Commit Graph
14 Commits
Author SHA1 Message Date
DottaandPaperclip 862a5758ba fix(agents): reduce hiring templates to role descriptions (#14985)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for work.
> - New agents receive role instructions from onboarding, the hiring skill, or a team package.
> - These sources repeat harness procedures and impose generic work policies.
> - They can crowd out the task and the harness instructions.
> - This pull request reduces those sources to short role descriptions.
> - It preserves configuration, skills, authentication, reporting lines, and approval controls.
> - The benefit is less repeated instruction text with explicit coverage for the default hiring path.

## Linked Issues or Issue Description

Refs #3307. The CEO template can impose a fixed delegation route instead of letting the agent choose how to fulfill the request. This change removes that route. It does not implement autonomous goal selection.

Related work: #14920 preserves stock Codex base instructions. #14948 reduces the generic manual and shared runtime prompts. #14961 improves native completion-tool descriptions. This PR is separate from those changes.

## What Changed

- Select only the short CEO `AGENTS.md` for new default CEO bundles. Keep the three former companion files as compatibility assets.
- Reduce the first-agent chief-of-staff prompt and coder, QA, UX, and security role examples.
- Reduce seven bundled team role bodies. Preserve their role, reporting, and skill metadata. Regenerate the catalog.
- Make hiring examples optional. Replace the long generic role manual with short role drafting guidance. Preserve explicit requester instructions.
- Add configuration and import coverage for native and legacy managed bundles, custom instructions, first-agent rendering, and catalog contents.
- Add an explicit-only hiring eval that starts from the production CEO default and checks one coder hire, independently computed JSON output, saved instructions, and worker reuse.
- Include the full prompt comparison and a separate three-request drafting simulation. Neither is a live provider comparison.

Prompt differences: [before and after](doc/plans/2026-10-02-hiring-template-prompt-diff.md). The CEO default falls from 1,897 to 20 words. The coder example falls from 652 to 18 words. Word counts describe instruction size, not outcome quality or billing.

## Verification

- PASS: 99 focused server tests and eight shipped-catalog tests.
- PASS: catalog generation and validation for four shipped teams.
- PASS: hiring skill validation.
- PASS: `pnpm -r typecheck`.
- PASS: `pnpm build`.
- INCOMPLETE: the full local `pnpm test:run` was stopped before rebase. Its original log is retained. This is not a completed full-suite pass. The full current-head GitHub CI workflow passed: https://github.com/paperclipai/paperclip/actions/runs/37073372419.
- PASS: `pnpm test:e2e:runner:typecheck` and `pnpm test:e2e:runner:unit` (63 files / 843 tests).
- PASS: discovery for the two new hiring cells, 50 existing everyday cells, and the full 438-cell catalog.
- PASS after rebase: 99 server tests, 11 catalog tests, 62 selected E2E support tests, and the E2E typecheck.
- PASS: all current-head PR checks at `57dcee147ed0b2d2e3cc657cd9e50fb16bf9ec25`: 51 successful check runs, two intentional Storybook skips, and successful Snyk status. Fresh Greptile is 5/5 with zero unresolved threads.
- PENDING follow-up: matched live hiring runs on frozen integration refs. No live outcome-quality or non-regression result is claimed from the configuration checks or this merge.

The new suite has two local native cells: Codex and ACPX Claude. It expects five provider turns per cell. It compares source-derived bundles, so the historical long templates remain admissible. Missing successful source-read receipts make a pair uncomparable. They do not establish a behavior regression or equivalence.

## Risks

- New default roles have fewer prescribed procedures. Live checks must determine whether a removed instruction was needed for an outcome.
- Existing custom and saved bundles keep their contents. The retained companion assets avoid a source-file compatibility break.
- Specialized Summarizer, Reflection Coach, and Wiki Maintainer prompts remain unchanged. Their product contracts need separate review.
- The generic non-CEO fallback reduction is in #14948. This PR alone does not provide its eight-word fallback.
- Configuration tests and drafting simulations do not establish live outcome quality. QA, UX, security, and chief-of-staff hiring behavior remain outside the new two-cell comparison.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`.

## Model Used

OpenAI Codex, based on GPT-6, with reasoning, code editing, shell tools, and delegated verification. The runtime does not expose the exact deployment model ID or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context to this change
- [x] I have specified the model used (with version and capability details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before requesting merge

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-02 17:56:42 -05:00
DottaandPaperclip 3c561642b4 fix(chat): resolve approvals and preserve unanswered questions (#14613)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents ask for decisions and optional details through cards in chat.
> - A clear approval in a message can leave the matching card pending.
> - An unanswered question can also block an unrelated later reply.
> - Decisions need a saved source message, while optional questions need
to remain answerable in history.
> - This pull request records conversational decisions and lets users
move on from questions and answer them later.

## Linked Issues or Issue Description

**What happened?**

Native Claude and Codex could act on approval in chat while the original
approval card stayed pending. Pending question forms stayed above the
composer, were absent from history, and could suppress later chat
replies. A late native question answer could wait for a finished run to
reconnect.

**Expected behavior**

The active agent records a clear approval or refusal against the exact
card and user message. Ambiguous replies do not grant consent. Users can
send another message without answering a question. The question remains
pending in history and can be reopened and answered later. The saved
answer reaches the agent.

**Steps to reproduce**

1. Ask an agent to propose work with a confirmation card, then approve
it in chat.
2. Check that the original card records that approval before work
starts.
3. Ask an interactive question, send an unrelated message, and reload.
4. Open the unanswered question from history and submit an answer.

Related work: #14408 added completion delivery. #14607 tests completion
reporting turns. Neither records conversational answers on approval
cards.

## What Changed

- Add a confirmation endpoint backed by a user comment, with schema
validation, OpenAPI discovery, and native Plan-mode access. Ask mode
remains read-only.
- Check company, active run, actor, current session, message provenance,
revision, and resolver policy. Save the decision and audit in one
transaction. Retries do not repeat effects. Emit resolution telemetry
after commit.
- Give fresh and resumed chat turns the actual pending confirmation
identities. Teach agents to save clear conversational decisions before
acting and to clarify ambiguity.
- Keep unanswered Agent Chat questions as compact history entries. A
newer user message closes the old form. Question cards never contribute
to composer pending counts or navigation, including after dismissing a
fresh form. The history card is the sole reminder; clicking it restores
that exact form and draft.
- Preserve Agent Chat questions when later messages or questions arrive.
Historical ordinary inputs no longer gate later chat replies.
Current-run requests, task execution, and governed approvals keep their
gates. Remove the special acknowledgement-publication proof helpers that
this rule replaces.
- Route answers to finished native runs through durable fresh-wake
delivery, with existing idempotency and source-question context. Settle
late replies against contiguous completed conversation turns and freeze
their history replay; failed, unhandled, and newly arriving messages
remain actionable.
- Add real-component Storybook scenarios, database and UI regressions,
and a three-turn native Claude/Codex E2E case. Capture distinct,
UI-ready screenshots and report the individual assertions.

## Verification

- Focused decision/publication/UI regressions after merging master: 288
passed; subsequent UI draft, failed-send, and conversation checks: 199
passed.
- Native question and durable delivery regressions: 106 passed,
including all four terminal run states and exactly-once late delivery.
Seven targeted regressions fail against the original implementation and
pass with the fix.
- Latest conversation/decision/native-delivery regressions after the
master merge: 121 passed. Covers completed progress, missing or failed
intervening turns, new messages during a late reply, stale sessions, and
frozen retry/replay boundaries. Four new assertions fail before the
ordering fix.
- E2E support suite after the master merge: 792 passed. Negative
controls reject expired cards, wrong questions/answers, stale or missing
replies, unrelated clarification forms, and unexpected tasks.
- The embedded-browser walkthrough caught one additional defect:
dismissing a fresh question still showed a composer badge. Both Cancel
and close-button regressions failed before the fix. The fix at
`65f2ade12` passes 170 chat-thread tests and 792 E2E support tests.
After merging master, 232 chat-thread/confirmation tests, server/UI
typechecks, and token gates pass. The preview and two-provider live E2E
pass at `e5512a206`; Greptile is 5/5 with zero unresolved threads at
that commit. All 55 checks are now successful at `e5512a206` (four
conditional checks skipped), including the aggregate verification gate
and clean-install canary test. The first attempt was interrupted by
simultaneous CI worker shutdowns; one failed-job rerun passed without
code changes.
- [Published
Storybook](https://d1p6rlowie26tp.cloudfront.net/storybook/branches/codex~2Fchat-approval-resolution/?path=/story/chat-comments-agent-chat-unanswered-questions--moved-on):
nine real-component scenarios. Manually exercised move on, reopen,
preserve draft, answer later, answer one of multiple questions, and a
custom mobile answer in the embedded browser. Retested fresh Cancel and
close-button dismissal in the updated build, then reopened and submitted
the preserved Green selection and inspected its answered receipt. Static
preview has no live model/backend; its callbacks are fixture responses.
- [First live
campaign](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36714504406-1/)
reproduced the late-answer completion-state defect on both providers
despite correct saved answers and acknowledgements. It also exposed a
valid imperative clarification rejected by the old oracle. Both issues
are fixed with regression controls; this failing run is retained as
evidence.
- [Four-cell
qualification](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36717804064-1/)
passed 4/4 at `2bf8a1009`: unanswered-question return and ambiguous
confirmation, each on native Claude and Codex. Inspected saved state,
source-message decisions, visible cards, and agent replies. Both
late-answer chats settled to waiting; no unrequested tasks were created.
[Final branch
rerun](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36719666238-1/)
passed 2/2 at `142630720`: the same unanswered-question journey after
merging master, plus an additional screenshot and browser assertion for
the actual late-answer acknowledgement.
- [Composer-reminder
E2E](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36727006818-1/)
passed 2/2 at `5b62c52d9`: native Claude and Codex, three turns each,
with explicit no-badge assertions before and after reload. Inspected
saved pending/answered state, both screenshots with a clear composer,
and actual Blue acknowledgements; all five behavioral matchers passed
per provider and neither created tasks. Cost coverage is partial; this
is bounded workflow qualification.
- [Fresh-dismissal
E2E](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36742773318-1/)
passed 2/2 at `e5512a206`: native Claude and Codex, including fresh
Cancel, clear composer, reopen, unrelated message, reload, late Blue
answer, and actual agent acknowledgement. All five behavioral matchers
pass per provider. Inspected the fresh-dismissal screenshots and saved
pending/answered identity; neither created tasks. Cost coverage is
partial (4/6 runs).
- Prior evidence remains available in [the earlier
campaign](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36642252725-1/).
Its early loading screenshot and overwritten final capture prompted the
UI-ready, distinct screenshot fixes.

## Risks

- The model interprets intent. The server verifies permission and
provenance; it does not infer consent from text. Ambiguous and unrelated
replies are not approvals.
- Historical questions can accumulate. They remain visible, pending, and
answerable; no automatic answer or expiry is invented.
- The change to completion gates is scoped to Agent Chat and ordinary
historical inputs. Current-turn and governed approvals retain their
existing controls.
- Live qualification is limited to the selected stories. Broader native
onboarding finalization remains separate work.
- No database migration. Telemetry adds no fields or values; the
contract and README document the commit boundary. Privacy review was
requested on the PR.

## Model Used

OpenAI Codex, GPT-6 family, with reasoning, repository tools, code
execution, and browser-test orchestration. The exact model ID and
context-window size are not exposed to this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 11:46:46 -05:00
DottaandPaperclip 83076d7e7c feat: return completed handoffs to Agent Chat (#14408)
Return completed Agent Chat handoffs through a durable outbox and scope each generated update to its supplied tasks. Add recovery, browser delivery, result access, and calibrated quality coverage.

Validated with two consecutive ten-case Claude/Codex campaigns, all CI checks, and a 5/5 review.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-29 10:25:52 -05:00
DottaandPaperclip 3f4f8b37ba fix: grade Codex clarification and refusal outcomes from evidence (#14570)
Grade clarification lists, obsolete unstarted wakes, and refusal cancellation from persisted evidence. Preserve execution and ownership assertions, add boundary regressions, and version the affected eval definitions.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-29 09:40:55 -05:00
0be2afcca6 feat(ui): improve task composer controls and pending input (#14322)
## Thinking Path

> - Paperclip lets operators assign tasks to AI agents and review their
work.
> - The task composer controls the next message and its assigned agent.
> - Operators needed a way to choose that agent's model and effort
without leaving the composer.
> - The old mode selector, upload button, and input cards made the
mobile composer crowded and hid normal messaging during a pending
decision.
> - Harnesses publish different model and effort capabilities, so the
picker must follow the selected agent.
> - This pull request adds one responsive composer flow, keeps pending
cards visible above it, and protects Codex ACP authentication in the
local test path.
> - Operators can choose run settings, send a message, and answer a
pending card as separate actions.

## Linked Issues or Issue Description

**Subsystem affected**

Task composer UI, issue thread interactions, Codex ACP credential
handling, and Storybook.

**Problem or motivation**

The composer did not expose model or effort for the selected agent.
Mobile actions wrapped poorly. Pending questions and confirmations
replaced the composer. A local Codex ACP test could also reuse host
authentication after the managed key was removed.

**Proposed solution**

Put assignee search, model search, exact model IDs, effort, and fast
mode in one picker. Use a mobile dialog. Replace the direct-upload plus
action and separate mode selector with an Add menu and removable Plan or
Ask chips. Place pending interaction cards above the usable composer.
Keep these cards pending after an ordinary message unless their creator
asks for comment superseding. Replace managed ACP auth files atomically
and isolate the test key from host credentials.

**Roadmap alignment**

ROADMAP.md does not list an overlapping composer milestone. This change
improves the existing task and review flows.

## What Changed

- Added the combined assignee, model, and effort picker to both task
composers. Search matches agent name, role, and harness. The server uses
a curated Codex list by default and honors instance-declared models.
Manual IDs remain available.
- Added an effort slider for known model capabilities, a conditional
Codex fast control, and reset. The picker opens in a modal on mobile.
- Added the Add menu for files, supported goals, Plan mode, and Ask
mode. Plan and Ask are exclusive removable chips. Keyboard mode cycling
remains available.
- Adjusted mobile spacing, avatars, wrapping, and Send placement.
Removed the composer divider.
- Moved pending question, confirmation, review, and related cards above
the composer. Ordinary comments now leave question and confirmation
cards pending by default. The onboarding prompt retains explicit comment
superseding.
- Updated the Storybook composer group with responsive states and the
production picker. Added UI, service, route, and browser regression
coverage.
- Isolated Codex ACP API-key authentication, skipped subscription auth
merge and shared-home copy-back for remote API-key runs, and replaced
the managed auth file atomically.

## Verification

- `pnpm -r typecheck` — passed on the final local head.
- `pnpm check:token-gates` — passed on the final local head.
- `pnpm exec vitest run server/src/__tests__/adapter-models.test.ts
ui/src/components/task-chat/ComposerRunSettingsPicker.test.tsx` — 31
tests passed, including role and harness search, declared Codex models,
and filtering general OpenAI models.
- `pnpm exec vitest run
server/src/__tests__/issue-thread-interactions-service.test.ts` — 74
tests passed.
- `pnpm exec vitest run
packages/adapters/codex-local/src/server/acp.test.ts` — 42 tests passed,
including remote API-key copy-back isolation.
- `pnpm test:run` — attempted locally; the embedded PostgreSQL test
database could not initialize on macOS. The isolated
`heartbeat-run-event-sequencing` suite reproduced that environment
failure. GitHub CI runs the full test matrix for this head.
- `pnpm build` — passed on the final head. `pnpm build-storybook` passed
after the last UI change; only server code, tests, and docs changed
afterward.
- Live local test drive — Codex ACP ran a task with a managed API key.
The test agent was restored to its default ACP configuration afterward.
- Review the interactive stories under the top-level Composer group with
`pnpm storybook`. Check a narrow desktop width and mobile Plan, Ask,
picker, and pending-question states.

## Risks

- A pending card stays open when an ordinary comment changes the
discussion. Its creator can set `supersedeOnUserComment: true` when a
new comment should replace it.
- Model and effort overrides persist on the task until reset or changed.
An unlisted manual model ID may fail when the provider runs it.
- Some harness catalogs do not report effort support. The picker hides
effort for those models.
- No database migration is required.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

OpenAI GPT-6 via Codex. This runtime does not expose the exact model ID
or context window to the task. The model used code execution and browser
tools.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: OpenAI Codex <codex@openai.com>
2026-09-28 23:02:39 +00:00
DottaandPaperclip cea8dda472 test: evaluate completion updates after native task handoffs (#13969)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Users can delegate work through onboarding and Agent Chat.
> - A completed task does not prove that its result reached the original
conversation.
> - Existing tests do not isolate completion after the source chat
becomes idle.
> - This pull request adds four explicit native-runner probes across
Claude and Codex.
> - The probes preserve the result and reply so we can separate delivery
failures from inaccurate answers.

## Linked Issues or Issue Description

Refs #13775. Refs #13813.

These evals extend native-runner qualification. They measure completion
updates before we choose a product change.

## What Changed

- Add the opt-in `completion-updates` suite with two stories for each
native provider.
- Test completion in the existing onboarding task flow and after an
Agent Chat handoff becomes idle.
- Gate the chat worker on a brief inside its managed project workspace.
Prove the source is idle before releasing the worker.
- Check durable task completion, saved output, a subsequent source
reply, and rendered access to the result.
- Preserve replies, task state, screenshots, run events, and a separate
semantic review rubric.
- Add grader regression tests and update the documented eval contract.
- Preserve the suites added on master and include four completion cases
in the 306-cell catalog.

Production behavior and prompts are unchanged.

## Verification

- Passed all 565 eval support tests across 45 files after merging
current master: `node node_modules/vitest/vitest.mjs run --config
tests/runner-e2e/vitest.config.ts`.
- Passed eval TypeScript: `node node_modules/typescript/bin/tsc -p
tests/runner-e2e/tsconfig.json`.
- Confirmed four selected cells: `node cli/node_modules/tsx/dist/cli.mjs
tests/runner-e2e/launch.ts --list --suite completion-updates`.
- Four-cell behavior campaign on source
`ad47cf1da2b1e36f19f4227cfeb53998720b0b5b`:
https://github.com/paperclipai/paperclip/actions/runs/36072337485.
- A screenshot-only follow-up waits for the restored source reply to
render after result-link navigation. Its one-cell Claude onboarding
verification passed on final head:
https://github.com/paperclipai/paperclip/actions/runs/36075716141.
Corrected report:
https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36075716141-1/.
The original four-cell onboarding screenshots caught navigation loading;
its saved reply evidence remains valid. The follow-up again found stale
wording: "That work will run next" was posted 38 seconds after the child
was Done. The four-cell campaign keeps its original source and
measurements.
- Published evidence:
https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36072337485-1/.
- Suite definition:
`afba4d85d6c53d9f64c08b37a2e9cc20481b78f5bd7e2fa012045e2c69444d9d`,
version 6. Models: native `gpt-5.6-sol` and `claude-sonnet-5`, local
execution, one attempt per cell. All four cleanup checks passed.
Onboarding billing coverage is partial; reported zero cost must not be
read as a free run.

| Story | Automated delivery/access | Separate semantic review |
| --- | --- | --- |
| Codex onboarding | Pass | Pass: accurate completion reply with an
accessible result |
| Claude onboarding | Pass | Fail: reply says it will save the note once
the task runs, after the note is already saved and the task is Done |
| Codex idle chat handoff | Fail | Worker completed and saved the note;
no completion reply during the full observation window |
| Claude idle chat handoff | Fail | Worker completed and saved the note;
no completion reply during the full observation window |

Both chat cases positively recorded the source waiting and the worker at
the brief gate before release. Both saved outputs include the brief-only
start time. The opt-in campaign is red because it exposes current
behavior. It is not a required merge gate. The PR does not fix that
product behavior. Semantic review is a recorded human/agent assessment
of retained evidence; it is not an automated prose-quality judge.
- Second campaign:
https://github.com/paperclipai/paperclip/actions/runs/36071065098. Codex
chat reached the idle boundary and completed its task, then received no
completion reply during the full window. Claude onboarding again
returned a stale handoff answer. Claude chat exceeded the prior
110-second handoff setup budget; this revision raises that bounded setup
window to 180 seconds.
- Retained baseline:
https://github.com/paperclipai/paperclip/actions/runs/36069427676.
Onboarding passed delivery/access for both providers, but Claude gave a
stale handoff answer. Chat cases stopped at fixture problems; they do
not establish a completion-delivery failure. This revision fixes the
workspace path and competing reference requirements.
- On the previous head `4023a2a3c28d45c9eb2c42d452ce99ffba5c7b73`, 54 PR
checks passed and two were skipped, including typecheck, tests, and
build. Broad checks ran in CI, not locally. That head received Greptile
5/5 with no unresolved findings. The unchanged mobile
repository-settings browser test passed on one targeted retry after a
detached/disabled Save-button timeout.

- Merged current master in `9b4491e1f` and resolved the catalog-count
conflict. Eval support tests and eval TypeScript pass locally. All
individual CI jobs passed on this merge commit, including build,
typecheck, server tests, runner checks, and browser shards. The final
aggregate check also passed: 54 checks passed and two were skipped.
Greptile reviewed this exact commit at 5/5 with no unresolved findings.

## Risks

- These explicit probes can expose current product failures. They do not
change the default paid test selection.
- Mechanical delivery and result access do not establish answer
accuracy. The preserved reply still requires semantic review.
- A fixture failure before the idle boundary or worker completion cannot
establish a completion-update failure.
- The handoff setup window lasts three minutes. The worker brief wait is
bounded at four minutes. The observation window lasts two minutes after
worker completion. It retains later replies without erasing earlier
accessible delivery.

## Model Used

OpenAI Codex, GPT-6 (`gpt-6-astra`), with reasoning, repository
inspection, code execution, and GitHub tool use. The runtime does not
expose the context window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 09:58:14 -05:00
DottaandPaperclip f56aee5423 fix(evals): capture complete durable run event streams (#13977)
## Thinking Path

> - Paperclip manages AI agents and records their durable work outcomes.
> - Product E2E checks those outcomes through the browser and public
API.
> - A successful long run can emit more than 1,000 durable events.
> - The harness read one page and missed the later completion evidence.
> - This pull request reads every page before it checks runtime
invariants.
> - Invalid or incomplete capture still fails. A missing page cannot
produce a pass.

## Linked Issues or Issue Description

Related: #13882 (Grok qualification). A duplicate search found no
existing event-pagination fix.

**What happened?**

The local structured-question case in [campaign
36071063537](https://github.com/paperclipai/paperclip/actions/runs/36071063537)
completed the task and passed its six outcome matchers. It failed native
runtime invariants because the capture contained exactly 1,000 events.
The last captured event preceded the run's completion by more than a
minute. The API caps each response at 1,000 rows. The harness did not
request the next page.

**Expected behavior**

Read the complete durable event stream through the public API before
checking semantic-result and terminal-event counts. Reject incomplete or
malformed evidence.

**Steps to reproduce**

1. Complete a native task that emits more than 1,000 durable events.
2. Place the semantic-result and terminal events after row 1,000.
3. Capture the run with the Product E2E harness.
4. Before this fix, the invariant checker sees only the first page.

**Paperclip version or commit**

Observed at `4196a4cd76db434854b679035e4146c7f69689ce`. The same
single-page capture exists on master. The original failed result remains
unchanged; missing historical tail evidence is not reconstructed or
graded as a pass.

## What Changed

- Add a bounded event collector that advances through the public
`afterSeq` cursor.
- Use it for task success/failure evidence and shared chat run evidence.
- Reject invalid pages, missing or non-increasing sequence numbers,
repeated cursors, failed later requests, and an exhausted page limit.
- Test completion events beyond the first page, exact page boundaries,
and malformed evidence.
- Correct the existing Everyday catalog test from 38 to the maintained
47 cells. The suite stays explicit-only.
- Document the complete-capture requirement and its bound.

## Verification

- Product harness typecheck passes.
- The full credential-free harness suite passed 477 tests. After adding
the chat integration regression, all 34 chat evidence tests pass.
- Seventeen pagination tests cover the valid tail and malformed-evidence
cases.
- `git diff --check` passes.
- [Full repository
CI](https://github.com/paperclipai/paperclip/actions/runs/36080683422)
passes at `aed6f79c089226be79e75dcf390969561ec4f787`: 52 successful
checks and two intentional skips, including Rust, typecheck, build,
server tests, and browser shards. Greptile gives this exact head 5/5
with no findings. No local Docker, Rust build, browser suite, or model
invocation was used for this change.
- [Live Grok
requalification](https://github.com/paperclipai/paperclip/actions/runs/36080870743)
is running on combined source
`1b0551bb7c8de3c54f4bee64dbe2c88328b3645e`, with the runner, scheduler,
and evidence fixes. Its full credential-free harness passes all 493
tests and typecheck. Live results are pending; the original campaign
remains a failed measurement.

## Risks

Long runs need more read-only API requests and larger private evidence
files. Capture stops with an explicit error after 100 full pages. This
changes neither production APIs nor provider behavior. It does not relax
an invariant or change a historical grade.

## Model Used

OpenAI GPT-6 through Codex, with repository tools and code execution.
The exact serving identifier and context-window size are not exposed in
this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-25 10:06:33 -05:00
DottaandPaperclip 5842185e4f fix: surface native cleanup quarantine and add chat qualification evals (#13775)
## Thinking Path

> - Paperclip lets people manage AI agents and their work.
> - Agent Chat needs reliable native execution before native runners
become the onboarding default.
> - Existing stories covered idle reassignment and controller restart,
but not an executing worker handoff or worker process loss.
> - Status answer tests also need to reject stale claims and invented
facts.
> - This pull request adds six opt-in full-stack cells with independent
state assertions and retained evidence.
> - The probes exposed a misleading Retry across server projection and
recovery-banner paths; the fix reports the blocked recovery honestly.
> - The tests preserve failures without changing recovery policy,
production prompts, or onboarding defaults.

## Linked Issues or Issue Description

Refs: #13762. Related: #13765 (Retry targets the latest failed attempt),
#13753 (task context ownership), #13746 (native recovery work).

## What Changed

- Add active reassignment with saved draft and plan preservation,
old-worker cancellation, and successor completion checks.
- Preserve recovery-needed projection when native cleanup fails before
its coordinator exists, refuse a generic retry that would immediately
fail again, and replace the recovery banner's misleading Retry with
Inspect run.
- Add verified local worker process loss with a required successful
continuation; retain a failing qualification result when recovery is
unavailable, while independently verifying the UI/API refuse doomed
retries.
- Add two-turn factual answer checks for current blockers, stale claims,
inactive backlog work, and unknown facts. Retain prose for separate
semantic review.
- Add positive and negative oracle calibration and document fault
isolation, cleanup, billing, and qualification limits.

## Verification

- Eval TypeScript check passes.
- All 442 eval support tests pass locally. The 89 focused server tests
and server typecheck pass. Six recovery-banner UI tests and token gates
pass.
- Initial new-cell campaign:
https://github.com/paperclipai/paperclip/actions/runs/35657128077. All
six results are retained; four failed on fixture-contract issues and two
exposed real worker cleanup quarantine.
- All 26 existing native onboarding cells:
https://github.com/paperclipai/paperclip/actions/runs/35656761484 (26/26
passed on master 846336e5a, all cleanup passed).
- Intermediate handoff/fault campaign:
https://github.com/paperclipai/paperclip/actions/runs/35657945095 (four
retained failures: two overly strict draft oracles, two real crash
quarantines).
- Final active handoff:
https://github.com/paperclipai/paperclip/actions/runs/35659014397 (2/2
passed on cf6d4ae3a; both cleanup passed).
- Clarified answer-quality fixtures:
https://github.com/paperclipai/paperclip/actions/runs/35658262695 (2/2
passed on 4a26f10be; both cleanup passed; all four answers semantically
reviewed).
- Quarantine guard regression campaign:
https://github.com/paperclipai/paperclip/actions/runs/35659580100 (both
API requests correctly refused with 409/no second run, but exposed a
separate misleading Retry in the recovery banner and a fixture wait on a
non-admitted run).
- Final quarantine guard verification:
https://github.com/paperclipai/paperclip/actions/runs/35661067305
(147e42f7e: both providers verify Inspect run/no Retry, HTTP 409, one
retained run, unchanged saved plan, and successful disposable cleanup.
Both evals intentionally remain red with
`worker_crash_recovery_unqualified`; no successful continuation exists).
The preceding campaign 35658772755 never ran provider cases because
GitHub artifact finalization returned HTTP 403.
- Full repository CI passes on 147e42f7e: typecheck, tests, build, and
browser gates. One unchanged local-service-supervisor readiness test
failed initially; its six-test file passed in isolation and the failed
shard passed on its single rerun. Latest-head rollup: 54 successful, 2
intentionally skipped, no failed or pending checks. Greptile is 5/5 with
zero unresolved findings.
- See tests/runner-e2e/QUALIFICATION-2026-09-21.md for retained attempts
and semantic review. Published reports:
[onboarding](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35656761484-1/),
[handoff](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35659014397-1/),
[grounded
answers](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35658262695-1/),
[crash guards and unqualified
recovery](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35661067305-1/).

## Risks

- Paid cells are explicit-only and local-only. The fault fixture signals
only the exact native run PID after checking its identity.
- Live worker-loss probes currently fail on cleanup quarantine for both
providers. The eval must remain red until there is a usable recovery,
even when preservation and refusal checks pass. Verified cleanup with a
fresh attempt versus exact-session resume remains a product decision.
- Structured facts alone do not qualify prose quality; semantic review
remains separate.
- Onboarding uses the existing runtime switch after the real wizard and
before provider execution. Native UI selection and public defaults
remain unchanged.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, tool use, and code
execution. The exact served snapshot and context-window size are not
exposed in this session.

## Checklist


- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 18:13:01 -05:00
DottaandPaperclip 846336e5a0 test: harden agent chat setup, interruptions and restart evals (#13762)
## Thinking Path

> - Paperclip lets people manage agents through ongoing conversations.
> - Chat users can change instructions while a provider is already
working.
> - Existing chat evals wait for each turn to settle before the next
message.
> - They cannot prove delivery during active work or the saved effect of
a correction.
> - Existing fixtures also enable Agent Chat through the API rather than
the settings UI.
> - This PR adds bounded browser workflows and checks their persisted
outcomes.

## Linked Issues or Issue Description

Refs #13741, #13752, #13750.

**What happened?**

The chat suites cover planning, delegation, status, and recovery. They
lack active-turn follow-ups and the experimental settings lifecycle. A
sequential conversation can pass even if messages sent during work are
lost.

**Expected behavior**

A follow-up submitted during a provider turn survives and affects the
final reply. A changed launch day appears in the saved plan. Disabling
Agent Chat rejects new messages while preserving history; re-enabling
resumes the same conversation.

**Steps to reproduce**

Run the explicit `agent-chat-stories` suite. It selects three local
cases for each native Claude and Codex profile. An ordinary provider
command waits for a fixture brief file so the browser can send the
follow-up at an observed active-run boundary.

## What Changed

- Add six opt-in Product E2E cells for settings, active follow-ups, and
plan corrections.
- Drive experimental settings through the UI and verify disabled sends
are rejected by the public API.
- Use a bounded file wait in the actual isolated agent workspace, with
provider-written readiness and an undisclosed brief reference.
- Grade persisted user messages, final replies, native run outcomes, and
exact saved plan fields.
- Accept active-turn steering or one queued successor; reject lost
input, duplicate input, and stale outputs.
- Allow one steered run or two sequential runs throughout the shared
harness, while preserving exact counts for other cases.
- Require a single marker-bearing response attributed to the final
provider run.
- Unload the development browser client before restarting the server,
avoiding reconnect/navigation races without weakening the post-restart
memory check.
- Add browser regressions for restart isolation and asynchronously saved
settings switches.
- Document prepared-agent setup, native onboarding limits, and the
separate API-tool rollout gate.

## Verification

- Eval TypeScript check passed.
- Eval support suite: 436 tests passed in 39 files.
- New oracle calibration: six tests passed, including plausible invalid
outcomes.
- Browser support regressions: seven tests passed; the restart
regression was observed failing before the fix.
- Catalog discovery selects exactly six local native cases and leaves
default paid selection unchanged.
- [Consolidated existing native chat
report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35643286055-1/):
master `b82661b56`, 33/34 passed, all cleanup passed. The failure was a
browser navigation timeout across restart; the page request returned 200
and the chat rendered.
- [Nine targeted restart/replay
cells](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35645850088-1/)
passed on `1fe2fe275`, including the original failure, across native
Claude/Codex and local/Daytona; all cleanup passed.
- [Initial six-story
campaign](https://github.com/paperclipai/paperclip/actions/runs/35644832817)
retained all six failures: asynchronous switch assertions, unavailable
fixture paths, and rich-text escaping in raw command comparisons. The
corrected fixtures preserve the same behavioral assertions.
- [Six-story campaign
v2](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35646270035-1/)
on `8232773a0`: 4/6 passed (both settings cases and both Claude
interruptions). Codex could not see the host-temp fixture outside its
workspace; this failed before follow-up delivery was exercised.
- [Four affected interruption
cases](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35647760635-1/)
all passed, including cleanup, on definition v3 / `ad6ac0545`. Files
live inside the actual agent workspace and the observed run workspace is
verified. Both providers saved Friday in the real plan with the
undisclosed brief reference; follow-ups persisted while the original run
was active. Together with both unchanged settings cases from v2, all six
new scenario variants have passing live evidence.
- Final head `ad6ac05456646c09d3452e320279457625353948`: 54 successful
checks, two intentional skips, zero pending/failing checks; mergeable
and clean. Fresh Greptile 5/5, zero unresolved findings.
- Full typecheck, tests, build, and browser CI passed remotely. One
earlier head encountered a signoff-policy browser timing failure; the
final head passed that shard.
- Local pnpm wrapper could not fetch its version/signature metadata in
the restricted environment; local eval checks used the installed Node
executables. Repo-wide validation was completed by GitHub Actions.

## Risks

These are eval-only changes. The file wait is a timing fixture in the
isolated agent workspace, not a production runner hook. Native Codex
host-filesystem isolation stays unchanged. It has a two-minute limit and
is released in `finally`. The prepared-agent settings case is not full
native onboarding: the wizard currently offers legacy adapters. The
disabled-entry assertion uses full document navigation, which clears the
prior React Query cache; preserved history is checked through the public
API and re-enabled chat. No production prompt, rollout default, adapter
behavior, or credential policy changes. Active-task reassignment and
worker-crash recovery remain outside these new cases.

## Model Used

OpenAI Codex, GPT-6, with repository tools and code execution. The exact
deployment model ID and context window are not exposed in this task.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 16:18:06 -05:00
DottaandPaperclip 9d19f98b50 fix: harden native chat recovery and add coordination evals (#13741)
## Thinking Path

> - Paperclip helps people manage AI agents and their work.
> - Agent chat uses native runner sessions to plan, delegate, and track
that work.
> - A user can press Stop while the native session is still starting.
> - The server can acknowledge that Stop without dispatching it, then
let the session submit a turn.
> - This leaves chat recovery waiting for an execution that the user
expected to stop.
> - This PR waits for the startup handle, dispatches cancellation, and
prevents a late startup from submitting a turn.
> - New full-stack evals check the resulting records and outputs across
Claude and Codex.
> - Those evals also exposed missing ACPX readiness fields, unbounded
polling, and an old-run identity check that rejected valid warm
handoffs.

## Linked Issues or Issue Description

**What happened?**

Stop during native startup could record an acknowledged cancellation
with `dispatched: false`. The provider could then begin work. A
subsequent `/new` stayed queued. A remote Claude follow-up also
exhausted the command journal while probing warm-session readiness: ACPX
never returned the readiness fields required by the shared transport.
Once readiness worked, attachment incorrectly compared the next run
descriptor against the old run ID. The 25 ms polling loop could issue
4,800 commands during its two-minute wait, beyond the 500-command bound.
The existing chat eval treated lifecycle logs as proof of an active
provider turn, so it did not distinguish startup cancellation from
active-turn cancellation.

**Expected behavior**

A Stop during startup must reach the pending session. A late session
must not submit a prompt after Stop. Recovery must retain control when
startup exceeds the bounded wait. Chat evals must check saved task
state, document contents, worker identity, account binding, and
duplicate effects.

**Steps to reproduce**

1. Start a native Claude or Codex chat turn.
2. Press Stop after process startup is requested but before the provider
turn starts.
3. Send `/new`, then send a fresh message.
4. On the affected base, cancellation can be acknowledged without
dispatch and the reset stays queued.

**Paperclip version or commit**

The live Claude baseline reproduced this on `29d6b3509`. The branch also
includes master commit `0f5fafe16`.

Related work: #13678, #13686, #13693, #13291, #13738. A separate runner
reliability branch also contains a startup-wait fix. Its overlap must be
reconciled before merging; this branch additionally prevents prompt
submission after a late startup.

## What Changed

- Wait for a pending native startup before acknowledging a run-scoped
Stop. Preserve the existing recovery error when that wait expires.
- Keep a Stop guard on startup. Cancel a late handle before it can
submit a provider turn.
- Add regression tests for normal handle publication and publication
after the Stop deadline.
- Back off blocked warm-attachment probes. Keep the fast two-snapshot
barrier, fail closed, and record changed blockers.
- Add red/green tests for delayed readiness, persistent blockers,
alternating readiness, and readiness near the deadline.
- Publish ACPX readiness and blockers. Preserve the old authority’s
event acknowledgement barrier; only settled sessions can proceed to
attachment.
- Bind warm ACPX descriptors to the validated next authority while
retaining old-run event correlation until activation. Preserve session
identity and provider profile checks.
- Exercise two consecutive run rotations through a qualified fake
sidecar, verifying checkpointing, provider identity, pre-activation
rejection, and new-run work admission.
- Separate startup and active-turn cancellation checkpoints in the
browser eval.
- Add 18 explicit native chat eval cells: 12 local and 6 Daytona cells
across Claude and Codex.
- Cover hiring and reuse through managed AI accounts, source-based
review, current blocked-task status, request replay after a lost HTTP
acknowledgement, server restart continuity, and Stop/reset continuity.
- Use ordinary production agent instructions. Enable API tools only for
the two coordination cases that need them.
- Calibrate the matchers with invalid records and outputs. Require
remembered context after restart and a structured status snapshot that
distinguishes the current blocker from history and task status from
active execution. Compare the public issue mutation contract and
relationships during read-only reporting. Preserve before/after source
records in failed eval evidence.
- Fix the lost-ack browser harness and verify it against a real HTTP
server. Check the chat composer after restart instead of waiting for an
unrelated document lifecycle event.
- Document the scope and limits of each case.

## Verification

- The startup regression failed on the unfixed executor and passed after
the fix.
- `pnpm test:e2e:runner:typecheck` passed.
- `pnpm test:e2e:runner:unit` passed: 424 tests in 37 files.
- `pnpm exec vitest run
server/src/services/native-runtime/native-session-executor.test.ts`
passed: 385 tests.
- [Baseline live
campaign](https://github.com/paperclipai/paperclip/actions/runs/35608208868):
Claude Stop reproduced the bug. Codex Stop and Claude hire/reuse passed.
Codex delegation was blocked by provider capacity.
- [Eval-only startup
campaign](https://github.com/paperclipai/paperclip/actions/runs/35609479786):
both providers failed as expected. Both persisted `dispatched: false`
and left `/new` queued.
- [First fixed
campaign](https://github.com/paperclipai/paperclip/actions/runs/35610533706)
on `c9e95797d`: 10/18 cells passed. Startup Stop passed for both
providers. Failed cases exposed eval harness defects and remote
continuity failures. All attempts remain available.
- [Original workflows and stronger memory
checks](https://github.com/paperclipai/paperclip/actions/runs/35611896649)
on `c04324fab`: 9/12 passed. Reassignment, local restart memory, and
startup Stop passed for both providers; Codex remote restart passed.
Claude remote restart exposed the missing readiness contract. Two Codex
planning cells hit provider capacity.
- [Unchanged-model
retry](https://github.com/paperclipai/paperclip/actions/runs/35613854548):
Codex planning and backlog creation both passed.
- [18-cell campaign with ACPX
readiness](https://github.com/paperclipai/paperclip/actions/runs/35614586963)
on `6a98ef743`: 16/18 passed, including all local/remote Stop and
committed-send cases. Claude remote continuity exposed the
next-authority check, now fixed. Codex hiring produced its checklist,
but the runner redacted the requested marker after it appeared as
“Tracking token: …”. That content-redaction policy is unchanged and
remains an explicit limitation.
- [Structured status
grading](https://github.com/paperclipai/paperclip/actions/runs/35614954725)
on `50448c228`: both providers passed on their first attempt, including
cleanup.
- [Complete read-only state
grading](https://github.com/paperclipai/paperclip/actions/runs/35616089011)
on `551e13892`: both providers passed.
- [Final ACPX handoff and hiring
retry](https://github.com/paperclipai/paperclip/actions/runs/35617045456)
on `cbd637587`: all three Claude Daytona cases passed (restart
continuity, active Stop/reset, and lost-ack replay). Codex hiring
reproduced the content-redaction failure: the saved checklist contained
`Tracking token: [REDACTED]` instead of the required business marker.
All four cases completed cleanup successfully. [Published
report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35617045456-1/).
The only subsequent commit adds the qualified-sidecar integration test;
production code is identical to this live proof.
- `pnpm test:e2e:runner:browser-support` passed: 5 browser tests without
paid models.
- Runner TypeScript typecheck passed. All 5 warm-readiness tests pass;
two failed with the prior fixed-rate loop, and the late-readiness test
failed before the pacing correction.
- ACPX readiness and warm-identity regressions each failed before their
fixes. All 292 runner-core Rust library tests passed. The
qualified-sidecar integration test passes. Rust formatting is checked.
- Status-grader regressions for misleading historical mentions and
previously unchecked mutations each failed before tightening the oracle
and pass now.
- [Latest-head
CI](https://github.com/paperclipai/paperclip/actions/runs/35617522307)
passed on `a4093c8f1`: full build, type checks, test partitions, browser
E2E, and native runner checks. Two unrelated tests initially failed
(Sentry fixture release attribution and local-service fixture
readiness); both passed locally together (35 passed, 5 optional SDK
tests skipped) and on the failed-job retry. No changes were made to
those tests.
- Greptile reviewed `a4093c8f1` at 5/5; both earlier findings are fixed
and all review threads are resolved.
- The paid live suite is not fully green: the reproducible
content-redaction case remains red. This is separate from the passing PR
merge checks. No production content-redaction, prompt, model, or
completion-policy change is included.
- Managed-account hiring and review cases explicitly enable API tools;
these do not qualify default new-user onboarding.

## Risks

- Stop can wait up to 30 seconds for startup, then use the existing
pending-recovery path. This does not prove that remote cleanup has
finished.
- Blocked warm readiness adds up to 750 ms between later probes with the
two-minute remote budget, or about 32 ms with the default five-second
budget. Ready sessions retain the short second barrier.
- Paid evals can fail because of provider capacity or agent decisions.
Each failure needs evidence-based classification.
- The HTTP request replay case checks comment idempotency and duplicate
effects. It does not prove replay safety for an ambiguous provider tool
call.
- The new suite is opt-in. It does not increase the default paid
campaign.
- No production prompts or model selection change. Review-handoff
behavior and content-redaction policy remain separate product decisions.
The latter can remove harmless business content that looks like
credential syntax; the failing attempt is retained.

## Model Used

OpenAI Codex, GPT-6, with repository tools and code execution. The exact
deployment model ID and context window are not exposed in this task.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 10:35:38 -05:00
DottaandPaperclip 7bc03e0acd feat(runner): default harnesses to full auto and support task reassignment (#13686)
## Thinking Path

> - Paperclip lets people manage AI agents and their work.
> - Agent Chat uses native runners to save plans and coordinate tasks.
> - Provider defaults differed across harnesses and could stop
unattended work at a second permission gate.
> - Agents also lacked a dedicated tool to move existing work to another
agent safely.
> - This change defaults native providers to full automatic permission
for provider tools and connected tools.
> - A guarded reassignment tool preserves task identity, stops the
previous run, and schedules the new owner once.
> - Codex and Claude chat acceptance tests now use production permission
defaults.

## Linked Issues or Issue Description

**Subsystem affected**

Native runner, ACPX Claude permission policy, task authority, and Agent
Chat acceptance tests.

**Problem or motivation**

A user can authorize an agent to save a plan or create a task, but
Claude's default provider gate can still stop that action. Reassignment
needs a dedicated operation that preserves context and avoids concurrent
owners or unintended recovery runs.

**Proposed solution**

Default Claude/ACPX to `approve-all`, OpenCode to `allow`, and Codex to
`never`. Apply the defaults at configuration, execution, fresh-session,
resume, driver, and proxy boundaries. Keep explicit permission settings
and server-side company, claim, task-mode, and approval checks. Add
`reassign_task` with version checks, durable idempotency, audited
cancellation, and guarded successor scheduling.

**Alternatives considered**

A Paperclip-only allowlist still blocks provider tools and other
connections during unattended work. Full automatic permission is the
requested product default. Recreating a task discards its identity and
history. Updating assignment without stopping the previous run can leave
two agents working on the same task.

**Roadmap alignment**

This extends the existing planning, delegated work, governed tool
access, and recovery features. It adds no new service or schema
migration. Recent related tasks and open PRs were checked for duplicate
work.

**Additional context**

Related: #13678 (Agent Chat tools and recovery), #13677 (remote runner
startup). The stacked legacy-adapter companion is #13693. This also
fixes the deployed-server artifact fallback needed to stage the current
runner binary.

## What Changed

- Default Claude/ACPX to `approve-all`, OpenCode to `allow`, and Codex
to `never`, including missing settings at direct driver and proxy entry
points. These defaults cover provider tools and connected tools.
Preserve explicitly configured restrictive modes.
- Include assigned approval reads using canonical side-effect
classifications, so verifying a recorded approval does not trigger
another provider gate. Paperclip approval decisions still enforce
controller authority.
- Carry the new permission mode through server configuration, execution
contracts, recovery identity, TypeScript, and Rust. Keep
`approve-paperclip` as an optional restricted mode, with exact SDK rules
and closed unknown requests. It is not a default.
- Add `reassign_task` to the semantic catalog, controller, mock
authority, and generated contracts.
- Guard reassignment with company authorization, expected owner and
version, protected-state checks, and durable retry receipts.
- Honor explicit backlog task creation atomically with the initial plan,
without scheduling a wake. Preserve backlog holds regardless of
dependency readiness.
- Stop active work before changing ownership. Restore the prior owner
through a guarded, idempotent wake if final handoff validation fails.
Keep intentional reassignment stops out of failure recovery. Preserve
backlog and blocked states without waking them early.
- Add authorization, concurrency, replay, stop, and permission boundary
regressions. Add Codex and Claude chat reassignment cases and run native
chat cases with production defaults.
- Clarify shared runner guidance: save plans and Paperclip documents
directly with `write_document`; create and register a local file only
when a downloadable file is requested.
- Document provider defaults and the operator choices for existing
agents.

## Verification

- Current head `d82fbb0f03546d27cecf072250e4172e0b1ee662`: **55 checks
passed**, with two intentional skips. [PR
checks](https://github.com/paperclipai/paperclip/pull/13686/checks).
- Greptile reviewed that exact head at **5/5**. The security reviewer
acknowledged the intended full-auto default, and the acknowledged
discussions are resolved.
- Full workspace `pnpm -r typecheck` and `pnpm build` passed locally
after rebasing onto current master. Targeted adapter/server, runner,
API, default/resume, and heartbeat configuration tests passed.
- **All six real-provider acceptance cases passed on their first
attempt, with cleanup passing:** plan handoff, task reassignment, and
backlog creation/status, each on native Claude and Codex. Evidence
records Claude's effective `approve-all` mode. [Campaign and
downloadable
evidence](https://github.com/paperclipai/paperclip/actions/runs/35469926548).
- The live campaign tested combined revision
`a37881c824dcd7170380fc4b788732fc743e5da7`. The final PR heads add only
a heartbeat test expectation correction; application code is unchanged
from that live-tested revision.
- The campaign's result-enforcement job passed. Its separate report
publisher failed because the trusted workflow's `patchedDependencies`
configuration differs from its frozen lockfile. All six results and
screenshots remain available as GitHub artifacts. The overall manual
workflow is red for this publishing failure.
- Full-suite coverage is supplied by the passing CI partitions. The
separate unsharded local run was stopped after the corresponding CI
partitions passed; it is not counted as a completed local run.
- Reassignment tests cover stale state, cross-company access, denied
authority, cancellation failure, compensating wake, and idempotent
retries. Backlog tests verify the original creation audit, saved plan,
exact task count, and absence of task-bound runs.

## Risks

- Agents with no explicit permission mode now receive full provider tool
permission, including connected tools. This is a deliberate broad
default. Existing explicit restrictive modes still apply. Controller
authorization, company isolation, workspace boundaries, and Paperclip
governance remain in force.
- Reassignment crosses run cancellation and task ownership transactions.
Durable stop intent, revalidation, audit receipts, and guarded queue
dispatch cover interruptions and retries.
- The new permission enum requires a current runner artifact. The remote
artifact fallback uses the same resolved controller binary for upload
and execution.
- Live provider behavior remains subject to the selected model. Targeted
live results do not qualify the full catalog.

## Model Used

OpenAI Codex, based on GPT-6, with code execution and repository tools.
The exact deployment model ID and context-window size are not exposed in
this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-19 16:49:18 -05:00
DottaandPaperclip 43acbcc398 fix(runner): preserve sessions and complete question and approval continuations (#13655)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The native runner connects task state to provider sessions.
> - Follow-up turns must retain provider memory and carry new user
direction.
> - Lost session IDs caused repeated context and extra input tokens.
> - Native question answers and approval races could leave valid work
blocked.
> - This pull request repairs those paths and adds regression coverage.
> - Agents can continue accepted work without repeating the conversation
or losing the user's answer.

## Linked Issues or Issue Description

Refs #13574. That merged PR shortened continuation prompts and moved
question instructions into tool documentation. This change preserves
sessions and fixes failures exposed by broader testing. Related runtime
work: #13408 and #13410.

**What happened?**

Native follow-up turns could lose the provider session ID. Completion
guidance could replace the original task with its latest comment. Claude
native questions could remain pending after the user answered. Approval
during a running tool call could suspend the run before the tool
response arrived. Onboarding and chat handoff instructions also caused
repeated planning or missing plan documents.

**Expected behavior**

Reuse a valid provider session. Send only new events when that session
already has the history. Preserve the task requirements and apply later
user direction. Store the question answer and deliver it to the waiting
run. Finish governed tool responses before suspending. Execute the
accepted plan without asking for the same approval again.

**Steps to reproduce**

Run the continuation, local-session-integrity, first-task, and
agent-chat suites with native Codex and Claude. Include
provider-question-bridge, accept-while-running, and plan-handoff.

**Paperclip version or commit**

This branch is based on master d54b75011. The active full catalog run
tests 4e75881db. Later review fixes have separate regression coverage.

**Deployment mode**

Isolated local instances and Daytona sandboxes in the existing Runner
full-stack E2E harness.

## What Changed

- Retain provider session identity across turns and late usage
snapshots. Send new continuation events on session reuse, with full
context available for a fresh session.
- Preserve task requirements and later direction in completion guidance.
Return the current contract revision after a stale completion
submission.
- Bridge native Claude questions to saved Paperclip cards. Submit
answers through the saved card and resume the same run.
- Delay governed suspension until tool results settle. Add a
deterministic test barrier for approval during an active run.
- Clarify free-text question examples, explicit onboarding plans, and
execution of accepted chat plans.
- Fix continuation readiness, verified output evidence, and declared
screenshot collection.
- Qualify the legacy Claude test CLI at 2.1.277. The old 2.1.19 CLI did
not discover mounted skills. Update the existing workflow pin and
isolated launcher together.
- Refresh the Daytona image lockfile integrity pin after reviewing
master patch updates.
- Carry continuation mode as runtime metadata instead of inferring it
from user-visible text. Install the test Claude CLI without lifecycle
scripts.

## Verification

- Targeted paid verification: 20/20 cases passed across
local-environment campaigns before the rebase.
[Report](https://pages.paperclip.ing/runner-e2e-seven-fixes-35397904249/).
- Harness checks: 379 unit tests passed; harness typecheck passed.
- Latest-head PR checks: 55 passed, two intentionally skipped. Greptile
is 5/5; the security scan passes.
- Review regressions: 350 executor tests and 204 session/driver tests
passed. A script-free Claude install was verified with the actual CLI.
- Full catalog, including the explicit-only everyday suite: [run
35417932353](https://github.com/paperclipai/paperclip/actions/runs/35417932353).
Completed: **164/205 passed; 41 failed**. [Full dashboard and failure
investigation](https://pages.paperclip.ing/runner-e2e-full-catalog-35417932353/).
Includes 204 case artifacts and one pre-case GitHub authorization
timeout; missing evidence is not scored as a pass. The full run tested
`4e75881db`; Final-head metadata/CLI smoke cases both passed. In the
separate [six infrastructure
retries](https://github.com/paperclipai/paperclip/actions/runs/35419769343),
the GitHub timeout case passed and all five Docker preflight failures
repeated. [Follow-up
dashboard](https://pages.paperclip.ing/runner-e2e-full-catalog-35417932353/follow-up/).
- Full local typecheck and build passed on the rebased branch. The full
local unit run completed with 657 passing files, two test timeouts and
one suite setup timeout. All three affected files passed when rerun in
isolation (84 tests). The first full local run was not clean.
- Focused regression coverage includes the live question bridge,
same-run response delivery, UI routing, stale revisions, approval
overlap, and session reuse.

## Full-catalog follow-ups

- Test infrastructure: 14 Claude everyday cells probe an absent host
CLI; six cells failed pre-task GitHub/Docker qualification (GitHub
passes on retry; all five Docker cases repeat; the workflow preflight
allowlist omits their case IDs); five ACPX Codex cells cannot create
sandbox namespaces.
- Runtime: four OpenCode completion-criteria mismatches masked by
shutdown errors, one service-approval suspension failure; three Daytona
recovery failures encounter existing skill files; one duplicate
completion wake.
- Confirmed test defects: question pagination and a noncanonical plan
document key.
- Product/behavior: mismatched visible/required question sets, an
attachment instead of the requested task document, one lone-option
onboarding question, early completion instead of review, and a Codex
Mini completion-schema failure.
- The report job itself fails on trusted master’s stale patch/lock
configuration. The linked report is rebuilt with the shared renderer
from original cell results and public fixture screenshots; it excludes
private snapshots, logs and traces.

These are investigated follow-ups, not silently regraded passes.
First-task passed 51/52. The PR checks are green independently of the
broader catalog’s behavioral/infrastructure failures.

## Risks

- Session reuse depends on a valid provider identity and context
coverage. Fresh-session fallback and reset tests cover this boundary.
- Native question delivery spans saved interaction state and a live
provider run. Tests cover duplicate events, closed runs, and same-run
answers.
- Provider behavior varies. The full paid catalog may expose failures
beyond these targeted fixes; those results will be reported without
relaxing valid approval or output checks.
- The legacy Claude version update is limited to test infrastructure. No
database migration is included.

## Model Used

OpenAI Codex, GPT-6 family, with repository inspection, code execution,
and browser/E2E tools. The exact runtime model identifier and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (targeted checks and all
three timeout-file reruns pass; full-run timeout caveat above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-19 07:42:57 -05:00
DottaandPaperclip 11921075a4 Add first-task onboarding skill and Runner E2E coverage (#13517)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The first task helps a new user define and approve useful work.
> - That workflow needs reusable instructions and tests against the
production experience.
> - Native Codex and Claude must load the assigned skill, including
after resume.
> - Maintainers need recorded conversations and precise failed checks to
judge regressions.
> - This pull request adds the first-task skill and a suite in the
shared Runner E2E harness.
> - It keeps behavior results separate from informational quality scores
and incomplete recordings.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

The first onboarding task and the Runner E2E report used to review it.

**Current behavior**

Onboarding embeds its policy in a hidden brief. Native Codex drops the
skill-instructions setting at the Rust boundary. The shared E2E harness
has no onboarding suite or full conversation view.

**Proposed behavior**

Assign and invoke `/first-task` for the onboarding task. Send selected
Codex skills as structured protocol inputs. Run twelve scenarios across
legacy Codex, legacy Claude, native Codex, and native ACPX Claude.
Include all 48 cells in full campaigns. Show recorded chat, question and
approval cards, exact checks, instructions, and billing in the shared
dashboard.

**Reason and benefit**

Measure the real onboarding experience before changing prompts.
Distinguish infrastructure failures, behavior failures, and unexercised
journey steps.

**Breaking changes**

No database migration or production API change. First-task instructions
now live in an assigned skill. The user-edited persona is preserved; the
skill includes the maintainer-approved proposal-mode mapping and
saved-plan requirement.

Related: #11043 is earlier onboarding work. #13422 already fixes native
Claude model pinning, context delivery, and read permissions on master;
this branch includes those fixes through its base. The new Claude
recovery test supplements them.

## What Changed

- Extract and assign the first-task skill while retaining the production
greeting and opening question.
- Carry the Codex skill-instructions flag through thread start and
resume. Resolve explicit task skill references only against assigned
skills and send native skill inputs.
- Invoke an unambiguously selected assigned skill through Claude ACPX’s
native slash-command parser on initial and resumed turns, retaining the
entire task/wake envelope as its argument. Do not carry that invocation
into ordinary tasks.
- Restore the saved single-task proposal modes: confirmation card, or
saved plan with revision-targeted checkbox approval. Explicit plan
requests also require a saved plan.
- Add first-response and complete-journey cases with fixed user facts,
acceptance checkpoints, durable outcome checks, and accounting for child
runs.
- Fail the eval when choice questions have fewer than two real options.
Recognize planning documents without treating them as completed work.
- Add optional, bounded quality judging as explicit post-processing.
- Render full conversations and static interaction cards in the shared
report. Conversations start folded. Show original and regraded results
and incomplete journeys distinctly.
- Keep credential-persistence scanning outside the first-task behavioral
suite; retain public evidence redaction.
- Refresh generated capability references after the API-reference edits.
- Correct shared native question guidance and tool schemas: choices need
at least two meaningful options; open-ended questions use canonical text
fields with the required compatibility payload. Verify both formats
through real tool-authority persistence.
- Disable announcements automatically for every isolated Runner E2E
process and label the gallery environment/provider/target explicitly.
- Remove CI races in the GitHub connection browser test and native
session recovery test by waiting for the actual async work before
asserting its results.

## Verification

- `pnpm exec vitest run
server/src/services/onboarding-first-task-assets.test.ts
server/src/__tests__/issue-onboarding-first-task-routes.test.ts`: 19
passed.
- `pnpm --dir packages/paperclip-runner exec vitest run
src/drivers/acpx/runtime-host.test.ts
src/drivers/acpx/native-skill-prompt.test.ts
src/cli/acpx-runtime-sidecar.test.ts`: 70 passed. Native command
forwarding and the 1 MiB input boundary both failed before their fixes
and passed afterward. Coverage includes changed skills on reopen,
approval context, and an ordinary subsequent task.
- Runner E2E unit suite: 306 passed. Harness typecheck passed. The 64
first-task fixture and grader tests also pass.
- Full repository typecheck and build passed locally. Server typecheck
and Runner build passed again after the native-command change.
- Full GitHub Actions CI passed on `23e56447b`: all
server/workspace/browser shards, Runner verification, typecheck/release
registry, build, canary, policy, and Docker checks. Greptile reviewed
this exact head at 5/5 with no unresolved threads. The earlier broad
local run had database startup/timing failures that passed isolated
retries; the complete remote suite is green.
- Merge verification against current master: 312 harness tests and 13
native recovery tests passed. Regenerated semantic contracts and fixture
hashes pass their consistency check. Full local typecheck and build also
passed on the stacked queue branch. After merging the latest master and
preserving the GitHub setup timing regression in the split browser
suite, both focused GitHub browser tests passed. Three CI timing/startup
flakes passed local verification and one remote retry; all latest-head
checks are green.
- Real pinned Claude SDK and Claude ACP JSON-RPC probes against a local
mock API confirmed that `/skill-name` expands the assigned skill body
before the model request and retains the task arguments. A prose mention
does not. The probes made no paid model calls. The ACP probe used the
current first-task skill body and retained the wake arguments.
- [Full 48-case campaign and
report](https://pages.paperclip.ing/runner-e2e-first-task-35053063880/):
44 passed after three interrupted Codex cases completed in targeted
reruns. Original results, regrades, and all 51 executions remain in the
report provenance.
- [Claude campaign after the shared-question
fix](https://pages.paperclip.ing/runner-e2e-first-task-claude-35099525201/):
10/12 passed with zero single-option failures. All 12 recorded the
current assigned skill and corrected guidance. The failures exposed
skipped skill invocation and a missing saved plan. This PR adds native
command invocation and explicit saved-plan instructions; the subsequent
report below still shows behavior failures.
- [Fresh 12-case Claude
report](https://pages.paperclip.ing/runner-e2e-first-task-claude-35102737804/)
at `78452129e`: 10/12 pass after correcting two false proposal-matcher
failures. The recordings said “Here is the task I will create and
run/complete” in approval cards; the old matcher missed that word order.
Regression tests failed before the fix and pass after it. Original
results and offline regrade provenance remain linked. No agent rerun was
needed. Zero single-option-question failures; two behavior failures
remain: direct work before acceptance on a plain first message, and an
explicit plan request without a saved plan. Neither check was relaxed.
The follow-up `82087ac7e` fixes command-prefix size accounting;
`94aefb1f3` fixes only that proposal matcher.
- Report browser checks confirm folded conversations, rendered cards,
explicit Local/Daytona labels, and no page errors. The published-object
audit scanned 1,306 text files across 2,154 objects with no
credential-format findings or prohibited files. Image pixels and unknown
token formats are outside that scan.

## Risks

- Model behavior is nondeterministic. One campaign is evidence, not a
guarantee. The two remaining Claude behavior failures are visible in the
report and require further product work; this PR does not claim all
onboarding scenarios pass.
- The suite checks persisted Paperclip effects. It cannot prove the
absence of arbitrary external effects.
- Historical recordings can miss later journey steps. These remain
incomplete, never passes.
- Native profiles switch runtime after the production onboarding wizard
because it does not yet expose a native option.
- Quality scores are informational and cannot override behavioral
failures.

## Model Used

OpenAI Codex, GPT-6, with reasoning, repository tools, and code
execution. The exact deployed model identifier and context-window size
are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-16 14:24:32 -05:00
DottaandPaperclip ab15aff390 feat: add experimental persistent agent chat (#13284)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Conversations must use the same tasks, controls, and execution
history.
> - Users need an ongoing chat with an agent without managing task
properties.
> - Agents should clarify and plan work, then hand execution to assigned
project tasks.
> - This pull request combines the reviewed Agent Chat stack for one
squash merge.
> - The benefit is persistent conversation with normal task governance
and shared UI.

## Linked Issues or Issue Description

**Subsystem affected**

Task lifecycle, agent runtime tools, shared task UI, and browser/paid
runner tests.

**Problem or motivation**

Users need one persistent conversation with each agent. A separate chat
store or renderer would duplicate task behavior and bypass existing
controls.

**Proposed solution**

Use a task-backed chat per company, user, and agent. Reuse the task
composer and transcript. Clarify and plan in chat, then create assigned
project tasks with the relevant plan. Keep Agent Chat behind its own
disabled-by-default experimental setting.

**Roadmap alignment**

This implements the task-backed direction in [CEO
Chat](https://github.com/paperclipai/paperclip/blob/master/ROADMAP.md#-ceo-chat).
Related proposals: #2504 and #9693. Related request: #7981. The
maintainer requested one squash merge of the complete stack.

Consolidates the reviewed runtime
[#13281](https://github.com/paperclipai/paperclip/pull/13281), backend
[#13282](https://github.com/paperclipai/paperclip/pull/13282), and UI
[#13283](https://github.com/paperclipai/paperclip/pull/13283) layers
with this PR's E2E coverage. All four layers passed CI and received
Greptile 5/5 before consolidation. This PR targets master and includes
the complete feature.

## What Changed

- Add personal canonical chat tasks with ordinary company visibility,
immutable identity, idempotent first sends, and an idle waiting state.
- Process `/new` in queue order. Preserve history, release a chat pause,
and fence old provider context and delayed writes.
- Keep chat lifecycle rules across recovery, finalization, assignment,
task lists, and rollups.
- Support research and plan revision in chat. Hand plans to ordinary
assigned project tasks before execution starts. Reject new chat
subtasks.
- Add repository-aware project creation and discovery tools, including
multiple repository IDs and GitHub URLs, authorization, idempotency, and
durable project-created cards.
- Reuse task UI components for chat, with starred/recent agent
navigation and a separate `enableAgentChat` experimental flag.
- Add deterministic browser tests and 24 paid chat cells across four
Codex/Claude profiles, with validated reports and screenshots.
- Integrate current master recovery, controller lease, queued-message,
and task UI changes. Gate chat interruption and deferred promotion on
ownership/feature policy. Guarantee lease renewal and active controls
are stopped even if teardown fails.
- Preserve master's migration 0273 and generate chat migration 0274 with
idempotent replay for development databases.

## Verification

- Prior exact heads of all four PRs passed Linux CI, including build,
typecheck, general/serialized tests, and browser E2E. Each had Greptile
5/5 and no unresolved findings.
- Integrated local verification passed: full repository typecheck and
production build, Storybook build, token gates, 340 focused UI tests,
all 20 deterministic chat browser tests, two migration replay tests, 88
focused chat/queue/native/controller tests, and provider/session
regressions including real lease expiry. These include the three
lifecycle regressions for the final admission/teardown fixes; server
typecheck also passes. Current head
`1268eda16cc2af892055917e7292f068820be135` has Greptile 5/5 with no
unresolved findings and passing security scans. All final-head CI gates
passed: build, full Runner verification, typecheck/release registry,
canary, all general/serialized test shards, and all browser E2E shards
([CI
run](https://github.com/paperclipai/paperclip/actions/runs/34696739927)).
Local PostgreSQL startup contention required serialized retries; skipped
fixtures do not count as passing coverage.
- The earlier paid campaign passed all 24 chat cells and retained 32
screenshots:
[report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-34648511170-1/index.html?report=agent-chat#suite-agent-chat).
It tested `abacbdfd2f660709ec37312cdb758284c8399d04`; it is prior
evidence, not a paid run of this integrated head.
- Manual check: enable Agent Chat in Experimental settings, open an
agent, clarify and revise a plan, then hand off to an assigned project
task. Stop a reply, send `/new`, and verify fresh context with retained
history. Disable the setting and verify agent shortcuts/new chat turns
are blocked.

## Risks

- Queue/session integration can affect retries and delayed writes. Tests
cover ownership, cancellation, reset boundaries, idle recovery, and
ordinary task behavior.
- Migration 0274 adds conversation fields and constraints. Replay is
idempotent and preserves existing development chat history.
- This combines the previously reviewed stack at the maintainer's
request. Agent Chat remains off by default and is separate from
Conference Room.

## Model Used

OpenAI Codex, GPT-6 Astra (`gpt-6-astra`), with reasoning, code
execution, browser tools, and parallel review. The exact context-window
size is not exposed in this session. Codex and Claude also ran as test
subjects in the linked paid campaign.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-12 08:56:04 -05:00