Preserve both profile-specific guards and complete terminal diagnostics; bind Copilot v4 to the regenerated ACPX patch.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
Return completed Agent Chat handoffs through a durable outbox and scope each generated update to its supplied tasks. Add recovery, browser delivery, result access, and calibrated quality coverage.
Validated with two consecutive ten-case Claude/Codex campaigns, all CI checks, and a 5/5 review.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - The shared ACP adapter engine records agent failures for operators.
> - ACP providers can report a failure category, title, and detailed
cause.
> - Our patch kept only the category in the saved error, so an operator
could not diagnose a failure when tracing was off.
> - This pull request preserves redacted provider diagnostics in the run
error, transcript, and structured run result.
> - Operators can now inspect the provider message and any supplied
request ID or stack trace after the run ends.
## Linked Issues or Issue Description
Refs #13889 (the diagnostic gap; this PR does not update the bundled
Claude version).
Refs #14484 (related model-refusal classification; this PR retains
diagnostics for all terminal failure categories).
**What happened?**
An ACP turn failed with only `ACP agent reported a terminal service
failure.` The provider's title and details were available in memory but
absent from the saved error and transcript.
**Expected behavior**
The run retains useful provider diagnostics even when raw tracing is
disabled. Credentials remain redacted. A size limit must report
truncation instead of silently removing the cause.
**Steps to reproduce**
1. Run an ACP agent that returns an error-severity typed session
failure.
2. Include an HTTP error, request ID, and stack text in its title and
details.
3. Inspect the failed run with tracing disabled. Before this change,
only the category survives.
## What Changed
- Both pinned ACPX patches pass complete error text to the in-memory
callback, so redaction happens before truncation.
- The shared engine retains the sanitized category, title, and details
in `resultJson.terminalSessionFailure` and includes the text in the run
error and error transcript.
- Diagnostics redact configured environment values even under arbitrary
names, unknown launch-environment values, connection URL passwords, run
credentials, and common credential syntax. Known boolean settings remain
readable, while credential values are redacted even when embedded in
other text. Diagnostics remove control characters and invalid Unicode.
- Title and detail limits keep escaped transcript JSON below the
server's chunk limit. Truncated fields include an omission count. The
safe run-result projection preserves a byte-bounded diagnostic preview
when the result exceeds its byte budget, with an explicit pointer to the
full adapter-bounded run error and transcript.
- The existing UI and CLI display the error. Diagnostics do not become
assistant output. Issue continuation summaries and session-compaction
prompts receive only the generic category, preventing provider text from
becoming handoff instructions. Existing quota classification, warnings,
timeout precedence, and control-channel failure precedence remain in
place.
- Regression tests cover real ACP child processes with both pinned
versions in one-shot and persistent modes, credential redaction, request
IDs after the old 4 KiB cutoff, transcript parsing, storage bounds, and
database retrieval of oversized multibyte diagnostics.
## Verification
- Full CI on `20ad4f5f1f66c46d2c260e6ad0339cbea607b4cf`: **54 passed, 2
intentionally skipped, no pending or failing checks**. Includes
typechecking, build, all Vitest shards, Runner checks, browser E2E, and
the canary packaging/public-install dry run.
- Greptile: **5/5** on this commit. Superagent security scan passes. All
review threads are resolved.
- Local verification passed: shared ACP engine suite (395 tests); real
Claude ACP child-process and diagnostic regressions across both pinned
runtimes and both execution modes; run retrieval and model-handoff
regressions (59 tests); ACPX patch packaging (16 tests); full typecheck
and build. Affected package typechecks and focused tests were rerun
after review fixes.
- The broad local `pnpm test:run` was stopped after review edits made
its cached imports stale. Fresh targeted runs pass, including both
affected server suites. Cold-build import failures were also rerun after
dependency builds: chat integration (1,063 tests) and tool access (351
tests) pass. The final commit's complete CI matrix is green.
## Risks
- Provider diagnostic text is untrusted. This change retains more of it
in company-scoped run records. Redaction and size bounds apply before
persistence.
- Diagnostics are limited to fields the provider supplies. Old runs
cannot recover discarded error text.
- No schema migration, recovery-policy change, or new Telemetry or
OpenTelemetry export.
## Model Used
- OpenAI GPT-6 through Codex, with reasoning, repository inspection,
code editing, and test execution. The exact serving model ID and
context-window size are not exposed in this session.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip manages AI agents that work on assigned tasks.
> - The task composer lets a person choose an assignee, model, and
effort for the next run.
> - A Paperclip Runner agent can use Codex as its provider.
> - The composer hid Codex effort for that agent because it checked only
the older Codex adapter.
> - The native Runner input also did not carry an effort choice to
Codex.
> - This pull request carries the chosen effort from the composer to
each Codex turn.
> - People can now select a supported effort and get the effort they
selected.
## Linked Issues or Issue Description
Refs #14322
**What happened?**
The composer showed a model but no effort slider when the assignee used
Paperclip Runner with the Codex provider. A task-level model override
also did not reach the native Runner input.
**Expected behavior**
The composer shows effort choices for a known Codex model. The next
native Codex turn uses the selected model and effort.
**Steps to reproduce**
1. Open a task composer.
2. Select an agent that uses Paperclip Runner with the Codex provider.
3. Select a known Codex model such as `gpt-6-astra`.
4. Open the assignee and model picker. The effort slider is missing
before this change.
## What Changed
- Show known Codex effort levels for Paperclip Runner Codex assignees.
- Save the task effort override in the native run input and send it to
Codex on each turn.
- Apply the task's merged model and effort overrides when the native run
starts.
- Apply a task model override for OpenCode Runner without changing the
agent's provider.
- Add Runner effort tests and desktop and mobile Storybook cases.
## Verification
- `pnpm -r typecheck` passed.
- `pnpm build` passed.
- `pnpm build-storybook` passed.
- `pnpm check:token-gates` passed.
- Focused UI, server, Runner contract, and Codex driver tests passed.
- The full CI test matrix, build, typecheck, and canary dry run passed
on the latest head.
## Risks
- Native Runner inputs add an optional Codex effort field to the current
v5 input. Older inputs keep their previous behavior.
- A known model rejects an effort that its catalog does not support.
Unknown models do not show a slider.
> This fixes an existing composer bug. I checked `ROADMAP.md`; it does
not describe this bug as planned work.
## Model Used
OpenAI Codex, GPT-6. The exact deployment ID and context window are not
exposed in this session. The model used reasoning, code execution, and
repository tools.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: OpenAI GPT-6 Astra <noreply@openai.com>
Grade clarification lists, obsolete unstarted wakes, and refusal cancellation from persisted evidence. Preserve execution and ownership assertions, add boundary regressions, and version the affected eval definitions.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - The Mine inbox shows work that needs the current user.
> - Failed-run rows used the latest run for every agent in the company.
> - A failure from another user therefore appeared in Mine and its
badge.
> - Run list responses also omitted the responsible user needed to
filter these rows.
> - This pull request uses run ownership for personal failure routing.
> - Users see their own failures and can still inspect company failures
in All.
## Linked Issues or Issue Description
**What happened?**
An agent run started for one user failed. Its row and failure badge
appeared in another user's Mine inbox.
**Expected behavior**
Mine and its badge include failed runs for the current responsible user.
Other users' failures remain available in All and run details.
**Steps to reproduce**
1. Use a company with two human users.
2. Create a failed or timed-out run attributed to the first user.
3. Open Mine as the second user. Before this fix, the failed run appears
there and increases the badge.
**Paperclip version or commit**
Reproduced in regression tests on master at `24beb0057`.
**Deployment mode**
Authenticated deployment with multiple users. Tests also cover the local
single-user board.
Related prior work: #933 addressed inbox dismissal and badge
consistency. No duplicate ownership fix was found.
## What Changed
- Return `responsibleUserId` in normal and summary run lists.
- Share one ownership rule across both inbox versions and client/server
badges.
- Select the latest run per agent before applying the ownership filter.
This prevents old failures from resurfacing on shared agents.
- Keep unattributed historical failures in the local board's Mine view.
Hide them from authenticated users with no matching owner.
- Keep company health alerts outside the personal badge, consistent with
the client.
- Document the routing contract and add page, badge, and database
regression coverage.
## Verification
- Red: the new badge cases failed with three company failures instead of
one personal failure; eight Mine page cases failed across both inbox
versions.
- Green: 113 focused tests pass in `ui/src/lib/inbox.test.ts`,
`ui/src/pages/Inbox.test.tsx`,
`server/src/__tests__/heartbeat-list.test.ts`, and
`server/src/__tests__/inbox-dismissals.test.ts`.
- `pnpm check:token-gates` passes.
- Agent calls on behalf of a user have two additional red-to-green API
regressions.
- Full `pnpm -r typecheck` and `pnpm build` pass. Server typecheck also
passes after the agent-call fix.
- All CI test shards and browser tests pass on `243bfa681`. The
duplicate local `pnpm test:run` was stopped after the CI test lanes
completed; it did not finish locally.
## Risks
- Authenticated users no longer receive unattributed legacy failures in
Mine. Those failures remain visible in All.
- The server badge no longer counts company health alerts, matching the
existing client badge.
- No migration, run state, retry behavior, or company access rules
change.
## Model Used
- OpenAI GPT-6 through Codex, with reasoning, terminal execution, and
browser tools. The exact deployment variant and context window size are
not exposed in this session.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-Authored-By: Paperclip <noreply@paperclip.ing>
* codex/runner-cursor-acp:
feat(runner): add rich ACP transport and durable interaction foundation (#14430)
fix: preserve terminal task owners during release (#14561)
fix(ui): recover gracefully during server restarts (#14560)
Add shared rich ACP transport, durable questions and permissions, verified provider packaging, and bounded activity and plan presentation. Keep Cursor, Copilot, and Pi pending their separate provider qualification.
Persist interaction settlement before publication, fence failed writes until fresh recovery, and preserve owned-process cleanup. Incorporate reviewed mainline integration with extended harness coverage.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Tasks record an assigned owner and separate checkout and execution
locks.
> - A completed task must retain its owner after execution ends.
> - The release endpoint currently clears that owner when it clears the
locks.
> - This pull request preserves the assignee of Done and Cancelled tasks
during release.
> - Unfinished tasks keep the existing relinquishment behavior.
> - The benefit is stable task attribution without retaining execution
locks.
## Linked Issues or Issue Description
**What happened?**
An agent completed an assigned task, then called the release endpoint.
The task stayed Done, but its assignee became null. The same defect
affects Cancelled tasks. It caused the legacy Claude clarification/reuse
and multiple-repository handoff E2E assertions to fail.
**Expected behavior**
Release must clear execution locks on terminal tasks and preserve their
assignee, final status, and disposition timestamps. Release of
unfinished tasks must still clear the agent assignee. Only In Progress
work returns to Todo.
**Steps to reproduce**
1. Create an assigned task with checkout and execution locks.
2. Complete or cancel the task.
3. Call `POST /api/issues/:id/release` as the assigned agent.
4. Read the saved task. Before this fix, its assignee is null.
**Paperclip version or commit**
Reproduced on master commit `d172197117a14b80a1eb2d2835a0e7cce2679656`.
**Deployment mode**
Local tests against real PostgreSQL through the production issue routes
and services. This is a core lifecycle defect, independent of the agent
adapter.
Refs: #11689, #6899, #7769. These are related open release proposals.
This is an independent fix limited to terminal task ownership. It does
not include timer scheduling changes.
## What Changed
- Preserve the current assignee when releasing Done or Cancelled tasks.
- Keep all execution-lock cleanup and existing unfinished-task behavior.
- Cover all seven task statuses through the release API and read back
saved state.
- Check disposition timestamps, activity attribution, and repeated board
cleanup.
- Update the API contract, agent reference, and CLI help.
## Verification
- Red commit `ddaabb754`: the two terminal-owner regressions failed with
`assigneeAgentId: null`; 12 other route tests passed.
- Green: all 14 route tests pass, plus the existing successor-checkout
race test (15 selected tests total).
- Command: `pnpm --filter @paperclipai/server exec vitest run
src/__tests__/issue-stale-execution-lock-routes.test.ts
src/__tests__/issues-service.test.ts -t 'stale issue execution lock
routes|does not let stale release clobber a successor checkout lock'`.
- The local host has exhausted its SysV semaphore pool. The red/green
runs used the existing test-provider hook to start disposable Docker
PostgreSQL 17 instances. Routes, services, migrations, and assertions
were unchanged. No database tests in the selected set were skipped. The
other 134 tests were excluded by the name filter.
- Capability contract and inventory drift checks pass.
- `pnpm build` and `pnpm -r typecheck` pass.
- The local `pnpm test:run` was interrupted after environment failures
while the complete sharded CI suite ran in parallel: native PostgreSQL
bootstrap fails under the host semaphore limit, and the large Git
fixture hits macOS `ENAMETOOLONG`. A focused rerun confirmed these
happen before the relevant assertions. The interrupted local run is not
counted as a full pass.
- Greptile completed on `1caeeb827e9cb658ddb71f16c2421ec20f80634e` with
**5/5**, a successful check, and no review threads.
- All CI gates pass on the current head: typecheck, build, general and
serialized tests, Runner checks, browser E2E, release packaging, and
security checks. Server shard 11 passed on one targeted retry; the first
attempt had 836 passing tests but an unhandled workspace-runtime startup
rejection caused by an existing timing window. All other successful jobs
were reused.
- No paid provider evaluations were run.
## Risks
- A caller that used release to erase ownership from terminal work will
now retain that owner. An explicit assignment update or the board
force-release option with `clearAssignee=true` can still clear it.
- No schema or migration changes. The transaction, company access,
assignee/run checks, and activity log remain in place.
## Model Used
- OpenAI GPT-6 through Codex, with tool use, code execution, and test
debugging. The session does not expose a more specific backend model
version or context-window size.
- OpenAI `gpt-6-luna` assisted with read-only test discovery and review.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
Preserve the extended harness and instruction-persistence suites, generated
contracts, and cleanup/collection ordering. Record the unqualified ACPX
agent-directory capability without broadening environment or path policy.
Co-Authored-By: Paperclip <noreply@paperclip.ing>