Verify and extract the pinned Copilot distribution, preserve native message IDs on the ordered ACP stream, and bind the guarded distribution to profile v7. Keep interim messages separate from final output.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
Carry the admitted native plan parent tool into canonical request identity and require its exact successful durable lifecycle before passive settlement. Preserve historical committed Cursor6 waits while versioning new admission to Cursor7.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
Keep current profile qualification mandatory when first creating a wait. Recovery uses its scoped committed receipt to verify the entire original proof, so future catalog/settings changes cannot authorize task work. Cover actual persisted finalization, catalog drift, original-profile tampering, and document explicit user continuation.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
Keep the failed full invocation and original paid cases explicit. Link the exact fixture corrections and independently reviewed v7 source proof.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
Retain the failed paid case and original proof gaps. Document the verified private dependency-lock overlay and bot-owned tracked lock restoration.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
Separate source415 and source81 controller evidence from the unchanged e822 runtime. Retain failed attempts, scope the native denial pass, document pending plan settlement and exact command correlation, and distinguish implemented Pi notices from remaining rich-field gaps.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
Record Cursor5, Copilot5 and Pi8 deterministic validation, retain all earlier failures and identify fresh qualification gates.
Co-authored-by: Paperclip <noreply@paperclip.ing>
Pin Pi profile 8 to native message provenance, preserve explicit empty starts and ends, and keep replay/history and native notices out of final assistant aggregation. Validate fresh, warm and loaded SDK sessions and exact native closure bytes.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
Record three passing workflows and the exact-response failure without promoting partial file success to qualification.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
Retain failed attempts, new local and Daytona passes, current runtime identities, and remaining production gates.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
Keep historical proofs distinct while recording current Copilot local and Pi Daytona passes and the remaining qualification gates.
Paperclip-Task: rich-acp-production
Co-Authored-By: Paperclip <noreply@paperclip.ing>
Co-Authored-By: Paperclip <noreply@paperclip.ing>
* commit '72ded1794':
test(runner): distinguish Pi blank transport from form support
fix(runner): admit Pi MCP names and bound native questions
Copilot passthrough does not identify the originating turn. Retain bounded session evidence and source identity without typed turn projection or invented turn provenance. Acquire smoke resource cleanup before every fallible setup operation.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
Preserve Pi v6 and bounded startup pools, admit current Cursor and Copilot v4 identities, and combine additive Product flows and catalog assertions.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
Preserve both profile-specific guards and complete terminal diagnostics; bind Copilot v4 to the regenerated ACPX patch.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
Record sanitized actual SDK model-request assertions for fresh and loaded sessions, with exact runtime provenance and explicit qualification limits.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
Retain exact checked file reads and traversal order, drain failed batches, and preserve the v6 closure identity. Add credential-free closed Runner startup regression and record both the paid timeout and diagnostic timings without claiming authenticated qualification.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
Return completed Agent Chat handoffs through a durable outbox and scope each generated update to its supplied tasks. Add recovery, browser delivery, result access, and calibrated quality coverage.
Validated with two consecutive ten-case Claude/Codex campaigns, all CI checks, and a 5/5 review.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - The shared ACP adapter engine records agent failures for operators.
> - ACP providers can report a failure category, title, and detailed
cause.
> - Our patch kept only the category in the saved error, so an operator
could not diagnose a failure when tracing was off.
> - This pull request preserves redacted provider diagnostics in the run
error, transcript, and structured run result.
> - Operators can now inspect the provider message and any supplied
request ID or stack trace after the run ends.
## Linked Issues or Issue Description
Refs #13889 (the diagnostic gap; this PR does not update the bundled
Claude version).
Refs #14484 (related model-refusal classification; this PR retains
diagnostics for all terminal failure categories).
**What happened?**
An ACP turn failed with only `ACP agent reported a terminal service
failure.` The provider's title and details were available in memory but
absent from the saved error and transcript.
**Expected behavior**
The run retains useful provider diagnostics even when raw tracing is
disabled. Credentials remain redacted. A size limit must report
truncation instead of silently removing the cause.
**Steps to reproduce**
1. Run an ACP agent that returns an error-severity typed session
failure.
2. Include an HTTP error, request ID, and stack text in its title and
details.
3. Inspect the failed run with tracing disabled. Before this change,
only the category survives.
## What Changed
- Both pinned ACPX patches pass complete error text to the in-memory
callback, so redaction happens before truncation.
- The shared engine retains the sanitized category, title, and details
in `resultJson.terminalSessionFailure` and includes the text in the run
error and error transcript.
- Diagnostics redact configured environment values even under arbitrary
names, unknown launch-environment values, connection URL passwords, run
credentials, and common credential syntax. Known boolean settings remain
readable, while credential values are redacted even when embedded in
other text. Diagnostics remove control characters and invalid Unicode.
- Title and detail limits keep escaped transcript JSON below the
server's chunk limit. Truncated fields include an omission count. The
safe run-result projection preserves a byte-bounded diagnostic preview
when the result exceeds its byte budget, with an explicit pointer to the
full adapter-bounded run error and transcript.
- The existing UI and CLI display the error. Diagnostics do not become
assistant output. Issue continuation summaries and session-compaction
prompts receive only the generic category, preventing provider text from
becoming handoff instructions. Existing quota classification, warnings,
timeout precedence, and control-channel failure precedence remain in
place.
- Regression tests cover real ACP child processes with both pinned
versions in one-shot and persistent modes, credential redaction, request
IDs after the old 4 KiB cutoff, transcript parsing, storage bounds, and
database retrieval of oversized multibyte diagnostics.
## Verification
- Full CI on `20ad4f5f1f66c46d2c260e6ad0339cbea607b4cf`: **54 passed, 2
intentionally skipped, no pending or failing checks**. Includes
typechecking, build, all Vitest shards, Runner checks, browser E2E, and
the canary packaging/public-install dry run.
- Greptile: **5/5** on this commit. Superagent security scan passes. All
review threads are resolved.
- Local verification passed: shared ACP engine suite (395 tests); real
Claude ACP child-process and diagnostic regressions across both pinned
runtimes and both execution modes; run retrieval and model-handoff
regressions (59 tests); ACPX patch packaging (16 tests); full typecheck
and build. Affected package typechecks and focused tests were rerun
after review fixes.
- The broad local `pnpm test:run` was stopped after review edits made
its cached imports stale. Fresh targeted runs pass, including both
affected server suites. Cold-build import failures were also rerun after
dependency builds: chat integration (1,063 tests) and tool access (351
tests) pass. The final commit's complete CI matrix is green.
## Risks
- Provider diagnostic text is untrusted. This change retains more of it
in company-scoped run records. Redaction and size bounds apply before
persistence.
- Diagnostics are limited to fields the provider supplies. Old runs
cannot recover discarded error text.
- No schema migration, recovery-policy change, or new Telemetry or
OpenTelemetry export.
## Model Used
- OpenAI GPT-6 through Codex, with reasoning, repository inspection,
code editing, and test execution. The exact serving model ID and
context-window size are not exposed in this session.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>