Commit Graph
384 Commits
Author SHA1 Message Date
DottaandPaperclip 71cd0a2621 fix(skills): honor the current run harness checkout (#15548)
## Thinking Path

> - Paperclip manages work for AI agents.
> - The runtime claims eligible assigned tasks before it starts an
agent.
> - The wake tells the agent when the runtime already holds that claim.
> - The legacy skill still requires another checkout in every case.
> - This PR makes the skill honor the current task and run claim.
> - Manual checkout and server ownership checks remain in place for
other cases.

## Linked Issues or Issue Description

**Where is the issue?**

`skills/paperclip/SKILL.md`, in the scoped wake procedure and Step 5.

**What's wrong?**

The wake can say that the harness already checked out the issue. The
skill still tells the agent that it must call checkout. These
instructions conflict.

**Suggested fix**

Skip the second checkout only when the runtime wake explicitly confirms
the claim for this issue and run. Retain manual checkout when that
statement is absent or the agent selects another task. Refs #14948 for
the existing shared prompt reduction.

## What Changed

- Honor the explicit runtime claim in the scoped wake procedure and Step
5.
- Keep context reads, status writes, deliverable handling and conflict
rules.
- Add checks for normal and resumed wake text and excluded automatic
claims.
- Retain successful checkout HTTP activity for legacy stock-task evals.
Bind each receipt to the exact company, task, agent and run. Keep this
observation separate from the original task grades.

## Verification

- Checkout observation calibration: nine tests pass.
- Focused skill, wake and database ownership tests: in progress.
- Full repository build, typecheck and tests: in progress.
- Planned live comparison: the existing assigned-skill document case on
legacy Codex and Claude. One attempt per variant and profile. No
automatic retries. The baseline and candidate share the observation code
and task oracle.
- Live results are pending. This draft does not claim behavioral
qualification.

## Risks

- Agents may misread prompt guidance. The API still enforces ownership;
the text grants no new authority.
- The exception is specific to the current issue and run. It does not
remove ordinary legacy completion writes or authorize another task.
- Activity measures successful checkout HTTP calls. Failed attempts
require separate run-log inspection. Missing or mismatched observations
cannot count as zero calls.
- One trial per profile cannot establish general reliability, speed or
cost trends.

## Model Used

OpenAI Codex, GPT-6 family. The exact model build and context window are
not exposed in this session. Used code editing, shell tools and test
execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-08 06:38:09 -05:00
DottaandPaperclip d66acb7ac1 feat: automate Slack bot app setup and installation (#15413)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Chat connectors give each agent a customer-owned bot and task-backed
conversations.
> - Manual Slack setup requires app creation and copying durable
credentials.
> - Operators need a shorter setup that an assisting agent can use
safely.
> - This pull request creates the app through Slack's Manifest API and
installs it through OAuth.
> - Durable registration state supports recovery without creating
another app.
> - A four-screen wizard, automatic avatar upload, and OAuth account
linking reduce setup work.
> - Connector settings and per-turn tool guidance support daily use
after installation.

## Linked Issues or Issue Description

**Subsystem affected**

Native Slack bot setup, company secret storage, chat connector
management, and agent tool guidance.

**Problem or motivation**

New Slack bots require manual app creation and copying a signing secret
and bot token. Interrupted setup can create duplicate apps. The setup
and management screens contain unnecessary controls. Agents also need
guidance for native questions, files, thread replies, and governed Slack
actions.

**Proposed solution**

Use a temporary app-configuration access token to create a
customer-owned app. Save durable secrets in the vault. Bind OAuth to the
initiating actor, company, endpoint, registration revision, scopes, and
configured origins. Preserve manual and existing-app recovery. Link the
installing user's account, send a welcome DM, and advance from saved
server evidence. Keep request-URL recovery instructions available if
automatic connection detection waits.

**Alternatives considered**

The Slack CLI adds installation requirements. Socket Mode changes
transport. A shared Paperclip-owned app changes app ownership. These
alternatives are outside this change.

**Roadmap alignment**

This extends existing chat connectors and secrets capabilities. Related
public work: #14037 and #13954 cover Slack MCP prerequisites and user
OAuth. No duplicate bot-registration PR was found.

## What Changed

- Share one reviewed manifest builder between automatic registration and
manual setup.
- Add replay-safe migration 0318 and company-bound registration state
with vault references and uncertain-creation recovery.
- Add registration, installation, callback, and resume APIs with
short-lived, single-use OAuth state.
- Save installation credentials before downstream checks and preserve
bot identity constraints.
- Reduce automatic setup to four screens. Keep advanced app details,
manual recovery, and existing-app setup.
- Upload the agent avatar with the Paperclip dark background. Link the
OAuth installer's account and send setup DMs.
- Show agent and connector-owner avatars. Simplify settings, access, and
conversation screens.
- Discover joined Slack channels and enable them by default. Start a
task from a bare mention and admit same-thread follow-ups.
- Refresh Slack tool guidance each turn. Add native-form, file,
approval, and delivery regressions plus manual model probe definitions
and sanitized acceptance records.
- Update deployment/database docs, OpenAPI, redaction, removal cleanup,
production Storybook stories, and provider browser tests.
- Merge current master and move the registration migration after its
latest migration without rewriting published commits.

The completed Slack success view intentionally has a single centered
**Done** action and no **Save & exit**, as explicitly requested by the
product owner. `DESIGN.md` records this exception; unfinished setup
steps retain the aligned wizard footer.

## Verification

- Passed after the master merge: repository typecheck, full build,
Storybook build, design-token gates, module-boundary gates, and
migration generation.
- Passed: all 352 focused Slack deterministic tests and all 14 affected
provider browser tests. Browser tests use controlled provider fixtures
and a separate throwaway instance.
- Passed on current head `c5d01e0e2`: the complete GitHub test matrix
(general server, chat, all workspaces, serialized server, and Runner),
all eight browser shards, typecheck/release registry, build, canary dry
run, security checks, and policy gates. There are 52 passing checks and
no pending or failing checks.
- Greptile completed on the exact current head with 5/5 and no
actionable findings or open review threads.
- Local repair verification passed 93 focused tests, including same-app
reinstall after revocation and rejection of consent started before
revocation, the AgentMail browser journey, and repository typecheck.
Local build and Storybook build also passed. The redundant local
full-suite rerun was stopped after the complete current-head CI matrix
passed.
- Real Slack setup and agent replies were exercised in the authorized
isolated test drive during the setup iteration.
- The ten additional model probes were attempted with legacy
`codex_local`, `gpt-5.6-sol`: five passed, two failed, and three were
partly verified. Native runtime is not qualified. See
`server/src/services/connectors/slack/evals/2026-10-08-acceptance.md`
for evidence and limits.
- Passing model probes cover native forms, downloaded file bytes, bare
mentions with thread replies, explicit posts/reactions, and saved
approval denial.
- The controlled uncertain-write probe found wrong delivery-check IDs.
The canvas fallback attempt used an invented tool name. Search
pagination/native search, a private-source denied-tool receipt, and
distinct board/webhook origins remain unqualified.

Reviewer path: enable Chat connectors, start Slack chat setup, select an
agent, enter an app-configuration access token, and approve Slack
installation. Send a message to the bot and confirm that setup advances
to success. Inspect settings and allowed channels. See
`doc/connections/SLACK-AUTOMATIC-SETUP.md` for deployment and recovery.

## Risks

- Slack app creation has no provider idempotency guarantee. A timeout
after dispatch stays uncertain until the operator checks Slack.
- OAuth needs a stable public HTTPS board origin. Webhook ingress may
use a separate configured HTTPS origin. Workspace policy can delay
installation.
- Migration 0318 can replay safely on instances that applied the earlier
development migration.
- OAuth installation now links the installer to the initiating Paperclip
user. Identity checks and company access rules still apply.
- Joined channels now enable bot responses by default. Linked-user
authorization and per-action approval rules still apply.
- Model behavior has the documented delivery-check and canvas fallback
failures. A passing CI run does not establish that every model probe
passed.
- Removing the connection does not delete the customer's Slack app. No
new first-party telemetry is added.

## Model Used

OpenAI Codex, GPT-6 family, with reasoning, repository tools, code
execution, and browser verification. The runtime does not expose a more
specific authoring model ID or context-window size. The live bot probes
used OpenAI `gpt-5.6-sol` through `codex_local` in legacy mode.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-08 06:33:31 -05:00
DottaandPaperclip fc6304dfe5 feat(runner): add experimental OpenAI Dot provider over MCP Events (#15402)
## Thinking Path

> - Paperclip manages AI agents, tasks, permissions, and execution
budgets.
> - Paperclip Runner gives each provider the same admitted task and tool
authority.
> - OpenAI Dot runs outside the local process tree and needs
asynchronous work delivery.
> - The merged MCP gateway supplies OAuth consent and signed event
delivery.
> - A personal assistant grant cannot safely stand in for an assigned
agent.
> - This pull request adds a separate Dot agent connection and a durable
Rust Runner bridge.
> - The operator can assign work to Dot and inspect its accepted work,
tool receipts, and result.

## Linked Issues or Issue Description

**Agent or provider**

OpenAI Dot, as an experimental provider of the existing Paperclip Runner
adapter.

**Why this adapter is useful**

An operator can assign normal Paperclip tasks to an existing Dot. Dot
can read its mailbox, request work on an assigned task, use admitted
task tools, and submit a result. Paperclip keeps company scope,
checkout, approvals, known budget limits, and activity attribution.

**How the agent is invoked**

A dedicated `/mcp/runner` OAuth resource pairs one Dot grant with one
agent. A signed MCP mailbox event wakes Dot. Dot explicitly accepts the
assignment. The Rust Runner owns the durable turn and operation
receipts. The first release supports self-hosted instances with a local
Runner controller.

**Additional context**

This extends the merged public MCP gateway from #14846 and the assistant
invitation and device-consent work from #14933. This also integrates the
merged assistant tool and configuration expansion in #15380. Dot retains
its dedicated agent resource and cannot receive personal configuration
permission. The public assistant connection remains a personal
connection.

## What Changed

- Add a durable Rust Dot provider and its TypeScript Runner driver.
- Add closed PRP v3 external-provider operations and native execution
input v6.
- Add company-scoped pairing, mailbox, assignment, and operation
records.
- Reuse merged browser/device consent, client metadata verification,
webhook admissions, refresh, secret rotation, and warm-standby gates.
- Keep Dot scopes, issuer, grants, event workers, and tool access
separate from personal assistant access.
- Add Dot configuration, pairing, readiness, and consent UI. Keep agent
grants out of the personal Connections entry.
- Regenerate the Dot-only migration after master. Preserve published
gateway migrations. Make the new migration safe to reapply.
- Document setup, recovery, accounting limits, evidence, and remaining
account qualification.
- Reverify reconnect callbacks and wake outstanding work with a fresh
mailbox reference; preserve the existing assignment and operation
receipts.
- Clean up Dot bindings and waiting runs on OAuth revoke and
refresh-token replay. Old grants cannot revoke replacement bindings.
- Restore the pairing reference when an unsaved agent form is reopened;
document board-only pairing routes in OpenAPI.
- Accept a clean Rust exit after the acknowledged shutdown receipt.
Unexpected exits still require recovery.
- Clear the cached binding after a successful revoke so a failed
connection refresh cannot restore it.
- Add production-component Storybook states and screenshots for pairing
and connection review. All preview account data is synthetic.
- Persist normalized completion, serialize Dot turns and durable work
admission, and poll subscription readiness.
- Serialize mailbox writes and cursor reads; retain paused fence
acknowledgement without task authority.
- Authorize admitted review runs without changing the worker assignee.
Include the fenced assignment ID in production stop notices.

## Verification

- This PR integrates master `4a8178e9c`. Dot migration
`0317_messy_famine.sql` follows the published history and is safe to
reapply. The merge preserves the reserved migration connection,
batch-commit handling, private task checks, task monitors, and native
accounting.
- Local workspace typecheck, full build, and UI token gates pass. The
server typecheck passes after the review fixes. Database and native
executor regressions pass.
- All twelve real Rust/PostgreSQL Dot integration tests and twelve Dot
driver tests pass. The tests cover native document writing and
finalization, durable replay, queue admission, mailbox ordering,
admitted reviews, stale authority, production stop references, and
paused acknowledgements.
- Current head `d0e7e0626` passes all 57 checks: 53 pass and four are
intentionally skipped. This includes full typecheck, build, tests, Rust
Runner verification, browser E2E, release verification, and Canary Dry
Run. Greptile rates this exact head 5/5. All review threads are
resolved.
- The full local root test run is slower than the sharded CI run and has
not completed. The full CI test gates pass on the current commit.
Focused local regressions pass.
- Real-account pairing and event delivery on this base commit remain
unqualified. Live account and setup proof are recorded in the follow-up
#15414.

The following screenshots use synthetic preview data. They show the
production pairing component and do not qualify a real account or the
full agent setup journey.

![Synthetic pairing
preview](https://github.com/paperclipai/paperclip/blob/codex/dot-events-prototype/doc/screenshots/openai-dot-runner/pairing.jpg?raw=true)

![Synthetic connected
preview](https://github.com/paperclipai/paperclip/blob/codex/dot-events-prototype/doc/screenshots/openai-dot-runner/connected.jpg?raw=true)

## Risks

- This base adapter uses `PAPERCLIP_ENABLE_OPENAI_DOT=1` plus Public MCP
and Paperclip Runner. The separate experimental-settings follow-up in
#15414 replaces this environment flag with saved operator settings.
- Dot does not expose provider token usage or cost. The operator must
acknowledge external billing. Known Paperclip budget gates still apply.
- Cancellation fences Paperclip authority. It does not confirm that Dot
stopped all external activity.
- Assigned skill files and third-party MCP bindings are unsupported and
reject admission. There is no mounted workspace, model selector, or
provider thread identifier.
- Hosted agent-broker and remote controller deployments are not
qualified.
- The new migration follows the merged master history. Existing
prototype databases still need the normal master migration history
before this Dot-only migration.

## Model Used

OpenAI Codex, based on GPT-6. The exact deployment ID and context window
size are not exposed in this session. Capabilities used: reasoning,
repository editing, code execution, and test inspection.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #123` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub references)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-07 19:11:50 -05:00
Devin FoleyandPaperclip cd40ebb95b Preserve ACP bridge terminal evidence before disconnect (#15478)
Drain terminal frames through backpressure and record bounded terminal evidence without waiting for log persistence. Guard later events and input failures against writes after socket end.

Validated with 217 focused tests, full typecheck/build, independent review and green CI with Greptile 5/5.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-07 13:01:22 -07:00
DottaandPaperclip a6306ba606 feat(runner): consolidate Cursor production integration (#15075)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The native Runner keeps provider sessions under company authority,
approvals, budgets and durable recovery.
> - Cursor work was spread across candidate branches. The published
branch lacked later plan, permission and cleanup fixes.
> - Production also needs public installation and matching runtime
assets for local and Daytona execution.
> - This pull request consolidates Cursor onto current mainline recovery
behavior and completes that installation path.
> - The installed v11 release passed focused local and Daytona
qualification after the generic mode and lifecycle cleanup. The later
model-selection correction and current mainline merge produce v14
artifacts that need matching release qualification.
> - Cursor admission is enabled in source; publish only an artifact
combination with matching qualification. Native AskQuestion and complete
per-run dollar accounting remain excluded.

## Linked Issues or Issue Description

Refs: #14435, #14631, #14669, #14699, #14724.

This completes the Cursor implementation by @cryppadotta from combined
source `22c78242a4e0c2369fecf0c2dc4e7600fbad6706`. It preserves newer
mainline recovery, completion and warm-directory behavior. Pi and
Copilot remain gated.

## What Changed

- Generate named Rust and TypeScript ACPX release profiles from one
manifest. Share runtime pins with packaging and server verification.
Preserve vendor runtime versions; bind the updated ACPX patch to Cursor
profile v14 and reject stale generated declarations at build/typecheck.
- Remove ACPX model allowlists, including the former Codex and Pi
restrictions and the duplicate developer test-drive gate. Send any
explicit model ID unchanged to its provider and verify the effective
selection before prompting. The bundled ACPX package forwards unlisted
IDs, rejects mismatched acknowledgements, and restores the exact
selection after session load. It does not expand Cursor model aliases.
Provider rejection, mismatch, or missing model controls fails without a
fallback. Model examples live in evaluation fixtures, outside runtime
declarations.

- Add pinned Cursor execution, contained instructions, exact model
verification and Agent/Plan/Ask modes.
- Carry an opaque generic `mode` identifier in shared native execution,
sidecar, Rust and recovery contracts. The provider adapter owns
supported modes, defaults, native translation and acknowledgement.
- Keep native RPC recognition, accepted-plan interpretation and
permission evidence behind provider adapters. Shared settlement and
recovery verify normalized facts and their committed evidence.
- Replace the Cursor-only warm-attachment branch with a runner-owned
capability. Only Cursor opts into it. Move profile compatibility and
optional usage parsing into provider metadata and adapters.
- Write generic plan-wait receipts. Read exact historical Cursor
receipts through a separate compatibility decoder. Reject mixed formats
and preserve existing authority checks.
- Carry native plans, semantic questions, todos, child activity,
permission identities and partial usage diagnostics through the Runner.
- Preserve durable response delivery, cancellation, warm ownership and
process retirement.
- Finish accepted planning runs successfully. Keep their tasks open for
explicit direction. Acceptance does not start implementation.
- Ship `paperclipai runtime setup cursor` and its provisioner through
the public package. npm installation does not download Cursor. Setup
uses the OS account's closure-keyed cache so system-wide npm packages
can remain read-only. Run it as the Paperclip service account.
- Include Cursor in normal provider packs and Daytona images for macOS
ARM64/x64 and Linux x64.
- Reject stale release packs by source revision and current ACPX/Cursor
pins before assembly writes files. Verify current Cursor
version/profile/closure again at runtime.
- Ship all three daemon targets and the expected Linux image-pack
identity. A macOS controller uses its packaged Linux daemon for Daytona.
Image mismatches fail before provider launch.
- Use the vendored Runner boundary for installed readiness probes.
Verify the actual installed Cursor probe.
- Verify compiled public Daytona plugins and their release versions in
installed smokes.
- Record exact artifacts, the acceptance matrix, retained failures,
supported capabilities and rollback behavior in the [readiness
report](https://github.com/paperclipai/paperclip/blob/codex/cursor-production-readiness/doc/plans/2026-10-03-cursor-production-readiness.md).

## Verification

- Current head `f7ec5cc1f0e30c62a829c016ff2013a2a9d79037` merges
mainline `faa8e452c73bae5e044dd6379179a00106abb131`. It keeps Cursor
plan and cancellation guards alongside mainline historical-question
filtering. The evaluation catalog includes both Cursor and expanded
adapter accounting cases (683 total). Recursive typecheck, full build,
696 lifecycle/recovery tests, 45 fixture tests and fixture typecheck
passed. Current-head CI passed: 56 successful checks, one neutral and
four skipped. [Complete
CI](https://github.com/paperclipai/paperclip/actions/runs/37557996535).
The fresh Base Greptile review is 5/5 on this exact head, with 304 files
reviewed, zero new comments and zero unresolved threads. The user
authorized overriding the CODEOWNER review gate after checks passed; no
failing checks are overridden. Prior results below retain their own head
identities.
- Corrective head `3d2b168366258036f6b6a6fccb382c49138cc601` fixes the
post-merge Apex finding. Automatic-review and new-evidence
reconciliation preserve pending child results and recheck delivery under
the status lock before completing. Account repair now excludes unrelated
secret consumers and requires the failed agent's identity. Regression
coverage includes the commit race, delivery statuses,
current-run/current-intent exclusions, repeated reconciliation, both
database reconciliation paths, and credential consumer boundaries. All
184 affected tests, server typecheck and server build passed.
Current-head Base Greptile review is 5/5, with 304 files reviewed, zero
new comments and zero unresolved threads. Current-head CI passed: 56
successful checks, one neutral and four skipped. [Complete
CI](https://github.com/paperclipai/paperclip/actions/runs/37535994724).
This Base review is distinct from the earlier Apex review.
- Merge head `5957c257a` reconciles mainline `b508a05c4`. It preserves
both accepted-plan waits and pending-child-completion checks, current
provider selectors, task-creation response identities, and mainline ACPX
missing-file handling. The combined patch is bound to Cursor profile
v14; historical records keep their original identities.
- Merge head `5957c257a` passed recursive typecheck, full build, 43
installed ACPX/package contracts, 107 provider UI and plan/recovery
tests, 593 database-backed lifecycle tests, 49 profile/native contract
tests, 45 Product E2E fixture tests, fixture typecheck, token gates,
three provider-free browser task-creation cases, and Runner
conformance/replay checks. Its complete CI passed (55 successful checks,
one neutral and four skipped), while Apex returned 2/5 with a
child-delivery finding addressed below.
- The local full-suite attempt again failed the unchanged Git streaming
test (360-second timeout) and was stopped. The concurrent local Rust
attempt failed four unchanged Codex process/deadline tests; all four
passed serially without code changes in 7.29 seconds after removing the
competing test load. These failed commands are retained and are not
reported as full-suite passes; the fresh Linux CI runs are tracked
separately.
- The previous head `907bdb2a2778c7ffeb4a662a91460c9d1ddfc9c5` earned
Apex 5/5 with zero comments after fixing all three findings: per-user
install cache, stale release-pack rejection, and public Linux smoke
account/home handling. Its real built installer passed from read-only
public packages on macOS ARM64 and Linux x64. All 137 release-registry
checks and 64 ACPX package contracts passed. That review does not cover
this mainline reconciliation.
- Prior `beadd3654` passed the full CI matrix; its one unchanged chat
test failure and successful single retry remain in the [CI
history](https://github.com/paperclipai/paperclip/actions/runs/37521449327).
Historical results below remain attributed to their original builds.

- Fixture follow-up `dd59d7e82b103a88b7cbd7d2c38b612c0fbbff7a` removes
provider-specific model choices from generic offline ACPX tests. The
fake sidecar preserves the model and session identity selected at open
through suspension. Affected verification passed: 106 Rust tests and 73
TypeScript tests. This commit changes test code only; the
production-code checks below retain their recorded identities. Its CI
and Greptile review later passed; those results belong to that
historical head.
- Model-selection cleanup `9a070808b48960a41fdfd369ae0636b95af82459`:
252 focused Runner tests passed (six platform skips), covering all six
ACPX agents, native model acknowledgement, rejected selections,
installation integrity and recovery identity. The merged branch passed
recursive typecheck, full build, token gates, server admission (19
tests), and the Product E2E catalog (45 tests). The acceptance catalog
passed all four tests. The full Rust suite passed: 643 tests, 2 ignored.
It verifies sidecar acknowledgement of unlisted models and rejection of
model mismatches. The final commits only update Rust tests; production
sources match the verified build at
`65ec3279ac50185e3cda109b5cfd9b4f56105de0`. No new paid provider calls
were made.
- The merge preserves both Cursor and the new mainline public-MCP
fixture cases. Auto-merge remains disabled; the latest follow-up status
is recorded above. The local `pnpm test:run` attempt hit the unchanged
Git streaming test's 300-second timeout and was interrupted before
merging mainline. The broad Runner attempt found obsolete single-model
assertions plus three macOS fixture-path failures caused by a
`/private/tmp` override. The assertions are corrected; affected
TypeScript checks passed with the standard macOS temporary directory,
and the complete Rust suite passed. Neither interrupted command is a
full-suite pass.
- Earlier declaration-cleanup head `6f4a5e9e2` passed recursive
typecheck, build, Rust and focused tests. Its CI later exposed a test
expecting duplicated Grok digest literals. The current source fixes that
assertion to compare launcher bytes with the shared manifest. Historical
successes and failed attempts are retained; no new live provider
qualification is claimed.
- Previous head `e75fde6098b0ddd8cec765bfb6ecaeecb88a26a6` passed
complete CI (56 successful checks, one neutral, four skipped) and
Greptile 5/5. [Historical complete
CI](https://github.com/paperclipai/paperclip/actions/runs/37489112305).
Those results are not claimed for the cleanup head.
- Frozen live application: `d7b696f9b8f79095233e9e3d56d23e6a6018dd48`.
Public package version: `0.0.0-cursor-verify.3d0c9b7761c6`. The
declaration cleanup preserves release pins and does not relabel that
tested artifact as a build of the new source. Mainline through
`e34abee670` was reconciled while preserving accepted-plan waits,
provider-capacity handling, and both Cursor and public-MCP fixtures.
- Clean normal installation, explicit Cursor setup and daemon resolution
passed on macOS ARM64, macOS x64 under Rosetta, and Linux x64. npm
lifecycle hooks ran without silently downloading Cursor.
- Historical v11 live matrix: **18/18 passed with cleanup** (nine local,
nine Daytona) after the generic mode and lifecycle cleanup. The campaign
has 23 attempts; all five failures and their diagnoses remain recorded.
Exact case identities, hashes and limits are in the readiness report.
All provider calls are real, use the explicit Luna model and
company-bound credentials, and run without qualification or
runtime-asset overrides.
- The immutable Daytona image is
`ghcr.io/paperclipai/paperclip-daytona-runner@sha256:d6259b6bba094702c13fc2283bd85550849c1c53145b656fb2746778f9fa1747`.
The public Daytona plugin is installed independently and its version is
checked.
- Recursive typecheck, full build, token gates and Runner
contract/conformance/replay checks passed on the frozen application. Its
complete Linux CI suite passed. The duplicate local full-suite command
was incomplete after timing failures; affected repeats passed, but that
command is not reported as a clean pass.
- Qualification fixtures passed typecheck, 1,675 Vitest tests (one
skip), 128 Node checks, three provider-free browser tests, and 150
focused lifecycle tests after the final diagnostic correction. The
affected legacy Cursor command file also passed all five tests after
removing its shorter 10-second override; it now inherits the suite’s
standard 15-second timeout. Greptile is 5/5 on `e75fde609` with no
unresolved review threads. CI results above are recorded separately from
historical build results.

## Risks

- Cursor v14 includes the updated ACPX dependency patch and release
identity. The v11 live matrix and image below remain historical
evidence. They do not certify new v14 package/image artifacts.

- ACPX accepts models beyond the qualification fixtures. Availability
and entitlement depend on the provider. Successful configuration is not
a claim of live qualification for every model.
- Shared mode is an opaque identifier. Provider adapters own its
meaning. Incompatible historical sessions remain fenced; exact committed
plan waits and task history remain inspectable.
- Native AskQuestion is excluded. Paperclip semantic questions are
supported. Authoritative per-run dollar accounting is unavailable;
partial counters remain diagnostics and unknown cost is not zero.
- Image input, detailed native diffs, deeper child transcripts and
native plan-file export remain follow-ups.
- macOS x64 has clean-install and daemon-startup proof under Rosetta,
not a separate live campaign on Intel hardware.
- Release only the tested package/image combination. Merging this PR
does not publish npm packages or deploy that image. Later builds need
their own release verification. Rollback disables new Cursor admission
while preserving records and recovery inspection.
- A model can fail an exact instruction: one cancelled-plan attempt
returned the wrong summary marker despite correct cancellation. The
unchanged repeat passed; both results remain in the report.

> ROADMAP.md was checked. This completes existing native Runner/Cursor
work; it does not add an independent core feature proposal.

## Model Used

OpenAI Codex, GPT-6. The exact serving variant and context window are
not exposed in this session. The agent used reasoning, repository
inspection, code execution, protocol tests and browser-backed Product
E2E tools. Cursor acceptance uses the explicit
`gpt-5.6-luna[context=272k,reasoning=medium,fast=false]` model. That is
the evaluated provider model.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass — affected suites passed;
full CI and the retained local failed attempts are recorded separately
above.
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green — 56 successful checks, one
neutral and four skipped on `f7ec5cc1f0e30c62a829c016ff2013a2a9d79037`
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups —
fresh Base review passed on `f7ec5cc1f0e30c62a829c016ff2013a2a9d79037`;
zero new comments and no unresolved threads. The earlier Apex finding
remains fixed.
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 20:48:15 -05:00
Devin FoleyandPaperclip eab93fd4a0 fix: checkpoint adapter usage and preserve unknown prices (#14991)
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-06 18:16:18 -07:00
DottaandPaperclip b508a05c43 feat: add internal agent complaints and suggestions (#15367)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents use legacy skills or native runner tools to work on tasks.
> - Those agents can encounter friction that does not belong in the task
thread.
> - A complaint should preserve the raw reaction. A suggestion should
describe an improvement.
> - This pull request adds attributed local storage and both submission
paths.
> - Agents can submit feedback once and continue their primary work.

## Linked Issues or Issue Description

**Subsystem affected**

Server, database, shared contracts, runtime skills, and native runner
tools.

**Problem or motivation**

Agents have no default internal channel for incidental complaints and
suggestions. Sending this feedback through task comments adds noise and
can alter task workflows.

**Proposed solution**

Store free-form feedback in the current instance database. Derive agent,
run, company, and task attribution from active authority. Provide
default legacy skills and provider-neutral native actions. Keep the
instructions close to Warp's MIT-licensed originals.

**Alternatives considered**

Task comments and external Slack delivery add unwanted side effects.
Mandatory suggestion fields and short editorial limits would discard
useful feedback. This release has no listing API, UI, read tool,
automatic triage, or external forwarding.

**Roadmap alignment**

This is a maintainer-requested addition to the existing runtime skills
and runner tool paths. It does not duplicate a listed roadmap milestone.
Searches for complaint tooling, suggestion-box, and agent commentary
found no overlapping public PR or issue.

## What Changed

- Add the company-scoped `agent_commentary` table, shared validation,
and idempotent migration `0310`.
- Add one transactional service and the agent-only POST route. Validate
active authority before writes or replay. Redact known credentials.
Commit a content-free audit with each new record.
- Add `submit_complaint` and `submit_suggestion` to standard, ask, and
planning modes. Keep review, revocation, and completion restrictions.
Store replay identity on the commentary row.
- Mount `complain` and `suggestion-box` by default for legacy agents.
Bundle a dependency-free Node.js stdin helper in the operational skill
and allow its POST through the sandbox bridge.
- Preserve Warp's complaint voice and suggestion guidance, with
attribution and local transport adaptations. Keep source attribution and
MIT notices in each skill's LICENSE, outside runtime instructions.
- Document custom-runtime HTTP use and database inspection. Add
real-database tests and a repeatable live Codex smoke for local and
Daytona execution.
- Pin the lagging-source migration fixture before the identity-repair
migration so later migrations preserve its regression coverage.

## Verification

- Personally ran real Codex submissions in all four environments on
2026-10-06. Local runs passed at 20:35 UTC. Daytona native passed at
20:31 UTC; Daytona legacy passed at 20:33 UTC. Each stored exactly two
rows with company, agent, run, and task attribution, wrote the
continuation marker, exited zero, created no task comments, and left
task status unchanged. Each recorded two content-free activity entries.
- Daytona used production provider hooks, real remote execution and file
transfer, the legacy queue callback bridge, and native private WebSocket
ingress. The current Linux runner was built from `abf47b595`, staged,
and verified against controller contracts. Both sandboxes were confirmed
deleted. This is a focused feedback transport smoke; it does not claim
full Runner E2E catalog or browser qualification.
- The immutable base image and Linux binary digest are recorded in [the
verification
documentation](https://github.com/paperclipai/paperclip/blob/codex/agent-commentary/doc/agent-commentary.md#verification).
The smoke script can save content-free JSON evidence. No credentials or
feedback bodies are in these reports.

| Environment | Runner | Complaint row | Suggestion row |
| --- | --- | --- | --- |
| local | legacy Codex | `59413a00-1de2-4bb1-bcc6-9c4b54c64aa6` |
`3db2364d-3e15-4f47-846f-875d3902999d` |
| local | native Codex | `5da22b5f-41df-4de5-8ba0-d9345ab01267` |
`2d5abe17-dd41-403c-a5ee-4729f2d58921` |
| daytona | legacy Codex | `27c9d0aa-8477-409f-9da0-e8ffa48dee50` |
`209681c9-d1e9-4ce1-999e-48fa07692389` |
| daytona | native Codex | `6eb001bb-4bcf-43f7-8717-f662f53dc7c3` |
`77c383d8-a997-49e5-a33e-25c70e15c0b2` |

- Run the local check with `node cli/node_modules/tsx/dist/cli.mjs
server/scripts/verify-agent-commentary-live.ts`. The documentation gives
the Daytona invocation. Both use disposable instance databases and
normal Codex provider usage.
- Repository `pnpm -r typecheck` and `pnpm build` passed after the test
extension. The build includes runner generation, contracts, and replay
checks. The smoke scripts also passed a separate TypeScript check. The
lagging-source migration regression passed. All equivalent current-head
Vitest CI shards passed. The local monolithic `pnpm test:run` invocation
was stopped after CI supplied that coverage; it did not complete
locally.
- Focused tests cover company isolation, spoofing, revoked credentials,
stale ownership, post-finish rejection, concurrent replay, conflicting
keys, atomic rollback, and deletion through existing services. Boundary
tests cover empty text, Unicode, text beyond 8,000 characters, and the
524,288-character ceiling without truncation. Mounting tests cover
Codex, Claude, and sandbox staging. Helper tests cover standalone Node
execution, stdin, invalid UTF-8, redirects, HTTP failure, and its
deadline. Privacy and bridge tests cover successful and rejected
requests.
- Instructions were compared with Warp's originals. MIT notices and
source credits live only in LICENSE files. Native tools preserve
truthful disclosure when asked, without routine announcements.
- [Full
CI](https://github.com/paperclipai/paperclip/actions/runs/37508559190)
and Greptile 5/5 passed on the earlier feature commit `5209c3501`. The
later head found the migration-fixture assumption fixed in this update.
On `8a4965164`, all 55 check contexts passed after one browser shard
rerun. Its initial reviewer signoff failure also passed an isolated
local browser run (1 test). Greptile scored that head 5/5 and identified
one smoke cleanup gap. `6ecbafb0b` fixes failed-acquisition cleanup with
four passing tests and a passing smoke-script typecheck. Fresh CI is
pending for this final test-only fix. No commentary production code
changed during verification.

## Risks

- Feedback is internally attributed. It is not anonymous. Existing
redaction removes known credentials, but agents must still omit
sensitive content. Normal provider transcripts can include their
submitted arguments.
- Default skill availability changes for existing legacy agents. Runtime
policy filtering still applies. The helper uses the existing Node.js
runtime with no extra dependencies; custom runtimes can call the HTTP
endpoint.
- Feedback is removed with its run, agent, or company. Task deletion
clears only the issue pointer. Normal database backups include the
table.
- The migration is additive and has no backfill. Writes serialize on the
active run for replay consistency. No server suggestion quota is
imposed.

## Model Used

OpenAI `gpt-6-astra` through Codex, with `xhigh` reasoning effort and a
reported 258,400-token context window. Capabilities used: repository
inspection, code execution, and live runtime verification. No subagents
were used.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 16:14:06 -05:00
DottaandPaperclip e38d6d16b6 feat(connections): add advanced provider setup and live browser qualification (#15341)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Users connect accounts and choose an agent harness and model.
> - The runtime change in #14970 supports custom providers on those
connections.
> - Normal setup must stay simple while advanced users can choose a
compatible gateway.
> - Shared connector rows and access controls keep these choices
consistent.
> - This pull request refines the agent setup UI and adds review stories
and repeatable browser qualification.
> - The qualification checks real tools and downloaded outputs, not only
a successful run status.

## Linked Issues or Issue Description

Refs #14970, #37, #13083, #14104, #14565, #12692.

The core implementation in #14970 is merged. This branch incorporates
its squash commit and targets `master`. Both PRs contain our
implementation. #14016 is a reference only and is not a dependency. This
PR has 96 changed files.

## What Changed

- Complete model-provider connector presentation beside other
connectors. Each row uses the existing Connect action and connection
list. Tags are stored without category UI. The base PR includes the
provider forms and routes.
- Show persistent Subscription, API Key, and Advanced choices. Label
Advanced as Custom Gateway. Reuse provider logos, connection lists, and
permissions controls. Default access to the organization and all agents
when permitted; keep narrowing controls under Advanced.
- Keep Configure reachable before subscription sign-in, so users can
select a supported environment when the default cannot sign in. Testing
and saving still require a connection. Show the execution environment in
Configure. Preserve the confirmed Connect choice. Editing a method,
credential, saved account, or advanced choice requires that current
choice to connect before testing or saving. Use matching model and
thinking-effort dropdowns and retain connection icons in selected
values.
- Preserve the new harness model default when switching an existing
OpenCode agent to Codex or Claude, and resolve user-selected model names
with the effective harness.
- Load popular OpenRouter models through the shared connection-model
discovery path. Keep explicit model lists and manual model entry
available.
- Group onboarding, connection setup, agent runtime, management,
recovery, and production-component stories under AI Connections /
Provider routing.
- Add an explicit-only provider-connections browser suite for managed
local or existing local/staging targets. Use private browser profiles
and credential handoffs. Support human-assisted subscription sign-in
without sharing passwords or tokens in reports.
- Verify persisted connection identity, runtime probes, tool execution,
exact artifact bytes, completion, and context-dependent follow-up.
Retain source/model provenance, cost bounds, closed error diagnostics,
original failures, and cleanup evidence.
- Add Gemini startup-model and skill-root fixes, Grok private-history
detection, ACP filesystem regression fixtures, selected-workspace
handling for local Hermes, and artifact-helper workspace fallback.
- Keep managed Grok runtime homes disposable. Remove host-side
transcript retention/restoration because private file modes do not
isolate same-user agent processes. Ignore earlier development archives
and use a fresh task handoff when history is unavailable. Verify the
absence of restored transcripts with a separate same-user process.
- Capture stopped-run diagnostics before deleting an attached-company
fixture agent. Track creation and owned sign-in receipts; revoke only
this attempt's accounts and never adopt a concurrent campaign's newly
created account. Preserve failure signals and final status through
cleanup.
- Require the requested environment in the saved agent and every run,
including follow-ups. Reject a forced incompatible target. Keep one
cancellation state through startup, every cell, reporting, and teardown
for SIGINT, SIGTERM, and SIGHUP. Stop further paid cells after
interruption. Document qualification limits.

## Verification

- Current head `b3bb3e94d577d43d9965a6b9daba039f599b2e49` includes
master `d9f600043`. The security fix in `a758fde31` passes full
workspace typecheck, production build, and 119 connection/Grok
regressions. The unchanged UI passes all 126
configuration/model-discovery tests and token gates. The final
published-guide correction passes Grok adapter typecheck. Earlier head
`eebd8225c` passed the complete deterministic runner suite (1,404 Vitest
tests and 128 Node tests) and all CI jobs. Current-head CI run
`37520147514` passed all 47 jobs, including the full sharded Vitest and
browser matrix, production build, and canary dry run. All 55 checks
completed: 53 successes and two expected skips. The current-head
security scan passed, Greptile is 5/5, and no review threads remain
open.
- A separate same-user process reproduced reading a restored Grok
transcript before the security fix. The regression now finds no
transcript. Existing fresh-session fallback and ordinary session
metadata behavior pass.
- The final account-choice and cleanup fixes pass 85 setup tests and 26
qualification-harness tests. Regressions verify that editing a
connection invalidates confirmation, Configure remains reachable before
sign-in, diagnostics are captured before fixture deletion, and
concurrent campaigns cannot adopt or revoke each other's accounts. UI
and E2E typechecks pass.
- The Storybook build and actual Chromium production-component stories
passed during this change. Review the neighboring AI Connections /
Provider routing stories, regular connector rows, three connection
modes, model discovery, and the single execution-environment control in
Configure.
- Cancellation smoke verified authenticated cleanup before browser close
for SIGINT, SIGTERM, and SIGHUP. Regressions cover interruption during
startup and reporting, missing-file ACP resource errors, and preserved
permission denials. Both ACP runtime versions and 54 ACPX/Grok
regressions passed. The deterministic connection-intent browser suite
passed two tests.
- Historical local qualification retained 43 passing API/gateway cells
out of 46, with downloaded outputs and follow-up receipts. These
attempts span earlier builds; they do not qualify this exact commit or
staging. Subscription combinations, Gemini overloads, and the unresolved
follow-up failure remain recorded rather than counted as passing.
- Use `pnpm test:e2e:runner -- --list --suite provider-connections` to
inspect the matrix. Follow `tests/runner-e2e/PROVIDER-CONNECTIONS.md`
for credentials, target URL, sign-in assistance, budget, evidence, and
cleanup. Paid live tests remain opt-in.

## Risks

- The core implementation in #14970 is merged. This PR adds no database
migration of its own.
- Subscription login needs an interactive provider session. Dedicated
accounts and staging qualification remain follow-up work; this PR does
not certify every login combination for production.
- Managed Grok transcript resume is deferred until provider history has
an OS isolation or authorized broker solution. Follow-ups start fresh
with Paperclip task context; earlier live Grok results do not qualify
this behavior.
- Gemini CLI 0.58.0 has an upstream ACP new-file error conversion
defect. Live overloads and one unresolved follow-up timeout remain
recorded. The stock CLI is unchanged, and those cases are not marked as
passing.
- Real-provider tests spend credits and use private credential/evidence
directories. The launcher requires explicit selection and checks target
ownership. It must not attach to a developer's database by accident.
- OpenClaw Gateway, Hermes Gateway, Claude Managed, AWS AgentCore,
Process, HTTP, and legacy ACPX local remain outside custom provider
setup.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, code
execution, and browser testing. The exact deployment model ID and
context window size were not exposed in this session.

## Checklist


- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 15:21:22 -05:00
DottaandPaperclip a590ab769d Give managed agents persistent cryptographic identities (#15352)
Give agents persistent Ed25519 identities encrypted with the existing instance master key. Create keys transactionally for new agents and lazily before supported managed runs, expose public identities in the API and agent UI, and protect private material during runtime delivery and output persistence.

Preserve identities in recovery backups while giving imported and development-cloned agents fresh keys.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-06 13:57:03 -05:00
Barış ÖZDEMİR bf14f803d5 fix(ssh): transport project repositories as their own git checkouts (#14782)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - A project can attach more than one repository. The task workspace
keeps the selected repository at its root and puts the other project
repositories under `.paperclip-repositories/<name>-<key>`, each with its
own `.git`
> - Agents can run on an SSH execution environment. Paperclip copies the
task workspace to the remote host before the run and restores it after
the run
> - The SSH copy excludes `.git` at every depth, but the restore
baseline excludes it only at the workspace root
> - So the other project repositories reach the remote host without Git,
and the restore then deletes their `.git` directories on the Paperclip
host
> - The next run of the same task fails during workspace setup, and the
agent cannot commit to those repositories on the remote host
> - This pull request transports each project repository as a Git
workspace of its own, the same way the sandbox path already handles them
> - The benefit is that multi-repository projects work on SSH
environments across consecutive runs

## Linked Issues or Issue Description

Refs #11632 (SSH workspace transfer exclude list). Related SSH workspace
PRs: #14233, #14428, #14472. I found no issue or PR for this bug.

**What happened?**

A project has two repositories and its agent runs on an SSH environment.
After the first run, the second repository under
`.paperclip-repositories/` has no `.git` directory on the Paperclip
host. The next run of the same task fails during setup with `Managed
workspace path "…/.paperclip-repositories/<repo>" already exists but is
not a git checkout.` On the remote host, `git` inside that repository
resolves to the parent repository.

**Expected behavior**

Each project repository reaches the remote host as a Git checkout with
its local changes. Remote commits and edits come back after the run. The
next run of the same task starts normally.

**Steps to reproduce**

1. Create a project with two repositories.
2. Configure an SSH execution environment and make it the agent's
default environment.
3. Assign a task to the agent and let it run once.
4. Look at `.paperclip-repositories/<repo>` in the task workspace:
`.git` is gone.
5. Wake the agent on the same task again: the run fails with
`setup_failed`.

**Paperclip version or commit**

Reproduced on `v2026.916.1` and on `master` (`5edf55d73`).

**Deployment mode**

Self-hosted (Docker), authenticated, with an SSH execution environment.

## What Changed

- `ssh.ts`: `prepareWorkspaceForSshExecution` lists the project
repositories under `.paperclip-repositories/`. It applies the discovery
rules of `readGitWorkspaceSnapshot`: each entry must be a directory with
a valid name and must be a Git repository root, else the prepare step
fails before any transfer.
- `ssh.ts`: the anchor copy leaves `.paperclip-repositories/` out. Each
project repository then gets the same import, sync, and deleted-path
steps as the anchor. The remote anchor repository ignores
`/.paperclip-repositories/`, as the local checkout does.
- `ssh.ts`: `prepareWorkspaceForSshExecution` returns the transported
repositories (the field is present only when there are repositories).
`restoreWorkspaceFromSshExecution` accepts them with their baselines. It
validates each path and baseline first, then restores the repositories
before the anchor and stops at the first failure, as the sandbox restore
does.
- `remote-managed-runtime.ts`: the anchor baseline excludes
`.paperclip-repositories/`, and each project repository gets its own
baseline for the restore merge.
- `ssh-fixture.test.ts`: regression tests for two consecutive managed
runs and for the direct restore path, on a workspace with a project
repository (commits, dirty edits, and a deleted file). Two tests for the
new validation.
-
`docs/guides/board-operator/execution-workspaces-and-runtime-services.md`:
one line about project repositories in the SSH round trip.

## Verification

- The new regression test fails on `master` (`expected 'backend
initial\n?? ../\n' to contain 'frontend initial'`) and passes with this
change.
- `PAPERCLIP_ENABLE_DARWIN_SSH_ENV_LAB=1 npx vitest run
packages/adapter-utils/src/ssh-fixture.test.ts
packages/adapter-utils/src/remote-managed-runtime.test.ts`: 32 passed,
with the sshd fixture running.
- `tsc --noEmit` passes for `packages/adapter-utils` and `server`, and
`pnpm -r typecheck` passes for the other workspaces. The Rust step of
`@paperclipai/paperclip-runner` did not run locally because `cargo` is
not installed.
- `node ./scripts/check-no-git-push.mjs` and `pnpm
check:module-boundaries` pass.
- `pnpm test:run` did not complete locally. Before it stopped, 5 tests
failed: 2 in `server/src/__tests__/workspace-runtime.test.ts` and 3 in
`server/src/__tests__/company-skills-service.test.ts`. The same 5 tests
also fail on the base commit `5edf55d73` without this change. CI runs
the full suite.
- `pnpm build` passes for all workspaces except
`@paperclipai/paperclip-runner` and `server`, because their build
compiles the Rust runner binary and `cargo` is not installed. `tsc
--noEmit` passes for `server`.
- Manual test on a self-hosted `v2026.916.1` instance with the same
change applied: a project with two repositories and an SSH environment.
Two runs on the same task passed. After each run, the second repository
keeps its `.git` on the host. On the remote host it is a Git checkout,
and the remote anchor ignores it.

## Risks

- Low risk. Workspaces without `.paperclip-repositories/` take the same
path as before, and the return value is unchanged for them.
- A workspace with an invalid entry under `.paperclip-repositories/` now
fails the SSH prepare step. The sandbox path already rejects such
entries.
- If one repository fails to restore, the restore stops, as in the
sandbox path. The remote run directory keeps the agent's work.
- Each project repository adds one bundle import and one restore per
run. The time grows with the number and size of the repositories.
- Out of scope: other nested `.git` directories (for example a vendored
checkout inside a repository) keep the existing SSH behavior.

## Model Used

- Provider and model: Anthropic Claude Opus 5.5 (`claude-opus-5-5`), in
Claude Code.
- Capabilities: extended thinking, tool use, and code execution. The
context window size was not recorded.
- Use: the model investigated the bug, wrote the change and the tests,
and ran the checks. A separate Claude Code agent reviewed the diff. The
author reviewed the change. The manual test ran on the author's
self-hosted instance.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-10-05 22:21:00 -07:00
Devin FoleyandPaperclip efac8ff2f4 Add bounded Git integration restore diagnostics (#15291)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Sandbox runs must restore their workspace before finalization can
succeed.
> - Restore diagnostics identify the failed phase and operation.
> - A Git integration exit code can still describe several different
failures.
> - This pull request adds fixed command labels and supported failure
classes.
> - Operators can distinguish these failures without collecting private
Git output.

## Linked Issues or Issue Description

Refs #15005. This branch includes merged #15268. It labels that change's
locked ref transaction and nested branch probe without changing their
behavior.

**What happened?**

A failed Git integration can report only `git_integration`, `unknown`,
and an exit code. That evidence does not identify the failed command.
Some Git versions also return exit 1 for both a merge conflict and an
invalid object.

**Expected behavior**

Record a fixed command family and a supported failure class. Keep
unknown cases as `unknown`. Exclude command arguments, process output,
paths, repository URLs, filenames, and ref names.

**Steps to reproduce**

The tests create local repositories with conflicting commits, a missing
object, and an expected-old ref mismatch. They call the real Git
operations and inspect the resulting diagnostic. No hosted workspace or
external provider is used.

**Paperclip version or commit**

Base: `e99854249c`.

**Deployment mode**

Built from source. The diagnostic applies to sandbox workspace restore.

## What Changed

- Label Git integration calls with a closed command enum. Preserve
arguments, options, errors, and retry behavior.
- Recognize supported object and ref errors and OS permission codes.
Require both exit 1 and completed tree output for a merge conflict.
- Carry the closed fields through the existing restore receipt, saved
adapter result, and Sentry projection. Revalidate saved metadata before
projection.
- Test nested wrappers, parallel failures, reused errors, handled
probes, result settlement, and privacy with the real Sentry SDK.
- Document the field contract and its limits.

## Verification

- Eight focused suites pass: 295 tests. These cover Git sync, restore
diagnostics, result settlement, teardown, sandbox runtime, failure
projection, the real Sentry SDK, and native warm-workspace Git history.
- `PAPERCLIP_REQUIRE_SENTRY_TEST_SDK=1 pnpm exec vitest run
packages/adapter-utils/src/workspace-restore-diagnostics.test.ts
packages/adapter-utils/src/workspace-restore-result.test.ts
packages/adapter-utils/src/git-workspace-sync.test.ts
packages/adapter-utils/src/workspace-restore-teardown.test.ts
packages/adapter-utils/src/sandbox-managed-runtime.test.ts
server/src/services/__tests__/run-failure-diagnostics.test.ts
server/src/__tests__/run-failure-sentry-real-sdk.test.ts
server/src/__tests__/native-workspace-sync-history.test.ts`
- Local checks use Node 24.21.0 and the pinned pnpm 9.15.4. The optional
Sentry SDK is pinned to the declared 10.71.0 and uses an in-memory
transport.
- `pnpm -r typecheck` and `pnpm build`: pass on the rebased head.
- The full local `pnpm test:run` was stopped before source changes when
#15268 merged and required a rebase. It reported three pre-existing
company-skill cache test failures on macOS. The same three cases fail on
clean bases `2c43b39167` and `e99854249c` and the earlier head with
`EACCES` when publishing a read-only cache directory. The relevant test,
service, and cache source blobs are identical. No cache changes are
included. Later local test groups were not reached.
- All 54 checks pass on exact head `8448896bc6`, including the
post-ready security scan, with two intentional Storybook skips. Greptile
scores this head 5/5 with no review threads. Full Linux CI covers the
local groups that were not reached.
- Independent review found no blocker. Its diagnostics, Git workspace,
and native history suites pass 115/115 on this head.
- A source comparison confirms that removing only diagnostic wrappers
yields the merged upstream Git integration code exactly, including ref
locks, transaction protocol, arguments, options, and retry decisions.
- The clean base has stale dependency overrides in its lockfile. Local
installation resolved them as the existing PR CI fallback does. No
manifest or lockfile change is included.
- `git diff --check` and a local secrets/PII review pass.

## Risks

A Git version or localized message may not match a known failure form.
Such cases remain `unknown`. Up to 16 KiB of stderr and the bounded
tree-ID prefix of stdout are inspected only in memory. The saved fields
contain enum values only. These diagnostics do not establish workspace
recovery or authorize retries. Restore decisions and Git mutations are
unchanged. No schema migration is required.

## Model Used

OpenAI Codex, GPT-6, with reasoning, tool use, and code execution. The
service does not expose the exact model deployment ID or context-window
size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-05 17:10:05 -07:00
Devin FoleyandPaperclip e99854249c fix(workspaces): restore rebased sandbox history against its starting snapshot (#15268)
## Thinking Path

> - Paperclip manages agents and preserves their work across runs.
> - Sandbox execution restores Git history and files to the host
workspace.
> - An agent can rebase or amend its branch before it finishes.
> - Restore currently treats the original and rewritten tips as
concurrent work.
> - That merge can conflict even when the host has not changed.
> - This change uses the starting Git snapshot to accept rewritten
history safely.

## Linked Issues or Issue Description

**What happened?**

A successful sandbox turn can end with `workspace_restore_failed` after
the agent rebases and pushes its branch. The restore step merges the
original host tip with the rewritten sandbox tip. This can recreate
conflicts that the agent already resolved.

**Expected behavior**

Accept the rewritten history when the host still has the recorded
starting branch and commit. Preserve concurrent host work through the
existing merge and recovery paths.

**Steps to reproduce**

1. Start a sandbox from a feature branch.
2. Rebase that branch onto an upstream commit that changes the same
file. Resolve the conflict in the sandbox.
3. Restore the sandbox while the host remains on the original commit.
4. The old implementation attempts a conflicting Git merge and fails the
run after the agent finishes.

**Paperclip version or commit**

Reproduced on `cab4263dc9` with a real Git rebase fixture.

**Deployment mode**

Self-hosted server with sandbox execution.

Related work: #15005 records restore failure stages. #11638 preserves
unrelated imported history with a graft. #10601 handles bundle
prerequisites. Those changes do not distinguish a rebase from a
concurrent host edit.

## What Changed

- Pass the run's starting Git branch and commit into sandbox history
integration.
- Adopt a related rewritten tip when the host still matches that
snapshot. Keep the expected-old-value ref update and bounded retry.
- Verify the branch attachment inside a prepared Git transaction while
Git holds its ref locks. Abort if a checkout changed the branch.
- Check host and sandbox Git identity before warm reuse, including
nested repositories. Restage when their tips or branches differ.
- Reject host branch changes and unrelated imports after a concurrent
host commit.
- Retain the existing unrelated-history graft for an unchanged host and
the conservative behavior for callers without a snapshot.
- Export a full bundle for an intentional reset to an ancestor so
restore receives the actual sandbox tip.
- Add real Git and sandbox restore regressions. Document the restore
contract.

## Verification

- Eight Git sync, sandbox restore, and native workspace suites pass: 255
tests.
- The new checkout-race and warm-reuse regressions failed before the
fixes. Real Git hooks verify that prepared transactions prevent a
concurrent HEAD change.
- `pnpm -r typecheck` passed after rebasing onto `984f092ddf` and
applying the Apex findings.
- `pnpm build` passed on `f5132603d6`.
- The earlier `pnpm test:run` attempt reported three
`company-skills-service.test.ts` failures on macOS (`EACCES` renaming a
read-only staging directory). The same failures reproduced on unchanged
master. The broader run was stopped after confirming that baseline
failure; it was not a full-suite pass. Those source and test files are
unchanged in the current base.
- [Apex
review](https://github.com/paperclipai/paperclip/pull/15268#issuecomment-6002549219):
5/5 on `f5132603d6`, requested with `@greptileai apex review`. All three
historical findings are addressed and all review threads are resolved.
- All CI gates passed on `f5132603d6` (53 successful checks, two
skipped). The signoff-policy browser fixture initially timed out waiting
for a local process-agent run; its one retry passed without code
changes. [CI
run](https://github.com/paperclipai/paperclip/actions/runs/37387511152)

## Risks

- A recorded starting snapshot now authorizes replacement of related
rewritten history. A stale or changed host tip retains the
concurrent-history path. Ref writes still compare the expected old
commit.
- Branch changes and unrelated rewrites after host advancement require
recovery instead of replacing host work.
- An intentional reset to an ancestor uses a full bundle. Large
histories can increase transfer time in that case.
- The existing directory merge rules remain in force. The fix does not
resolve an earlier failed restore or replay its external actions.

## Model Used

OpenAI GPT-6 via Codex, with reasoning, repository analysis, code
editing, and local test execution. The exact deployment model ID and
context window were not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass — 255 affected tests; the
broader-suite baseline failure is documented above
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-05 16:43:15 -07:00
DottaandPaperclip 4857799a88 feat(connections): deliver saved instructions to authorized agent turns (#15216)
Persist optional connection instructions and deliver authorized snapshots to agent execution prompts. Keep provider templates with each app definition, preserve edits and opt-outs, and replace sessions when guidance or access changes.

Use shared production settings across setup and Permissions, with source visibility in agent Instructions. Add the initial memory-provider defaults and managed Honcho workspace configuration. Include migration 0298 and regression coverage for generic providers, runtime delivery, authorization, and catalog regeneration.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-05 18:16:00 -05:00
Devin FoleyandPaperclip 3b47a6befd Record bounded workspace restore failure stages (#15005)
Capture allowlisted restore substep and failure metadata for future diagnosis while preserving workspace recovery behavior, error identity, cleanup ordering, and privacy.

Validation: exact-head Greptile 5/5, passing CI, no unresolved review threads, and a clean merge. Detailed verification and limitations are recorded in the pull request.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-05 10:58:32 -07:00
DottaandPaperclip b17019e14d fix(agents): reduce default instructions and qualify stock harnesses (#14948)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Its adapters supply task context and access to Paperclip skills and
tools.
> - The default hire manual and shared prompts also repeat general work
procedures.
> - Those procedures overlap with stock provider instructions and the
Paperclip skill.
> - Existing E2E fixtures supply a QA manual, so they do not qualify the
production default.
> - This pull request reduces the generic instructions and adds real
default-hire coverage.
> - The benefit is less competing guidance, with inspectable evidence
for preserved skills and task context.

## Linked Issues or Issue Description

Refs: #14920. That merged change preserves native Codex base
instructions. This PR covers the default manual, shared legacy prompts,
operational skill guidance, and the narrowly approved ACP
skill-discovery/session-environment repair for measured delivery and
credential-persistence failures.

**What existing behavior does this improve?**

New non-CEO hires without a custom bundle and legacy task/chat startup
and continuation prompts.

**Current behavior**

The shipped default manual contains 602 words. Generic task/chat prompts
and ordinary resume deltas repeat work procedures already available
through the harness and Paperclip skill.

**Proposed behavior**

The default manual contains only the eight-word company identity. Shared
startup prompts retain identity and connection guidance. Ordinary resume
deltas retain current work context without the generic execution
contract.

**Reason and benefit**

Let the stock harness guide general work. Keep Paperclip-specific
capabilities and independently test default hires, skills, ordered
comments, and chat restart.

**Breaking changes**

New default hires receive less guidance. Existing saved manuals,
explicit custom bundles, CEO templates, and specialized wake contracts
retain their behavior. The obsolete includeExecutionContract option
remains accepted for source compatibility.

## What Changed

- Reduce the default hire manual to one sentence.
- Reduce shared task/chat defaults and remove the generic
ordinary-resume contract.
- Keep connection guidance, auth, skills, custom prompts, and
specialized wake context.
- Add credential-free instruction-boundary gates and 26 explicit Product
E2E cells across eight legacy/native profiles, including two focused
Paperclip-storage cases.
- Capture public hire receipts before providers run, then grade
delivered prompts and independent task/chat outcomes.
- Add an early legacy skill API recipe for saving a task document,
checking the saved revision receipt and linking the document. Improve
stock task/heartbeat skill-selection metadata and show a clickable
Markdown UI-link example. Keep native tool completion separate.
- Advertise bounded routing descriptions and exact successfully staged
SKILL.md paths in legacy ACP Claude; keep full bodies on demand and
preserve remote path rebasing.
- Remove only the provider environment from copied persisted ACP session
records, while loading current run credentials and preserving all other
options/conversation state.
- Regenerate both capability metadata inventories and reject stale
manifests/inventories before provider admission.
- Publish the original reduction and focused skill-repair comparisons,
preserving all failures, automatic recovery, cost coverage and
limitations.

## Verification

**Behavioral qualification remains pending.** Original legacy ACP Claude
loses the issue document only in the reduced cohort beneath an unchanged
credential failure. A source-backed diagnosis finds that neither
ordinary assignment reads the staged operational skill, while the
runtime persists provider environment in session state. The new common
repairs expose skill metadata/path and omit persisted env; strict
document and credential guards stay intact. [Inspectable diagnosis and
retained
hashes](https://github.com/paperclipai/paperclip/blob/9f654db4541d3d002769c988f6e51fc0b08dadbd/doc/plans/2026-10-03-legacy-acp-claude-readiness.md).

Current repair head `de0965984ff3edf611ae6d0e7ca5c7d5ae3947bb`
incorporates master `569c7203aa24b95440682983ce7940ba1d4247bd` (merged
#14961/#15007). All 222 affected adapter tests, adapter-utils/E2E
typechecks, and final 96 variant/grader/retry calibrations pass. The
frozen historical comparator is
`c25697f4260b6f3adfea143c3ae9932e2f42986d`: 8,280 of 8,291 paths
identical, exactly two production instruction paths plus nine declared
unit expectations differ. The operational skill/discovery/environment
repairs, selected model/profile/task/core grader/auth/permissions/retry
policy are identical. Both actual launcher prepare→verify admissions
pass with zero providers. [Immutable manifest and exact
receipts](https://github.com/paperclipai/paperclip/blob/9f654db4541d3d002769c988f6e51fc0b08dadbd/doc/plans/2026-10-03-legacy-acp-claude-evidence/manifest.json).

One original legacy ACP Claude cell per variant is authorized, with
enforced single campaign attempts, 12-minute deadlines and company/agent
1,000-cent hard stops; every product recovery run/cost is counted.
Actual live outcomes are pending. Current normal CI has one failed
server shard and failed aggregate verify under diagnosis; other normal
gates including typecheck/build/Rust/all eight browser shards pass.
Fresh review completed successfully; the valid historical startup/resume
masking finding was fixed with per-invocation task/chat checks and
strict complete-snapshot capture, calibrated and resolved. Prior heads,
failures and campaigns below remain historical evidence, not checks on
this repair head.

- Prior head `36aa4d81c49a1a8f6f04b1a068fae19aa901955f` is replayed on
merged hiring master `862a5758ba0e88a33232c1f1fa645e85c38a3113`. All 52
current-head checks pass with two intentional Storybook skips, including
repository typecheck/test/build and the browser shard. Fresh Greptile is
5/5 with zero unresolved review threads. Exact-head stock prerequisites
pass 599 assertions (598 TypeScript + 1 Rust), all six gates and
retained receipt verification, zero providers/source errors. Fingerprint
`a7f5a22d860a88fe20cce213c6d5e0004788f32c8363930729aea4fd740ad16d`.
Combined catalog/hiring calibrations pass 67 assertions, E2E typecheck
and 26-cell stock discovery pass. Canonical contract/inventory checks
and the later issue-derived reference calibration are retained; that
reference-only follow-up is not live-qualified by earlier frozen runs.
- Prior full repository typecheck/build passed. The complete local
Vitest run executed 14,956 tests: 14,870 passed, 83 skipped, three
timing failures. All three affected files passed unchanged narrow
reruns; original failures remain retained. Current-head CI now passes
the full general checks; the original local failures remain retained.
- The original 24-pair default-manual/shared-prompt comparison has two
new overall classic Claude/OpenCode document-delivery failures plus an
additional legacy ACP Claude document loss beneath an unchanged
credential-guard failure (not closed by later runs), two newly passing
OpenCode ordered cases, seven unchanged failures and 13 unchanged
passes. Equal 15/24 totals do not establish behavioral equivalence.
[Complete original
report](https://github.com/paperclipai/paperclip/blob/875f4c397d9e8c3f12f39dedd59abaf1eaf5236e/doc/plans/2026-10-02-stock-harness-live-comparison.md).
- The skill-only repair holds the eight-word manual/shared prompts and
merged #14920 fixed. All four matched profile configurations and 203
fixture/behavior files match. Candidate
`abd0b628ca642c09a54a4edc56a5227402f6686e` varies only the two skill
sources against baseline `bc83fe030234439ac51279502a28803958963e2e`.
[Candidate
workflow](https://github.com/paperclipai/paperclip/actions/runs/37060885547)
and [baseline
workflow](https://github.com/paperclipai/paperclip/actions/runs/37060888047)
each pass 571 exact-source prerequisites before providers; all eight
cells clean up successfully. Failed campaigns publish successfully and
remain failed.
- Repair pairs: Claude original Fail → Pass; Claude explicit Pass →
Pass; both OpenCode cases Fail → Fail. Explicit OpenCode's handoff
worsens beneath the unchanged failing UI-link grade: baseline gives a
clickable API URL, candidate gives a code-formatted path without an
anchor. The request's usable-link wording is narrower in the UI-only
oracle. [Complete repair report and safe
projection](https://github.com/paperclipai/paperclip/blob/875f4c397d9e8c3f12f39dedd59abaf1eaf5236e/doc/plans/2026-10-02-legacy-document-skill-repair.md).
- The subsequent narrow stock metadata/link correction has two matched
Pass → Pass cases, zero new machine failures/passes and no pending
pairs. Both original-case handoff links remain deficient: candidate uses
a wrong PAP prefix, baseline supplies a bare prefix-less slug path; the
preserved original oracle only requires a durable document. Both
explicit clickable UI-link cases pass revision/content/link grading. All
four exact-source 587-check gates, single assignment runs and cleanup
pass. This does not establish fix causality because baseline also
succeeds. [Candidate
workflow](https://github.com/paperclipai/paperclip/actions/runs/37069547401)
freezes `fe9dc1e3c518825242ed889ab9c8352986f8c2ed`; [matched
baseline](https://github.com/paperclipai/paperclip/actions/runs/37069552374)
freezes `0d7ecfa96d72fba79b7f0a25052b42c0686c0488`. This is a skill-only
comparison with reduced manuals/shared prompts held constant, not a
repeat of the historical-manual comparison. Only original and clarified
explicit classic OpenCode cases are selected, two per variant/four
expected turns. 8,242 other tracked files and both profile hashes match;
protected workflows admit each exact source before credentials.
[Complete qualification
report](https://github.com/paperclipai/paperclip/blob/74d0d3d945f4c52d0814b5a845ab5bd09f33cd6b/doc/plans/2026-10-02-opencode-skill-routing-link-qualification.md).
Candidate original loads Paperclip/reference before saving publicly;
baseline original loads it after writing locally, then saves publicly
within the same assignment. Reported cost totals are $0.0107824490
candidate / $0.0107909015 baseline, with unmetered runtime. The later
reference-only issue-derived link correction is provider-free calibrated
and **not live-qualified** by these frozen runs; no further paid runs.
- Retained tool calls show the repaired original OpenCode assignment
loads only its assigned output skill before writing locally. Operational
Paperclip is first loaded during automatic disposition recovery; its
early recipe is visible then, but it never saves the missing document.
Explicit candidate loads Paperclip and reads the new reference before
saving successfully. All nine actual runs are counted. Reported LLM
totals are $0.3802537209 baseline and $0.4918990161 candidate; local
runtime is unmetered.
- Initial setup, packaging, cancelled/missing-cell recovery, callback
test and relative-output attempts remain retained. No completed provider
failure was rerun. Frozen measurement branches are unchanged by later
canonical metadata maintenance.
- Run `pnpm test:e2e:runner:stock-harness`, `pnpm test:e2e:runner:unit`,
and `pnpm test:e2e:runner:typecheck`. Select `stock-harness` explicitly
for paid execution; it is excluded from `--all`.

Prior-head integration: `36aa4d81c49a1a8f6f04b1a068fae19aa901955f`
replays this PR on merged hiring #14985
(`862a5758ba0e88a33232c1f1fa645e85c38a3113`), preserving the four
explicit custom-CEO-bundle checks, minimal generic manual boundary, and
both suites. The combined fixture catalog and hiring calibrations pass
67 assertions; exact-head stock prerequisites pass 599 assertions (598
TypeScript + 1 Rust), all six gates and retained-receipt verification,
zero providers/source errors, fingerprint
`a7f5a22d860a88fe20cce213c6d5e0004788f32c8363930729aea4fd740ad16d`. E2E
typecheck and 26-cell stock discovery pass. Fresh current-head CI passes
all 52 checks with two intentional skips, and fresh Greptile is 5/5 with
zero unresolved review threads.

The prior source-plan browser failure is retained: a deterministic
process fixture replayed its last `fixture:plan` command on
`chat_task_completed`, writing revision 2 with identical body after the
approval handoff. This was not paid provider execution. Rebased
current-head CI passes the same assertion without an old-head retry or a
change to that browser fixture.

The merged hiring change was measured separately on immutable matched
unions, with this reduced/shared/operational context and native
completion guidance held constant. [Complete original two-profile
report](https://github.com/paperclipai/paperclip/blob/f0512647656be78e48abd8c22a3078db8bf6bcd2/doc/plans/2026-10-02-hiring-template-live-comparison.md):
[candidate](https://github.com/paperclipai/paperclip/actions/runs/37075466208)
/ [historical
baseline](https://github.com/paperclipai/paperclip/actions/runs/37075469463),
705 provider-free prerequisites each. Both pairs are unchanged Fail →
Fail on the exact-five count, with six core delivery checks passing all
four cells; 28 actual successful runs include eight automatic completion
wakes, zero retries, four successful cleanups. Source-read coverage is
uncomparable, actual model charges unknown. Separately versioned
provider-free accounting remains analytical work; original verdicts are
preserved. This does not rerun or qualify the completed default-manual
or native campaigns.

## Risks

- Legacy ACP Claude's additional delivery loss is not closed by any
later matched run and blocks the no-extra-failing-behavior merge
criterion. Legacy document delivery may have relied on the prior
manual/shared prompts. The early skill repair improves Claude in one
trial; the later OpenCode pairs pass in both variants and cannot
establish causality or robust recovery. Both original-case links remain
deficient beneath the storage-only grade. The later issue-derived
reference correction has only provider-free validation. Native
finish/block descriptions must not be supplied to legacy agents.
- The comparison holds merged native Codex fix #14920 constant; it
cannot measure that fix's before/after task performance.
- These bounded skill/context/chat workflows do not measure general
coding quality. Unrepresented providers remain unqualified.
- Saved manuals and old Codex sessions are not automatically migrated.
Codex through ACP still has a separate base-instruction follow-up.

## Model Used

OpenAI Codex, GPT-6 family as identified by this session. The exact
deployment ID and context-window size are not exposed. The assistant
used reasoning, repository tools, code execution, and delegated PR/eval
work.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` /
`Closes: #` / `Refs: #` OR (b) described the issue in-PR following the
relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (relevant suites and all
three unchanged narrow reruns pass; complete-run timing failures
retained in Verification)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green on the new repair head
(prior-head checks retained above)
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
on the new repair head
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-03 12:32:42 -05:00
Devin FoleyandPaperclip 1815474597 fix: report pending execution phase at Stop timeout (#14990)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The server owns each adapter execution and waits for it to settle
after Stop.
> - A Stop timeout reports that termination remains unverified.
> - Existing phase timings arrive only after their work completes, so a
stalled await has no timing.
> - This pull request samples the pending phase when the Stop timer
expires.
> - Operators can identify the pending operation without treating
diagnostics as stop proof.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

The opt-in Sentry context for an unconfirmed adapter Stop timeout.

**Current behavior**

The timeout includes execution identity but no pending phase. A session
close, instruction collection, workspace restore, or diagnostic write
can remain pending without producing its completion timing.

**Proposed behavior**

Add a closed-list phase and elapsed milliseconds from the exact live
execution control. Sample them when the timeout fires. Report `unknown`
and a null age when the control or attribution is unavailable.

**Reason and benefit**

The next timeout can identify which operation is still pending. It does
not require task text, paths, provider output, or additional database
writes.

**Breaking changes**

No API or execution behavior change. The existing opt-in error context
gains two fields. Related public work: #14639 added Stop identity
diagnostics; #14866 and #14945 cover instruction cleanup and teardown
outcomes. This change adds pending attribution to those paths. The
native Stop work in #14802 remains separate.

## What Changed

- Add a bounded tracker per execution control. Token scopes support
nested and overlapping awaits. A late release cannot clear a newer
scope.
- Track adapter execution, ACP cancellation and settlement, diagnostic
writes, and host cleanup. Keep a coarse host scope until the executor
finishes.
- Sample only the matching current control and settlement promise at
timeout. Freeze the sanitized result. Use a monotonic clock and cap
elapsed time at one day.
- Test stalled operations, repeated Stop calls, stale and wrong-run
controls, callback failures, scope bounds, and the real Sentry SDK
context.

## Verification

- `pnpm -r typecheck` passed.
- `pnpm build` passed.
- Six focused suites passed: 79 tests. They cover pending scopes, Stop
control ownership, real ACP settlement stalls, and Sentry context
isolation.
- The real Sentry SDK contract ran with the audited optional peer
`@sentry/node@10.71.0` installed outside the workspace. Valid phase and
elapsed values were exported; arbitrary labels and nonfinite elapsed
values were rejected.
- An independent agent reviewed the production diff and ran the focused
tests without blockers.
- `git diff --check` and a redacted Gitleaks scan passed. The local full
`pnpm test:run` was stopped during its large serial server batch to
avoid duplicating the sharded CI suite. No complete local broad-suite
pass is claimed.
- All CI gates passed on `c223237b58aa8d479d1c66d7d399de30270bac4f`: 54
successful checks and two expected Storybook skips. This includes the
full sharded test suite, typecheck, build, real Sentry SDK isolation,
browser tests, canary dry run, and the security scan after the PR became
ready for review.
- Greptile scored the same commit 5/5 with no actionable findings or
unresolved review threads.

## Risks

This is diagnostic instrumentation. Cancellation, deadlines, teardown
order, termination proof, and file recovery proof remain unchanged.
Unsupported or uninstrumented work uses a coarse phase. Tracker overflow
fails closed to `unknown`. The tracker emits no new run-log or Telemetry
event. The existing Sentry opt-in gate remains in place.

## Model Used

OpenAI Codex, GPT-6. The agent used code inspection, local command
execution, automated tests, and an independent agent review. The runtime
did not expose a more specific model identifier or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-02 16:55:38 -07:00
Devin FoleyandPaperclip 4abff286c2 fix: retain resolved ACP execution timeout metadata (#14986)
## Thinking Path

> - Paperclip manages agent work through adapters and heartbeat runs.
> - ACP adapters resolve a timeout for the selected execution target.
> - An untouched sandbox timeout uses a four-hour default.
> - Heartbeat finalization rebuilt the metadata from the stored zero.
> - This made a timed-out sandbox run report an effective timeout of
zero.
> - This change retains the adapter's resolved policy for accurate run
diagnostics.

## Linked Issues or Issue Description

**What happened?**

ACP sandbox runs with `timeoutSec: 0` use the four-hour default. Their
terminal metadata reports `effectiveTimeoutSec: 0` and `timeoutSource:
config` because heartbeat finalization only reads the stored agent
configuration.

**Expected behavior**

The result must report the policy the adapter used: `14400` and
`sandbox_default`. Explicit limits, fractional limits, explicit
unlimited overrides, and local defaults must keep their resolved values.

**Steps to reproduce**

Run an ACP adapter on a sandbox target with `timeoutSec: 0`. Compare the
start log's four-hour policy with the terminal result's effective
timeout. The new tests exercise the adapter result and the heartbeat
metadata merge without waiting four hours.

**Paperclip version or commit**

Reproduced from `6eaf218924f0a89faf1c02eb0d6877a6c5c8a2cb`.

**Deployment mode**

Self-hosted server with an ACP sandbox execution target.

Searched open timeout and metadata issues and PRs. Related #14804
exposes timeout configuration in forms; #14496 proposes a default
policy; #14833 addresses CLI session retention. This PR changes only ACP
result metadata.

## What Changed

- Retain the resolved timeout in the ACP result after settlement.
- Use validated adapter resolution when merging terminal timeout
metadata. Preserve config fallbacks for older adapters and the HTTP
millisecond policy.
- Test sandbox defaults, explicit and fractional limits, explicit
unlimited overrides, local defaults, malformed metadata, and unchanged
cancellation fields.
- Document the result fields and their meaning.

## Verification

- `pnpm -r typecheck` passed.
- `pnpm build` passed.
- Stop metadata tests: 23 passed.
- Reporter and diagnostic suites: 84 passed; two real-Sentry-SDK tests
skipped because the optional SDK is not installed.
- ACP engine suite: all 206 tests passed with a deterministic local
`gemini --version` shim. Ambient host CLI probes made the existing
Gemini session-resume fixture intermittent (one assertion failure in
each of two broad runs); an isolated 15-case rerun and the clean-base
206-test suite also passed. No assertion or timeout was changed.
- `pnpm test:run` is in progress. This PR does not claim a complete
local suite pass.
- Independent review found no blocking issues. The tests cover real
adapter emission and the real metadata merge separately.

## Risks

Low runtime risk: this changes result diagnostics. It does not change
timeout values, cancellation acknowledgement, cleanup, checkpoint
safety, retries, or provider operations. It does not fix the cause of a
quiet or long-running tool. Older stored results are not rewritten. The
new source values apply only when an adapter returns a valid resolution.

## Model Used

OpenAI GPT-6 with reasoning, repository inspection, code editing, and
test execution. The deployment-specific model ID and context-window size
are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-02 16:09:36 -07:00
Devin FoleyandPaperclip 22cea6b2e6 fix: bound sandbox bridge waits and flag silent runs sooner (#14979)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Sandbox agents exchange input and output through bridge control
commands.
> - A provider can stop responding to a command even when it receives a
timeout.
> - These small commands can inherit a four-hour agent lifetime and
block input or teardown.
> - The board also calls a silent run healthy for the first hour.
> - This pull request bounds bridge control waits and surfaces silence
sooner.

## Linked Issues or Issue Description

**What happened?**

A sandbox run can remain active when a bridge control command never
returns. The shared helper passes a timeout to the provider but does not
enforce it on the host. It also accepts the agent's hours-long timeout.
Output silence remains `ok` for an hour and becomes `critical` only
after four hours.

**Expected behavior**

Bound short bridge operations even if the provider never settles. Report
failed input delivery through the existing shutdown path. Warn after
five silent minutes and escalate after fifteen. Keep normal agent
command limits and require verified termination before releasing
execution ownership.

**Steps to reproduce**

1. Use a sandbox runner whose bridge read or input-upload promise never
settles.
2. Set its configured timeout to four hours.
3. Observe that the old queue client never returns or rejects.
4. Inspect a running task with 35 minutes of output silence. The old
summary still reports `ok`.

**Paperclip version or commit**

Base commit `d6d88b9de2`.

**Deployment mode**

Self-hosted server with sandbox execution.

**Agent adapter(s) involved**

Shared command-managed sandbox bridge, including Codex ACP sessions. The
informational silence thresholds apply to active runs across adapters.

Related: #14889 recovers stalled Daytona output streams; #14485 retries
explicit gateway failures during input delivery. This change bounds
short control operations whose provider promises never settle. It does
not add tool replay or automatic cancellation for output silence. #6297
proposes configurable per-agent silence thresholds; this patch only
changes the existing defaults.

## What Changed

- Enforce at most 30 seconds per bridge control shell command on the
host and provider, including callback startup and shutdown,
process-session launch, and payload setup. Preserve shorter configured
deadlines and launch environments.
- Keep the long-lived agent command outside this deadline. Use a fixed
timeout diagnostic without command payloads.
- Surface suspicious output silence after five minutes and critical
silence after fifteen minutes.
- Decouple the shared-workspace holder cutoff from warning thresholds
and preserve its existing one-hour value.
- Add regressions for hung reads, a late upload response, failed input
delivery, exact warning boundaries, and fresh output clearing warnings.
- Update the adapter guide and execution contract.

## Verification

- The three new queue-client regressions fail on the unchanged base and
pass with this patch.
- Final callback bridge and sandbox session suites: 214 passed. These
cover hung reads, writes, startup, shutdown, process-session launch,
payload setup, and the separate long-running agent limit.
- Stdin ordering and shutdown suite: 56 passed after the lifecycle
change.
- Daytona and watchdog coverage passed in the earlier focused runs.
Across the focused suites, 602 distinct tests pass.
- `pnpm -r typecheck` and `pnpm build`: passed. Server and adapter
typecheck/build also passed after their respective follow-up changes.
- `pnpm test:run`: attempted and stopped after known local failures.
Four chat/email cases used an external ancestor skill path, three
skill-cache cases failed on macOS, and one wakeup case timed out. The
wakeup case passes alone (1 passed, 27 skipped). This run spanned the
workspace-cutoff follow-up and also failed its new holder case; a fresh
final-head workspace suite passes all 19 tests. The interrupted run is
not a full local-suite pass or final-head verification.
- A filesystem queue-drain test failed once during the lifecycle rerun
and passed on the complete two-suite rerun. It uses the filesystem
client, outside the changed command-runner path.
- Complete CI on `ff2212c235`: 53 successful checks and two expected
skips, including the full test suite and canary packaging dry run. No
failed or pending checks.
- Greptile reviewed `ff2212c235` at 5/5. All review findings are
addressed, no threads remain unresolved, and the branch has no merge
conflicts with `master`.
- `git diff --check` and a scan of added text for secrets and private
identifiers passed.

## Risks

- A bridge control operation that needs more than 30 seconds now fails,
even if the caller selected a longer run lifetime. Agent commands retain
their own limits.
- Timing out a provider promise does not cancel the remote operation or
prove it stopped. Existing execution settlement still owns termination
verification. No uncertain tool action is replayed.
- Quiet healthy runs display warnings sooner. Existing snooze, continue,
and false-positive dismissal controls still apply. Silence alone does
not cancel a run, create review work, or change assignments.
- No schema or API shape change.

## Model Used

OpenAI GPT-6 through Codex, with tool use and code execution. The exact
serving model ID and context window are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run change-specific tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-02 15:29:41 -07:00
DottaandPaperclip 2ec82c5774 fix(runner): preserve task context when tool connections change (#14963)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents use tools through company-scoped connections and provider
sessions.
> - Resolving a tool connection currently forces a fresh session even
when the provider can load new tools into the existing conversation.
> - A fresh provider conversation can receive too little history to
continue the task.
> - This pull request adds explicit tool-refresh capabilities and uses
them in both runner paths.
> - Fresh attempts receive bounded task history with source IDs and
retrieval instructions.
> - The benefit is that agents can continue the same task after a
connection changes.

## Linked Issues or Issue Description

**What happened?**

A resolved tool connection forced a fresh provider conversation. The new
conversation could lose the original goal and prior answers. Claude also
rejected resume when only the MCP server set changed.

**Expected behavior**

Resume the provider conversation when its harness can refresh tools.
When a fresh session is required, supply enough bounded history to
continue the task. Preserve company, agent, task, workspace, model,
instruction, and skill checks.

**Steps to reproduce**

1. Start a conversation and agree on a task and its constraints.
2. Request and connect a tool needed for the task.
3. Continue the conversation after the connection resolves.
4. Check that the agent remembers the task and can use the new tool.

**Paperclip version or commit**

The bug was reproduced on master at `c46e41e81`. This branch is rebased
on current master.

**Deployment mode**

Self-hosted server. Both legacy adapters and the native runner are
affected.

Related public work: Refs #13282 for task-backed conversations. Refs
#13057 for the broader session-compaction proposal. Refs #14659 for
another report about local CLI session continuity. This change fixes
tool-connection continuation. Provider authentication repairs keep their
existing recovery behavior.

## What Changed

- Expose tool-refresh support in native harness descriptors and legacy
adapter metadata.
- Request tool refresh after connection resolution. Keep provider
authentication repair as a fresh-session wake.
- Reload current tools and credentials while retaining supported Claude,
Codex, Grok, and other provider conversations.
- Allow MCP-only changes during qualified native recovery. Keep all
other compatibility checks.
- Refresh managed-provider and ACPX tool bindings when attaching a new
run.
- Add a fresh-session handoff for both runner paths. Bound database
reads, excerpts, and the final packet to 24,000 bytes.
- Include the original request, recent messages, decisions, plans, prior
answers, and source IDs. Mark omitted content. Apply reset boundaries,
wake cutoffs, quarantine, and secret redaction.
- Add regression tests and document the capabilities and handoff
behavior.

## Verification

- `pnpm -r typecheck` and `pnpm build` passed. Rust formatting passed.
- Final review fixes passed 314 server tests, 333 adapter utility tests,
139 native-session runtime tests, and 12 managed-provider Rust tests.
They verify historical quarantine, raised budgets across attachment, no
history reads on successful resume, and handoff delivery on fresh retry.
- Broader branch verification also passed 1,401 adapter utility tests,
1,047 runner TypeScript tests, 43 Grok adapter tests, and 311 Rust core
tests.
- Live Claude CLI and Grok ACP probes preserved the provider session ID,
recalled a prior task constraint, and called a newly added read-only MCP
tool.
- GitHub CI passed on `b21486d18084a7aa4cafbe8e012f7cad6585d9cc`: 55
successful checks and 4 skipped checks. This includes all test shards,
all eight browser shards, runner checks, and the Grok clean public npm
install canary. [CI
run](https://github.com/paperclipai/paperclip/actions/runs/37057514976).
- A full local test attempt encountered a separate Git snapshot timeout.
All affected local suites passed after the final edits, and the full CI
test gates passed.
- Review the capability matrix in `packages/paperclip-runner/README.md`.
Repeat the four reproduction steps with a supported provider and with an
unsupported harness.

## Risks

- Provider tool refresh can fail. Existing recovery falls back to a
fresh conversation where policy permits it.
- A new transport can replace an old process while preserving the
provider conversation. Tests cover current credentials and unchanged
identity.
- Long history can omit older context. Explicit markers and source IDs
let the agent retrieve needed context within task scope.
- Unknown and unqualified harnesses use the fresh-session path. No
database migration is required.

## Model Used

OpenAI Codex, GPT-6, with reasoning, tool use, code execution, and live
provider testing. The exact model ID and context-window size are not
exposed in this session. Claude and Grok also ran as test subjects.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-02 15:11:44 -05:00
DottaandPaperclip 9786f6df56 fix(runner): preserve credential content in document saves (#14937)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The Runner sends authorized tool calls to the control plane.
> - Agents use these calls to save plans and instruction files.
> - The Runner used diagnostic secret detection to reject execution
arguments.
> - Ordinary credential-related prose could reject a document save
before persistence.
> - This pull request forwards the original arguments and leaves
credential policy to the provider harness.
> - The benefit is reliable saves with useful diagnostic records.

## Linked Issues or Issue Description

Related foundation: Refs #12415 and #14430. No duplicate save-policy fix
was found.

**What happened?**

A `write_document` call failed before the server saved its plan. The
Runner reported `semantic tool input contains credential material;
refusing to execute altered arguments`. The detector also masked
ordinary phrases such as `secret manager` and `credential handling` in
diagnostics. Both TypeScript dispatchers had equivalent execution gates.
One dispatcher also rewrote structured approval and question payloads
before execution.

**Expected behavior**

Paperclip forwards authorized arguments unchanged. The provider harness
decides credential-content policy. Log and audit redaction does not
reject or rewrite save input.

**Steps to reproduce**

1. Send an authorized `write_document` call with a plan that discusses
credential handling.
2. Include an intentional credential value in the body to exercise
harness-owned policy.
3. The old Runner rejects the call. With this change, the document
service stores the exact body.
4. Diagnostic records still mask explicit credential values. Qualified
credential fields, short bearer values, opaque diagnostic pairs, and
valid encoded JSON token headers have regression coverage.

**Paperclip version or commit**

Reproduced at `c46e41e81c03cd3c8b64cf993615b604d7fe8c62`. The branch is
based on current `master`.

**Deployment mode**

Server deployment with the native Paperclip Runner. Local regression
tests use the real document service and an embedded test database.

## What Changed

- Remove credential-content vetoes from Rust admission and both
TypeScript semantic dispatchers.
- Preserve original structured approval and question arguments during
execution.
- Keep transport bounds, schema checks, authorization, idempotency, and
audit masking.
- Require explicit credential syntax or recognized formats for
diagnostic masking. Preserve ordinary prose, metadata, and dotted
identifiers.
- Test exact document persistence, replay, nested argument identities,
and masked audit copies.
- Remove obsolete retry guidance and document harness-owned credential
policy.

## Verification

- `cargo test --manifest-path
packages/paperclip-runner/runner/Cargo.toml --locked -p
paperclip-runner-core --lib --test acpx_event_payload --test
acpx_provider_state --test acpx_provider_turns`: 355 tests passed.
- Focused server and adapter tests: 189 tests passed after rebase. These
include the real document save and the complete tool-gateway suite.
- Semantic dispatcher and conformance tests: 34 tests passed.
- Diagnostic redaction and MCP tests: 46 tests passed, including all six
review examples.
- `pnpm -r typecheck` and `pnpm build` passed on the repaired branch.
- The broad local root suite was interrupted after database fixture
setup failures. The focused database suites passed. CI runs the complete
configured test lanes.

## Risks

- Authorized tool arguments can intentionally contain credentials. The
harness must enforce its content policy.
- Diagnostic detection is narrower. Explicit assignments, credential
fields, and recognized credential formats remain masked.
- The change does not add a database migration or change company
authorization.

## Model Used

- OpenAI GPT-6 through Codex. The session exposes the GPT-6 model
family; its exact runtime model identifier and context window size are
not exposed. Used reasoning, tool use, and code execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-02 14:19:41 -05:00
DottaandPaperclip 7a52dcdc74 fix: repair MCP validation and cancelled execution recovery (#14951)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The tool gateway gives agents access to connected services. Recovery
controls what happens when a run stops.
> - Generated tool names can exceed the provider limit after the MCP
client adds its prefix.
> - The same invalid definition can fail each automatic retry. A
cancelled run can also hold saved messages without showing its cause.
> - This pull request bounds tool names, stops configuration retries,
and retains cancellation evidence.
> - It shows the stopped run and admits saved input only after the
existing safety checks pass.
> - The benefit is a clear recovery path that preserves operator Stop
and prevents duplicate message delivery.

## Linked Issues or Issue Description

**What happened?**

A long connected MCP tool name makes the provider reject the entire
request. Automatic recovery repeats the invalid request. Separately,
unexpected legacy cancellations can leave saved input behind a recovery
hold. The notice does not identify the stopped run or its cause.

**Expected behavior**

Complete MCP names fit the provider limit. Tool-definition errors
require configuration repair. Cancelled runs retain their source and
reason. The recovery notice shows the cause and saved-message count.
Verified unexpected cancellations can start a fresh turn through the
existing admission checks.

**Steps to reproduce**

1. Assign an App gallery connection with a long application key and tool
name to a Claude agent.
2. Start a run. The provider rejects a name over 128 characters,
including its MCP prefix.
3. For cancellation recovery, stop a legacy provider turn without an
operator Stop request and send a user message while the recovery hold is
active.
4. Inspect the recovery notice and the deferred message queue.

**Paperclip version or commit**

Rebased onto master at `cf8ad63c806685bfd7c48e3ed4a919d61a7c55f1`.

**Deployment mode**

Hosted or self-hosted server with legacy Claude or Codex execution.

Related public work:

- Refs #14017. That PR caps name segments. This PR preserves existing
short names and uses stable hash aliases for long complete names. It
also covers classification and recovery.
- Refs #4510. That PR adds a cancellation-source column. This PR records
bounded evidence in the existing run result, without a migration.
- Refs #12552 and #4506. Those PRs suppress recovery after operator
cancellation. This PR preserves operator intent and uses the existing
continuation gates.

## What Changed

- Bound gateway names with the full provider prefix in the 128-character
budget. Retain the original upstream tool name for dispatch and
permissions.
- Classify invalid tool definitions as configuration failures before
diagnostic redaction. Stop automatic retries and continuation attempts
for that error code.
- Persist cancellation source, expectedness, initiator, reason, and
time. Preserve recorded Stop intent when adapter results arrive. Report
unexpected started cancellations with closed diagnostic labels.
- Show the run cause, saved-message count, and Inspect run link. Offer
Continue for eligible unexpected cancellations. Require verified
provider stop, empty tool inventory, ownership, and the existing pause,
budget, approval, and dependency gates. Use the existing queue for
single delivery.
- Add regression coverage and update the execution, MCP gateway, and
run-log documentation.

## Verification

- `pnpm -r typecheck` and `pnpm build` passed.
- `pnpm check:token-gates` passed.
- Ran `pnpm test:run` and completed its workspace and serialized groups.
Initial resource and timing failures passed on isolated reruns. All 149
serialized route suites passed.
- Reran the changed server, adapter, and UI suites after the rebase.
Coverage includes long-name upstream dispatch, configuration retry
suppression, cancellation evidence retention, privacy labels, oversized
run projection, and concurrent saved-message delivery.
- `pnpm test:e2e tests/e2e/legacy-failure-continuation.spec.ts` passed
all six browser scenarios. The recovery notice shows the run cause and
inspection link, and each recovery entry point reaches one new response.
- Added database-backed checks for active, removed, paused, unavailable,
and disabled chat connections. The final continuation and
recovery-notice suites passed 167 tests. Externally bound chats hide
board Continue and show a usable next action.
- All 55 GitHub checks passed on
`42afbf1371dcaeb72646e3d8f65c19ff7cddf8de`. Two unrelated Storybook jobs
were skipped by their normal conditions. Greptile reviewed that commit
at 5/5 with no findings and no open review threads.

## Risks

- Long tool names change to aliases. Existing short names stay
compatible. The original connection and upstream name remain the
dispatch authority.
- Invalid tool definitions no longer get automatic retries. An operator
must repair the configuration before a new attempt.
- Continuation changes apply only to positively identified unexpected
legacy cancellations with complete empty tool inventory. Operator Stop,
unknown historical cancellations, outstanding tools, and unverified
provider termination keep their holds.
- No database migration. The added projection fields are optional.
Cancellation reason and initiator IDs remain local run evidence; Sentry
receives only closed source and initiator-type labels and expectedness.

## Model Used

- OpenAI GPT-6 through Codex, with reasoning, repository editing, shell
execution, and GitHub tool use. The runtime does not expose the exact
model variant or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-02 13:47:59 -05:00
Nicky LeachandPaperclip 9d0f7e2ddd fix(adapter-utils): make the directory merge lock crash test deterministic (#14881)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - A workspace restore merges a directory, and a cross-process lock
serializes that merge
> - The lock must recover after the process that holds it crashes
> - One test proves that recovery: it kills the holder process and then
acquires the lock
> - That test failed intermittently for two independent reasons, and
this pull request removes both
> - First, it spawned the holder through the tsx command-line entry
point, which re-spawns the evaluated code in a further child process, so
the kill signal reached only the wrapper and the real holder kept the
lock
> - Second, it replaced the global clock to force a timeout, which left
the acquisition with zero real retries, so a single transient busy
result failed the test
> - The benefit is a deterministic crash-recovery test and a reliable
continuous-integration signal

## Linked Issues or Issue Description

**What happened?**

The test `recovers a killed holder even when its recorded PID has been
reused` in `packages/adapter-utils/src/directory-merge-lock.test.ts`
failed intermittently in continuous integration. The failure reported
`ERR_WORKSPACE_RESTORE_LOCK_TIMEOUT` with `waitMs: 3` and
`knownLocalHolder: false`. A rerun of the same job on the same commit
passed.

**Expected behavior**

The test must pass every run. It must acquire the lock after the holder
process dies.

**Steps to reproduce**

1. Check out `master`.
2. Run `npx vitest run
packages/adapter-utils/src/directory-merge-lock.test.ts`.
3. Repeat the run. The named test fails intermittently.

**Paperclip version or commit**

`32e9f3ba0ec000578936731990d23bb0e77493fa`

**Deployment mode**

Built from source. The failure appears in the general test job of
continuous integration.

**Agent adapter(s) involved**

Not adapter-specific (core bug).

**Relevant logs or output**

```
ERR_WORKSPACE_RESTORE_LOCK_TIMEOUT
workspaceRestoreLock: { ownerState: 'alive', knownLocalHolder: false, waitMs: 3,
                        ownerSameProcess: true, ownerAgeMs: 60030, ownerPredatesProcess: true }
```

## What Changed

The test file had two independent defects. This pull request removes
both.

**1. The kill signal did not reach the real lock holder.**

The test spawned its holder through the tsx command-line entry point.
That entry point re-spawns the evaluated code in a further child
process. `SIGKILL` therefore killed only the wrapper, and the process
that had opened the lock database survived as an orphan that still held
the lock. The test now loads tsx as an `--import` hook, so the spawned
process is the real holder and the kill releases the lock at once. This
also stops the test from leaking an orphan process.

**2. The forced clock left the acquisition with zero retries.**

A test helper replaced `Date.now` to force a timeout. The implementation
reads `Date.now()` one time, to compute its deadline, so that single
read consumed the forced value and every later read returned a time
already past the deadline. The retry loop therefore got one attempt and
no retries. That is correct for a test that asserts a timeout, but the
crash-recovery test asserts a *successful* acquisition, so any transient
busy result on the first attempt failed it.

The fix removes the clock replacement from the whole file and gives each
test a real, short, explicit wait budget:

- `withDirectoryMergeLock` takes a new optional wait-budget parameter.
It threads through to the lock acquisition function. The production
default is the existing 30-second budget, and no production call site
changed.
- The five tests that assert a timeout pass a real 200-millisecond
budget. Each one still times out for the real reason, because the lock
is genuinely held or the legacy lock directory genuinely exists. Each
one now exercises at least four real retries of the 50-millisecond retry
interval.
- The crash-recovery test passes a real 5-second budget. A failure now
reports the structured `ERR_WORKSPACE_RESTORE_LOCK_TIMEOUT` diagnostic
well inside the test timeout, instead of a bare test timeout.

**No test timeout increased.** Every `it(..., N)` timeout in the file
equals its value on `master`.

## Verification

- Measured the first cause rather than assumed it: the spawned wrapper
process reported one process id, and the process that opened the lock
database reported a different process id and named the wrapper as its
parent. The real holder kept the lock for about 50 to 60 milliseconds
after the kill.
- Reproduced the failure deterministically before the change, with no
artificial processor load: 15 of 15 runs failed. Confirmed the fix: 15
of 15 runs passed.
- Ran the lock test file 15 times in series: 12 of 12 tests passed every
time.
- Confirmed the clock replacement is gone: a search for a `Date.now`
override in the file returns nothing.
- Confirmed the production default is unchanged at 30 seconds, and that
the diff touches no production call site.
- Proved the diagnostic still surfaces: with a temporary edit that held
the lock with a genuine live holder, the test failed with
`ERR_WORKSPACE_RESTORE_LOCK_TIMEOUT` and the full `workspaceRestoreLock`
block at about 5 seconds, inside the 15-second test timeout. The
temporary edit was reverted.
- `workspace-restore-merge.test.ts` passed 56 of 56. The adapter test
files that cover every production caller passed 116 of 116 and 52 of 52.
`agent-directory-working-copies.test.ts` passed 70 of 70.
- The `adapter-utils` and `server` type-checks passed with no error.
- Confirmed that no spawned process survives the test run.

## Risks

Low risk. The production change is one optional parameter with the
existing default, so every production caller keeps the real 30-second
budget and no production call site changed. The remaining change is
limited to one test file. The `--import` form of the tsx hook is already
used elsewhere in this repository, in the container image command and in
an end-to-end test configuration. Test coverage does not drop: the owner
record is diagnostic only, the SQLite reserved lock remains the
authority that the tests exercise, and the timeout-asserting tests now
exercise the real retry loop instead of a replaced clock. The file costs
about 0.5 to 0.9 seconds more wall clock than `master`, which is the
cost of the short real waits that replace the instant forced timeout.

## Model Used

Claude Sonnet 5 (`claude-sonnet-5`), used with extended thinking and
tool use for the diagnosis, the measurement, and the change.

## Checklist

Check every box that the state of the pull request satisfies. The local
test runs and the type checks are complete. Reconcile the
continuous-integration and review boxes after the checks reach their
terminal state.

## Test plan

- [x] Continuous integration is green on every check, including the
general test job.
- [x] The general test job passes the file
`packages/adapter-utils/src/directory-merge-lock.test.ts`.
- [x] Greptile returns 5 of 5 with no open item.
- [x] `mergeable: MERGEABLE` is terminal.

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-02 08:17:03 -07:00
Devin FoleyandPaperclip f2e0f19630 Defer agent directory cleanup until stop proof is available (#14866)
## Thinking Path

> - Paperclip manages agents and their persistent files.
> - Each run owns a temporary agent directory and a save receipt.
> - Cleanup needs independent proof that the owning process stopped.
> - A cleanup call without that proof currently waits for the directory
lock anyway.
> - A second lock failure can prevent environment release after the run
already reported a failed save.
> - This change skips cleanup that has no authority and retries
unavailable remote copies after exact destruction proof.
> - The save failure stays visible. Existing lock owners remain
protected.

## Linked Issues or Issue Description

Related work: Refs #14787 (lock diagnostics), #14695 (warm instruction
ownership), #9667 (stale lock proposal), and #9872 (control-plane
ownership proposal). I checked open PRs and issues. This change leaves
the shared filesystem lock protocol in place and does not duplicate the
warm-retention work in #14695.

**What happened?**

Heartbeat cleanup records an explicit unavailable instruction-save
warning, then calls directory release before releasing the environment
lease. Release can wait for a lock even though the copy has no
process-stop proof and cannot be removed. That secondary timeout
prevents the following lease-release step. If destruction proof arrives
later, the unavailable copy is excluded from both recovery queries.

**Expected behavior**

Skip a release that cannot remove anything. Preserve the failed-save
receipt and candidate fields. Once exact remote destruction is recorded,
recover remote cleanup without running a provider command. Unavailable
local copies retain their potentially uncollected edits even if local
stop proof arrives later. A blocked cleanup must not prevent cleanup for
other agents.

**Steps to reproduce**

1. Prepare an agent directory, report its save unavailable, and leave
process-stop proof absent.
2. Hold the shared directory lock and call release. Before this change,
release waits and fails although removal is not authorized.
3. Record destruction of the copy's exact remote lease. Before this
change, neither recovery sweep selects the unavailable copy.

**Paperclip version or commit**

Reproduced against `efc2e6810e9bc0dc8cb412b0e7647c0db9821caa`.

**Deployment mode**

Local and remote execution with persistent agent directories. Tests use
an isolated embedded PostgreSQL database and fixture transports.

## What Changed

- Re-read receipts and skip release before lock acquisition when stop
proof is absent, the copy is superseded, or cleanup is complete. Keep
the same checks inside the lock.
- Recover unavailable remote copies only after exact destruction proof.
Preserve their unavailable state, errors, candidate hash, and candidate
bytes. Keep unavailable local copies and their uncollected edits
unchanged.
- Store destruction-only cleanup authority with the stop proof. Later
cleanup honors it after a lost database response or restart, including
when a transport remains cached.
- Defer failed or unproven cleanup with bounded batches and a retry
delay. Keep failed cleanup visible in logs and its receipt.
- Serialize preparation of an existing run with cleanup. Fresh run
preparation keeps its existing admission path.
- Cover held locks, receipt scope, delayed proof, batch fairness, lost
update responses, cached transports, and concurrent same-run preparation
with database regressions.

## Verification

- Focused directory, legacy instruction-copy, shared lock, and bounded
diagnostic suites: 169 tests passed across four files.
- `pnpm -r typecheck`: passed on the final source.
- `pnpm build`: passed on the final source.
- Completed all selected local `pnpm test:run` groups: 733 general
server suites, 149 serialized suites, and 14 workspace projects. There
are 13 known macOS `EACCES` failures in the unchanged runtime skill
cache tests. Their exact signatures match earlier clean-base results,
and the cache source and test blobs match both that base and this PR
base (existing fix: #14290). One CLI import test timed out under
concurrent load; its full file passed separately (17 tests). Broad
coverage began before the review corrections; the final source has the
focused 169-test run, typecheck, and build. This is a local verification
limit, not a passing full local suite.
- `git diff --check` and local Gitleaks plus private-identifier/PII diff
scans passed.
- Independent review of the final source found no remaining actionable
issue. Its 17 targeted tests cover crash recovery, cached transports,
same-run preparation, real local edit preservation, proof scope, and
batch fairness. The main focused run also covers contained scheduling
failures.
- Final commit `35a24085f7`: Greptile 5/5 with no recommendations and
zero unresolved review threads.
- Final commit `35a24085f7`: all 53 checks passed, including Canary Dry
Run and the security scan; two visual checks were intentionally skipped.
The workspace shard passed on retry after GitHub reported that its first
runner lost communication. An earlier Canary runner shut down after the
release dry run passed. Neither interruption recorded an application
assertion failure; the exact final-head checks are now green.

## Risks

- This repairs cleanup ordering and recovery eligibility. It does not
repair an ambiguous legacy lock owner or restore unsaved files. Actual
collection still fails visibly when its lock cannot be acquired.
- An unavailable remote copy is recovered only after exact destruction
proof. A stopped but retained environment stays protected; recovery does
not execute a command that could restart it.
- Unavailable local copies with later stop proof still retain
potentially uncollected edits. A general local recollection or
reclamation policy remains outside this change.
- Existing-run preparation now waits for the same lock as cleanup. The
fresh-run path is unchanged.
- The cleanup mode is stored in the existing private receipt JSON. No
schema migration or public API change is required.
- No deployment, task replay, or runtime lock deletion was performed.

## Model Used

OpenAI GPT-6 (Codex), with reasoning, repository tools, and test
execution. The runtime does not expose a more specific model suffix or
context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub references)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-01 14:28:47 -07:00
Michael NguyenandClaude Opus 5.5 b721d24cac fix(adapter-utils): retry GitHub broker transport failures before falling back (#14856)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Agents run `git` and `gh` through a managed launcher. The launcher
gets a GitHub credential from the Paperclip control plane
> - The launcher sends one request to the credential broker for each
command
> - If that request fails at the transport level, for example after a
10-second timeout, the launcher continues without managed credentials
> - So a slow or restarting control plane removes the managed GitHub
identity from that command. Some agents then use other GitHub identities
that do not have the necessary permissions
> - This pull request retries a failed broker request two more times,
with a short backoff, before the launcher gives up
> - The benefit is that a short control-plane delay does not remove the
managed identity from an agent's GitHub operation

## Linked Issues or Issue Description

Refs #14175. That pull request changes the same broker request loop for
a different failure: sandbox network denials. The pull request that
merges second must rebase.

**What happened**
A Codex agent ran `git` and `gh` through the managed launcher while the
control plane was under heavy memory pressure. Each command printed
`Paperclip: GitHub broker_transport_unavailable; continuing without
managed credentials.` The agent then tried to open the pull request
through a different GitHub integration. GitHub rejected the request with
`403 Resource not accessible by integration`.

**Expected behavior**
A short broker delay or a short transport failure must not remove the
managed GitHub identity from the command. The launcher must try the
broker again before it continues without credentials.

**Steps to reproduce**
1. Set `PAPERCLIP_GITHUB_BROKER_URL` to a closed port.
2. Start a broker on that port after about 300 ms.
3. Run `gh` through the launcher.
4. Before this change, the launcher prints
`broker_transport_unavailable` and runs `gh` without the managed token.

**Version or commit**
`4ac374103` on master. Commit `3166e93a7` has the same code.

**Deployment mode**
Local trusted instance that runs as a launchd service, with
`codex_local` agents.

## What Changed

- `packages/adapter-utils/src/github-launcher.ts`: the broker request
loop now catches transport errors and retries up to two more times,
after 0.5 s and then after 1 s. The loop reads the response body inside
the retry, so a failed or slow body read is also retried. Busy (409)
responses keep their own budget of 30 attempts, separate from transport
retries. After the third transport failure, the launcher prints
`broker_transport_unavailable` as before.
- `packages/adapter-utils/src/github-launcher.test.ts`: two new tests
make the broker fail the first request and answer the second. In one,
the connection drops before the response. In the other, the connection
drops in the middle of the body. Each test checks that `gh` gets the
managed token, that the broker receives exactly two requests, and that
no `broker_transport_unavailable` message appears.
- The existing `broker-offline` test now has a 15-second timeout,
because each command now retries twice before it falls back.

## Verification

- `npx vitest run packages/adapter-utils/src/github-launcher.test.ts`: 9
of 9 tests pass.
- The body-read test fails on the first commit of this pull request and
passes with the second commit.
- `pnpm --filter @paperclipai/adapter-utils typecheck`: passes.
- The existing `broker-offline` test confirms that the launcher still
falls back after the retries, and that local Git still works.

## Risks

- When the broker is unreachable, each `git` or `gh` command now waits
about 1.5 s more before it continues without credentials. When the
broker times out, the worst case is about 31.5 s instead of 10 s.
- The change only adds retries. It does not change which credentials the
launcher accepts or which environment variables it copies.
- #14175 changes the same loop. The pull request that merges second
needs a small rebase.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- Anthropic Claude Opus 5.5 (`claude-opus-5-5`), used through Claude
Code with tool use: shell commands, file edits and test runs. The model
wrote the change, the test and this description. The repository owner
approved the change before it was made.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass — *targeted tests and the
package typecheck; see Verification*
- [x] I have added or updated tests where applicable
- [ ] I have updated relevant documentation to reflect my changes — *no
documentation describes the broker retry*
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green — *CI has not run yet*
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups —
*Greptile has not reviewed yet*
- [x] I will address all Greptile and reviewer comments before
requesting merge

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 12:57:41 -07:00
Devin FoleyandPaperclip dd9983b894 fix(adapter-utils): release restore locks when a process crashes (#14869)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agent runs restore workspace files and collect instruction-file
changes.
> - Writers to the same target directory must wait for each other.
> - The current lock records a PID, which a new process can reuse after
a crash.
> - A reused PID can keep an orphaned lock alive and make each later run
fail.
> - This pull request makes a SQLite file lock decide ownership. The OS
releases it when the process exits.
> - Later runs can proceed after a crash, and concurrent live writers
remain protected.

## Linked Issues or Issue Description

Refs #10914. This addresses crash recovery. It does not cancel a stalled
operation in a process that is still alive.

Related work: #9667, #14787, and #12187. The earlier attempt in #9667
assumes one live server per lock root. This implementation uses an
OS-backed lock to support concurrent writers without treating a
different process token or an old timestamp as proof of a dead owner. It
retains the private lock root and bounded timeout diagnostics from the
merged changes.

After a process dies while holding a restore lock, a replacement process
can reuse its PID. The existing `process.kill(pid, 0)` check then
reports a live owner forever. Later runs can complete their model turn
but fail during file collection or restore.

## What Changed

- Hold a SQLite `BEGIN IMMEDIATE` transaction for each directory write.
Use the existing built-in `node:sqlite` dependency.
- Keep each lock database on a stable inode. Keep PID and time metadata
only for diagnostics.
- Retain the 30-second asynchronous wait and existing timeout error code
and diagnostic fields.
- Fail closed when an old directory lock exists. Document a
stopped-writer upgrade and rollback procedure.
- Add real child-process tests for crashes, PID reuse, live owners, and
connection cleanup. Cover callback failures, independent targets, stable
inodes, invalid lock files, and ambiguous legacy records.

## Verification

- Before the fix, the crash/PID-reuse test and the live-owner test both
failed. Both pass with this change.
- `pnpm exec vitest run
packages/adapter-utils/src/directory-merge-lock.test.ts
packages/adapter-utils/src/workspace-restore-merge.test.ts`: 56 tests
passed.
- Restore and agent-file working-copy integration tests: 118 tests
passed before the additional connection-cleanup test.
- `pnpm -r typecheck`: passed.
- `pnpm build`: passed.
- Full GitHub CI: all checks passed, including Linux workspace tests,
server test shards, build, typecheck, and browser tests.
- Greptile: 5/5, with no review threads or unresolved comments.
- `pnpm test:run`: started locally, then stopped with SIGINT (exit 130)
after full CI passed. The local serial run did not complete and is not
counted as a full local pass. The completed CI shards provide the
full-suite result.

## Risks

- **Upgrade and rollback require a drain.** Stop every old writer that
shares an instance root before switching protocols. Old and new versions
must not write concurrently.
- Existing legacy `.lock/` directories remain blocking. After all
writers stop, preserve run evidence and move those directories to an
operator scratch directory. The new code does not infer that they are
abandoned from PID or age.
- Never delete or replace a `.lock.sqlite` file while writers can run.
These small files remain after release.
- The shared filesystem must support reliable SQLite locking. Broken
network-filesystem locking is unsupported.
- This change prevents new orphaned ownership. It does not recover file
changes lost during earlier failed collections, or interrupt a live
operation that stalls.
- No application database migration or new native dependency is
required. See `doc/workspace-restore-locks.md` for the procedure.

## Model Used

OpenAI Codex based on GPT-6, with code execution and repository tools.
The exact model variant and context window are not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused regression and
integration suites; see the full-suite note above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-01 12:14:48 -07:00
Devin FoleyandPaperclip 4b9a6000f7 Add bounded evidence for directory lock timeouts (#14787)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agent files use directory locks during collection and cleanup.
> - A lock timeout can fail finalization after the model turn completes.
> - The timeout currently identifies no owner state or waiting
operation.
> - This pull request adds bounded evidence to the existing run failure
report.
> - Operators can distinguish a known local holder from a possible old
lock without changing lock safety.

## Linked Issues or Issue Description

**What happened?**

A directory lock timeout does not distinguish active local work from an
owner record left by an earlier process. The stored execution stage can
also precede the cleanup operation that failed.

**Expected behavior**

The failure report should identify the waiting operation and expose
bounded ownership clues. It must preserve the timeout and keep unknown
ownership protected.

**Steps to reproduce**

Hold a directory merge lock while a second caller reaches its
acquisition deadline. The regression tests exercise a live holder and an
older owner record with a live PID.

Related: #9667 proposes stale-lock recovery under a single-server
assumption. This change only adds evidence and does not adopt that
assumption. #14575 and #14665 add other run failure diagnostics.

## What Changed

- Record lock owner state, capped age and wait duration, same-process
and process-age comparisons, and whether this module holds the lock.
- Label agent-directory release, collection, checkpoint, and warm
handoff timeouts with a fixed operation code.
- Validate each field before the existing event-local Sentry report
accepts it. Exclude owner records, PIDs, paths, and absolute timestamps.
- Limit the extra diagnostic owner read to 100 ms with best-effort
abort; malformed JSON is `invalid` and unreadable owner records remain
`unknown`.
- Document the diagnostic limits and verify that contenders never
reclaim protected locks.

## Verification

- Focused lock, diagnostic, real Sentry SDK, and database-backed
agent-directory tests: 126 passed, including stalled-read and
malformed/missing/unreadable-owner regression coverage.
- Final revision `0691613dcc`: all 54 reported checks successful, with
two intentionally skipped Storybook checks. Greptile: 5/5, zero
unresolved review threads; no merge conflicts.
- `pnpm -r typecheck`: passed.
- `pnpm build`: passed.
- `pnpm test:run`: complete suite coverage ran with the existing
repository shard flags: four general-server shards, four serialized
shards, two general-workspaces-a shards, and general-workspaces-b. The
full run is not green because of the base failures below.
- The broad run found 13 failures in the unchanged macOS skill-cache
tests. All 13 reproduce on the clean base revision. Open PR #14290
covers that existing failure.
- Two unchanged CLI archive tests hit their five-second limits during
the broad run; all 17 tests in that file pass on recheck. A CLI auth
socket error also cleared on recheck (19 tests), and its full serialized
shard passed on rerun.

## Risks

This is a diagnostic change, not a stale-lock fix. Owner observations
can race with release. Wall-clock shifts can affect the age comparison.
A local-holder flag covers only this module instance. None of these
fields authorizes reclamation or proves a file save. Lock acquisition,
release, retries, task status, and recovery guards retain their current
behavior. No schema change or deployment action is required.

## Model Used

OpenAI Codex, based on GPT-6, with code execution and repository tools.
The exact model build and context window were not exposed to this agent.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass (focused checks pass;
existing base failures are documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 18:35:12 -07:00
Devin FoleyandPaperclip d6d67b00d3 Prevent background workspace scans from refreshing the Git index (#14666)
## Thinking Path

Paperclip runs background workspace scans alongside real Git writers.
`git status` can refresh the index as an optional side effect, taking a
lock that makes another operation fail. Disable optional locking in the
shared scan subprocess so background observation does not compete with
workspace updates.

## Linked Issues or Issue Description

**What existing behavior does this improve?**
Workspace Git scans used by changed-file browsing, cleanliness guards,
and sandbox snapshots.

**Current behavior**
The scan process inherits Git's default optional-lock behavior. Even a
clean `status` can rewrite stale stat-cache entries in the index and
contend with a concurrent writer.

**Proposed behavior**
Always set `GIT_OPTIONAL_LOCKS=0` for the shared scan subprocess while
preserving the selected environment and Git's required write locks.

**Reason and benefit**
Background reads stop creating avoidable index contention. Git documents
this behavior and recommends disabling optional locks for background
status: [background
refresh](https://git-scm.com/docs/git-status#_background_refresh).

**Breaking changes**
None to scan results or required write locking. Later scans may repeat
stat checks that would otherwise have been cached in the index.

Related scan implementation: #11572, #14253. This avoids one known
contention source; it does not identify every historical lock owner or
repair abandoned locks.

## What Changed

- Disable optional locking at the shared scan subprocess boundary,
including explicit caller environments.
- Test a clean status against a deliberately stale index and prove an
ordinary status would rewrite it.
- Test tracked/untracked results with an existing index lock,
preservation of that lock and working files, and continued rejection of
a mandatory-lock write.
- Document the scan behavior and performance tradeoff.

## Verification

- Focused stream, workspace-sync, and scheduler suites: 63 tests passed.
- Full `pnpm -r typecheck` and `pnpm build` passed locally. Final head
`effe6420c77bd18d36af6b093db3a564c04b8b38` passed all 53 CI checks,
including complete test coverage, typecheck, build, and browser/runner
gates; two unrelated checks intentionally skipped.
- Full local test attempts initially had missing embedded-Postgres
library symlinks; the dependency setup was repaired. Duplicate local
full-suite runs were stopped after full CI completed. This PR does not
claim a completed full local suite.
- Review regression: real Git honors the supplied `GIT_CONFIG_*` setting
and the input environment remains unchanged; all four direct subprocess
cases passed.
- Reviewed the diff for secrets, customer data, and internal references.

## Risks

Low risk. Disabling optional index refresh can repeat filesystem stat
work on later scans. Required locks remain enforced; no lock is removed,
no failed reset is retried, and workspace mutation guards are unchanged.
No schema changes.

## Model Used

OpenAI GPT-6 via Codex, with repository inspection, code execution, and
tests. Exact model build identifier is not exposed by this session.

## Checklist

- [x] Thinking path and model are specified
- [x] Checked ROADMAP.md; this is a maintenance correction, not planned
feature work
- [x] Searched for duplicate and related PRs
- [x] Described the issue using the enhancement template
- [x] No internal issue references, customer data, or private instance
links
- [x] Descriptive branch name
- [x] Focused regression tests pass
- [x] Added tests and updated documentation
- [x] Risks documented
- [x] Required validation and CI gates are green (full suite validated
in CI; local scope documented above)
- [x] Greptile is 5/5 with no unresolved findings
- [x] I will address review comments before requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 14:28:07 -07:00
Devin FoleyandPaperclip 0ea6b10967 Record ACP activity and workspace restore failure evidence (#14665)
## Thinking Path

Paperclip records terminal run failures for operators. A timeout's own
log and cleanup output update the run's last-output timestamp, so that
timestamp can make a long-silent provider look active. Snapshot runtime
activity before finalization and include the saved workspace restore
classification to make the next failure actionable without copying tool
payloads.

## Linked Issues or Issue Description

**What existing behavior does this improve?**
Terminal run diagnostics in the existing opt-in Sentry integration.

**Current behavior**
Reports cannot distinguish runtime events from finalization logging and
omit the already-persisted workspace restore code. A later successful
run also does not establish that earlier workspace files were restored.

**Proposed behavior**
Record runtime-event age/count and pending-tool inventory at
finalization, before status reads and cleanup. Forward only finite
counts, a completeness boolean, and known restore codes through the
existing reporter.

**Reason and benefit**
Operators can distinguish a silent turn with unfinished tools from
recent runtime activity and see restore failures without retrieving
private run output. Neither signal certifies productive work or
successful recovery.

**Breaking changes**
None. Error grouping, execution deadlines, cancellation, recovery
policy, and the Sentry opt-in remain unchanged.

Related diagnostic work: #14573, #14575, #14639.

## What Changed

- Snapshot ACP activity before success/failure finalization, including
thrown relay failures.
- Forward bounded numeric/boolean evidence and shared workspace restore
codes; exclude commands, tool IDs, paths, and arbitrary result data.
- Document limitations and test silence, empty streams, timeout, cleanup
delay, incomplete tool inventory, and privacy.

## Verification

- `pnpm -r typecheck` passed after the final implementation.
- Changed suites: 252 tests passed; all 29 database reporter tests
subsequently passed after restoring the embedded-Postgres package
library symlinks. The migration test also passed (30 database cases
total).
- `pnpm build` passed during implementation. Final head
`fe78dba6f592b1abccac7cdbf341bd2e0b0d30cb` passed all 53 CI checks,
including complete test coverage, typecheck, build, and browser/runner
gates; two unrelated checks intentionally skipped.
- Full local test attempts initially hit missing embedded-Postgres
library symlinks; the dependency setup was repaired and database tests
passed. Duplicate local full-suite runs were stopped after full CI
completed. This PR does not claim a completed full local suite.
- Review regression: completed, failed, and cancelled tools are excluded
from the pending count; focused activity/timeout tests and adapter-utils
typecheck passed.
- Reviewed the diff for secrets, customer data, and internal references.

## Risks

Low risk, diagnostic-only. The existing tool inventory is incomplete for
some runtime events, so the report carries its completeness flag. Event
age is measured at finalization and does not prove useful work or
identify the underlying provider failure. No schema changes or new
capture gate.

## Model Used

OpenAI GPT-6 via Codex, with repository inspection, code execution, and
tests. Exact model build identifier is not exposed by this session.

## Checklist

- [x] Thinking path and model are specified
- [x] Checked ROADMAP.md; this is a maintenance correction, not planned
feature work
- [x] Searched for duplicate and related PRs
- [x] Described the issue using the enhancement template
- [x] No internal issue references, customer data, or private instance
links
- [x] Descriptive branch name
- [x] Focused regression tests pass
- [x] Added tests and updated documentation
- [x] Risks documented
- [x] Required validation and CI gates are green (full suite validated
in CI; local scope documented above)
- [x] Greptile is 5/5 with no unresolved findings
- [x] I will address review comments before requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 14:27:53 -07:00
DottaandPaperclip b54b2dc35c fix: preserve warm Codex turns with incremental managed file checkpoints (#14735)
## Thinking Path

> - Paperclip manages AI agents and keeps their instructions and files
durable.
> - Native Codex runners can keep a process alive between compatible
turns.
> - Managed file collection stopped that process after each turn, which
defeated warm reuse.
> - Agent folders can contain large images and other files, so full
copies on every turn are expensive.
> - This change keeps one managed directory for the live session and
saves only file changes after each turn.
> - Ownership, authorization, instruction changes, and process
retirement still control when reuse is safe.

## Linked Issues or Issue Description

Related: #13710 introduced native warm session reuse. This fixes managed
file collection that still forced those sessions to stop. No duplicate
open PR or issue was found.

**What happened?**

With managed instructions and warm native Codex enabled, consecutive
turns reused a Daytona sandbox but started a new runner process each
time. The managed directory collector required process termination
before saving files.

**Expected behavior**

Compatible turns keep the same process and managed `AGENT_HOME`. Each
completed turn saves added, changed, and deleted files before the next
turn starts. Unchanged large files do not transfer again.

**Steps to reproduce**

1. Use a native Codex agent with managed instructions and a reusable
Daytona environment.
2. Enable warm session reuse and run three turns on the same task.
3. Write a large binary on the first turn, edit a small note on each
turn, and delete a file on the second turn.
4. Compare process identity across turns and read the canonical files
through the public agent-files API.

**Paperclip version or commit**

Reproduced on `d30b03bd8c17604cdab1533eeeeb087aba30e8b1`.

**Deployment mode**

Local server with remote Daytona execution; cloud native runner uses the
same path.

## What Changed

- Retain the managed directory only for the verified owner of a live
native Codex session.
- Checkpoint each completed turn before releasing the session for reuse.
Retry unstable captures, then stop and collect when a warm checkpoint
cannot be validated.
- Compare metadata and cached hashes, stream only changed file payloads,
record deletions, and validate path, content, quota, and authorization
before saving.
- Rotate sessions when canonical files, loaded instructions,
credentials, or launch policy change. Fence stale collection and cleanup
callbacks from later owners.
- Keep cleanup and recovery aware of the current session owner. Recheck
canonical files under the writer lock at handoff, attach the successor
collector before fallible bookkeeping, and emit one final save receipt
on checkpoint fallback. Preserve storage warnings across unchanged
checkpoints.
- Add regression coverage and a three-turn Daytona test with independent
public API file checks, an unchanged 8 MiB binary, deletion checks, and
strict process identity checks.
- Document checkpoint consistency, lifecycle behavior, and local run-log
counters.
- Replace a timing assumption in the Daytona teardown test with explicit
transfer-arrival gates after CI exposed an unset release callback.

## Verification

- Full local `pnpm -r typecheck` and `pnpm build` passed. Server checks
were repeated after the final storage-warning fix.
- Runner E2E typecheck and 749 runner E2E unit tests passed.
- Focused file checkpoint, directory ownership, instruction collection,
native session, and merge tests passed. After review fixes, the
managed-directory and native-session suites passed 550 tests, including
intervening canonical edits, same-run fresh restore, failed handoff
collection, and one-call fallback collection. Server typecheck passed
again. The Daytona plugin suite passed 218 tests. The quota-warning
regression failed before the fix and passed afterward.
- Three real Daytona campaigns passed before the final handoff review
fixes. The latest kept PID 547 across all three turns. The first
checkpoint copied 8,388,635 bytes; the next two copied 36 and 54 bytes.
Public API reads verified the binary, note contents, and deletion after
every turn. Test cleanup deleted the sandbox.
- The final head was also deployed to an isolated cloud staging instance
and passed three UI-triggered native Codex turns with managed
instructions. All three retained the same process ID/start time, native
session, provider session, runner instance, and Daytona sandbox.
Checkpoints copied 8,388,643 bytes on turn 1, then only 52 and 78 bytes
on turns 2 and 3; those warm captures also hashed only 52 and 78 bytes.
Independent canonical API reads verified every byte of the unchanged 8
MiB binary and the exact note contents after every turn; the deleted
file returned 404 after turns 2 and 3. After restoring the original
lifecycle and agent-auth configuration, removing the temporary secret,
pausing the test agent, and deleting both test sandboxes, independent
canonical API reads still verified the entire binary, the final 78-byte
three-line note, and the deletion. The native runner flag remained
enabled and the final serving revision remained the PR head.
- Two earlier staging attempts are preserved as failures and are
excluded from the acceptance result: a saved ChatGPT login failed with a
provider routing 401, and its subsequent stopped-sandbox retry failed
before provider startup with a closed-lease admission error. The
successful campaign used a fresh sandbox and a temporary encrypted
API-key binding. The stopped-lease retry remains unexplained; this
campaign does not establish recovery of that failed sandbox.
- All [Paperclip CI
gates](https://github.com/paperclipai/paperclip/actions/runs/36750397355)
pass on `26ef2ef56a389259246809805c0b34a4747eb86b`, including full test
partitions, build, typecheck, runner verification, E2E shards, and the
Canary clean public-npm install. Greptile reviewed that exact head at
5/5 with no unresolved review threads or outstanding findings.
- Full local repository coverage used the existing CI partitions, but
the 40,000-file Git streaming stress test timed out and its local retry
was interrupted by macOS thermal emergency sleep; this is not a green
full local suite claim. The exact stress test passed on the final head
in [CI server shard
2/12](https://github.com/paperclipai/paperclip/actions/runs/36750397355/job/110008294290),
in 111.9 seconds.
- Repeat the live test with configured credentials and a Linux runner
artifact: `pnpm test:e2e:runner -- --id
daytona-warm-continuity.runner-codex.daytona.warm-three-turn`.

## Risks

- This is a file-level checkpoint, not an atomic snapshot of the whole
folder. Background writes after a capture are saved by the next
checkpoint or final stopped collection.
- Metadata scans still visit all paths. Modified files transfer in full;
unchanged files do not rehash or transfer.
- Incorrect ownership or reuse could collect the wrong directory. Run
ownership fences, current authorization, stable capture validation, and
stopped collection fallbacks are covered by tests.
- Warm reuse remains opt-in. No database migration or fleet default
changes.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, code editing, tool use, and
test execution. The exact serving model ID and context-window size are
not exposed by this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 13:33:37 -05:00
DottaandPaperclip 3b4b270650 fix(adapters): preserve ACP terminal failure diagnostics (#14573)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The shared ACP adapter engine records agent failures for operators.
> - ACP providers can report a failure category, title, and detailed
cause.
> - Our patch kept only the category in the saved error, so an operator
could not diagnose a failure when tracing was off.
> - This pull request preserves redacted provider diagnostics in the run
error, transcript, and structured run result.
> - Operators can now inspect the provider message and any supplied
request ID or stack trace after the run ends.

## Linked Issues or Issue Description

Refs #13889 (the diagnostic gap; this PR does not update the bundled
Claude version).
Refs #14484 (related model-refusal classification; this PR retains
diagnostics for all terminal failure categories).

**What happened?**
An ACP turn failed with only `ACP agent reported a terminal service
failure.` The provider's title and details were available in memory but
absent from the saved error and transcript.

**Expected behavior**
The run retains useful provider diagnostics even when raw tracing is
disabled. Credentials remain redacted. A size limit must report
truncation instead of silently removing the cause.

**Steps to reproduce**
1. Run an ACP agent that returns an error-severity typed session
failure.
2. Include an HTTP error, request ID, and stack text in its title and
details.
3. Inspect the failed run with tracing disabled. Before this change,
only the category survives.

## What Changed

- Both pinned ACPX patches pass complete error text to the in-memory
callback, so redaction happens before truncation.
- The shared engine retains the sanitized category, title, and details
in `resultJson.terminalSessionFailure` and includes the text in the run
error and error transcript.
- Diagnostics redact configured environment values even under arbitrary
names, unknown launch-environment values, connection URL passwords, run
credentials, and common credential syntax. Known boolean settings remain
readable, while credential values are redacted even when embedded in
other text. Diagnostics remove control characters and invalid Unicode.
- Title and detail limits keep escaped transcript JSON below the
server's chunk limit. Truncated fields include an omission count. The
safe run-result projection preserves a byte-bounded diagnostic preview
when the result exceeds its byte budget, with an explicit pointer to the
full adapter-bounded run error and transcript.
- The existing UI and CLI display the error. Diagnostics do not become
assistant output. Issue continuation summaries and session-compaction
prompts receive only the generic category, preventing provider text from
becoming handoff instructions. Existing quota classification, warnings,
timeout precedence, and control-channel failure precedence remain in
place.
- Regression tests cover real ACP child processes with both pinned
versions in one-shot and persistent modes, credential redaction, request
IDs after the old 4 KiB cutoff, transcript parsing, storage bounds, and
database retrieval of oversized multibyte diagnostics.

## Verification

- Full CI on `20ad4f5f1f66c46d2c260e6ad0339cbea607b4cf`: **54 passed, 2
intentionally skipped, no pending or failing checks**. Includes
typechecking, build, all Vitest shards, Runner checks, browser E2E, and
the canary packaging/public-install dry run.
- Greptile: **5/5** on this commit. Superagent security scan passes. All
review threads are resolved.
- Local verification passed: shared ACP engine suite (395 tests); real
Claude ACP child-process and diagnostic regressions across both pinned
runtimes and both execution modes; run retrieval and model-handoff
regressions (59 tests); ACPX patch packaging (16 tests); full typecheck
and build. Affected package typechecks and focused tests were rerun
after review fixes.
- The broad local `pnpm test:run` was stopped after review edits made
its cached imports stale. Fresh targeted runs pass, including both
affected server suites. Cold-build import failures were also rerun after
dependency builds: chat integration (1,063 tests) and tool access (351
tests) pass. The final commit's complete CI matrix is green.

## Risks

- Provider diagnostic text is untrusted. This change retains more of it
in company-scoped run records. Redaction and size bounds apply before
persistence.
- Diagnostics are limited to fields the provider supplies. Old runs
cannot recover discarded error text.
- No schema migration, recovery-policy change, or new Telemetry or
OpenTelemetry export.

## Model Used

- OpenAI GPT-6 through Codex, with reasoning, repository inspection,
code editing, and test execution. The exact serving model ID and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-29 10:18:30 -05:00
DottaandPaperclip 24beb00575 feat(runner): add rich ACP transport and durable interaction foundation (#14430)
Add shared rich ACP transport, durable questions and permissions, verified provider packaging, and bounded activity and plan presentation. Keep Cursor, Copilot, and Pi pending their separate provider qualification.

Persist interaction settlement before publication, fence failed writes until fresh recovery, and preserve owned-process cleanup. Incorporate reviewed mainline integration with extended harness coverage.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-29 08:56:21 -05:00
DottaandFry 3ca196b0a6 feat(agents): persist agent files across tasks without revision history (#14420)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - An agent needs personal files across tasks and sessions.
> - AGENTS.md is one file in that directory. Supporting files need the
same persistence.
> - The Instructions Editor and agent runs must share one current
directory.
> - Concurrent runs should apply only the files they change. The last
sync of the same file wins.
> - This pull request uses existing file transport and removes temporary
copies after sync.
> - Old instruction-only sessions keep their restore contract. New saves
do not create revision history.

## Linked Issues or Issue Description

Refs #14325. This replaces its revision-oriented design with persistent
agent files. Keep #14325 unmerged.

Transport prerequisite #14416 merged first at
`d172197117a14b80a1eb2d2835a0e7cce2679656`. This PR now targets master
and remains below 100 changed files.

Related work: #4513 and #8798 cover instruction tooling. This change
handles run synchronization, cross-task personal files, browser editing,
and old-session restoration.

## What Changed

- Keep one current directory per company and agent. Point AGENT_HOME at
a temporary working copy for each active run. Keep task files and
provider HOME separate.
- Restore text, binary files, and nested folders through workspace
transport. Exclude remote agent files from task Git snapshots with a
self-ignoring file inside the reserved runtime directory; never write
through repository-controlled Git metadata.
- Collect after the provider and child processes have stopped. Keep
resumable conversation state.
- Apply changed and deleted files under the agent lock. The last sync
wins for the same file. Unrelated concurrent changes survive.
- Remove temporary copies after successful sync, rejected sync, and
staging failure. Register ownership before copying so restart recovery
can remove interrupted preparation. Retry transient synchronization up
to three times. Preserve the original remote lease reference until
deletion succeeds; restart cleanup never acquires a replacement sandbox.
Do not create captured directories or a conflict-review queue for new
runs.
- Keep browser editing, stale-draft protection, and streaming binary
downloads. Keep the instruction entry and text editor limited to 1 MiB.
- Keep historical agent-folder sync failures on their affected runs
instead of repeating them above current saved instructions. Preserve
legacy candidate review and current browser-save errors. Avoid duplicate
quota warnings while retaining separate sync failures when they describe
a different problem.
- Require target-scoped caller grants for peer instruction access, while
preserving self edits, responsible-user checks, and protected-change
consent.
- Treat full storage as a nonblocking run warning, never an agent pause
or run-admission failure. Restore already-over-quota saved folders so
ordinary agent cleanup can recover; warn on each run until cleanup. The
run detail view shows the warning.
- Allow 256 MiB per file, 2 GiB per directory, and 100,000 entries. Hash
large files as streams. Check editor-save quotas with metadata instead
of hashing unrelated files.
- Preserve old native inputs, instruction-only copies, paths, digests,
and pending legacy candidates. Adopt old revision heads once. New writes
do not append history rows.
- Add idempotent migration 0287 and verify upgrades from the preview
tables and receipts.
- Add nine interactive stories under **Agents / Persistent files**,
including automatic incoming edits, stale browser drafts, and
storage-limit diagnostics.

## Verification

- Merge candidate: `4f5390107ec6ffd80a76d1d2e85530e66f21d079`, after
merging current master and the landed transport prerequisite.
Integration required no manual conflict resolution; the feature remains
99 changed files. Full workspace typecheck, production build, token
gates, and 715 focused tests passed on this merge candidate. Fresh
Greptile review is 5/5 with no unresolved findings. All 55 checks
passed, with four conditional skips, including the build, typecheck,
browser E2E, and canary dry run. A single retry recovered four jobs
interrupted by runner shutdowns; no source changes were required.
- Historical-warning UI fix: all 6,834 UI tests across 640 files passed,
including regression coverage for three old failures, legacy preserved
edits, and warnings scoped to the affected run. Full workspace
typecheck, production build, Storybook build, and token gates passed.
Browser-verified Storybook playtests passed for Historical Failures
After Successful Save, Storage Limit, and Full Storage Run Warning.
- Review follow-ups at `4e20c9fb2`: all 18 focused tests passed,
including external Git directories, linked worktrees, symlinks,
hardlinks, and distinct I/O failures alongside storage warnings. Server
and UI typechecks, token gates, and the production build passed.
- Storage warning regressions at `0724f3012`: all 33 directory tests and
all five heartbeat-list tests passed, with no skips in their successful
runs. They cover repeated runs while full, an already-over-quota saved
folder, cleanup, warnings retained after unrelated save failures, and
bounded warnings in large result JSON. Server typecheck passed after the
final warning fixes.
- Full workspace typecheck, production build, and token gates passed
during this follow-up. Product E2E harness: 631 tests passed across 52
files; harness typecheck passed. Earlier native session/context and
directory/legacy collection suites passed 537 tests; Runner
unit/transport suites passed 329 tests.
- **Real E2E at `0724f3012` (before this follow-up):** legacy local
Codex and native Daytona Codex each passed six tasks, one server
restart, seven independent assertions, and cleanup verification. Both
prove browser-to-agent edits, agent-to-browser edits, nested/binary
restoration, per-file last-sync-wins, a successful run after an
oversized save rejection, and cleanup clearing the warning.
- Native local Codex also passed the six-task quota flow before the
final warning-retention fixes. That pass began at `918d1ed02` while the
bounded-result warning fix was being edited, so it is not claimed as
exact-final-head evidence. Its final-head rerun failed during embedded
PostgreSQL bootstrap before any provider run: the macOS host had 87,365
of 87,381 SysV semaphores occupied. No unrelated services or kernel
limits were changed.
- The final-source report intentionally records **2/3 cells passed**,
preserving the blocked native-local attempt:
`tests/runner-e2e/results/agent-files-quota-final-20260928-report/`.
Earlier failed attempts and provenance notes remain under
`tests/runner-e2e/results/agent-files-quota-final-20260928-input/` and
the original campaign directories.
- Daytona used immutable image
`ghcr.io/paperclipai/paperclip-daytona-runner@sha256:5643f0d801417cae3581833a1a3bc6715b325e028602738d2652c44cac5dc6bf`
and its exact Linux runner binary. Controller source is `0724f3012`;
image source is recorded separately.
- Legacy-session compatibility and all three ACP Stop/resume browser
regressions passed on the prior validated feature head
`169fab46d5af21caa2269b4c1b29b69c933a6951`. They assert the same
provider session is retained and interrupted writes are not replayed.
Migration upgrade tests also passed earlier.
- Nine interactive stories are under **Agents / Persistent files**,
including **Full Storage Run Warning**. Its playtest and visual browser
inspection passed; the warning states that runs continue and the editor
remains available.
- Prior-head checks on `4e20c9fb2`: 55 passed, two conditional jobs
skipped, no failures or pending checks. All eight browser E2E shards and
their aggregate passed. Fresh Greptile review is 5/5 with no findings;
all review threads are resolved, the security scan passed, and GitHub
reports no merge conflicts.
- The broad local follow-up test run was interrupted after host
semaphore exhaustion affected isolated PostgreSQL instances. It also
encountered the existing macOS long-path fixture failure and two timeout
failures. This is not a claim that the full local suite passed. Logs are
retained; focused storage/warning tests passed.

## Risks

- A later sync can overwrite an earlier edit to the same file, including
a saved browser edit. There is no text merge or retained version. This
is the intended last-sync-wins policy.
- A save that exceeds a storage limit is rejected and its temporary copy
is discarded. The run itself continues normally, and later runs restore
the last saved files with a warning until cleanup. Transient sync
failures get bounded retries. An I/O failure partway through a sync can
leave some files updated; a failed receipt does not claim whole-folder
success.
- Larger folders increase copy time, network traffic, and temporary disk
usage. Active runs still need working copies. Terminal runs do not
accumulate archives. Operators must provision disk for agents and
configured concurrency; these limits are not company-wide quotas.
- A restored old native session remains instruction-only until a fresh
session starts. Its original conflict fence and existing pending
candidates remain compatible.
- Provider processes close at the collection boundary. Conversation
resume remains available, but warm process reuse is lost.
- Backups must include the instance filesystem and database. External
bundles keep their existing behavior until explicitly moved to managed
storage.

## Model Used

OpenAI Codex, GPT-6 family. The session does not expose a more specific
model ID or context-window size. Reasoning, code execution, and browser
tools assisted this change. Real provider E2E uses `gpt-5.6-sol`.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Fry (Paperclip) <noreply@paperclip.ing>
2026-09-29 08:25:56 -05:00
DottaandPaperclip d172197117 feat(storage): add plain directory sync with conflict preflight (#14416)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agent runs use workspace transport to restore and collect files.
> - Some directories belong to the agent across tasks.
> - Those directories need plain file transport without task Git state.
> - A concurrent file edit must be detected before a merge changes any
file.
> - This pull request adds optional plain-directory sync and conflict
preflight.
> - Existing task workspace sync keeps its defaults.

## Linked Issues or Issue Description

Refs #14325. This is the transport prerequisite for a replacement of its
instruction revision design with current agent files.

## What Changed

- Add an opt-in plain-directory transport mode to command and sandbox
runtimes.
- Add strict merge preflight for file edits, deletions, and directory
changes.
- Accept identical replay after an interrupted merge. Preserve competing
changes.
- Set the compiled OpenCode test executable to 0755, independent of the
CI host’s file-creation mask. Preserve the original startup error if
cleanup also fails.

## Verification

- At `ced53ae532ce6966cad5a83575d83db61af98126`, all 212 targeted
transport tests passed across workspace restore, remote managed runtime,
SSH fixture, and execution-target sandbox suites. Adapter-utils
typecheck passed.
- Real isolated SSH retry fixture previously passed with
`PAPERCLIP_ENABLE_DARWIN_SSH_ENV_LAB=1`; stale deleted files remain
absent while gitignored binary bytes survive.
- Dependent PR #14420 passed real native local, legacy local, and native
Daytona persistence E2E at `169fab46d5af21caa2269b4c1b29b69c933a6951`,
which includes all production transport changes through `ced53ae53`; the
subsequent two commits only fix the OpenCode test fixture. Nine tasks,
three server restarts, and all cleanup checks passed.
- A hosted OpenCode fixture failed twice at `ced53ae53`. Reproduced the
failure locally and in Linux with `umask 0002`: the compiler created a
group-writable executable, correctly rejected by the qualified launch
boundary. Explicit 0755 permissions fix the test without weakening the
production guard. The focused test and non-root Linux reproduction now
pass under that same mask.
- Before rebase, head `69e97de0475d34aac5d532e559a405eaf015fd2b`
includes the deterministic fixture permission fix and preserves original
bootstrap diagnostics. All production transport code is unchanged since
the 212-test validation. Fresh Greptile review is 5/5 on this exact head
with no unresolved findings. All 54 current-head checks passed, with two
conditional skips. The full CI run completed successfully, including the
previously failing OpenCode runner shard.

- Merge validation on rebased head
`c509d79dd190c5cb00dc65edfde209097ff21465`: all five commits are
patch-identical to the reviewed branch. All 54 checks passed with two
conditional skips, and fresh Greptile review is 5/5 with no findings.
One retry cleared an npm archive 404 and a Cursor fixture timeout.

## Risks

- New behavior is opt-in. Existing task snapshot behavior retains its
defaults.
- Generic strict merge preflight remains opt-in. The dependent
agent-folder feature rebases changed paths before applying them to
provide per-file last-sync-wins; it does not create a conflict-review
queue.
- This change adds no database migration, dependency, or UI.

## Model Used

OpenAI Codex, GPT-6 family. The session does not expose a more specific
model ID or context-window size. Reasoning, code execution, and tool use
assisted this change. Live provider validation used `gpt-5.6-sol`.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-29 07:59:11 -05:00
Devin FoleyandPaperclip 53aad90b9e fix: retry sandbox ACP input delivery after gateway failures (#14485)
## Thinking Path

> - Paperclip coordinates agent work through execution adapters.
> - Sandbox ACP sessions send ordered input through a remote file queue.
> - A temporary provider 502 currently closes the session during an
input upload.
> - A lost response can occur after the sandbox has consumed the
message, so a blind retry can duplicate input.
> - This pull request retries gateway failures with the same sequence
and drops consumed sequences at the receiver.
> - The session can continue through a brief provider failure without
repeating a tool call.

## Linked Issues or Issue Description

**What happened?**

A sandbox ACP run can fail with `ACP agent disconnected during request
(connection_close, exit=null, signal=null)` when a provider input upload
returns HTTP 502. The bridge destroys its local socket on the first
failure and can discard the diagnostic before the proxy reads it.

**Expected behavior**

A temporary gateway failure should get a bounded retry. A lost response
after successful delivery must not duplicate input or reorder later
messages. Permanent failures must still close the session.

**Steps to reproduce**

1. Run the real sandbox process bridge with an echo child and a local
test runner.
2. Inject a provider 502 before preparation, after a chunk upload, or
after final publication and consumption.
3. Send the next input message. Before this change, the connection
closes instead of delivering it.

Searched open and closed PRs for `ACP disconnect`, `bridge retry`, and
`502 sandbox`. Related work: #13287 covers shutdown after bridge loss;
#13793 covers large launch envelopes. This change covers ordered input
delivery within a running legacy ACP session.

## What Changed

- Retry input uploads up to three times for recognized Daytona and
Cloudflare HTTP 502, 503, and 504 diagnostics, with 250 ms and 500 ms
delays.
- Give each upload separate temporary paths and discard already-consumed
input sequences, including late publication from an earlier attempt.
Clean failed attempts in the background without removing a published
message or another attempt’s files. Cleanup cannot delay retries or
shutdown.
- Keep later input behind the retry. Stop queued input on permanent
failure and flush a fixed diagnostic before closing the socket. Neither
failure-diagnostic persistence nor shutdown-warning persistence can
block teardown.
- Add real-process regression tests for lost responses, late
publication, retry exhaustion, immediate permanent failure, and
diagnostic redaction.
- Give accepted run-log file appends up to three seconds to drain before
finalization computes the size, hash, and durable copy. Close the run
handle to later appends. This waits only for file writes, independently
of later DB progress or live-event persistence. If writes remain
stalled, return null size/hash metadata and skip the final durable copy
so the run can settle. Late writes cannot restart mirroring.
- Preserve legacy comment attribution when final log size is unknown by
reading existing entries within the unchanged 2 MB scan limit. Storage
errors or a three-second read deadline return the evidence already read
instead of failing the comment listing; pagination stops at the
deadline. The deadline requests cancellation of the underlying local
stream or S3 HEAD, GET, and response stream. A separate response timeout
returns partial evidence even when filesystem I/O delays cancellation;
late reads cannot append evidence or start another page. Each listing
retains its existing batches of eight reads, without a shared admission
cap that skips readable logs under contention.
- Document the retry and log-finalization boundaries in the development
guide.

## Verification

- Final commit `347daa564b`: [Linux
CI](https://github.com/paperclipai/paperclip/actions/runs/36506995168/attempts/2)
passed. Greptile Apex review 13 scored this commit 5/5 with no new
findings; all 12 review threads are resolved.
- The final CI run initially hit a Cursor test timeout and four Discord
credential-lock contention failures. All five cases passed in isolation.
The two failed shards and their aggregate gate passed on retry without a
code change. Those intermittent failures are not claimed fixed by this
PR.
- `pnpm --filter @paperclipai/adapter-utils typecheck` passed.
- `pnpm exec vitest run
packages/adapter-utils/src/execution-target-stdin-race.test.ts
packages/adapter-utils/src/execution-target-sandbox.test.ts
packages/adapter-utils/src/sandbox-callback-bridge.test.ts`: 262 tests
passed on the final implementation, including 21 new regressions. The
original three fault-injection cases failed before the fix.
- The regressions cover failed and indefinitely stalled cleanup,
Cloudflare gateway responses and retry exhaustion, permanent errors that
must not retry, and teardown while failure logging remains indefinitely
stalled. Seven Apex regression cases failed before the review fixes.
Adapter-utils typecheck and build passed again after the final review
change.
- `pnpm exec vitest run server/src/services/run-log-store.test.ts
server/src/services/run-log-store-cancellation.test.ts`: all 25 tests
passed, including four new regressions that failed before the
finalization fix. They cover delayed and failed appends, late-write
admission, agreement between the local bytes/summary/durable copy, and a
stalled append that exhausts the three-second budget. The timeout case
verifies unknown metadata, no final upload, and no mirror restart after
late completion. New cancellation tests use the real AWS SDK against a
local HTTP server. They verify that stalled HEAD, GET, and response-body
connections close on abort and that a subsequent read succeeds. Local
range and already-aborted read cases also pass.
- `pnpm exec vitest run server/src/__tests__/issues-service.test.ts -t
'readIssueCommentRunLogText|deriveIssueCommentRunLogAttribution'`: 14
targeted tests passed. The null-size reader case, both storage-error
cases, the stalled-read case, and the cancellation/concurrent-listing
cases failed before their fixes. The new regressions verify that
timed-out reads are cancelled, subsequent listings recover, and two
concurrent listings both retain their attribution markers. A read that
ignores cancellation still returns partial evidence at three seconds and
cannot resume pagination when it finishes; this regression failed before
the response-timeout fix.
- `pnpm --filter @paperclipai/server typecheck` and `pnpm --filter
@paperclipai/server build` passed after the response-timeout change.
- Full `pnpm -r typecheck` and `pnpm build` passed earlier in this PR;
the affected packages were rechecked after review fixes.
- Full local `pnpm test:run` failed in the general-server group: 511
files passed, 40 failed, and 158 were skipped. Failures include embedded
PostgreSQL initialization, read-only cache directory renames, a macOS
long-path fixture, and a workspace exposure assertion. The PostgreSQL,
cache-permission, and long-path failures also reproduce with both
changed implementation files restored to baseline commit `24c58e479a`.
The exposure suite passes in isolation both on baseline and the fixed
branch (28 passed, 3 skipped). CI runs the full suite on Linux. Later
local test groups were not reached.
- An earlier CI run hit the Telegram retry-timing failure fixed upstream
in #14501. The branch includes that master fix. The selected recovery
test passed against a fresh, migrated PostgreSQL 16 database. The
embedded PostgreSQL runner is unavailable on this Mac; the isolated
database was stopped and removed afterward.
- No live agent turn was replayed. The tests use local child processes
and injected provider failures.

## Risks

Retries are restricted to recognized Daytona SDK and Cloudflare bridge
gateway-error messages, which survive plugin RPC serialization. Other
errors fail immediately. Temporary upload paths are now unique for all
command-managed queue writes. Receiver sequence checks prevent duplicate
input; retries do not restart an agent turn. Cleanup and failure logging
are nonblocking and best effort; session teardown remains the final
cleanup boundary. Log finalization now drains accepted local file writes
for at most three seconds and ignores later appends on the closed run
handle. A timeout leaves final size/hash unknown and skips the final
durable upload; an existing partial mirror may remain available, but it
is not claimed as a verified final snapshot. It does not wait for later
DB progress or live-event persistence. Optional attribution keeps
partial evidence when a read fails or times out. Cancellation closes S3
requests and response streams. Local filesystem I/O may finish after the
caller deadline, but a late read cannot change the returned evidence or
continue pagination. Later listings can retry after storage recovers.
There is no schema, authentication, or permission change. Revert this
commit to restore the previous behavior.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository inspection, code
editing, and local test execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally; targeted tests pass and full-suite
limitations are documented above
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 19:15:33 -07:00
0be2afcca6 feat(ui): improve task composer controls and pending input (#14322)
## Thinking Path

> - Paperclip lets operators assign tasks to AI agents and review their
work.
> - The task composer controls the next message and its assigned agent.
> - Operators needed a way to choose that agent's model and effort
without leaving the composer.
> - The old mode selector, upload button, and input cards made the
mobile composer crowded and hid normal messaging during a pending
decision.
> - Harnesses publish different model and effort capabilities, so the
picker must follow the selected agent.
> - This pull request adds one responsive composer flow, keeps pending
cards visible above it, and protects Codex ACP authentication in the
local test path.
> - Operators can choose run settings, send a message, and answer a
pending card as separate actions.

## Linked Issues or Issue Description

**Subsystem affected**

Task composer UI, issue thread interactions, Codex ACP credential
handling, and Storybook.

**Problem or motivation**

The composer did not expose model or effort for the selected agent.
Mobile actions wrapped poorly. Pending questions and confirmations
replaced the composer. A local Codex ACP test could also reuse host
authentication after the managed key was removed.

**Proposed solution**

Put assignee search, model search, exact model IDs, effort, and fast
mode in one picker. Use a mobile dialog. Replace the direct-upload plus
action and separate mode selector with an Add menu and removable Plan or
Ask chips. Place pending interaction cards above the usable composer.
Keep these cards pending after an ordinary message unless their creator
asks for comment superseding. Replace managed ACP auth files atomically
and isolate the test key from host credentials.

**Roadmap alignment**

ROADMAP.md does not list an overlapping composer milestone. This change
improves the existing task and review flows.

## What Changed

- Added the combined assignee, model, and effort picker to both task
composers. Search matches agent name, role, and harness. The server uses
a curated Codex list by default and honors instance-declared models.
Manual IDs remain available.
- Added an effort slider for known model capabilities, a conditional
Codex fast control, and reset. The picker opens in a modal on mobile.
- Added the Add menu for files, supported goals, Plan mode, and Ask
mode. Plan and Ask are exclusive removable chips. Keyboard mode cycling
remains available.
- Adjusted mobile spacing, avatars, wrapping, and Send placement.
Removed the composer divider.
- Moved pending question, confirmation, review, and related cards above
the composer. Ordinary comments now leave question and confirmation
cards pending by default. The onboarding prompt retains explicit comment
superseding.
- Updated the Storybook composer group with responsive states and the
production picker. Added UI, service, route, and browser regression
coverage.
- Isolated Codex ACP API-key authentication, skipped subscription auth
merge and shared-home copy-back for remote API-key runs, and replaced
the managed auth file atomically.

## Verification

- `pnpm -r typecheck` — passed on the final local head.
- `pnpm check:token-gates` — passed on the final local head.
- `pnpm exec vitest run server/src/__tests__/adapter-models.test.ts
ui/src/components/task-chat/ComposerRunSettingsPicker.test.tsx` — 31
tests passed, including role and harness search, declared Codex models,
and filtering general OpenAI models.
- `pnpm exec vitest run
server/src/__tests__/issue-thread-interactions-service.test.ts` — 74
tests passed.
- `pnpm exec vitest run
packages/adapters/codex-local/src/server/acp.test.ts` — 42 tests passed,
including remote API-key copy-back isolation.
- `pnpm test:run` — attempted locally; the embedded PostgreSQL test
database could not initialize on macOS. The isolated
`heartbeat-run-event-sequencing` suite reproduced that environment
failure. GitHub CI runs the full test matrix for this head.
- `pnpm build` — passed on the final head. `pnpm build-storybook` passed
after the last UI change; only server code, tests, and docs changed
afterward.
- Live local test drive — Codex ACP ran a task with a managed API key.
The test agent was restored to its default ACP configuration afterward.
- Review the interactive stories under the top-level Composer group with
`pnpm storybook`. Check a narrow desktop width and mobile Plan, Ask,
picker, and pending-question states.

## Risks

- A pending card stays open when an ordinary comment changes the
discussion. Its creator can set `supersedeOnUserComment: true` when a
new comment should replace it.
- Model and effort overrides persist on the task until reset or changed.
An unlisted manual model ID may fail when the provider runs it.
- Some harness catalogs do not report effort support. The picker hides
effort for those models.
- No database migration is required.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

OpenAI GPT-6 via Codex. This runtime does not expose the exact model ID
or context window to the task. The model used code execution and browser
tools.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: OpenAI Codex <codex@openai.com>
2026-09-28 23:02:39 +00:00
DottaandPaperclip 992f720262 fix: make runner task context ownership explicit (#13753)
<!-- Write all pull request text in Simplified Technical English
(ASD-STE100). -->

## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Task descriptions, comments, continuation data, skills, and
execution rules enter several agent adapters.
> - The same source can be rendered by more than one automatic input
carrier.
> - Failed resumes can also rebuild input from stale or compact context.
> - This pull request gives each Paperclip-owned source one delivery
owner and preserves the required transport boundaries.
> - It adds deterministic adapter, interaction, runner, and browser
tests for these boundaries.
> - The benefit is more predictable context delivery with explicit
evidence for later live qualification.

## Linked Issues or Issue Description

Related: #13144 removes a duplicate environment payload and bounds wake
lists. Related: #11360 addresses Hermes resume behavior. This pull
request preserves compatible active-session formats while repairing
context ownership and stale question creation.

**What happened?**

Task descriptions and comments could enter more than one automatic
context block. Native transports could wrap a complete model input in a
second task envelope. Some legacy and gateway adapters could omit the
owned assignment on ordinary tasks or rebuild a failed resume with stale
compact context. A continuation could also request a question after
newer human comments had arrived.

**Expected behavior**

Each task or comment source has one automatic model-facing owner.
Distinct comment IDs and repeated wording remain distinct. Fresh
fallback attempts rebuild the required full context. A question request
is rejected when newer queued human direction makes it stale. Harness
access policy remains owned by execution configuration.

**Steps to reproduce**

1. Build a task with a description and current comments.
2. Capture the actual adapter or runner input.
3. Compare source ownership and task-envelope nesting.
4. Queue a human comment before a continuation requests a question.
5. Trigger a failed resume and inspect the fresh retry input.
6. Run the focused adapter, interaction, runner, and browser checks.

## What Changed

- Add shared prompt-section selection at the provider-attempt boundary.
- Deliver owned assignment context through native, legacy CLI, ACP,
gateway, cloud, Pi, Kimi, Grok, Gemini, OpenCode, Cursor, OpenClaw, and
Hermes paths.
- Rebuild full or compact context after resume recovery changes the
attempt. Add native and Claude ACP tests of actual recovery requests.
- Preserve custom templates, loaded instruction files, execution
policies, and older active-session formats.
- Record continuation source metadata and reject stale question creation
under the issue-row lock.
- Add explicit Product E2E context-integrity profiles, prerequisite
gates, credential-isolation checks, and report fixtures.
- Bypass service-worker forwarding for same-origin Vite development
modules. A real Chromium test fails with resource exhaustion before the
repair and passes after it. Production asset caching keeps its existing
policy.
- Add browser diagnostics and service-worker module-loading regressions.
- Add an explicit zero-retry eval option. The default retry behavior
remains unchanged. Each campaign records its effective policy.
- Remove the model-facing working-directory sentence from four prompt
builders. Existing workspace, sandbox, permission, and custom-template
configuration remains unchanged.
- Align the everyday workflow assertion with the current 47-entry
catalog.

Compared with current upstream master, the branch carries the
context-ownership implementation and its tests, the explicit
context-integrity catalog and evidence harness, and the focused browser
regression checks.

## Verification

**Merge assessment:** focused regression evidence supports merge. This
is not full completion of the original broad qualification matrix. The
maintainer has authorized merge after fresh verification of the master
integration.

- Current head: `bbd52f82114eabf09bc7b1a7e97d54a5b43bbc00`. This
integrates current master `2f585ef26a1814fa209715242d1ca791b63e4c4e`.
All 14 conflicts are resolved. Cancellation checks, workspace
finalization, native Grok support, and both sets of tests are retained.
- Current-head Greptile: **5/5**, with no blocking findings. The review
names this exact commit. All **59 reported checks are terminal: 55
successful, 4 skipped, zero pending or failing**. This includes the full
root general and serialized suites, separate runner checks, typecheck,
build, canary, browser E2E, Docker, and security checks. The successful
legacy security status is included in that total.
- After integration: workspace typecheck and full build passed. Separate
runner checks passed: **2,160 TypeScript tests (10 skipped), 582 Rust
tests, and 39 preparation checks**. Other passing checks include 621
Product E2E harness units, 376 focused shared/adapter tests, 160
real-database/API tests, 86 Hermes tests, 18 browser-support checks, and
Product E2E typechecking. The complete root suite passed in CI. The
duplicate local monolithic root run was stopped after that CI result; it
is not counted as a completed local pass.
- New native recovery coverage retains full assignment, completion
contract, and explicit skill selection after safe replacement, for old
and prepared input formats. Full native session test file: **136/136
passed**.
- New Claude ACP coverage captures actual fresh, resumed, and
missing-session fallback requests. It verifies one assignment copy,
comment order, identical text under distinct comment IDs, and full
fallback context. Full file: **33/33 passed**. Both affected TypeScript
checks passed.
- Existing deterministic tests cover source revisions, approval and
trust boundaries, completion validation, custom templates, compatible
sessions, standalone driver wrapping, and maintained adapter transport
requests.
- Provider-free browser support: **17/17 passed** after the master
merge. Service-worker unit tests: **33/33 passed**. The module-overload
regression failed before the repair and passed after it in real
Chromium.

### Fresh live comparisons

The new batch ran exactly four Product E2E attempts. **All four passed
on the first attempt; no retries.** Each has six terminal matchers plus
the existing browser lifecycle and invariant checks.

| Exact case ID | Control | Candidate |
|---|---|---|
| `core-compatibility.runner-codex.local.plan-revise-accept` | Passed |
Passed |
|
`local-session-integrity.runner-acpx-claude.local.structured-question-restart-resume`
| Passed | Passed |

The plan case checks a revised canonical plan and revision-bound
approval before completion. The question case restarts the server before
submitting the answer, then verifies the continuation completes.

Control source is `dfa4e1bda8d50a1a01746603251a9128dbe9d0d6`. Candidate
source is `79fcdb5dece501d28064ea9da306603881b46f0c`. They use identical
frozen definitions and provider versions: Codex `0.156.0` with
`gpt-5.6-sol`; ACPX `0.13.1` / Claude ACP `0.73.0` with
`claude-sonnet-5`. The September 24 head added master browser recovery
and test-only changes. The September 28 head also integrates newer
master changes, including cancellation, workspace finalization, and
native Grok. These are frozen-source live results, not exact-head live
runs.

The candidate received one description copy where the control initially
received three. The submitted initial plan envelopes were 7,969 versus
19,097 characters. Question envelopes were 7,592 versus 18,919. These
are structural measurements, not whole-provider token or dollar savings.

### Earlier evidence and failed attempts

- The preceding fresh batch has four effective passing pairs: OpenCode
comment continuation and assigned skill, native Codex comment
continuation, and native Claude comment continuation. It retains **11
attempts: eight passed and three failed**.
- Original failures remain recorded: missing local PostgreSQL library
links before task creation; host-sleep cleanup after task/page checks
passed; and a Claude **control** session-open rejection before a model
turn. Setup was repaired identically on both worktrees. The permitted
unchanged infrastructure retries passed. The underlying Claude provider
startup error was not retained and remains unknown.
- Older R2 retains **17 passes and one failure** across 18 attempts,
including eight both-pass native/legacy Codex/Claude pairs. Its OpenCode
blank-page failure led to the service-worker repair. R2 is historical
evidence: master changed the native fixed prompt and removed duplicate
wake environment data afterward.
- The September 24 CI run initially failed one unrelated preview
readiness test (`ECONNREFUSED` on its local fixture). Its test and
production code match master. Isolated local verification passed **28
tests, 3 skipped**. One unchanged CI retry passed the full shard: **831
passed, 1 skipped**, including all **31 preview-exposure tests**. The
aggregate CI gate passed afterward. The precise startup cause remains
unknown; a port race is a hypothesis, not a proved cause.

### Limits

The original wider profile/workflow matrix, repeated trials, and remote
Daytona qualification are incomplete. These results support a focused
merge recommendation, not statistical equivalence or universal harness
qualification. Some usage receipts are missing in both variants, so no
token or dollar savings are claimed. The $500 ceiling was preserved
using conservative allowances; failed attempts and unknown charges
remain in the ledger.

Reproduce the focused additions with `pnpm exec vitest run
packages/adapters/claude-local/src/server/acp.test.ts` and `pnpm
--filter @paperclipai/paperclip-runner exec vitest run
src/native-session-runtime.test.ts`. Full checks use `pnpm -r
typecheck`, `pnpm test:run`, `pnpm build`, and the separate runner
checks. Paid evals require the frozen definitions, profiles, and
credentials; do not use `--all` as a substitute for the selected cases.

## Risks

- Context placement changes can affect model behavior. Deterministic
checks cover the selected paths, but live qualification remains
incomplete.
- The stale-question guard can reject a request when queued human
comments arrived during the run. This is intended.
- New stored inputs and model envelopes retain compatibility readers for
older active sessions.
- Custom templates may intentionally repeat content.
- Removing a model-facing working-directory sentence does not change
filesystem, command, sandbox, or permission configuration.
- The worker bypass applies only to same-origin development module
paths. Cache-policy tests preserve private-response handling and
production asset caching. Mounted HTTP fixture changes remain test-only.
- This PR does not claim measured token savings or statistical
equivalence across every harness.

## Model Used

OpenAI Codex, exact model gpt-6-astra, with repository tools and code
execution. Bounded supporting work used gpt-5.6-luna and gpt-6-luna. The
serving context-window size is not exposed in this task.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have described the issue in-PR using the required issue fields
- [x] I have not referenced internal/instance-local Paperclip issues or
links
- [x] My branch name describes the change and contains no internal
ticket id
- [x] I have run the focused local checks and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect these changes
- [x] I have considered and documented risks above
- [x] All current-head Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
for the current head
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 14:49:14 -05:00
DottaandPaperclip cbc5132e6c fix(adapters): expose a verified provider stop before workspace restoration (#14311)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agent adapters own local or remote provider processes.
> - Run cleanup must know when the final provider process has stopped.
> - A remote timeout or a lost transport does not prove that the
provider stopped.
> - Retried provider invocations also make an earlier stop signal stale.
> - This pull request adds a verified final-invocation stop callback
before workspace restoration.
> - The dependent instruction revision change uses that boundary to
preserve private instruction edits safely.

## Linked Issues or Issue Description

**What happened?**

Adapter completion did not expose a reliable point between provider
shutdown and workspace restoration. Cleanup could lose provider-written
files, or treat a remote timeout as proof that a process stopped.

**Expected behavior**

Cleanup runs once after the final provider invocation has a verified
stop receipt and before workspace restoration. An incomplete remote
command keeps collection pending.

**Steps to reproduce**

1. Run a remote provider that returns a timeout without a numeric exit
code.
2. Let the adapter return or retry the provider.
3. Attempt to collect provider-written files during cleanup. The adapter
has no verified final-process boundary to use.

This is the prerequisite for the stacked canonical instruction revision
pull request. It has no database or UI dependency. Related #13291
verifies remote termination for later recovery; this change exposes the
earlier adapter-owned stop boundary before workspace restoration. The
scopes do not duplicate each other.

## What Changed

- Add a stop callback to adapter execution context and a
final-invocation fence.
- Confirm local child closure and complete remote exit receipts. Reject
remote timeouts, missing exits, transport failures, and SSH exit 255 as
stop proof.
- Invoke collection once before workspace restoration in eight CLI
adapters and at the confirmed ACP stop boundary.
- Preserve the stop observation when a later log flush fails.
- Run bridge and workspace cleanup in `finally` even when collection
rejects. ACP records a safe error without exposing a raw filesystem
path. Grok keeps collection errors separate from workspace restore
failures and preserves completed provider results when both cleanup
steps fail.
- Correct the existing Cursor test shell fixture so bounded remote file
reads run against real fixture files.

## Verification

- Three stop-boundary regressions failed before the callback
implementation and passed after it.
- Independent prerequisite branch: 377 tests passed across 24 adapter,
process-target, and ACP suites. Two added collector-rejection tests
failed before the cleanup fix and passed after it.
- A third regression reproduced Grok misclassifying a collection failure
as failed workspace restoration. Two additional cases covered completed
and failed provider turns when collection and restore both fail. The
Grok and restore-classifier suites passed 48 tests.
- All nine affected package typechecks and affected package builds
passed; Grok checks passed again after its classification fix.
- Integrated instruction branch: native local, legacy local, and native
Daytona each passed three browser tasks with exact persisted bytes,
fresh-task readback, history/restore, and explicit conflict resolution.
Unchanged warm Daytona passed three turns. Legacy Codex passed all five
checks again after the exception-safe cleanup fix.
- Alternate staging passed the same three-task native Daytona flow:
exact stopped-run save, independent downloaded readback, browser
history/restore, and explicit resolution of a real concurrent edit. The
deployed source was `14c3d810c9e05625121b3d27767aea9317b03125`, which
covers the initial adapter callback. Later Grok cleanup failures are
qualified by the adapter tests above.

- The final combined native Daytona flow passed again on deployed
`2bedd0f23bf4698b1f8b818f6796900647030427`: three fresh tasks proved
ordinary instruction edits, independent readback, History/Restore,
concurrent board conflict, and explicit candidate resolution. This
native staging flow does not claim to exercise the Grok adapter.

## Risks

- An unverified remote stop intentionally does not trigger collection. A
later controller with verified stop evidence must recover it or report
the copy unavailable.
- The callback is optional. Callers that do not register it retain their
existing behavior.
- The callback runs before workspace restoration and can delay cleanup
if its caller does not bound its own work. The dependent instruction
collector uses bounded reads and retries. A rejected callback still
permits bridge and workspace cleanup; it cannot claim an instruction
save.

## Model Used

OpenAI Codex, GPT-6, with tool use and code execution. The runtime does
not expose the exact deployment variant or context-window size. Multiple
Codex agents implemented and verified the change.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 08:45:53 -05:00
DottaandPaperclip 4b38db9622 fix(runtime): stream workspace Git snapshots through disk manifests (#14253)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Managed runs copy a selected workspace to an execution environment
and restore its changes.
> - Git snapshots select the files for that copy and for later recovery.
> - A fixed output limit stops large generated trees before the run can
start.
> - Increasing the limit still keeps the complete filename lists in
memory.
> - This pull request stores those lists and merge baselines in disk
manifests.
> - Large snapshots can now complete with bounded filename buffers and
explicit failure handling.

## Linked Issues or Issue Description

Refs #14194. This is the streaming follow-up to the merged 32 MiB limit
fix.

Related: #13619 and #11621 cover workspace scan admission and demand.
This change keeps the shared scheduler and changes the snapshot data
path.

## What Changed

- Stream changed, untracked, deleted, and ignored paths through the
shared scheduler and the standalone adapter path.
- Use SQLite manifests for file selection, duplicate removal,
ignored-path lookup, baseline capture, and merge lookup.
- Set a configurable 30-minute snapshot deadline. Keep the existing
interactive scan deadlines.
- Wait for each child process and pending sink write before removing
temporary storage after failure or cancellation.
- Use NUL archive lists and bounded deletion batches. Preserve unusual
names, explicit selection, nested repositories, and source root checks.
- Store manifest references in native recovery format v2. Check their
location and digest before recovery reads. Keep v1 descriptors readable.
- Remove temporary manifests at lifecycle completion. Use fixed-size
temporary copy names for long basenames.
- Admit each manifest with a SQLite page allowance based on current disk
capacity. Keep a configurable free-space reserve and fail explicitly
when either limit is reached.
- Preserve a host file that replaces a directory deleted by the sandbox,
and continue the rest of the restore.

## Verification

- Current head `5ab622ec43cd16d35429d79dedee6a5d8e3d2df2` has 54
successful checks/statuses and two skipped Storybook jobs. No checks
failed or remain pending.
- [CI
passed](https://github.com/paperclipai/paperclip/actions/runs/36318966368):
typecheck, build, all test shards, E2E, Rust checks, and the aggregate
verify job.
- [Greptile is
5/5](https://github.com/paperclipai/paperclip/pull/14253#issuecomment-5855670338)
on the current head. All four review threads are resolved. Security
checks passed.
- 229 focused tests passed across Git sync, runtime staging, merge,
manifest integrity, native recovery, and the scheduler (214
adapter/runtime tests and 15 scheduler tests).
- A real 40,000-file fixture produces 43,428,890 filename bytes. The
original standalone and scheduled scans fail. The new test passes all
four filename paths, complete staging, exclusion of late files, unusual
names, and deletion replay.
- Recovery tests reject changed bytes, symlinks, and paths outside the
controller state directory. Adapter-utils typecheck passed.
- A test executor returned buffered output and caused two retry
integration failures. The fixture now uses the shared streaming
scheduler. All 13 tests passed with `corepack pnpm exec vitest run
server/src/__tests__/heartbeat-project-repositories.test.ts`. The same
CI shard now passes.
- Ran `pnpm -r typecheck`, `pnpm test:run`, and `pnpm build` locally.
Each full local command hit SIGKILL/exit 137 in the 4 GiB container.
These local commands did not pass. The current-head CI gates above
provide the full verification.
- A real-Git disk-capacity regression confirms a typed failure and
removal of the incomplete manifest. Repeated writer attempts cannot
exceed the permitted page count.

- Follow-up real Daytona and separate staging qualification passed with
the related archive validator (#14315) and exact-owner finalization fix
(#14314). Three successive turns copied back all 60,000 files with
39,828,890 filename bytes and five unusual names. Independent host
inventories verified every file and the pinned Git HEAD. Native,
provider, session, and process identities stayed fixed; no retry
remained. The task reached Done, and its browser-downloaded final proof
matched exactly. The reusable regression is #14316, including an
assertion of the effective environment idle policy.

## Risks

- SQLite manifests use disk space. Each receives one quarter of the
available capacity above the host reserve at creation. The reserve
defaults to 256 MiB and has a 64 MiB configuration minimum. Disk
capacity, filesystem quotas, per-path limits, Git resource use, and
execution deadlines remain limits.
- Each path and sink chunk has a 64 KiB limit. SQLite connections use a
1 MiB page cache. Invalid or incomplete records fail explicitly.
- Restore transport keeps fixed and configured archive exclusions. A
remotely created Git-ignored file can be transferred, but the host merge
excludes it through the manifest.
- Provider archive buffers, Git and tar memory, repository metadata,
legacy v1 arrays, and the separate referenced-source resolver retain
their own limits. Existing provider safety validators still buffer
textual tar listings: Daytona allows 32 MiB and Kubernetes allows 64
MiB. These separate transport limits can stop a sufficiently large
restore before merge. This change does not claim bounded total process
memory or unlimited transport size.
- New descriptors use v2. Existing v1 recovery remains supported; a
downgrade cannot read v2 descriptors.

## Model Used

OpenAI GPT-6 through Codex. The exact deployment ID and context limit
are not exposed in this run. The agent used code editing, terminal
execution, tests, and GitHub tools.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 08:40:36 -05:00
DottaandPaperclip 0f14d26123 fix(runtime): allow bounded large untracked workspace snapshots (#14194)
## Thinking Path

> - Paperclip manages AI agents and their work.
> - Remote runs need a snapshot of the task workspace before the agent
starts.
> - The snapshot lists untracked filenames through the shared Git scan
scheduler.
> - A generated directory with a few thousand long filenames can exceed
the 1 MiB output limit.
> - This stops setup and prevents the agent from continuing its task.
> - This change gives that listing a 32 MiB bound and keeps the explicit
file snapshot.
> - Normal generated trees can now pass setup, while larger snapshots
still fail at a finite limit.

## Linked Issues or Issue Description

Refs #11572 and #12214 for the existing bounded scan and ignore-scan
protections.

**What happened?**

An agent continuation failed during workspace setup with `Workspace Git
scan exceeded its output limit`. A real Git fixture reproduces the
untracked-file path: 5,000 long filenames in one generated directory
exceed its 1 MiB output limit.

**Expected behavior**

The workspace snapshot must support ordinary generated trees with
thousands of files. It must keep a finite output bound and select
explicit files before staging.

**Steps to reproduce**

1. Commit a base file in a Git repository.
2. Add 5,000 untracked files with long names in one new directory.
3. Call `readGitWorkspaceSnapshot` with the normal scan limits.
4. Observe the output-limit error before this change.

**Paperclip version or commit**

Base commit: `640dee1`.

**Deployment mode**

Remote sandbox execution from source.

## What Changed

- Increase the untracked-file snapshot output bound from 1 MiB to 32
MiB.
- Keep explicit file selection, the shared scheduler, the timeout, and
the other scan bounds.
- Test that 40,000 long filenames above the old 8 MiB bound reach the
snapshot. Reuse those files with deeper paths to prove that output above
32 MiB still fails.
- Test that files created after the snapshot, including an ignored
secret, stay out of the overlay archive.
- Check the workspace root identity and reject selected paths with
symlinked parent directories before upload.
- Test root replacement after path resolution, including root-level
selected files.
- Preserve existing workspace root aliases by capturing the resolved
root before snapshot selection. Test that later alias retargeting cannot
change the archive contents.
- Stop staging on permission and I/O errors; continue to allow missing
files.
- Test these failures and preserve selected symlink entries.
- Accept valid case-renamed directories by checking ancestor file types.
A modeled case-insensitive regression failed before this correction and
now passes.
- Give the 40,000-file fixture enough time to remove its files.
- Document the larger bound and staging behavior.

## Verification

- Red: the 5,000-file regression failed with `stdout maxBuffer length
exceeded` before the fix.
- A separate check through the real server scheduler reproduced
`workspace_git_scan_output_limit` on the original code. The revised code
selected all 5,000 files.
- The late-file regression failed against the first PR revision because
the archive contained `drafts/late.secret`. It passes with the final
explicit-file approach.
- The staging regressions failed before the review fix: a substituted
parent directory and permission/I/O errors were accepted. All three
cases now stop before upload.
- Green: 134 tests passed across `git-workspace-sync.test.ts` and
`sandbox-managed-runtime.test.ts` with Vitest 4.1.11. This includes
complete selection above 8 MiB and rejection above 32 MiB.
- `pnpm --filter @paperclipai/adapter-utils... typecheck` passed after
the revision.
- The module-boundary check and `git diff --check` passed.
- `pnpm -r typecheck` and `pnpm build` stopped in the Rust runner steps
because this environment has no `cargo` executable.
- The full `pnpm test:run` attempt ended with `SIGKILL` during the
general server suite. It did not finish. That full-suite result belongs
to the earlier revision. Fresh checks are required for this revision.

- The new regression fails at the old 8 MiB bound with `stdout maxBuffer
length exceeded`. All 134 focused tests pass with the 32 MiB change.
- The affected typechecks and module-boundary check pass. The full local
typecheck requires Cargo, which is absent in this environment.
- Final verification for `b43e9bc95e55382c6a9bfe200487c164770c8be8`: 54
successful checks/statuses and two skipped Storybook checks. No checks
remain pending or failed.
- The [CI
run](https://github.com/paperclipai/paperclip/actions/runs/36315569674)
passes on attempt 2. The first attempt had one unrelated preview-fixture
readiness timeout. That exact test passed locally; its CI shard passed
on the single rerun.
- [Greptile reports
5/5](https://github.com/paperclipai/paperclip/pull/14194#issuecomment-5852045501)
on this revision. All review threads are resolved.
- [Security review accepts the documented memory
tradeoff](https://github.com/paperclipai/paperclip/pull/14194#discussion_r4115149465)
for this finite mitigation. The separate streaming follow-up will remove
full-list buffering.
- This revision also passes affected local typechecks and the
module-boundary check. Full local typecheck/build stop because Cargo is
absent. The full local Vitest attempt was stopped after about 18 minutes
once all remote gates passed; it did not complete locally.

## Risks

- Each untracked-file scan can buffer up to 32 MiB instead of 1 MiB. The
scheduler still limits concurrent scans and execution time.
- Snapshots above 32 MiB still fail with the existing error. Tracked and
ignored-file scan bounds stay unchanged.
- A workspace that replaces a selected path’s parent with a symlink now
fails staging.
- Detailed logs from the reported host were unavailable. The exact
command that exceeded its limit on that host is unconfirmed.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, tool use, and code
execution. The runtime does not expose a more specific model build or
context window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub references)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-27 06:42:19 -05:00
Devin FoleyandPaperclip b2e9e82f05 fix: stop remote Grok runs before continuing queued messages (#14100)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The execution service owns each run and saves messages sent while it
runs.
> - Interrupt must stop the current executor before it delivers those
messages.
> - Remote Grok commands did not register the host cancellation control.
> - A cancelled task run could still write Done and prevent queue
recovery.
> - This pull request connects remote cancellation and revokes cancelled
run writes.
> - Saved input can use the existing queue admission rules after
verified cleanup.

## Linked Issues or Issue Description

**What happened?**

Interrupting a queued message marked a remote Grok run cancelled before
its sandbox stopped. The old run could still post a reply and mark the
task Done. Its saved follow-up remained deferred behind execution
recovery.

**Expected behavior**

Stop revokes run write authority and waits for verified termination.
Saved messages remain durable and enter one successor through normal
admission after cleanup.

**Steps to reproduce**

1. Run a task with `grok_local` in a remote sandbox.
2. Send a follow-up and use Interrupt while the command runs.
3. Let the old command attempt a task status update after cancellation.
4. Observe the task disposition and the saved message queue.

**Paperclip version or commit**

The gap is present in master at `d3e0f0a238`.

**Deployment mode**

Authenticated server with a Daytona sandbox.

Related work: #14028 and #14046 handle bounded continuation. #13291
covers infrastructure interruption and verified remote cleanup. #13332
addresses atomic recovery holds. This change handles direct Grok
operator cancellation and stale task writes.

## What Changed

- Register remote Grok cancellation before preparation. Keep command
ownership until the host confirms sandbox termination.
- Reuse the sandbox cancellation boundary for the direct CLI invocation.
Reject fresh attempts after cancellation and preserve workspace restore
failure evidence.
- Reject writes from cancelled task JWTs and runs with a pending stop.
Preserve diagnostic reads and existing conversation error codes.
- Recheck run authority under a database lock before task updates and
interaction responses commit.
- Preserve authorized handoffs that stop their own run. Only the
server-issued stop receipt for that request permits the final task
update.
- Add tests for hung commands, unverified stops, early cancellation,
copy-back failures, late Done, late interaction responses, authorized
handoffs, exact lease receipts, and one queue successor across
concurrent restart sweeps.
- Document the cancellation and write-authority contract.

## Verification

- Targeted adapter, cancellation-boundary, authentication,
queued-message, interaction-service, and activity-route tests passed.
The expanded run passed 214 tests; one new test had an incomplete
fixture. After correcting the fixture, all 8 selected follow-up cases
passed.
- `pnpm -r typecheck`: passed on
`179c86caf1bf0d89914a503d46e24af7e4b8c557`.
- `pnpm build`: passed on the same commit.
- `pnpm test:run`: the general-server group completed with 13,521
passed, 99 skipped, and 18 failed tests. It then stopped, so the
remaining local groups did not run. Five Slack, email, and wake-batching
failures passed on focused reruns after correcting the local
environment. The remaining 13 failures reproduce as `EACCES` on rename
in unchanged skill-cache code on macOS. Two custom-image suite setup
hooks also failed to start embedded PostgreSQL after the machine
exhausted shared-memory slots; all 31 tests in that file passed on rerun
after the local resource issue was resolved. CI covers all test groups.
- CI: 53 checks passed and 2 were skipped on the latest commit,
including the aggregate verification gate. The last server shard passed
on its single rerun after a preview-server startup timeout. The affected
file also passed locally with 28 passed and 3 skipped.
- Greptile: 5/5 on the latest commit. Both review threads are resolved.
- No live deployment or staging task mutation has been performed.

## Risks

- Stopping the sandbox can prevent file copy-back. The result preserves
workspace restore failure evidence; termination does not imply restored
files.
- If provider termination fails, the adapter keeps ownership of its
outstanding command and does not acknowledge Stop.
- The write restriction now applies to ordinary cancelled tasks. Reads
remain allowed. Task and interaction checks add a shared run-row lock to
agent mutations. An exact server-issued receipt permits the task request
that stopped its own run to complete its handoff.
- Existing terminal tasks are not reopened automatically. An operator
must correct a historical late Done before its saved queue can continue.
- No schema migration or UI change.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository inspection, code
execution, and test tools. The precise backend revision and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-26 17:07:07 -07:00
Devin FoleyandPaperclip 4ca404b49a fix: record safe sandbox restore failure diagnostics (#14064)
Record a bounded diagnostic for failed workspace and staged-asset restores.
Preserve the original error, retry policy, and archive safety checks. Never
copy raw provider messages, credentials, paths, or asset names into the log.
Nested failures log once; safe fields survive throwing property getters.

Verified 129 focused restore/Claude tests, typecheck/build, and green full
PR CI. Greptile 5/5 with all review threads resolved.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-25 17:16:13 -07:00
DottaandPaperclip 96bf004a79 fix: use persisted state for lifecycle continuation and retry budgets (#13888)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Its control plane decides when a task can continue, wait, stop, or
complete.
> - Legacy continuation could change when an agent changed its wording
without changing task state.
> - Shared attempt counts also let repair and infrastructure retries
affect each other's limits.
> - This pull request uses persisted state and separate, bounded
allowances for these decisions.
> - If automatic repair stops, the task explains what happened and
offers a guarded retry.
> - Paired tests and real-provider evaluations verify that Stop,
approvals, ownership, and spending limits remain authoritative.

## Linked Issues or Issue Description

Related work: Refs #13761, Refs #11126, Refs #13610. These cover
obsolete continuation dispatch and retry storms. Open and closed issues
and PRs were searched for related lifecycle, continuation, and retry
work.

**What happened?**
Legacy continuation depended on English wording and progress heuristics.
Repair, failure retry, and productive continuation could consume shared
counts. When bounded repair stopped, the task showed a technical
recovery message without a clear next action.

**Expected behavior**
Persisted disposition and owned execution paths determine the next
action. Missing disposition prompts bounded agent repair. Explicit work
mode determines planning mode. Narrative changes and raw activity counts
cannot replenish allowances. An exhausted repair shows a readable
notice. An explicit retry checks current controls and preserves the
assigned agent.

**Steps to reproduce**
Run `pnpm test:lifecycle-baseline`. The paired probes keep structured
state constant while varying completion, planning, blocker, and progress
prose. Run the explicit `lifecycle-baseline` and
`continuation-accounting` Product E2E suites for real-provider coverage.
In Storybook, open **Design previews / Recovery notice** to inspect the
production component's normal, pending, acknowledged, unavailable,
failure, and mobile states.

## What Changed

- Hide the image attachment button, icon, and drop/paste hint in answer
composers. Image paste and drop support remains available.
- Merge current master and retain both browser regression sets. Use a
production-stamped service worker in the offline recovery browser
fixture.
- Share one state-based legacy continuation decision across immediate,
delayed, and recovered dispatch. Bind bounded repairs to their source
run and episode.
- Remove title and description wording from work-mode authority. Agents
can still write requested plans in execution mode.
- Persist separate failure-retry and productive-continuation counters.
Disposition repair and resource waits cannot consume or reset those
allowances.
- Validate delayed repair identity, then recheck current gates before
provider dispatch. Fence native startup cancellation.
- Show **Agent needs attention**, a plain-language explanation, **Retry
agent**, and expandable details in both task interfaces. Report request
progress, acknowledgement, and errors inline.
- Store typed recovery notice metadata. Recognize older active notices
only through exact stored action and run IDs. Notice text never grants
retry authority.
- Use the existing recovery-action endpoint for retry. Recheck current
action, status, owner, agent availability, dependencies, active runs,
pending questions and confirmations, approvals, pause controls, and
budget. Duplicate requests do not wake twice.
- Add component, page, route, database, contract, and Storybook
coverage. Keep the scenario inventory and executable evals here.
Historical reports and snapshots live in the [commit-pinned
paperclip-evals
archive](https://github.com/paperclipai/paperclip-evals/blob/ce3e5afcd4a1184650f586a2b5b8be5874c66c8b/experiments/2026-09-lifecycle-authority/README.md).
- Preserve unsaved project fields while the same project URL changes to
its canonical alias. Do not reuse data across projects or companies.
This separate fix addresses the repeated repository-editor browser
failure without changing the browser test.
- Keep the development service worker from intercepting Vite module
reloads. Update the connection-intent browser fixture to record progress
and completion through the agent API.

## Verification

Merge preparation on September 25, commit
`c1e8e4b7ddd9fbc4913ed55ce21b8e12906c2f97`:

- Merged master `bd2030932` and resolved the browser test-list conflict
by keeping both sets of regressions.
- Deterministic lifecycle baseline: 1,090/1,090 assertions passed; no
failures, skips, or missing selected evidence. Unit 423, runner 184,
database integration 397, grading 86.
- Browser support: 17/17 passed. The offline recovery test first failed
with an unstamped development worker, then passed with the production
stamp. Its assertions are unchanged.
- Focused interaction UI and offline fallback tests: 19/19 passed.
Verified the custom-answer composer in Storybook: no attachment controls
or hint; entering an answer enables Next.
- Recursive typecheck, production build, token gates, and diff checks
passed. The worktree is clean. No new real-provider campaign was run.
- Current CI and review: [Current PR CI
passed](https://github.com/paperclipai/paperclip/actions/runs/36166011243):
55 successful checks and two optional Storybook skips. Greptile scored
this exact commit 5/5. Hiding the question attachment controls is an
intentional UI change; paste/drop remains available.

Earlier recovery UI verification, commit
`21be0fec0e90e86b6d662b8ee4831847cd041cdb`:

- Recursive typecheck, production build, token gates, and diff checks
passed.
- Focused UI coverage: 338 tests passed across six suites (336 before
the interaction guard, with the two affected suites rerun at 149 passed
after it). Covers both task interfaces, the real page mutation,
pending/error acknowledgement, stale state, and unavailable controls.
- Recovery database integration: 352 tests passed before the interaction
guard. The complete recovery-action and mutation-route suites passed 181
tests after it. The two new pending question/confirmation regressions
failed before the fix and passed afterward, including
resolved-interaction controls. Shared validator suite: 31 passed. E2E
catalog suites: 34 passed.
- Browser inspection passed for light/dark themes, mobile layout,
expandable details, pending retry, acknowledgement, failure, and
disabled retry. Storybook renders the production component; its request
is simulated.
- The broad local run hit two chat callback-order wait failures and was
stopped after all CI unit/database/runner shards passed. Both local
failures passed when rerun without the competing full-suite process.
- CI exposed a repeated project-repository draft-loss race during
canonical redirects. A new unit regression failed before the fix; all
nine project-page tests now pass, including controls for other projects
and companies. Both unchanged repository browser tests passed against a
fresh local server. UI typecheck, production UI build, and token gates
passed after this fix.
- [Earlier PR CI
passed](https://github.com/paperclipai/paperclip/actions/runs/36072486798)
on `21be0fec0e90e86b6d662b8ee4831847cd041cdb`: 55 successful checks, two
optional Storybook skips, and no failed or pending checks. The
repository browser shard passed with the production fix. Greptile is 5/5
on this exact commit with no unresolved review threads. The PR is
mergeable.

Historical, source-qualified lifecycle evidence:

- Lifecycle baseline: 1,074 assertions. Native session coverage: 447
tests. Product E2E support: 515 tests. Browser support: 11 tests. Full
earlier verification is retained in the archive.
- [Real-provider campaign: 8/8 passed, zero
retries](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35881382080-1/index.html),
source `e88d210417280140b44a36449027290adcb1aeaa`. Evidence and cleanup
checks passed. This includes deliberately exhausted repair cases that
correctly remain blocked; it does not mean every task finished Done.
This campaign predates the recovery UI change.
- Archive migration verified all 16 original JSON files byte-for-byte
and all 24 checksum entries. App tests do not need private archive
access. [Archive PR
#27](https://github.com/paperclipai/paperclip-evals/pull/27) is merged.

## Risks

- Agents that omit durable disposition receive at most two repair
attempts by default. Prose-only completion exposes missing state rather
than silently changing scheduling.
- A retry is an explicit board action. The server rechecks current
controls. A successful response confirms the task returned to To do; it
does not claim that the provider has already started.
- Existing notice metadata remains valid. Only older active notices with
matching structured evidence receive the new UI. Historical notices
without that evidence keep their existing rendering. No schema migration
is required.
- Old run records require conservative retry accounting. Tests cover old
counters, alternating retry lanes, restarts, and exhausted repairs.
- Historical snapshots require private `paperclip-evals` access. The app
index retains public campaign links. Live campaigns qualify specific
sources and scenarios; no new real-provider campaign has run for the
recovery UI commit.

> This fixes existing lifecycle and recovery behavior and does not
duplicate planned core work.

## Model Used

OpenAI GPT-6 through Codex assisted implementation, reasoning, code
execution, and review. The exact serving model ID and context window are
not exposed in this task. Historical real-provider evaluations used
Codex model `gpt-5.6-sol`, separately from the implementation assistant.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-25 15:28:11 -07:00
Devin FoleyandPaperclip bd6caf51bb fix: preserve restore failure results and stop unsafe retries (#14035)
Preserve agent output and earlier execution errors when workspace restore fails. Report the restore phase and confirmed saved-plan links. Require verified repair before retrying unsafe archives, while preserving approval states and the retry budget.

Verified with full CI, 506 focused regression tests, and Greptile 5/5 with all review threads resolved.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-25 11:34:31 -07:00
Devin FoleyandPaperclip f3a214fe77 Fix relative symlinks in secondary sandbox repositories (#13953)
## Thinking Path

> - Paperclip manages agents and their task workspaces.
> - A task can use several independent Git repositories.
> - Sandbox staging copies secondary repositories from temporary clones.
> - The copy changed relative symlinks into absolute host paths.
> - Those links broke skill discovery and made workspace restore fail.
> - This change preserves link targets while keeping extraction checks
intact.

## Linked Issues or Issue Description

**What happened?**

Staging a secondary repository rewrites a link such as
`.claude/skills/demo -> ../../skills/demo` to an absolute path in a
temporary Git clone. That clone is then removed. The link is broken in
the sandbox, and Daytona refuses the outbound archive during workspace
restore. An agent can finish its turn but still have its run fail during
restore.

**Expected behavior**

Repository links keep their original targets after staging. Links within
a repository remain usable, and changes return to the local checkout.
Unsafe outbound archive links still fail before extraction.

**Steps to reproduce**

1. Create a project with a primary repository and a secondary
repository.
2. Commit a relative skill directory link in the secondary repository.
3. Stage the workspace for sandbox execution and inspect the copied
link.
4. Restore that repository through Daytona. Before this fix, the link
points at a removed host temporary directory and restore rejects it.

**Paperclip version/commit**

Reproduced on `0f8750627f11d855552abce9a837d7f3b67c9ddf` in the
multi-repository sandbox path.

Related: #13442 introduced multi-repository provisioning. #13882 adds
native Grok but keeps legacy adapters; #12991 addresses Grok instruction
isolation and leaves skill staging unchanged. Searches of open/closed
PRs and open issues found no direct fix for this copy behavior.

## What Changed

- Set `verbatimSymlinks: true` when copying secondary Git clones. This
[Node
option](https://nodejs.org/api/fs.html#fspromisescpsrc-dest-options)
preserves the stored link target instead of resolving it against the
temporary source.
- Test directory, file, chained and dangling links after temporary-clone
cleanup. Assert the copied Git checkout remains clean.
- Cover skill-link reads, edits and restore in fresh, warm-adoption and
durable-seed workspace modes.
- Extend Daytona checks for valid relative directory links and rejected
absolute targets.
- Document the staging behavior.

## Verification

- Four regression cases fail without the source fix: one clone test and
three staging modes.
- `pnpm exec vitest run
packages/adapter-utils/src/git-workspace-sync.test.ts
packages/adapter-utils/src/sandbox-managed-runtime.test.ts
packages/plugins/sandbox-providers/daytona/src/plugin.test.ts`: 301
passed.
- `pnpm -r typecheck` and `pnpm build`: passed.
- `pnpm test:run` was started locally, then stopped after the full Linux
CI suite passed. It did not finish locally; this is not a full
local-suite pass.
- Full PR CI: 53 checks passed, two conditional skips, on
`07b30ff298baab327a4e60708ade899935688a3f`. Three server shards were
interrupted by runner shutdowns; the unchanged mobile repository test
timed out waiting for a disabled Save changes button. One same-commit
failed-job rerun passed. Original attempts remain in [run
36036538369](https://github.com/paperclipai/paperclip/actions/runs/36036538369).
- Greptile: 5/5 on the same head, no review threads or actionable
findings.
- Diff scanned for secrets and private identifiers; no matches.

## Risks

Low risk: the production change is one copy option. It preserves
symlinks instead of following or materializing their targets. Daytona
extraction guards, workspace exclusions, authentication and database
behavior do not change.

This prevents corruption in newly staged snapshots. It does not rewrite
an already corrupted warm workspace or durable seed; those need fresh
staging from the source checkout. No live provider run or customer-task
replay was performed. The tests use real Git, filesystem and tar
operations with mocked provider transport.

## Model Used

OpenAI GPT-6 through Codex, with tool use and code execution. The exact
serving model identifier and context-window size are not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-24 10:58:49 -07:00
Devin FoleyandPaperclip 4b8ec588f3 Stop duplicating wake context in adapter environments (#13891)
## Thinking Path

> - Paperclip manages agent work and preserves task context.
> - Built-in adapters already include wake context in the agent prompt.
> - They also copy the full wake JSON into a process environment
variable.
> - A large environment entry can prevent the agent from starting with
`spawn E2BIG`.
> - This change removes the duplicate environment entry and uses the
existing prompt delivery.
> - The agent keeps its context without extra file transport or new
history limits.

## Linked Issues or Issue Description

Refs #13144, #13860, #13872, #13793.

Large wake payloads can exceed the operating system limit for one
environment entry. The launch-envelope fix in #13793 handles the outer
transport but leaves that child environment entry intact.

Credit to @nickyleach for the prompt-only approach in #13144. This PR
applies that part on current master. It does not include that PR's
30-item history limits or recovery-history endpoint. Those behavior
changes can be reviewed separately from the process launch fix.

## What Changed

- Stop exporting `PAPERCLIP_WAKE_PAYLOAD_JSON` in the shared ACP engine
and all ten built-in adapter writers.
- Ignore configured values of the retired variable so saved adapter
settings cannot restore the oversized entry. Also drop inherited copies
in Hermes, which builds its environment directly.
- Keep scalar runtime variables, existing prompt rendering, continuation
history, resume deltas, gateway bodies, and Hermes JSON template
variables.
- Document the prompt delivery contract and the migration for custom
instructions that read the retired variable.
- Test large local and sandbox child-process launches, fresh and resumed
ACP turns, SDK delivery, and configured-variable filtering.

## Verification

- `pnpm -r typecheck` passed.
- Focused adapter utility, ACP, Codex child-process, and Cursor Cloud
suites: 340 tests passed.
- Hermes execution and prompt tests: 18 tests passed using its package
Vitest configuration.
- The child-process tests deliver over 128 KB of context through stdin
and check the complete text. The ACP test retains 50 complete messages
and 50 completed actions, then checks the resumed delta.
- `pnpm build` passed.
- `pnpm test:run` was attempted, then stopped after it reproduced ten
macOS runtime-skill-cache permission failures (also reproduced on
unchanged master) and one HTTPS backfill test failure. The HTTPS test
passed when rerun unchanged on this branch and master. The complete
local suite was not completed; Linux CI provides the full-suite gate.
- CI is green on c3b8191f97: 53 checks passed and two skipped. Attempt 3
passed the remaining shards without source changes. Earlier attempts hit
a lost runner, preview readiness, and chat browser failures. The preview
test passed unchanged locally. Local chat tests passed 3/4; the
remaining sidebar focus-style assertion also fails on unchanged master.
No tests were weakened or skipped to obtain the passing rerun.
- The staged diff passed `gitleaks stdin --redact` and `git diff
--check`.

## Risks

- Custom instructions or scripts that read `PAPERCLIP_WAKE_PAYLOAD_JSON`
must use the wake payload in the prompt instead. The variable is absent
even for small wakes.
- This removes one cause of `E2BIG`. Legacy CLI paths for Gemini, Grok,
Kimi, Pi, and Hermes still pass prompts as arguments and retain their
existing argument-size limits. Other large environment variables also
remain subject to OS limits.
- No new context truncation, API route, database migration,
authorization change, or production rollout is part of this PR. Existing
prompt windows and resume rendering remain unchanged.

## Model Used

OpenAI GPT-6 through Codex, with code execution and repository tools.
The exact model variant and context window size are not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-23 13:50:10 -07:00
Devin FoleyandPaperclip 8326e33ada Fix oversized sandbox process launch payloads (#13793)
## Thinking Path

> - Paperclip manages agent work and preserves context across retries.
> - Sandbox ACP runs encode their command and environment into one
launch value.
> - A long continuation can make that encoded value exceed Linux's exec
limit.
> - The launch shell then exits before the agent can initialize.
> - This PR transfers large command envelopes through a private
temporary file.
> - The agent receives the complete environment and can start normally.

## Linked Issues or Issue Description

Refs #13777.

**Bug description**

Sandbox tasks with long retry context fail during ACP initialization
with exit 127. The streamed bridge also drops the shell error that
explains the failure.

**Steps to reproduce**

Launch the streamed sandbox process bridge on Linux with one valid
110,000-byte environment value. Its base64 command envelope exceeds the
limit on one exec argument or environment string. Daytona reports
`argument list too long: env`, then exit 127.

**Expected behavior**

The command envelope must not make a valid child environment too large
to launch. A shell startup failure must retain its diagnostic in the run
log.

## What Changed

- Keep envelopes up to 64 KiB on the existing launch path. Upload larger
envelopes in bounded chunks inside a mode-0700 session directory. Set
the final payload file to mode 0600.
- Read the file without following symlinks and delete it before spawning
the child. Remove incomplete uploads on failure. Both streamed and
polled bridges use the same envelope.
- Preserve stderr when the launch shell fails before the wrapper emits a
terminal event. Emit a fixed terminal error and shutdown acknowledgement
if a payload cannot be read or parsed, without exposing its contents.
- Cover large environments, file permissions, payload deletion,
interrupted uploads, missing or malformed payloads, and startup
diagnostics. Document the transfer and cleanup behavior.

## Verification

- Both large-envelope regressions fail before the fix and pass after it.
- Targeted bridge, ACP engine, real-spawn, and stdin-race checks pass:
370 tests.
- `pnpm --filter @paperclipai/adapter-utils typecheck` passes.
- A gated live Daytona probe on the current sandbox image reproduced
exit 127 with the old bridge. The fixed bridge launched the same command
successfully. A second probe completed real Claude ACP initialization
with a 110,000-byte context value. It did not run an agent task. All
temporary sandboxes were deleted.
- `pnpm -r typecheck` and `pnpm build` reach the unchanged Rust runner
step and stop because this machine has no `cargo` executable.
- Full CI passes on `ea816cf58d`: [run
35683853754](https://github.com/paperclipai/paperclip/actions/runs/35683853754).
All 53 checks pass; two optional checks are skipped. The local
full-suite run was stopped after equivalent CI suites passed; it has no
final local result.
- Greptile is 5/5 on `ea816cf58d`, with no unresolved review threads.
The branch is mergeable.

## Risks

Large envelopes require extra upload calls during startup. The temporary
data stays inside the private session directory and is removed before
child startup or during failure cleanup. Individual child environment
values still obey the operating system's native limits. No migration or
configuration change is required.

## Model Used

OpenAI GPT-6 (Codex), with reasoning, repository tools, and code
execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (targeted checks;
full-workspace limits described above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 20:54:07 -07:00
Devin FoleyandPaperclip c221589b69 fix: close sandbox process proxies after remote exit (#13777)
## Thinking Path

> - Paperclip manages AI agents and their tasks.
> - Sandbox agents use a local proxy to exchange ACP messages with a
remote process.
> - ACP keeps the proxy input stream open while it waits for a reply.
> - The proxy received a remote exit but kept that input stream open.
> - This left the proxy alive and could block later task retries.
> - This pull request closes the input handle after a terminal remote
event and lets final output drain.

## Linked Issues or Issue Description

**What happened?**

A sandbox process could exit while its local proxy stayed alive. During
startup, this could appear as a handshake timeout. A later task retry
could then stop because the previous process was still alive.

**Expected behavior**

The proxy should exit with the remote process status, even when ACP has
not closed its input stream. Final output and diagnostics should arrive
before the proxy closes.

**Steps to reproduce**

1. Start a sandbox process-session bridge with a child that writes
output and exits.
2. Start the generated local proxy and keep its input stream open, as
ACP does during startup.
3. Observe that the proxy stays alive after the remote process exits.

**Paperclip version or commit**

Reproduced on master at `846336e5a0`.

Related: #13272 covers legacy sandbox startup recovery. #13765 covers
the task retry UI. This change fixes the proxy process lifecycle.

## What Changed

- Close the proxy input handle on remote exit or error, and preserve the
exit status.
- Stop forwarding input after the terminal event.
- Wait for process close in the test helper so output is fully drained.
- Test successful exit, failed exit, and spawn failure with input held
open in both output modes. Check complete delivery of 208 KiB of final
output.

## Verification

- Before the fix, all four remote-exit regression cases timed out.
- After the fix, all six terminal-event cases pass.
- Targeted runtime suite: 362 tests pass across the sandbox bridge,
stdin queue races, real ACP spawn, and ACP engine suites.
- `pnpm --filter @paperclipai/adapter-utils typecheck` passes.
- Full workspace tests hit local platform failures: embedded PostgreSQL
fails during initialization, and runtime skill-cache tests fail with
`EACCES` while renaming read-only directories on macOS. The affected
server code is unchanged in this PR. Those suites pass in CI.
- `pnpm -r typecheck` and `pnpm build` stop at the existing Rust runner
step because this machine has no `cargo` executable. The CI typecheck
and build gates pass.
- [CI run
35658635907](https://github.com/paperclipai/paperclip/actions/runs/35658635907)
passes all required gates. The first browser shard run passed all 12
tests but hit a GitHub 403 during report upload; the rerun passed,
including upload.
- Greptile scored commit `9c3c89148e` 5/5 with no review threads. The
branch is mergeable.

## Risks

The proxy now closes its input immediately after a terminal remote
event. Tests cover final output delivery and preserve nonzero exit
codes. Existing orphan processes still require cleanup or a server
restart. Live sandbox startup needs validation after deployment; this
change addresses the confirmed proxy hang and does not establish why an
earlier remote process stopped.

## Model Used

OpenAI GPT-6 (Codex), with reasoning, repository tools, and code
execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (targeted checks;
full-workspace limits described above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
(existing behavior restored; no documentation change needed)
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 16:08:23 -07:00
DottaandPaperclip d9b3a5653e feat(chat): add initial Slack communication guidance and connection menus (#13760)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Chat connectors let people use the same tasks and agent tools from
external conversations.
> - Agents need communication guidance that fits the conversation
medium.
> - That guidance belongs in the original task context, without repeated
instructions on each turn.
> - Connection owners also need clear settings and a consistent way to
remove a connection.
> - This pull request adds initial Slack guidance, optional connection
instructions, and chat connection menus.
> - The benefit is clearer Slack replies with the existing Paperclip
workflow and permissions.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

Agent replies in Slack and chat connection management in the Apps
catalog.

**Current behavior**

Slack tasks do not carry a saved communication profile. The catalog
shows a separate Manage button and does not offer removal on every chat
connection row.

**Proposed behavior**

Save Slack guidance when a new conversation creates a task. Restore that
original guidance when a model session is rebuilt. Do not append it to
ordinary follow-ups. Expose optional additional instructions in Slack
Settings. Put Manage and Remove connection in a three-dot menu for all
chat providers. Keep Finish setup visible for drafts.

**Reason and benefit**

Small answers fit in Slack. Substantial deliverables use ordinary
document or artifact tools with a useful Slack summary. Connection
settings apply to new tasks and cannot change permissions. Users can
remove both active and unfinished chat connections from the catalog.

**Breaking changes**

Two additive database columns store endpoint preferences and the initial
conversation snapshot. Existing endpoints default to empty preferences.
Existing conversations keep their original behavior. Non-Slack guidance
is unchanged.

Related public context:
https://github.com/paperclipai/paperclip/pull/13741 improves native chat
recovery. This change adds communication context to those existing
execution paths. A search found no duplicate communication-guidance PR.

## What Changed

- Add a provider-guidance registry, enabled for Slack first.
- Persist optional endpoint communication instructions and capture an
immutable snapshot when a conversation creates a task.
- Resolve guidance from the verified company-scoped connection. Restore
it for fresh native and legacy sessions without per-turn reminders,
extra model calls, or extra context queries.
- Add the Slack Settings field, validation, audit coverage, and
Storybook save/error states.
- Add Manage and Remove connection menus for all seven chat providers.
Keep the draft setup button. Require removal confirmation and allow
retry after failure.
- Add regression coverage, an active/draft menu story, and connector
documentation.

## Verification

All CI checks are green for 5f48df4e0. Greptile scored this head 5/5
with no actionable findings. No review threads remain unresolved.

- Passed `pnpm -r typecheck` and `pnpm build` on PR head 5f48df4e0.
- Passed design-token checks, UI typecheck, and all 20 catalog tests
after rebase. Tests cover all seven providers, active/draft removal,
confirmation, cache refresh, errors, and cancellation.
- Verified the active/draft menu in Storybook. The interaction test runs
without browser console errors.
- Passed focused guidance, endpoint persistence/isolation, heartbeat
trust, native context, ACPX, adapter utility, and CLI recovery tests.
Full UI and CLI groups passed (6,512 and 502 tests).
- Tested real Slack conversations on staging: concise updates with
public links, a planning question with buttons, a saved plan, a saved
report, task creation and assignment, and explicit detailed output. Old
tasks retained original preferences after an edit; a new task used the
changed preferences. Restored the staging setting afterward.
- Existing safe progress remained visible without duplicate final
replies or private reasoning.
- Broad local tests found resource/time-sensitive failures that passed
targeted reruns. One Cursor archive-download fixture failed on both this
branch and the unchanged main checkout. The full local suite is not
claimed clean. All PR-head CI test shards passed, including general,
serialized, Runner, and browser suites. The redundant local full-suite
rerun was stopped after CI completed successfully.
- Live delegation was not tested because the staging company has only
one agent. Live testing also found separate latency and runner
task-editing capability gaps; this PR does not add connector-specific
workflow behavior to hide them.

## Risks

- Prompt guidance changes the form of new Slack replies. Explicit
requests for detail still take precedence.
- The additive migration is idempotent. Conversation snapshots remain
fixed when connection settings change.
- Native and legacy recovery must preserve the initial context without
duplicates; targeted tests cover these paths.
- Removing a connection stops new work through the existing lifecycle
action. It retains Paperclip task history and does not delete the
external app or bot.

## Model Used

OpenAI GPT-6 through Codex, with repository editing, shell tools, and
browser testing. The host does not expose a more specific model ID or
context-window size. No separate model calls were added to the product.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (targeted suites; broad
local limitations are listed above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 14:14:25 -05:00
DottaandPaperclip e8c8ba3c19 feat(apps): add experimental MCP aggregator connectors (#13755)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Its tool gateway applies company access rules and approval controls
to connected apps.
> - MCP aggregators expose many apps through one provider endpoint.
> - Each aggregator needs its own credential, catalog, grants, and
lifecycle in Paperclip.
> - This pull request adds independent Zapier, Arcade, Composio Connect,
and Executor setup with a common Access → Connect layout.
> - A default-off MCP aggregators flag lets operators opt in while we
complete provider acceptance tests.
> - Agents use the normal Paperclip permissions, Test screen, and
gateway after setup.

## Linked Issues or Issue Description

**Subsystem affected**

Apps, connection setup, shared contracts, and the remote MCP gateway.

**Problem or motivation**

Aggregator endpoints need clear provider setup and correct MCP sessions.
Generic setup does not explain each provider's authentication or broad
execution tools. Provider approval must preserve the original execution
instead of replaying a write.

**Proposed solution**

Add four separate connectors behind Settings → Experimental → MCP
aggregators. Start with human and agent access, then connect the
endpoint and read its tools. Enable tools by default. Use the existing
Permissions and Test screens after setup. Keep legacy Composio API-key
and child connections intact.

**Alternatives considered**

A shared connection for all providers would mix credentials and access
rules. Separate provider-specific permission and test screens would
duplicate existing controls. Vercel Connect is outside this change.

**Roadmap alignment**

Extends the existing MCP Tool Gateway & Apps capability and the
Connected Apps roadmap area. This work was requested and reviewed by the
maintainer.

Related work: #11894, #12630, #12632, #12634, and #12906 concern the
legacy Composio broker. #13102 also covers remote MCP pagination. This
change preserves the broker path and adds initialized sessions, response
matching, and provider resume handling alongside pagination.

## What Changed

- Add branded setup and interactive Storybooks for Zapier, Arcade,
Composio Connect, and Executor. Use the existing access controls and
normal action tests. Do not request a connection name or action choices
during setup.
- Add the default-off `enableMcpAggregators` flag to settings, managed
feature metadata, the catalog, and setup guards. Hidden connections keep
running. Legacy Composio connections remain unchanged.
- Reuse the vault, grants, policy, and catalog models. Support OAuth
discovery, bearer tokens, custom headers, and credential-bearing URLs.
Add no database tables or migrations.
- Initialize and retain Streamable HTTP sessions by connection and
effective credentials. Read paginated catalogs and match streaming
responses to request IDs.
- Classify unfamiliar aggregator tools as writes despite upstream
read-only hints; only exact reviewed read capabilities enter the
read-only allowlist. Legacy Composio child behavior is preserved.
- Preserve provider authorization links and execution IDs. Support
Executor approve/resume, decline, and cancel without automatic replay of
uncertain writes.
- Preserve Off and Ask first choices during refresh and reconnect. Allow
new tools and retire removed tools. Keep agent access updates atomic and
preserve an empty agent selection.
- Document connector UX rules, provider branding sources, and live
acceptance results.
- Stabilize the existing Sentry release fixture after its repeated CI
failure by reusing one module mock; production Sentry behavior is
unchanged.

## Verification

- Final head `d11781970`: [CI
run](https://github.com/paperclipai/paperclip/actions/runs/35633534900)
passed, including broad typecheck, test shards, build, and E2E. All 54
checks pass; 2 optional checks are skipped. Greptile is 5/5, Security
Scan passes, and all review threads are resolved.

- Passed 27 focused connector Vitest checks and 18 connector-only
Storybook browser checks before the flag change. All 85 stories rendered
at desktop and narrow widths.
- Passed 5 connector lifecycle/server checks and 7 selected flag checks
after adding the flag. The latter cover settings, managed defaults,
cached catalog visibility, and all four setup routes.
- Review fixes passed 13 risk/handoff/lifecycle checks, dedicated
session-expiration and transport regressions, 13 selected
connector/gateway CI cases, and 10 selected setup/reconnect UI cases. A
real Composio connection-list call also succeeded through the refreshed
UI on `9ab115f71`.
- UI and server TypeScript checks passed. UI build, Storybook build,
token gates, and diff whitespace checks passed during implementation.
- Real browser and real Paperclip agent tests passed for Arcade,
Composio, and Executor. Tested action permissions, denied agent access,
reconnect, disconnect, and isolation. Tested Arcade catalog
additions/removal and Executor provider approve/resume, decline, and
cancel.
- Zapier live acceptance is incomplete. Its dedicated provider server is
configured, but its credential-copy dialog returned an empty clipboard
through browser automation. No live Zapier action is claimed.
- The three isolated Sentry release cases pass after the CI fixture fix.
- Local verification is deliberately narrow at the maintainer's request.
The full local suite, recursive typecheck, and repository-wide build
were not run. CI provides the broader checks.

## Risks

- Shared MCP transport changes affect other remote MCP servers. Protocol
fixtures cover initialized sessions, streaming response matching,
pagination, and isolation.
- Broad execution tools remain broad permissions. The provider governs
actions inside those tools.
- Provider handoff links are retained briefly in memory. After a server
restart, a one-time link may require reopening the provider dashboard.
Paperclip does not replay the original call.
- Zapier remains unproven live. Custom-header imports and self-hosted
endpoints have fixture coverage rather than a separate live account for
every variant.
- Turning the experimental flag off hides setup; it does not revoke
existing credentials or stop existing connections.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, shell
execution, and browser automation. The exact runtime model ID and
context-window size are not exposed in this session. A separate
Anthropic-backed Paperclip agent performed live gateway acceptance
tasks.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 12:53:11 -05:00
DottaandPaperclip 45c99a0d06 fix(adapters): default legacy harnesses and connected tools to full auto (#13693)
## Thinking Path

> - Paperclip lets people manage AI agents and their work.
> - Legacy adapters launch provider CLIs and expose connected tools.
> - Existing defaults did not consistently grant full automatic
permission.
> - Remote Claude used a fixed tool list that omitted MCP tools and
future tools.
> - Direct Codex launches and OpenCode configuration also used narrower
defaults.
> - This change gives all these paths the same full-auto default as
native runners.
> - Explicit restrictive settings continue to work.

## Linked Issues or Issue Description

Refs #13686. This PR is stacked on that native-runner and
task-reassignment PR. Merge #13686 first.

Related: #831 (constructed Claude agents), #1935 (adapter-switching
permission defaults).

## What Changed

- Use actual Claude permission bypass for local and remote runs and
probes. Remove the fixed tool list so MCP and future provider tools are
included.
- Identify actual managed sandbox targets to Claude with `IS_SANDBOX=1`.
Do not mark ordinary host execution as a sandbox.
- Default direct Codex execution to approval and sandbox bypass,
matching agent creation. Preserve explicit false, CLI profiles, sandbox
modes, approval policy, and network restrictions.
- Set OpenCode's full-auto runtime permission to `allow` for every tool
and connection. Preserve the existing explicit opt-out.
- Default Gemini probes to the same YOLO mode as execution. Make the
legacy ACP `default` alias use `approve-all` for fresh and resumed
sessions.
- Add default, opt-out, remote, probe, connected-tool, and resume
regression tests. Update adapter configuration documentation.
- Other adapter paths already request full automatic permission or have
no provider approval gate.

## Verification

- Full workspace `pnpm -r typecheck` and `pnpm build` passed locally
after rebasing onto current master. Targeted adapter/server and legacy
ACP tests passed, including defaults, explicit opt-outs, remote
launches, connected tools, and fresh/resumed sessions.
- Greptile reviewed current head
`8ca135eaffcf9cfdba6f1368e896a781a0891d50` at **5/5**. The security
reviewer acknowledged the documented full-auto requirement. Acknowledged
discussions are resolved.
- Current head has **54 passing checks**. [PR
checks](https://github.com/paperclipai/paperclip/pull/13693/checks). The
process-adapter signoff browser shard passed on one retry after its
first attempt exceeded a three-second issue-run wait.
- **Six native Claude/Codex real-provider cases passed on their first
attempt, with cleanup passing**, against the combined branch: plans,
reassignment, and backlog creation/status. [Campaign and downloadable
evidence](https://github.com/paperclipai/paperclip/actions/runs/35469926548).
This does not claim a real-provider run of every legacy adapter.
- The live-tested revision is
`a37881c824dcd7170380fc4b788732fc743e5da7`. The current head differs
only in the corrected heartbeat test expectation; application code is
identical.
- The campaign result-enforcement job passed. The separate report
publisher failed during frozen dependency installation because the
trusted workflow's patched-dependency configuration does not match its
lockfile. Passing case evidence remains downloadable from the workflow.
- Full-suite coverage comes from CI partitions. The separate unsharded
local run was stopped after the corresponding CI partitions passed; it
is not counted as a completed local run.

## Risks

- Missing permission settings now grant all provider operations,
including connected tools. OpenCode full-auto also overrides ambient
provider permission rules. An explicit Paperclip permission opt-out
preserves restrictive behavior.
- Claude refuses full bypass as root outside an identified sandbox.
Ordinary host deployments must run Claude as a non-root user. Managed
sandbox launches include the required marker.
- These defaults do not grant additional Paperclip roles, connections,
or company access. Existing controller authorization and governance
still apply.
- This PR depends on #13686. Retarget it to master after that PR merges.

## Model Used

OpenAI Codex, based on GPT-6, with code execution and repository tools.
The exact deployment model ID and context-window size are not exposed in
this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-19 16:49:18 -05:00