Recheck retry eligibility under the coordinator lock before creating a cancellation intent. Preserve later run and coordinator outcomes with a failed-only acknowledged-result CAS, while retaining same-intent recovery and NOWAIT conflict handling.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
(cherry picked from commit 29230e2000)
Check failed-run retry eligibility under the run and coordinator locks, fail closed on concurrent coordinator claims, and preserve same-caller recovery after the audited intent disables a retry. Keep terminal failures and foreign actors or intents rejected.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
(cherry picked from commit 79db1c2a6f)
Reserve an optional board request UUID under the run lock and retain it
through default Stop joins and native dispatch. Reject prior or competing
intents and require the same actor for idempotent retries.
Bind the Copilot denial fixture to its exact request and audited intent,
with versioned causal receipts and negative controls for earlier Stop.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
Include the production tsx name helper in generated Node programs so unchanged instruction copies remain warm. Exercise both generated programs through the actual tsx loader while retaining mutation and containment checks.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
Mark missing and invalid child run ordinals as partial observations. Exercise the pinned vendor child construction and reuse methods on all three platforms and regenerate the Cursor-only execution identity.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
Stop immediate collection retries when ownership is deferred or attempts stop advancing. Keep unconfirmed agent-file retirement recoverable through generic cleanup, with later-owner and restart-proof regressions.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
Co-Authored-By: Paperclip <noreply@paperclip.ing>
* commit '4d6e2c00b88af9902facb20ad9177fb70a3a26e0':
Keep the next Cursor profile fixture within the typed contract
Retain stable projectless scope, transfer collection to the exact successor lease, and preserve stopped remote files without restarting their sandbox. Validate complete remote directories and cover the composed DB/session lifecycle, retirement, and fresh preparation.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
Mark lost remote unchanged-turn copies unavailable only after exact termination proof, then perform owned cleanup. Preserve the live unchanged-turn guard and document warm ownership and crash preservation boundaries.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
Probe complete materializations in bounded children, atomically transfer current-run collection authority, and preserve the exact agent home across warm turns. Retire before collecting changed or uncertain copies, fence stale claims and cleanup, and retain unresolved retirement ownership.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
Reuse the production completion-contract envelope hash and exercise real contract creation and reuse in accepted-plan persistence tests.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
Retain the original Cursor6 event filter when verifying an exact committed wait. Progress rows remain part of Cursor7 lifecycle proof without consuming the historical query budget.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
Carry the admitted native plan parent tool into canonical request identity and require its exact successful durable lifecycle before passive settlement. Preserve historical committed Cursor6 waits while versioning new admission to Cursor7.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
Keep current profile qualification mandatory when first creating a wait. Recovery uses its scoped committed receipt to verify the entire original proof, so future catalog/settings changes cannot authorize task work. Cover actual persisted finalization, catalog drift, original-profile tampering, and document explicit user continuation.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
Bind normal native plan acceptance to its durable request, delivered answer, admitted run and completion contract. Commit a passive in-progress result with a visible next-message summary, preserve semantic finish priority, and suppress recovery until task-specific continuation without changing modes.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
Load route modules during fixture setup so cold transforms cannot leave a timed-out request running against the next test mock state. Dismiss delayed product announcements through the normal Board UI in the attachment receipt fixture.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
Keep authorized residue scrubbing and unsafe-path checks while leaving untouched workspaces unchanged when no wake attachments were selected.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
Fence the credential-reclamation hook used by the global recovery sweep so it cannot leave synthetic lease owners on earlier companies. Exercise a paused foreign receipt after the fault is armed and verify it settles without a leaked lease.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
Keep both live and catalog schema assertions through public runner exports while removing the standalone suite dependency on App source.
Paperclip-Task: rich-acp-production
Co-Authored-By: Paperclip <noreply@paperclip.ing>
Complete the dialog fixture context and wait for the actual setup render. Observe the post-spawn PID update before checking the starting service identity, preserving provisioning and readiness assertions.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
Return completed Agent Chat handoffs through a durable outbox and scope each generated update to its supplied tasks. Add recovery, browser delivery, result access, and calibrated quality coverage.
Validated with two consecutive ten-case Claude/Codex campaigns, all CI checks, and a 5/5 review.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - The shared ACP adapter engine records agent failures for operators.
> - ACP providers can report a failure category, title, and detailed
cause.
> - Our patch kept only the category in the saved error, so an operator
could not diagnose a failure when tracing was off.
> - This pull request preserves redacted provider diagnostics in the run
error, transcript, and structured run result.
> - Operators can now inspect the provider message and any supplied
request ID or stack trace after the run ends.
## Linked Issues or Issue Description
Refs #13889 (the diagnostic gap; this PR does not update the bundled
Claude version).
Refs #14484 (related model-refusal classification; this PR retains
diagnostics for all terminal failure categories).
**What happened?**
An ACP turn failed with only `ACP agent reported a terminal service
failure.` The provider's title and details were available in memory but
absent from the saved error and transcript.
**Expected behavior**
The run retains useful provider diagnostics even when raw tracing is
disabled. Credentials remain redacted. A size limit must report
truncation instead of silently removing the cause.
**Steps to reproduce**
1. Run an ACP agent that returns an error-severity typed session
failure.
2. Include an HTTP error, request ID, and stack text in its title and
details.
3. Inspect the failed run with tracing disabled. Before this change,
only the category survives.
## What Changed
- Both pinned ACPX patches pass complete error text to the in-memory
callback, so redaction happens before truncation.
- The shared engine retains the sanitized category, title, and details
in `resultJson.terminalSessionFailure` and includes the text in the run
error and error transcript.
- Diagnostics redact configured environment values even under arbitrary
names, unknown launch-environment values, connection URL passwords, run
credentials, and common credential syntax. Known boolean settings remain
readable, while credential values are redacted even when embedded in
other text. Diagnostics remove control characters and invalid Unicode.
- Title and detail limits keep escaped transcript JSON below the
server's chunk limit. Truncated fields include an omission count. The
safe run-result projection preserves a byte-bounded diagnostic preview
when the result exceeds its byte budget, with an explicit pointer to the
full adapter-bounded run error and transcript.
- The existing UI and CLI display the error. Diagnostics do not become
assistant output. Issue continuation summaries and session-compaction
prompts receive only the generic category, preventing provider text from
becoming handoff instructions. Existing quota classification, warnings,
timeout precedence, and control-channel failure precedence remain in
place.
- Regression tests cover real ACP child processes with both pinned
versions in one-shot and persistent modes, credential redaction, request
IDs after the old 4 KiB cutoff, transcript parsing, storage bounds, and
database retrieval of oversized multibyte diagnostics.
## Verification
- Full CI on `20ad4f5f1f66c46d2c260e6ad0339cbea607b4cf`: **54 passed, 2
intentionally skipped, no pending or failing checks**. Includes
typechecking, build, all Vitest shards, Runner checks, browser E2E, and
the canary packaging/public-install dry run.
- Greptile: **5/5** on this commit. Superagent security scan passes. All
review threads are resolved.
- Local verification passed: shared ACP engine suite (395 tests); real
Claude ACP child-process and diagnostic regressions across both pinned
runtimes and both execution modes; run retrieval and model-handoff
regressions (59 tests); ACPX patch packaging (16 tests); full typecheck
and build. Affected package typechecks and focused tests were rerun
after review fixes.
- The broad local `pnpm test:run` was stopped after review edits made
its cached imports stale. Fresh targeted runs pass, including both
affected server suites. Cold-build import failures were also rerun after
dependency builds: chat integration (1,063 tests) and tool access (351
tests) pass. The final commit's complete CI matrix is green.
## Risks
- Provider diagnostic text is untrusted. This change retains more of it
in company-scoped run records. Redaction and size bounds apply before
persistence.
- Diagnostics are limited to fields the provider supplies. Old runs
cannot recover discarded error text.
- No schema migration, recovery-policy change, or new Telemetry or
OpenTelemetry export.
## Model Used
- OpenAI GPT-6 through Codex, with reasoning, repository inspection,
code editing, and test execution. The exact serving model ID and
context-window size are not exposed in this session.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip manages AI agents that work on assigned tasks.
> - The task composer lets a person choose an assignee, model, and
effort for the next run.
> - A Paperclip Runner agent can use Codex as its provider.
> - The composer hid Codex effort for that agent because it checked only
the older Codex adapter.
> - The native Runner input also did not carry an effort choice to
Codex.
> - This pull request carries the chosen effort from the composer to
each Codex turn.
> - People can now select a supported effort and get the effort they
selected.
## Linked Issues or Issue Description
Refs #14322
**What happened?**
The composer showed a model but no effort slider when the assignee used
Paperclip Runner with the Codex provider. A task-level model override
also did not reach the native Runner input.
**Expected behavior**
The composer shows effort choices for a known Codex model. The next
native Codex turn uses the selected model and effort.
**Steps to reproduce**
1. Open a task composer.
2. Select an agent that uses Paperclip Runner with the Codex provider.
3. Select a known Codex model such as `gpt-6-astra`.
4. Open the assignee and model picker. The effort slider is missing
before this change.
## What Changed
- Show known Codex effort levels for Paperclip Runner Codex assignees.
- Save the task effort override in the native run input and send it to
Codex on each turn.
- Apply the task's merged model and effort overrides when the native run
starts.
- Apply a task model override for OpenCode Runner without changing the
agent's provider.
- Add Runner effort tests and desktop and mobile Storybook cases.
## Verification
- `pnpm -r typecheck` passed.
- `pnpm build` passed.
- `pnpm build-storybook` passed.
- `pnpm check:token-gates` passed.
- Focused UI, server, Runner contract, and Codex driver tests passed.
- The full CI test matrix, build, typecheck, and canary dry run passed
on the latest head.
## Risks
- Native Runner inputs add an optional Codex effort field to the current
v5 input. Older inputs keep their previous behavior.
- A known model rejects an effort that its catalog does not support.
Unknown models do not show a slider.
> This fixes an existing composer bug. I checked `ROADMAP.md`; it does
not describe this bug as planned work.
## Model Used
OpenAI Codex, GPT-6. The exact deployment ID and context window are not
exposed in this session. The model used reasoning, code execution, and
repository tools.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: OpenAI GPT-6 Astra <noreply@openai.com>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - The Mine inbox shows work that needs the current user.
> - Failed-run rows used the latest run for every agent in the company.
> - A failure from another user therefore appeared in Mine and its
badge.
> - Run list responses also omitted the responsible user needed to
filter these rows.
> - This pull request uses run ownership for personal failure routing.
> - Users see their own failures and can still inspect company failures
in All.
## Linked Issues or Issue Description
**What happened?**
An agent run started for one user failed. Its row and failure badge
appeared in another user's Mine inbox.
**Expected behavior**
Mine and its badge include failed runs for the current responsible user.
Other users' failures remain available in All and run details.
**Steps to reproduce**
1. Use a company with two human users.
2. Create a failed or timed-out run attributed to the first user.
3. Open Mine as the second user. Before this fix, the failed run appears
there and increases the badge.
**Paperclip version or commit**
Reproduced in regression tests on master at `24beb0057`.
**Deployment mode**
Authenticated deployment with multiple users. Tests also cover the local
single-user board.
Related prior work: #933 addressed inbox dismissal and badge
consistency. No duplicate ownership fix was found.
## What Changed
- Return `responsibleUserId` in normal and summary run lists.
- Share one ownership rule across both inbox versions and client/server
badges.
- Select the latest run per agent before applying the ownership filter.
This prevents old failures from resurfacing on shared agents.
- Keep unattributed historical failures in the local board's Mine view.
Hide them from authenticated users with no matching owner.
- Keep company health alerts outside the personal badge, consistent with
the client.
- Document the routing contract and add page, badge, and database
regression coverage.
## Verification
- Red: the new badge cases failed with three company failures instead of
one personal failure; eight Mine page cases failed across both inbox
versions.
- Green: 113 focused tests pass in `ui/src/lib/inbox.test.ts`,
`ui/src/pages/Inbox.test.tsx`,
`server/src/__tests__/heartbeat-list.test.ts`, and
`server/src/__tests__/inbox-dismissals.test.ts`.
- `pnpm check:token-gates` passes.
- Agent calls on behalf of a user have two additional red-to-green API
regressions.
- Full `pnpm -r typecheck` and `pnpm build` pass. Server typecheck also
passes after the agent-call fix.
- All CI test shards and browser tests pass on `243bfa681`. The
duplicate local `pnpm test:run` was stopped after the CI test lanes
completed; it did not finish locally.
## Risks
- Authenticated users no longer receive unattributed legacy failures in
Mine. Those failures remain visible in All.
- The server badge no longer counts company health alerts, matching the
existing client badge.
- No migration, run state, retry behavior, or company access rules
change.
## Model Used
- OpenAI GPT-6 through Codex, with reasoning, terminal execution, and
browser tools. The exact deployment variant and context window size are
not exposed in this session.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
Add shared rich ACP transport, durable questions and permissions, verified provider packaging, and bounded activity and plan presentation. Keep Cursor, Copilot, and Pi pending their separate provider qualification.
Persist interaction settlement before publication, fence failed writes until fresh recovery, and preserve owned-process cleanup. Incorporate reviewed mainline integration with extended harness coverage.
Co-Authored-By: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Tasks record an assigned owner and separate checkout and execution
locks.
> - A completed task must retain its owner after execution ends.
> - The release endpoint currently clears that owner when it clears the
locks.
> - This pull request preserves the assignee of Done and Cancelled tasks
during release.
> - Unfinished tasks keep the existing relinquishment behavior.
> - The benefit is stable task attribution without retaining execution
locks.
## Linked Issues or Issue Description
**What happened?**
An agent completed an assigned task, then called the release endpoint.
The task stayed Done, but its assignee became null. The same defect
affects Cancelled tasks. It caused the legacy Claude clarification/reuse
and multiple-repository handoff E2E assertions to fail.
**Expected behavior**
Release must clear execution locks on terminal tasks and preserve their
assignee, final status, and disposition timestamps. Release of
unfinished tasks must still clear the agent assignee. Only In Progress
work returns to Todo.
**Steps to reproduce**
1. Create an assigned task with checkout and execution locks.
2. Complete or cancel the task.
3. Call `POST /api/issues/:id/release` as the assigned agent.
4. Read the saved task. Before this fix, its assignee is null.
**Paperclip version or commit**
Reproduced on master commit `d172197117a14b80a1eb2d2835a0e7cce2679656`.
**Deployment mode**
Local tests against real PostgreSQL through the production issue routes
and services. This is a core lifecycle defect, independent of the agent
adapter.
Refs: #11689, #6899, #7769. These are related open release proposals.
This is an independent fix limited to terminal task ownership. It does
not include timer scheduling changes.
## What Changed
- Preserve the current assignee when releasing Done or Cancelled tasks.
- Keep all execution-lock cleanup and existing unfinished-task behavior.
- Cover all seven task statuses through the release API and read back
saved state.
- Check disposition timestamps, activity attribution, and repeated board
cleanup.
- Update the API contract, agent reference, and CLI help.
## Verification
- Red commit `ddaabb754`: the two terminal-owner regressions failed with
`assigneeAgentId: null`; 12 other route tests passed.
- Green: all 14 route tests pass, plus the existing successor-checkout
race test (15 selected tests total).
- Command: `pnpm --filter @paperclipai/server exec vitest run
src/__tests__/issue-stale-execution-lock-routes.test.ts
src/__tests__/issues-service.test.ts -t 'stale issue execution lock
routes|does not let stale release clobber a successor checkout lock'`.
- The local host has exhausted its SysV semaphore pool. The red/green
runs used the existing test-provider hook to start disposable Docker
PostgreSQL 17 instances. Routes, services, migrations, and assertions
were unchanged. No database tests in the selected set were skipped. The
other 134 tests were excluded by the name filter.
- Capability contract and inventory drift checks pass.
- `pnpm build` and `pnpm -r typecheck` pass.
- The local `pnpm test:run` was interrupted after environment failures
while the complete sharded CI suite ran in parallel: native PostgreSQL
bootstrap fails under the host semaphore limit, and the large Git
fixture hits macOS `ENAMETOOLONG`. A focused rerun confirmed these
happen before the relevant assertions. The interrupted local run is not
counted as a full pass.
- Greptile completed on `1caeeb827e9cb658ddb71f16c2421ec20f80634e` with
**5/5**, a successful check, and no review threads.
- All CI gates pass on the current head: typecheck, build, general and
serialized tests, Runner checks, browser E2E, release packaging, and
security checks. Server shard 11 passed on one targeted retry; the first
attempt had 836 passing tests but an unhandled workspace-runtime startup
rejection caused by an existing timing window. All other successful jobs
were reused.
- No paid provider evaluations were run.
## Risks
- A caller that used release to erase ownership from terminal work will
now retain that owner. An explicit assignment update or the board
force-release option with `clearAssignee=true` can still clear it.
- No schema or migration changes. The transaction, company access,
assignee/run checks, and activity log remain in place.
## Model Used
- OpenAI GPT-6 through Codex, with tool use, code execution, and test
debugging. The session does not expose a more specific backend model
version or context-window size.
- OpenAI `gpt-6-luna` assisted with read-only test discovery and review.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - An agent needs personal files across tasks and sessions.
> - AGENTS.md is one file in that directory. Supporting files need the
same persistence.
> - The Instructions Editor and agent runs must share one current
directory.
> - Concurrent runs should apply only the files they change. The last
sync of the same file wins.
> - This pull request uses existing file transport and removes temporary
copies after sync.
> - Old instruction-only sessions keep their restore contract. New saves
do not create revision history.
## Linked Issues or Issue Description
Refs #14325. This replaces its revision-oriented design with persistent
agent files. Keep #14325 unmerged.
Transport prerequisite #14416 merged first at
`d172197117a14b80a1eb2d2835a0e7cce2679656`. This PR now targets master
and remains below 100 changed files.
Related work: #4513 and #8798 cover instruction tooling. This change
handles run synchronization, cross-task personal files, browser editing,
and old-session restoration.
## What Changed
- Keep one current directory per company and agent. Point AGENT_HOME at
a temporary working copy for each active run. Keep task files and
provider HOME separate.
- Restore text, binary files, and nested folders through workspace
transport. Exclude remote agent files from task Git snapshots with a
self-ignoring file inside the reserved runtime directory; never write
through repository-controlled Git metadata.
- Collect after the provider and child processes have stopped. Keep
resumable conversation state.
- Apply changed and deleted files under the agent lock. The last sync
wins for the same file. Unrelated concurrent changes survive.
- Remove temporary copies after successful sync, rejected sync, and
staging failure. Register ownership before copying so restart recovery
can remove interrupted preparation. Retry transient synchronization up
to three times. Preserve the original remote lease reference until
deletion succeeds; restart cleanup never acquires a replacement sandbox.
Do not create captured directories or a conflict-review queue for new
runs.
- Keep browser editing, stale-draft protection, and streaming binary
downloads. Keep the instruction entry and text editor limited to 1 MiB.
- Keep historical agent-folder sync failures on their affected runs
instead of repeating them above current saved instructions. Preserve
legacy candidate review and current browser-save errors. Avoid duplicate
quota warnings while retaining separate sync failures when they describe
a different problem.
- Require target-scoped caller grants for peer instruction access, while
preserving self edits, responsible-user checks, and protected-change
consent.
- Treat full storage as a nonblocking run warning, never an agent pause
or run-admission failure. Restore already-over-quota saved folders so
ordinary agent cleanup can recover; warn on each run until cleanup. The
run detail view shows the warning.
- Allow 256 MiB per file, 2 GiB per directory, and 100,000 entries. Hash
large files as streams. Check editor-save quotas with metadata instead
of hashing unrelated files.
- Preserve old native inputs, instruction-only copies, paths, digests,
and pending legacy candidates. Adopt old revision heads once. New writes
do not append history rows.
- Add idempotent migration 0287 and verify upgrades from the preview
tables and receipts.
- Add nine interactive stories under **Agents / Persistent files**,
including automatic incoming edits, stale browser drafts, and
storage-limit diagnostics.
## Verification
- Merge candidate: `4f5390107ec6ffd80a76d1d2e85530e66f21d079`, after
merging current master and the landed transport prerequisite.
Integration required no manual conflict resolution; the feature remains
99 changed files. Full workspace typecheck, production build, token
gates, and 715 focused tests passed on this merge candidate. Fresh
Greptile review is 5/5 with no unresolved findings. All 55 checks
passed, with four conditional skips, including the build, typecheck,
browser E2E, and canary dry run. A single retry recovered four jobs
interrupted by runner shutdowns; no source changes were required.
- Historical-warning UI fix: all 6,834 UI tests across 640 files passed,
including regression coverage for three old failures, legacy preserved
edits, and warnings scoped to the affected run. Full workspace
typecheck, production build, Storybook build, and token gates passed.
Browser-verified Storybook playtests passed for Historical Failures
After Successful Save, Storage Limit, and Full Storage Run Warning.
- Review follow-ups at `4e20c9fb2`: all 18 focused tests passed,
including external Git directories, linked worktrees, symlinks,
hardlinks, and distinct I/O failures alongside storage warnings. Server
and UI typechecks, token gates, and the production build passed.
- Storage warning regressions at `0724f3012`: all 33 directory tests and
all five heartbeat-list tests passed, with no skips in their successful
runs. They cover repeated runs while full, an already-over-quota saved
folder, cleanup, warnings retained after unrelated save failures, and
bounded warnings in large result JSON. Server typecheck passed after the
final warning fixes.
- Full workspace typecheck, production build, and token gates passed
during this follow-up. Product E2E harness: 631 tests passed across 52
files; harness typecheck passed. Earlier native session/context and
directory/legacy collection suites passed 537 tests; Runner
unit/transport suites passed 329 tests.
- **Real E2E at `0724f3012` (before this follow-up):** legacy local
Codex and native Daytona Codex each passed six tasks, one server
restart, seven independent assertions, and cleanup verification. Both
prove browser-to-agent edits, agent-to-browser edits, nested/binary
restoration, per-file last-sync-wins, a successful run after an
oversized save rejection, and cleanup clearing the warning.
- Native local Codex also passed the six-task quota flow before the
final warning-retention fixes. That pass began at `918d1ed02` while the
bounded-result warning fix was being edited, so it is not claimed as
exact-final-head evidence. Its final-head rerun failed during embedded
PostgreSQL bootstrap before any provider run: the macOS host had 87,365
of 87,381 SysV semaphores occupied. No unrelated services or kernel
limits were changed.
- The final-source report intentionally records **2/3 cells passed**,
preserving the blocked native-local attempt:
`tests/runner-e2e/results/agent-files-quota-final-20260928-report/`.
Earlier failed attempts and provenance notes remain under
`tests/runner-e2e/results/agent-files-quota-final-20260928-input/` and
the original campaign directories.
- Daytona used immutable image
`ghcr.io/paperclipai/paperclip-daytona-runner@sha256:5643f0d801417cae3581833a1a3bc6715b325e028602738d2652c44cac5dc6bf`
and its exact Linux runner binary. Controller source is `0724f3012`;
image source is recorded separately.
- Legacy-session compatibility and all three ACP Stop/resume browser
regressions passed on the prior validated feature head
`169fab46d5af21caa2269b4c1b29b69c933a6951`. They assert the same
provider session is retained and interrupted writes are not replayed.
Migration upgrade tests also passed earlier.
- Nine interactive stories are under **Agents / Persistent files**,
including **Full Storage Run Warning**. Its playtest and visual browser
inspection passed; the warning states that runs continue and the editor
remains available.
- Prior-head checks on `4e20c9fb2`: 55 passed, two conditional jobs
skipped, no failures or pending checks. All eight browser E2E shards and
their aggregate passed. Fresh Greptile review is 5/5 with no findings;
all review threads are resolved, the security scan passed, and GitHub
reports no merge conflicts.
- The broad local follow-up test run was interrupted after host
semaphore exhaustion affected isolated PostgreSQL instances. It also
encountered the existing macOS long-path fixture failure and two timeout
failures. This is not a claim that the full local suite passed. Logs are
retained; focused storage/warning tests passed.
## Risks
- A later sync can overwrite an earlier edit to the same file, including
a saved browser edit. There is no text merge or retained version. This
is the intended last-sync-wins policy.
- A save that exceeds a storage limit is rejected and its temporary copy
is discarded. The run itself continues normally, and later runs restore
the last saved files with a warning until cleanup. Transient sync
failures get bounded retries. An I/O failure partway through a sync can
leave some files updated; a failed receipt does not claim whole-folder
success.
- Larger folders increase copy time, network traffic, and temporary disk
usage. Active runs still need working copies. Terminal runs do not
accumulate archives. Operators must provision disk for agents and
configured concurrency; these limits are not company-wide quotas.
- A restored old native session remains instruction-only until a fresh
session starts. Its original conflict fence and existing pending
candidates remain compatible.
- Provider processes close at the collection boundary. Conversation
resume remains available, but warm process reuse is lost.
- Backups must include the instance filesystem and database. External
bundles keep their existing behavior until explicitly moved to managed
storage.
## Model Used
OpenAI Codex, GPT-6 family. The session does not expose a more specific
model ID or context-window size. Reasoning, code execution, and browser
tools assisted this change. Real provider E2E uses `gpt-5.6-sol`.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Fry (Paperclip) <noreply@paperclip.ing>
## Thinking Path
> - Paperclip coordinates agent work through execution adapters.
> - Sandbox ACP sessions send ordered input through a remote file queue.
> - A temporary provider 502 currently closes the session during an
input upload.
> - A lost response can occur after the sandbox has consumed the
message, so a blind retry can duplicate input.
> - This pull request retries gateway failures with the same sequence
and drops consumed sequences at the receiver.
> - The session can continue through a brief provider failure without
repeating a tool call.
## Linked Issues or Issue Description
**What happened?**
A sandbox ACP run can fail with `ACP agent disconnected during request
(connection_close, exit=null, signal=null)` when a provider input upload
returns HTTP 502. The bridge destroys its local socket on the first
failure and can discard the diagnostic before the proxy reads it.
**Expected behavior**
A temporary gateway failure should get a bounded retry. A lost response
after successful delivery must not duplicate input or reorder later
messages. Permanent failures must still close the session.
**Steps to reproduce**
1. Run the real sandbox process bridge with an echo child and a local
test runner.
2. Inject a provider 502 before preparation, after a chunk upload, or
after final publication and consumption.
3. Send the next input message. Before this change, the connection
closes instead of delivering it.
Searched open and closed PRs for `ACP disconnect`, `bridge retry`, and
`502 sandbox`. Related work: #13287 covers shutdown after bridge loss;
#13793 covers large launch envelopes. This change covers ordered input
delivery within a running legacy ACP session.
## What Changed
- Retry input uploads up to three times for recognized Daytona and
Cloudflare HTTP 502, 503, and 504 diagnostics, with 250 ms and 500 ms
delays.
- Give each upload separate temporary paths and discard already-consumed
input sequences, including late publication from an earlier attempt.
Clean failed attempts in the background without removing a published
message or another attempt’s files. Cleanup cannot delay retries or
shutdown.
- Keep later input behind the retry. Stop queued input on permanent
failure and flush a fixed diagnostic before closing the socket. Neither
failure-diagnostic persistence nor shutdown-warning persistence can
block teardown.
- Add real-process regression tests for lost responses, late
publication, retry exhaustion, immediate permanent failure, and
diagnostic redaction.
- Give accepted run-log file appends up to three seconds to drain before
finalization computes the size, hash, and durable copy. Close the run
handle to later appends. This waits only for file writes, independently
of later DB progress or live-event persistence. If writes remain
stalled, return null size/hash metadata and skip the final durable copy
so the run can settle. Late writes cannot restart mirroring.
- Preserve legacy comment attribution when final log size is unknown by
reading existing entries within the unchanged 2 MB scan limit. Storage
errors or a three-second read deadline return the evidence already read
instead of failing the comment listing; pagination stops at the
deadline. The deadline requests cancellation of the underlying local
stream or S3 HEAD, GET, and response stream. A separate response timeout
returns partial evidence even when filesystem I/O delays cancellation;
late reads cannot append evidence or start another page. Each listing
retains its existing batches of eight reads, without a shared admission
cap that skips readable logs under contention.
- Document the retry and log-finalization boundaries in the development
guide.
## Verification
- Final commit `347daa564b`: [Linux
CI](https://github.com/paperclipai/paperclip/actions/runs/36506995168/attempts/2)
passed. Greptile Apex review 13 scored this commit 5/5 with no new
findings; all 12 review threads are resolved.
- The final CI run initially hit a Cursor test timeout and four Discord
credential-lock contention failures. All five cases passed in isolation.
The two failed shards and their aggregate gate passed on retry without a
code change. Those intermittent failures are not claimed fixed by this
PR.
- `pnpm --filter @paperclipai/adapter-utils typecheck` passed.
- `pnpm exec vitest run
packages/adapter-utils/src/execution-target-stdin-race.test.ts
packages/adapter-utils/src/execution-target-sandbox.test.ts
packages/adapter-utils/src/sandbox-callback-bridge.test.ts`: 262 tests
passed on the final implementation, including 21 new regressions. The
original three fault-injection cases failed before the fix.
- The regressions cover failed and indefinitely stalled cleanup,
Cloudflare gateway responses and retry exhaustion, permanent errors that
must not retry, and teardown while failure logging remains indefinitely
stalled. Seven Apex regression cases failed before the review fixes.
Adapter-utils typecheck and build passed again after the final review
change.
- `pnpm exec vitest run server/src/services/run-log-store.test.ts
server/src/services/run-log-store-cancellation.test.ts`: all 25 tests
passed, including four new regressions that failed before the
finalization fix. They cover delayed and failed appends, late-write
admission, agreement between the local bytes/summary/durable copy, and a
stalled append that exhausts the three-second budget. The timeout case
verifies unknown metadata, no final upload, and no mirror restart after
late completion. New cancellation tests use the real AWS SDK against a
local HTTP server. They verify that stalled HEAD, GET, and response-body
connections close on abort and that a subsequent read succeeds. Local
range and already-aborted read cases also pass.
- `pnpm exec vitest run server/src/__tests__/issues-service.test.ts -t
'readIssueCommentRunLogText|deriveIssueCommentRunLogAttribution'`: 14
targeted tests passed. The null-size reader case, both storage-error
cases, the stalled-read case, and the cancellation/concurrent-listing
cases failed before their fixes. The new regressions verify that
timed-out reads are cancelled, subsequent listings recover, and two
concurrent listings both retain their attribution markers. A read that
ignores cancellation still returns partial evidence at three seconds and
cannot resume pagination when it finishes; this regression failed before
the response-timeout fix.
- `pnpm --filter @paperclipai/server typecheck` and `pnpm --filter
@paperclipai/server build` passed after the response-timeout change.
- Full `pnpm -r typecheck` and `pnpm build` passed earlier in this PR;
the affected packages were rechecked after review fixes.
- Full local `pnpm test:run` failed in the general-server group: 511
files passed, 40 failed, and 158 were skipped. Failures include embedded
PostgreSQL initialization, read-only cache directory renames, a macOS
long-path fixture, and a workspace exposure assertion. The PostgreSQL,
cache-permission, and long-path failures also reproduce with both
changed implementation files restored to baseline commit `24c58e479a`.
The exposure suite passes in isolation both on baseline and the fixed
branch (28 passed, 3 skipped). CI runs the full suite on Linux. Later
local test groups were not reached.
- An earlier CI run hit the Telegram retry-timing failure fixed upstream
in #14501. The branch includes that master fix. The selected recovery
test passed against a fresh, migrated PostgreSQL 16 database. The
embedded PostgreSQL runner is unavailable on this Mac; the isolated
database was stopped and removed afterward.
- No live agent turn was replayed. The tests use local child processes
and injected provider failures.
## Risks
Retries are restricted to recognized Daytona SDK and Cloudflare bridge
gateway-error messages, which survive plugin RPC serialization. Other
errors fail immediately. Temporary upload paths are now unique for all
command-managed queue writes. Receiver sequence checks prevent duplicate
input; retries do not restart an agent turn. Cleanup and failure logging
are nonblocking and best effort; session teardown remains the final
cleanup boundary. Log finalization now drains accepted local file writes
for at most three seconds and ignores later appends on the closed run
handle. A timeout leaves final size/hash unknown and skips the final
durable upload; an existing partial mirror may remain available, but it
is not claimed as a verified final snapshot. It does not wait for later
DB progress or live-event persistence. Optional attribution keeps
partial evidence when a read fails or times out. Cancellation closes S3
requests and response streams. Local filesystem I/O may finish after the
caller deadline, but a late read cannot change the returned evidence or
continue pagination. Later listings can retry after storage recovers.
There is no schema, authentication, or permission change. Revert this
commit to restore the previous behavior.
## Model Used
OpenAI GPT-6 through Codex, with reasoning, repository inspection, code
editing, and local test execution.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally; targeted tests pass and full-suite
limitations are documented above
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip manages AI agents and their work.
> - Project workspace policies define how isolated worktrees are set up.
> - Tasks can override a branch without providing every setup field.
> - The resolver currently replaces the entire project strategy with
that partial override.
> - Losing an explicit setup command can run the repository fallback
script and block the task.
> - This pull request keeps enabled project defaults when the task uses
the same strategy type.
## Linked Issues or Issue Description
**What happened?**
A project uses `git_worktree` with `provisionCommand: "true"`. A task
overrides only `baseRef`. The resolver drops the command. Worktree
creation then invokes `scripts/provision-worktree.sh`, which can fail
because its required setup is absent.
**Expected behavior**
A branch override keeps the project's provision, runtime provision, and
teardown commands unless the task explicitly overrides them. A different
strategy type must not inherit those commands.
**Steps to reproduce**
Configure the project with an enabled `git_worktree` strategy and
`provisionCommand: "true"`. Give the task an isolated workspace with a
`git_worktree` strategy and a different `baseRef`. Add a failing
repository fallback provisioner. Before this change, worktree creation
invokes that script. After this change, it uses the project's explicit
command and succeeds.
Related: #4968 concerns agent strategy and working-directory fallback.
#13903 concerns gated API fields and reusable-workspace updates. #11091
concerns provision hooks on workspace reuse. None fixes partial task
overrides discarding project defaults.
## What Changed
- Merge a partial task strategy over the enabled project's strategy only
when their types match.
- Preserve explicit null values when parsing nullable strategy fields,
so they can clear project values.
- Keep explicit empty-string overrides and agent fallback behavior.
- Exclude disabled project strategies and avoid an inherited branch
template when a task pins an existing branch.
- Add policy regression coverage and a real Git worktree test with a
failing fallback script.
- Document inheritance, explicit clearing, and no-op provisioning in the
development guide.
## Verification
- Policy regression: eight failures before the fix; all 41 policy tests
pass after it.
- Real worktree regression: passes and creates a worktree using the
task's base branch without invoking the failing fallback provisioner.
- `pnpm -r typecheck`: passed.
- `pnpm build`: passed.
- `pnpm test:run`: general-server phase completed with 13,873 passed, 86
skipped, and 14 failures in unchanged macOS skills-cache and Git
long-path tests. The same failures reproduce on unmodified base code.
The command stops at that phase, so no full local pass is claimed.
- CI initially failed the existing Telegram subscription recovery test
on a 15-second timeout. The separate fix and investigation are in
#14501. A serialized job also lost its runner; GitHub reported lost
communication, and that job was rerun without source changes. All 52
final-commit checks pass, with two intentional skips. Greptile is 5/5,
with no unresolved comments or merge conflicts. The chat shard passed on
one unchanged rerun. The timeout cause remains unproven; #14501 adds
phase diagnostics for a recurrence.
## Risks
Tasks that specify a partial strategy now retain the project's omitted
fields, including setup and teardown hooks. This is the intended
behavior change. Inheritance requires an enabled project policy and
matching strategy types. Explicit task values still win. Null and empty
commands restore existing runtime defaults; they do not guarantee that
no script runs. Use `"true"` for an explicit no-op provision command. No
migration, live configuration change, or task replay is included.
## Model Used
OpenAI GPT-6 (Codex), with reasoning, terminal tools, and code
execution. The context window size is not exposed in this session.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes:` / `Closes`
/ `Refs` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub references)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run focused tests locally and they pass; full-suite status
is recorded above
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
Co-authored-by: Paperclip <noreply@paperclip.ing>
## Thinking Path
> - Paperclip manages AI agents and their work.
> - Telegram delivery recovers subscription changes after a restart.
> - The recovery test leaves a failed action on the real one-second
retry timer.
> - A slow runner can cross that deadline before the test checks that no
retry occurred.
> - This pull request holds the fixture deadline until the explicit
restart transition.
> - The test still checks that recovery uses fresh provider options.
## Linked Issues or Issue Description
**What happened?**
The Telegram subscription recovery test expected one `setWebhook`
request but saw two. A 1.5-second delay after the first failed request
reproduces the failure.
**Expected behavior**
The test controls when the failed request becomes eligible for retry.
Host speed does not change its result.
**Steps to reproduce**
Run the test named `retries an unknown subscription mutation after
restart` with a 1.5-second delay after the first failed-action
assertion. The old fixture retries too early. The updated fixture passes
with the same delay.
Related: #13952 fixes a separate Telegram fixture cleanup problem. This
change addresses retry timing.
## What Changed
- Set the stored retry deadline to 2099 before the pre-restart
assertions.
- Keep the existing explicit epoch deadline after restart and all
provider request assertions.
- Report the current test phase only when this test fails, to diagnose
an observed intermittent CI timeout.
- Leave production retry code unchanged. Temporary delay and per-step
console tracing are not included.
## Verification
- Delayed regression: failed before the change with two requests instead
of one; passed after the change.
- Focused Telegram durable private draft Stop group: 34 tests passed.
- `pnpm -r typecheck`: passed.
- `pnpm build`: passed.
- Full local chat shard: 355 tests passed twice.
- `pnpm test:run`: general-server phase completed with 13,861 passed, 86
skipped, and 14 failures in unchanged macOS skills-cache and Git
long-path tests. The same failures reproduce on unmodified base code.
The command stops at that phase, so no full local pass is claimed.
- CI exposed a separate 15-second timeout. A diagnostic run passed all
355 shard tests, with the affected test completing in under one second.
Its cause remains unproven. Normal step logging is removed; a
failure-only phase report remains for a recurrence. The final commit
also passes the 355-test chat shard, and Greptile rates it 5/5. All 52
final-commit checks pass, with two intentional skips. There are no
unresolved review comments or merge conflicts.
## Risks
Low risk. This changes only the fixture deadline. It does not disable a
test, extend a timeout, or change production retry behavior. Existing
assertions still verify the failed action, the pending state, restart
recovery, and fresh provider options. The intermittent CI timeout is not
claimed fixed; phase diagnostics narrow the next occurrence without
changing the timeout. No documentation change is needed for a test
fixture correction.
## Model Used
OpenAI GPT-6 (Codex), with reasoning, terminal tools, and code
execution. The context window size is not exposed in this session.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes:` / `Closes`
/ `Refs` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub references)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run focused tests locally and they pass; full-suite status
is recorded above
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes, or
explained why none is needed
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
Retain bounded reasons for runner identity, harness recovery, and provider-pack read failures. Preserve existing ownership and cleanup proofs and compatibility with receipt-gated chat recovery.
Verified with executor, recovery, diagnostic privacy, typecheck, build, and full CI checks.
Co-Authored-By: Paperclip <noreply@paperclip.ing>