Commit Graph
27 Commits
Author SHA1 Message Date
DottaandPaperclip bf9dbd18a8 fix(runner): recover incompatible Codex models before launch (#15519)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The Runner prepares and saves each task's provider configuration
before launch.
> - A sandbox image can have a supported Codex CLI that is too old for
the selected model.
> - The current version check stops that task even when the image can
run a similar model.
> - This pull request selects a compatible model before it saves a fresh
execution.
> - The task continues with a visible warning, and recovery uses the
saved effective model.

## Linked Issues or Issue Description

- Refs #15053. This fixes the same CLI and model mismatch for fresh
remote Runner tasks. It does not change subscription onboarding.
- Refs #14721. This adds a bounded startup choice for Codex CLI
compatibility. It does not add user-defined fallback chains or quota
failover.
- Related PR: #15399 added the model-specific CLI checks that this
change preserves at launch.

## What Changed

- Check the preinstalled remote Codex CLI before saving a fresh
execution that uses a model with a verified CLI minimum.
- Select the closest compatible older model in the same class, then the
stable Runner default. Consider each candidate once.
- Save the effective model before checkpoint selection. Keep the
requested agent and task settings unchanged.
- Add a task warning and a local `runner.model_fallback` run-log event
with both models and the CLI version.
- Share executable discovery with the launch verifier. Preserve explicit
artifact and install settings, saved executions, and existing artifact
checks.
- Add regression tests and document the selection and warning behavior.

## Verification

- Red: the three new provider-configuration regression cases failed
before the fix. They kept the incompatible requested model.
- Red: both rejected-probe cases failed before the review fix. They now
defer to launch verification.
- Green: targeted provider configuration, remote preflight, task
warning, and launch-verifier tests pass (680 tests).
- `pnpm -r typecheck` passes.
- [Full
CI](https://github.com/paperclipai/paperclip/actions/runs/37713759744)
passes on `eb273fc32661526d90a864b42c73a7da61dd62da`: 47 successful
jobs, including all general and serialized Vitest groups, browser
shards, Runner verification, typecheck, build, and the canary dry run.
- Local extended verification: three serialized shards pass (115
suites). HTTP route tests hit intermittent 15-second timeouts. The
authorization suite passes on rerun (130 tests). The document suite
passes on both master and the PR head in the same isolated setup (6
tests). The local general run was stopped after the equivalent full CI
matrix passed.
- `pnpm build` passes.
- Server typecheck and build pass again after the probe-error review
fix.
- `git diff --check` and `pnpm check:module-boundaries` pass.
- Example: request `gpt-6.1-sol` on Codex 0.158.0 to use `gpt-6-sol`; on
0.156.0, use `gpt-5.6-sol`. The warning names the requested model,
effective model, and CLI version.
- This PR has not been deployed to staging.

## Risks

- A fallback can have different capabilities. The task warning makes the
substitution visible. Later runs can use the requested model after the
image CLI is updated.
- Preparation adds two remote commands for models with a verified CLI
minimum. Failed or invalid version probes retain the existing launch
checks.
- This only handles known CLI and model mismatches before a fresh
provider launch. Authentication, capacity, and artifact failures keep
their existing behavior.
- No database migration is required.

## Model Used

OpenAI GPT-6 through Codex, with tool use and code execution. This
session does not expose the exact deployment model ID, context window
size, or reasoning setting.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-08 04:57:26 -05:00
Devin FoleyandPaperclip 465140596f Preserve bounded workspace sync diagnostics across RPC (#15481)
Preserve bounded error codes and HTTP/exit statuses across the environmentSyncOut worker RPC boundary, and revalidate that method-scoped envelope before attaching host restore diagnostics. Keep the original error and all recovery policy unchanged; do not transmit provider payloads or credentials.

Verified real RPC roundtrip and privacy regressions, 95 focused tests, 68 independent tests, full typecheck/build and all exact-head CI. Greptile 5/5 with no unresolved comments.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-07 12:22:53 -07:00
Devin FoleyandPaperclip 06484b3c41 fix: preserve conversation retries through execution cleanup (#15463)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Recovery schedules bounded retries after a provider disconnects.
> - A stopped run can still hold its environment lease while cleanup
runs.
> - Retrying before that lease is released cancels the new run before it
starts and spends another retry.
> - Restoring the task can also send the worker repair instructions from
an already resolved recovery action.
> - This pull request preserves the waiting retry and removes settled
recovery instructions from later wakes.
> - The task can continue after cleanup without an operator repairing
the same incident again.

## Linked Issues or Issue Description

**What happened?**

A legacy conversation run disconnected while it was doing ordinary work.
Its environment cleanup took longer than the retry delay. Two retries
were cancelled before dispatch with `execution_reconciliation_required`.
Those cancellations exhausted the failure budget. After an operator
restored the task, the wake still told the original worker to repair the
runtime and hand the task back to itself.

**Expected behavior**

Cleanup waits preserve the pending attempt. The same retry can continue
after ownership is released, subject to all current gates. Once a
recovery action is resolved or cancelled, subsequent task wakes omit its
repair instructions.

**Steps to reproduce**

1. Fail a legacy conversation run while its environment lease remains in
`pending_cleanup`.
2. Schedule a bounded retry and run promotion before cleanup releases
that lease.
3. Repeat the scheduler sweep. Before this fix, retries promote and then
cancel without starting.
4. Resolve a stranded-task recovery action and build the restored task
wake with that action ID. Before this fix, the wake still includes the
settled repair instructions.

**Paperclip version or commit**

Reproduced with database regressions against `ceabc3bc880` on master.

Related: #15019 restores a skipped assignment handoff after lease
release. #15235 filters stale handoff evidence in the recovery sweep.
This change preserves an existing scheduled conversation retry and
corrects restored wake content. It does not create a new handoff wake.

## What Changed

- Keep an unstarted legacy conversation retry on the same durable row
while prior execution ownership remains active. Recheck after 30 seconds
without increasing retry accounting.
- Return a queued retry to scheduled state if it encounters that hold at
the claim gate. Retain its issue claim and publish the status change.
- Record one local lifecycle diagnostic per blocking run. Remove that
wait marker on promotion.
- Include recovery action metadata only while the referenced action is
active or escalated.
- Return an explicit `waiting` response and the saved schedule when
Retry now meets cleanup. Show the wait inline without a false success or
disabled button.
- Add database, rendered-prompt, route, and UI regressions. Document the
execution and run-log contracts.

## Verification

- Five cleanup and restored-wake regressions fail against the original
production code. The Retry now route and UI regressions also fail before
their correction.
- Related retry, dispatch, stale-queue, and recovery suites: 321 tests
pass across seven files. All ten focused cleanup/restored-wake cases
pass after rebase. The final dispatch adjustment passes all 46 adapter
tests.
- Retry now routes and affected UI suites: all 46 tests pass. The tests
cover repeated clicks, the saved schedule, unchanged accounting,
promotion after release, and no false success or error state.
- `pnpm -r typecheck`, `pnpm build`, `pnpm build-storybook`, and `pnpm
check:token-gates` pass.
- Full local `pnpm test:run` was started and then stopped after the
final commit passed all GitHub CI test shards. No complete local
full-suite result is claimed; CI supplies the complete test result for
the final commit.
- Final head `6f1058a332c039e33c4f002b296d20a5554e760e`: all 55 GitHub
checks are green or intentionally skipped. Apex review is 5/5 after two
reviews, with no unresolved threads. The PR has no merge conflicts.

## Risks

The wait applies only to unstarted legacy conversation retries. Native
runs and non-conversation execution keep their existing recovery rules.
Cleanup must actually release ownership before execution can resume. The
wait does not fix a cleanup service that never finishes. Promotion and
dispatch still enforce cancellation, reassignment, pause, budget, and
reconciliation gates. No schema change is required. The Retry now
response adds a `waiting` outcome; the shared contract and all three UI
controls handle it.

## Model Used

OpenAI Codex based on GPT-6, with repository analysis, tool use, and
local code execution. The runtime does not expose the precise serving
model ID or context window.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused suites; complete
final-head test coverage in CI)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-07 11:15:49 -07:00
DottaandPaperclip 63f3aa2dbf fix(runner): continue restart-interrupted Codex turns (#15297)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Native runs retain their provider conversation across server
restarts.
> - A dead runner can restore a Codex conversation after its active turn
is lost.
> - The old recovery path synthesized a failed task result from that
interruption.
> - The task then required operator action even though its conversation
and workspace were available.
> - This pull request preserves the interruption cause and uses the
admitted restart attempt for one continuation in the same conversation.
> - The agent can reconcile unfinished actions and complete the current
request without resending the original task.

## Linked Issues or Issue Description

Related: #12845 added native restart recovery. #15042 covers admission
during shutdown. #14796 covers legacy shutdown recovery. This change
covers a lost native Codex turn after successful conversation
restoration.

**What happened?**

After a server restart killed the local runner, Paperclip restored the
saved Codex thread. Runnerd found that the old turn was no longer
active. It synthesized a failed terminal and a needs-review result from
the last progress message. The transport also discarded the terminal
error when it reconstructed thread history. The task failed instead of
continuing.

**Expected behavior**

After proving that the old process stopped and admitting a bounded
recovery attempt, resume the current request in the same conversation.
Preserve the workspace. Inspect unfinished actions before proceeding.
Keep real provider failures, accepted results, intentional stops,
unknown unreconciled effects, and exhausted attempts subject to their
existing rules.

**Steps to reproduce**

1. Start a local native Codex run and leave its turn active.
2. Kill the isolated runner and provider processes, as can happen during
a server restart.
3. Restore the same provider thread with no active turn.
4. Observe the synthetic task failure. The new real-process regression
reproduces this boundary with a scripted provider.

## What Changed

- Record an explicit recoverable process-loss cause without inventing a
task result.
- Preserve terminal errors and prior turns in reconstructed provider
history. Recover the authoritative saved result when adopting an
accepted continuation.
- Send one continuation in the same conversation for an admitted
dead-runner recovery. Require reconciliation of unfinished commands and
external actions.
- Persist the interrupted terminal before submission and retain the
existing recovery marker across another controller loss.
- Keep provider attempt limits, terminal failures, and intentional
cancellation behavior.
- Add red/green regressions, real process-kill coverage, restart
checkpoint coverage, and retry-budget coverage. Document the behavior
and run-log evidence.

## Verification

- Red: the new native runtime regression rejected with
`NativeProviderTerminalFailure` on the original code; the Rust restore
regression found a missing recovery cause.
- Red/green: if restoring the conversation fails and replacement is
allowed, the replacement receives the full task and fresh-session
handoff. Both prepared and legacy execution inputs are covered.
- Red: a second controller crash after the provider accepted the
continuation caused an extra `turn/start`. The regression now proves
there are exactly two submissions total: the original and its
continuation.
- Green: focused runtime, backend, driver recovery, and real-process
restart suites (207 tests). After the final history/result changes,
driver recovery and real-process restart suites passed again (36 tests).
- Green: complete Codex transport suite (186 tests), server restart
classification/database integration suites (34 tests), and Rust Codex
provider suite (92 passed, 2 ignored).
- Full `pnpm -r typecheck` and `pnpm build` passed on `60141e649`. The
subsequent replacement-prompt guard passed the Runner TypeScript check
and the complete runtime plus process-restart suites (148 tests).
- Local full-suite attempt: `pnpm test:run` reported two failures in the
untouched chat integration suite. Both passed individually, and the
complete chat suite passed on rerun (1,063 tests). After all remote test
shards passed, the duplicate serial local run was stopped with SIGINT;
it is not claimed as a full local-suite pass.
- Latest-head CI (`b1297dcd4`): 55 successful checks and 4 intentionally
skipped checks, including all test shards, typecheck, build, native
Runner verification, end-to-end tests, and canary dry run. [CI
run](https://github.com/paperclipai/paperclip/actions/runs/37405346303).
- Greptile: 5/5 on the latest head, with no unresolved review threads.
The PR is mergeable.
- The process tests use the real runner binary and a scripted Codex
provider. They do not call a live model service.

## Risks

- This changes local Codex recovery after process loss. A continuation
can execute more work in the retained conversation. Its prompt requires
state inspection before repeating an uncertain action; the system does
not replay tool calls.
- Recovery shares the existing three-attempt budget and one-shot
continuation marker. Real failures and older unmarked failed checkpoints
are not reopened.
- No database migration or API change is required.

## Model Used

- OpenAI Codex, based on GPT-6. The exact model ID and context-window
size are not exposed in this session.
- Capabilities: reasoning, source inspection, tool use, code editing,
code execution, and test analysis.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 07:49:47 -05:00
Devin FoleyandPaperclip efac8ff2f4 Add bounded Git integration restore diagnostics (#15291)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Sandbox runs must restore their workspace before finalization can
succeed.
> - Restore diagnostics identify the failed phase and operation.
> - A Git integration exit code can still describe several different
failures.
> - This pull request adds fixed command labels and supported failure
classes.
> - Operators can distinguish these failures without collecting private
Git output.

## Linked Issues or Issue Description

Refs #15005. This branch includes merged #15268. It labels that change's
locked ref transaction and nested branch probe without changing their
behavior.

**What happened?**

A failed Git integration can report only `git_integration`, `unknown`,
and an exit code. That evidence does not identify the failed command.
Some Git versions also return exit 1 for both a merge conflict and an
invalid object.

**Expected behavior**

Record a fixed command family and a supported failure class. Keep
unknown cases as `unknown`. Exclude command arguments, process output,
paths, repository URLs, filenames, and ref names.

**Steps to reproduce**

The tests create local repositories with conflicting commits, a missing
object, and an expected-old ref mismatch. They call the real Git
operations and inspect the resulting diagnostic. No hosted workspace or
external provider is used.

**Paperclip version or commit**

Base: `e99854249c`.

**Deployment mode**

Built from source. The diagnostic applies to sandbox workspace restore.

## What Changed

- Label Git integration calls with a closed command enum. Preserve
arguments, options, errors, and retry behavior.
- Recognize supported object and ref errors and OS permission codes.
Require both exit 1 and completed tree output for a merge conflict.
- Carry the closed fields through the existing restore receipt, saved
adapter result, and Sentry projection. Revalidate saved metadata before
projection.
- Test nested wrappers, parallel failures, reused errors, handled
probes, result settlement, and privacy with the real Sentry SDK.
- Document the field contract and its limits.

## Verification

- Eight focused suites pass: 295 tests. These cover Git sync, restore
diagnostics, result settlement, teardown, sandbox runtime, failure
projection, the real Sentry SDK, and native warm-workspace Git history.
- `PAPERCLIP_REQUIRE_SENTRY_TEST_SDK=1 pnpm exec vitest run
packages/adapter-utils/src/workspace-restore-diagnostics.test.ts
packages/adapter-utils/src/workspace-restore-result.test.ts
packages/adapter-utils/src/git-workspace-sync.test.ts
packages/adapter-utils/src/workspace-restore-teardown.test.ts
packages/adapter-utils/src/sandbox-managed-runtime.test.ts
server/src/services/__tests__/run-failure-diagnostics.test.ts
server/src/__tests__/run-failure-sentry-real-sdk.test.ts
server/src/__tests__/native-workspace-sync-history.test.ts`
- Local checks use Node 24.21.0 and the pinned pnpm 9.15.4. The optional
Sentry SDK is pinned to the declared 10.71.0 and uses an in-memory
transport.
- `pnpm -r typecheck` and `pnpm build`: pass on the rebased head.
- The full local `pnpm test:run` was stopped before source changes when
#15268 merged and required a rebase. It reported three pre-existing
company-skill cache test failures on macOS. The same three cases fail on
clean bases `2c43b39167` and `e99854249c` and the earlier head with
`EACCES` when publishing a read-only cache directory. The relevant test,
service, and cache source blobs are identical. No cache changes are
included. Later local test groups were not reached.
- All 54 checks pass on exact head `8448896bc6`, including the
post-ready security scan, with two intentional Storybook skips. Greptile
scores this head 5/5 with no review threads. Full Linux CI covers the
local groups that were not reached.
- Independent review found no blocker. Its diagnostics, Git workspace,
and native history suites pass 115/115 on this head.
- A source comparison confirms that removing only diagnostic wrappers
yields the merged upstream Git integration code exactly, including ref
locks, transaction protocol, arguments, options, and retry decisions.
- The clean base has stale dependency overrides in its lockfile. Local
installation resolved them as the existing PR CI fallback does. No
manifest or lockfile change is included.
- `git diff --check` and a local secrets/PII review pass.

## Risks

A Git version or localized message may not match a known failure form.
Such cases remain `unknown`. Up to 16 KiB of stderr and the bounded
tree-ID prefix of stdout are inspected only in memory. The saved fields
contain enum values only. These diagnostics do not establish workspace
recovery or authorize retries. Restore decisions and Git mutations are
unchanged. No schema migration is required.

## Model Used

OpenAI Codex, GPT-6, with reasoning, tool use, and code execution. The
service does not expose the exact model deployment ID or context-window
size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-05 17:10:05 -07:00
DottaandPaperclip a65ca09508 fix(runner): settle accepted results after shutdown failures (#15217)
## Thinking Path

> - Paperclip manages AI agents and the tasks they perform.
> - The native runner saves tool results and completion reports before
it releases a session.
> - Large project discovery responses can exceed the durable command
limit.
> - A shutdown failure can leave a saved answer waiting for workspace
repair.
> - Recovery reused the old assessment for a different status decision,
which violated a database constraint.
> - This pull request bounds discovery responses and lets recovery
commit the saved result after workspace repair.
> - The benefit is a task that reaches its correct final status without
another provider turn.

## Linked Issues or Issue Description

**What happened?**

A native run can save its final answer, fail during shutdown, and leave
the task In Progress after workspace repair succeeds. Reconciliation
tries to reuse the failed-workspace assessment for a new decision. The
one-decision-per-assessment constraint rejects the write. Replaying the
old decision can also retain a fresh coordinator lease. Separately,
retained session cleanup only recognizes the old `adapter_failed` label.

The project-list tool returns full project records, including large
descriptions and workspace configuration. A large response exceeds the
runner's durable command limit. The settlement diagnostic previously
recorded only a failure flag.

**Expected behavior**

Project discovery stays within the command limit. Recovery finishes the
saved result after workspace repair, releases its lease, and preserves
the original error for inspection. It does not repeat provider work or
relax session ownership checks.

**Steps to reproduce**

1. Return large project records from `list_projects` and observe an
oversized semantic result.
2. Persist an accepted native completion result, then record a shutdown
failure.
3. Finalize with a failed workspace, repeat that attempt, then record
successful workspace repair.
4. Reconcile the run. Before this fix, the issue stays In Progress.

**Paperclip version or commit**

Reproduced on `a386a599983519eb1d399f8b770bfccdb2a74762`.

**Deployment mode**

Self-hosted server with Paperclip Runner.

Related transport work: #12208 drains queued events; #12241 resumes
interrupted semantic calls. This change addresses bounded project
discovery and accepted-result finalization.

## What Changed

- Read bounded project summary projections from the database and return
at most 50 authorized summaries with a continuation cursor and explicit
description truncation. Agent and run trust boundaries narrow the
database candidates; project-specific policies still receive full
authorization. Only visible projects determine continuations. The
default project-list API remains unchanged.
- Record bounded, content-free settlement failure causes for command
limits, storage errors, and rejected dispatch.
- Include workspace state in assessment identity. Preserve the initial
assessment for interrupted finalization, and commit replacement
assessment and decision references together.
- Release the coordinator lease when an existing decision is replayed,
without repeating its effects.
- Clear stale errors when recovery succeeds and retain them in
`recoveredExecutionFailure`.
- Accept both current and legacy close-failure labels in the existing
exact-state cleanup path.
- Add regression tests and update the tool contracts and recovery
documentation.

## Verification

- Red: the new project paging, settlement diagnostic, current cleanup
label, and repaired-workspace regressions failed on the original
implementation.
- Green: protocol/catalog/tool checks (117 tests), the full project-tool
and finalizer suites (47 tests), cleanup ownership cases (70 tests), and
cleanup sweep cases (4 tests) pass. Database-backed pagination covers
large descriptions/configuration, complete enumeration, uppercase
cursors, agent/run restrictions, project-policy scope contributions, and
identical results/cursors when hidden projects are added.
- The initial CI failures in interrupted Board waits, contended endpoint
proof, and semantic schema validation were reproduced and fixed; all
affected cases pass locally.
- `pnpm -r typecheck` — passed.
- `pnpm build` — passed.
- `pnpm test:run` — started before the review corrections; it spanned
several source revisions and was stopped after reporting old-behavior
and timing failures. It is not claimed green. Fresh project and recovery
suites pass; the recovery suite also passes all 25 cases with the broad
runner’s isolated home/config. The Slack timing case passed in
isolation. Latest-head CI is the authoritative complete test matrix.
- Existing authorization suite — 66 tests passed.
- Latest-head CI on `8e5763d915aa6f375bdab6601996899ea01496fc` — 55
checks passed, 4 intentionally skipped, no pending or failing checks.
- Greptile — 5/5, zero unresolved threads on the same commit.
- `git diff --check` — passed.

## Risks

- `list_projects` now returns summaries. Callers must follow
`nextCursor` and use the authorized project API for full records.
Candidate narrowing is only an optimization: project policy and
responsible-user authorization remain authoritative.
- Assessment identity changes for workspace finalization. Existing
evidence remains intact; no migration is needed.
- Cleanup still requires matching identities, settled tool evidence, and
verified process ownership. Unknown tool outcomes remain blocked from
session reuse.

## Model Used

OpenAI Codex, based on GPT-6, with reasoning, code execution, and GitHub
tools. The exact serving model ID and context window are not exposed in
this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-05 10:31:11 -05:00
Devin FoleyandPaperclip d6d88b9de2 fix: preserve run outcomes when agent file cleanup is deferred (#14945)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The heartbeat service records each agent turn and releases its
working files.
> - A turn can save its work and finish before instruction-copy cleanup
runs.
> - A cleanup exception can replace that completed result with an
adapter failure.
> - This also loses result accounting and can prevent environment lease
release.
> - This pull request records a cleanup warning and keeps the original
run outcome.
> - The existing recovery sweep retries cleanup from the durable
working-copy record.

## Linked Issues or Issue Description

Related cleanup and lock work: #14866 and #14869. Related, distinct
work: #14695 retains warm-process files; #12021 handles provider-process
SIGTERM after a terminal result.

**What happened?**

An agent saved its plan, posted a comment, and requested approval. The
provider completed its turn. Instruction-copy cleanup then timed out on
a directory lock. Its exception escaped a `finally` block and replaced
the provider result, so the completed turn showed `Run failed`.

**Expected behavior**

Keep the provider outcome, usage, cost, saved work, and pending
approval. Record a cleanup warning and let the existing recovery sweep
retry. A real provider failure must keep its original error. A failed
file save must keep its failed-save receipt.

**Steps to reproduce**

1. Complete a legacy adapter turn that saves work and requests approval.
2. Make instruction-copy release throw a directory-lock timeout.
3. Read the run result. Before this change, the cleanup error replaces
the provider outcome.

The new heartbeat tests reproduce the failure without a live provider or
external service.

**Paperclip version or commit**

The regression reproduces on `c83df091b1a5207375eaf23466bb5c62e4e1518e`.
This branch is rebased onto `cf8ad63c80`.

**Deployment mode**

Server-managed agent execution with persistent instruction working
copies.

## What Changed

- Catch instruction-copy release failures in both heartbeat teardown
paths. Stop repeating a failed cleanup attempt within the same run.
- Write a sanitized `instruction_cleanup` warning. A warning-write
failure also preserves the run result.
- Test successful, failed, and throwing providers; both teardown paths;
warning-write failure; accounting; approval state; and execution-control
release.
- Extend the held-lock test to prove a fresh recovery worker removes the
deferred copy and preserves its failed-save receipt.
- Document deferred cleanup and the run-log event.

## Verification

- Red proof: all five new heartbeat regression cases fail with the
original release calls.
- At head `20bea4f431c916d2f5db1970213aab85f5daa34c`, all 417 tests
passed across heartbeat process recovery, agent directory working
copies, and directory merge locks.
- Full local `pnpm -r typecheck`, `pnpm build`, and `git diff --check`
passed.
- [GitHub
CI](https://github.com/paperclipai/paperclip/actions/runs/37031119795)
passed at this head. All 53 reported checks passed; the two Storybook
checks were correctly skipped. This includes general and serialized
tests, browser shards, runner verification, build, typecheck, and the
canary dry run.
- Greptile reviewed this head with 5/5, no code comments, and no
unresolved review threads. The branch has no merge conflicts.
- The local `pnpm test:run` attempt was stopped after it reported eight
failures in unchanged suites. Four Slack/AgentMail cases selected an
unrelated ancestor skills directory and failed with `ENOENT`; the two
Slack cases passed with a temporary local skill-root link, which was
then removed. Three company-skill cases reproduced macOS `EACCES` errors
when renaming read-only cache directories. One gateway case passed when
rerun alone. No full local-suite pass is claimed; the complete CI test
jobs passed.

## Risks

- Cleanup errors now leave recovery work pending. The durable
working-copy record remains available for the existing retry sweep.
- This change preserves provider failures and failed-save receipts. It
does not claim that unsaved file edits were saved.
- Native instruction reservation errors retain their existing behavior
because they guard process ownership.
- No schema, lockfile, workflow, API, or UI changes.

## Model Used

OpenAI Codex, based on GPT-6, with reasoning, repository tools, and code
execution. The exact backend model ID and context-window size are not
exposed by this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-02 14:16:18 -07:00
DottaandPaperclip 7a52dcdc74 fix: repair MCP validation and cancelled execution recovery (#14951)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The tool gateway gives agents access to connected services. Recovery
controls what happens when a run stops.
> - Generated tool names can exceed the provider limit after the MCP
client adds its prefix.
> - The same invalid definition can fail each automatic retry. A
cancelled run can also hold saved messages without showing its cause.
> - This pull request bounds tool names, stops configuration retries,
and retains cancellation evidence.
> - It shows the stopped run and admits saved input only after the
existing safety checks pass.
> - The benefit is a clear recovery path that preserves operator Stop
and prevents duplicate message delivery.

## Linked Issues or Issue Description

**What happened?**

A long connected MCP tool name makes the provider reject the entire
request. Automatic recovery repeats the invalid request. Separately,
unexpected legacy cancellations can leave saved input behind a recovery
hold. The notice does not identify the stopped run or its cause.

**Expected behavior**

Complete MCP names fit the provider limit. Tool-definition errors
require configuration repair. Cancelled runs retain their source and
reason. The recovery notice shows the cause and saved-message count.
Verified unexpected cancellations can start a fresh turn through the
existing admission checks.

**Steps to reproduce**

1. Assign an App gallery connection with a long application key and tool
name to a Claude agent.
2. Start a run. The provider rejects a name over 128 characters,
including its MCP prefix.
3. For cancellation recovery, stop a legacy provider turn without an
operator Stop request and send a user message while the recovery hold is
active.
4. Inspect the recovery notice and the deferred message queue.

**Paperclip version or commit**

Rebased onto master at `cf8ad63c806685bfd7c48e3ed4a919d61a7c55f1`.

**Deployment mode**

Hosted or self-hosted server with legacy Claude or Codex execution.

Related public work:

- Refs #14017. That PR caps name segments. This PR preserves existing
short names and uses stable hash aliases for long complete names. It
also covers classification and recovery.
- Refs #4510. That PR adds a cancellation-source column. This PR records
bounded evidence in the existing run result, without a migration.
- Refs #12552 and #4506. Those PRs suppress recovery after operator
cancellation. This PR preserves operator intent and uses the existing
continuation gates.

## What Changed

- Bound gateway names with the full provider prefix in the 128-character
budget. Retain the original upstream tool name for dispatch and
permissions.
- Classify invalid tool definitions as configuration failures before
diagnostic redaction. Stop automatic retries and continuation attempts
for that error code.
- Persist cancellation source, expectedness, initiator, reason, and
time. Preserve recorded Stop intent when adapter results arrive. Report
unexpected started cancellations with closed diagnostic labels.
- Show the run cause, saved-message count, and Inspect run link. Offer
Continue for eligible unexpected cancellations. Require verified
provider stop, empty tool inventory, ownership, and the existing pause,
budget, approval, and dependency gates. Use the existing queue for
single delivery.
- Add regression coverage and update the execution, MCP gateway, and
run-log documentation.

## Verification

- `pnpm -r typecheck` and `pnpm build` passed.
- `pnpm check:token-gates` passed.
- Ran `pnpm test:run` and completed its workspace and serialized groups.
Initial resource and timing failures passed on isolated reruns. All 149
serialized route suites passed.
- Reran the changed server, adapter, and UI suites after the rebase.
Coverage includes long-name upstream dispatch, configuration retry
suppression, cancellation evidence retention, privacy labels, oversized
run projection, and concurrent saved-message delivery.
- `pnpm test:e2e tests/e2e/legacy-failure-continuation.spec.ts` passed
all six browser scenarios. The recovery notice shows the run cause and
inspection link, and each recovery entry point reaches one new response.
- Added database-backed checks for active, removed, paused, unavailable,
and disabled chat connections. The final continuation and
recovery-notice suites passed 167 tests. Externally bound chats hide
board Continue and show a usable next action.
- All 55 GitHub checks passed on
`42afbf1371dcaeb72646e3d8f65c19ff7cddf8de`. Two unrelated Storybook jobs
were skipped by their normal conditions. Greptile reviewed that commit
at 5/5 with no findings and no open review threads.

## Risks

- Long tool names change to aliases. Existing short names stay
compatible. The original connection and upstream name remain the
dispatch authority.
- Invalid tool definitions no longer get automatic retries. An operator
must repair the configuration before a new attempt.
- Continuation changes apply only to positively identified unexpected
legacy cancellations with complete empty tool inventory. Operator Stop,
unknown historical cancellations, outstanding tools, and unverified
provider termination keep their holds.
- No database migration. The added projection fields are optional.
Cancellation reason and initiator IDs remain local run evidence; Sentry
receives only closed source and initiator-type labels and expectedness.

## Model Used

- OpenAI GPT-6 through Codex, with reasoning, repository editing, shell
execution, and GitHub tool use. The runtime does not expose the exact
model variant or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-02 13:47:59 -05:00
DottaandPaperclip ec3bacc9bd fix(chat): hide ignored provider information (#14929)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Task and agent chats show agent progress and problems that need
attention.
> - Codex also sends account, skill, and unrelated thread notifications.
> - The runner correctly ignores that information but reports it as a
warning.
> - Chat then shows an internal diagnostic as an actionable provider
notice.
> - This pull request keeps the diagnostic in run logs and removes it
from chat.
> - Real provider warnings, errors, and agent replies remain visible.

## Linked Issues or Issue Description

**What happened?**

Chat showed “Received a provider update” and a warning with the text
“ignored
unrelated provider information”. Its details said “User Actionable: Yes”
even
though no user action was needed. Saved conversations retained the same
noise.

**Expected behavior**

Keep ignored provider information in the run log. Do not show it as chat
activity
or a user warning. Preserve real warnings and errors.

**Steps to reproduce**

1. Start a conversation with the native Codex runner.
2. Have the provider send an account update, skill change, or unrelated
thread
   notification during the turn.
3. Inspect live chat and reload its saved history.

The regression tests also reproduce the old stored notice without a live
account.

**Paperclip version or commit**

Source implementation on master at `e00d10d5d`. The duplicate search
found no
open PR for this fix. Related prior work: #13109 improved
provider-notice
presentation. #12367 added Codex thread normalization. This change
addresses
the internal information that those paths still projected as chat
warnings.

**Deployment mode**

Native Paperclip Runner with the Codex app-server provider. The issue
was seen
in hosted chat and can be reproduced with local provider fixtures.

## What Changed

- Map ignored unrelated Codex information to `harness.diagnostic` in the
Rust
  and TypeScript normalizers.
- Retain a bounded allowlist of redacted provider method and thread/turn
identifiers.
- Use the same Unicode character limit and truncation marker in both
normalizers.
- Share the text redactor through a pure helper. Keep provider
connection code
  out of the standalone demo's source closure.
- Omit that diagnostic and the matching legacy notice from live chat.
- Omit the matching legacy notice from saved chat history.
- Test diagnostic retention, account-notification integration, live and
saved
  chat, and continued visibility of real warnings, errors, and replies.
- Document the local run-log event and historical display behavior.

## Verification

- Passed: 68 tests in the two affected UI transcript suites.
- Passed: 60 TypeScript tests across provider events, transport
behavior, and
  the standalone demo boundary.
- Passed: 13 Rust provider-event tests and the Codex
account-notification
  integration test.
- Passed: `pnpm check:token-gates` and Cargo formatting checks.
- Passed: full `pnpm build` and `pnpm -r typecheck`. After the review
fix,
the provider package build, typecheck, and both provider-event suites
passed again.
- Full local `pnpm test:run` failed: 608 files / 10,904 tests passed, 30
server
suites failed, and 104 files / 4,012 tests were skipped. Most failures
were
  embedded PostgreSQL startup errors. Two tests timed out in
`heartbeat-comment-wake-batching` and
`workspace-git-snapshot-streaming`.
  PostgreSQL startup also failed in `heartbeat-run-event-sequencing` and
`native-finalization-migration`. These server files are unchanged by
this PR.
Isolated heartbeat reruns were skipped locally. The stable test script
stopped
  after this general-server group, so later groups did not run locally.
- The original review thread is resolved. Greptile is 5/5 on current
head
  `683dab7cce57187c57e84c83f5e9da4ad75c9c04`.
- All current-head CI gates passed, including the full
server/chat/workspace
test matrix, Rust and TypeScript runner suites, browser E2E, build,
typecheck,
and release canary. [CI
run](https://github.com/paperclipai/paperclip/actions/runs/37021330663).
- Replay the exact old warning in either transcript adapter. It must
produce
no chat row. A genuine provider warning or error must still produce a
row.

## Risks

- Low risk. The display filter matches one diagnostic code or the
complete
  legacy warning shape. Other provider notices remain visible.
- New ignored-information events use the existing harness-diagnostic
event
type. They retain diagnostic evidence without original account payloads.
- No database migration, API permission, provider execution, or recovery
  behavior changes. This affects the local run log, not Telemetry or
  OpenTelemetry exports.

## Model Used

OpenAI Codex, GPT-6. The exact backend model ID and context-window size
are
not exposed in this session. Used reasoning, repository inspection, code
editing, tool use, and test execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run the affected tests locally and they pass (the broad
local run has PostgreSQL startup errors and timeouts documented above;
the full CI matrix passed)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-02 09:59:34 -05:00
DottaandPaperclip b54b2dc35c fix: preserve warm Codex turns with incremental managed file checkpoints (#14735)
## Thinking Path

> - Paperclip manages AI agents and keeps their instructions and files
durable.
> - Native Codex runners can keep a process alive between compatible
turns.
> - Managed file collection stopped that process after each turn, which
defeated warm reuse.
> - Agent folders can contain large images and other files, so full
copies on every turn are expensive.
> - This change keeps one managed directory for the live session and
saves only file changes after each turn.
> - Ownership, authorization, instruction changes, and process
retirement still control when reuse is safe.

## Linked Issues or Issue Description

Related: #13710 introduced native warm session reuse. This fixes managed
file collection that still forced those sessions to stop. No duplicate
open PR or issue was found.

**What happened?**

With managed instructions and warm native Codex enabled, consecutive
turns reused a Daytona sandbox but started a new runner process each
time. The managed directory collector required process termination
before saving files.

**Expected behavior**

Compatible turns keep the same process and managed `AGENT_HOME`. Each
completed turn saves added, changed, and deleted files before the next
turn starts. Unchanged large files do not transfer again.

**Steps to reproduce**

1. Use a native Codex agent with managed instructions and a reusable
Daytona environment.
2. Enable warm session reuse and run three turns on the same task.
3. Write a large binary on the first turn, edit a small note on each
turn, and delete a file on the second turn.
4. Compare process identity across turns and read the canonical files
through the public agent-files API.

**Paperclip version or commit**

Reproduced on `d30b03bd8c17604cdab1533eeeeb087aba30e8b1`.

**Deployment mode**

Local server with remote Daytona execution; cloud native runner uses the
same path.

## What Changed

- Retain the managed directory only for the verified owner of a live
native Codex session.
- Checkpoint each completed turn before releasing the session for reuse.
Retry unstable captures, then stop and collect when a warm checkpoint
cannot be validated.
- Compare metadata and cached hashes, stream only changed file payloads,
record deletions, and validate path, content, quota, and authorization
before saving.
- Rotate sessions when canonical files, loaded instructions,
credentials, or launch policy change. Fence stale collection and cleanup
callbacks from later owners.
- Keep cleanup and recovery aware of the current session owner. Recheck
canonical files under the writer lock at handoff, attach the successor
collector before fallible bookkeeping, and emit one final save receipt
on checkpoint fallback. Preserve storage warnings across unchanged
checkpoints.
- Add regression coverage and a three-turn Daytona test with independent
public API file checks, an unchanged 8 MiB binary, deletion checks, and
strict process identity checks.
- Document checkpoint consistency, lifecycle behavior, and local run-log
counters.
- Replace a timing assumption in the Daytona teardown test with explicit
transfer-arrival gates after CI exposed an unset release callback.

## Verification

- Full local `pnpm -r typecheck` and `pnpm build` passed. Server checks
were repeated after the final storage-warning fix.
- Runner E2E typecheck and 749 runner E2E unit tests passed.
- Focused file checkpoint, directory ownership, instruction collection,
native session, and merge tests passed. After review fixes, the
managed-directory and native-session suites passed 550 tests, including
intervening canonical edits, same-run fresh restore, failed handoff
collection, and one-call fallback collection. Server typecheck passed
again. The Daytona plugin suite passed 218 tests. The quota-warning
regression failed before the fix and passed afterward.
- Three real Daytona campaigns passed before the final handoff review
fixes. The latest kept PID 547 across all three turns. The first
checkpoint copied 8,388,635 bytes; the next two copied 36 and 54 bytes.
Public API reads verified the binary, note contents, and deletion after
every turn. Test cleanup deleted the sandbox.
- The final head was also deployed to an isolated cloud staging instance
and passed three UI-triggered native Codex turns with managed
instructions. All three retained the same process ID/start time, native
session, provider session, runner instance, and Daytona sandbox.
Checkpoints copied 8,388,643 bytes on turn 1, then only 52 and 78 bytes
on turns 2 and 3; those warm captures also hashed only 52 and 78 bytes.
Independent canonical API reads verified every byte of the unchanged 8
MiB binary and the exact note contents after every turn; the deleted
file returned 404 after turns 2 and 3. After restoring the original
lifecycle and agent-auth configuration, removing the temporary secret,
pausing the test agent, and deleting both test sandboxes, independent
canonical API reads still verified the entire binary, the final 78-byte
three-line note, and the deletion. The native runner flag remained
enabled and the final serving revision remained the PR head.
- Two earlier staging attempts are preserved as failures and are
excluded from the acceptance result: a saved ChatGPT login failed with a
provider routing 401, and its subsequent stopped-sandbox retry failed
before provider startup with a closed-lease admission error. The
successful campaign used a fresh sandbox and a temporary encrypted
API-key binding. The stopped-lease retry remains unexplained; this
campaign does not establish recovery of that failed sandbox.
- All [Paperclip CI
gates](https://github.com/paperclipai/paperclip/actions/runs/36750397355)
pass on `26ef2ef56a389259246809805c0b34a4747eb86b`, including full test
partitions, build, typecheck, runner verification, E2E shards, and the
Canary clean public-npm install. Greptile reviewed that exact head at
5/5 with no unresolved review threads or outstanding findings.
- Full local repository coverage used the existing CI partitions, but
the 40,000-file Git streaming stress test timed out and its local retry
was interrupted by macOS thermal emergency sleep; this is not a green
full local suite claim. The exact stress test passed on the final head
in [CI server shard
2/12](https://github.com/paperclipai/paperclip/actions/runs/36750397355/job/110008294290),
in 111.9 seconds.
- Repeat the live test with configured credentials and a Linux runner
artifact: `pnpm test:e2e:runner -- --id
daytona-warm-continuity.runner-codex.daytona.warm-three-turn`.

## Risks

- This is a file-level checkpoint, not an atomic snapshot of the whole
folder. Background writes after a capture are saved by the next
checkpoint or final stopped collection.
- Metadata scans still visit all paths. Modified files transfer in full;
unchanged files do not rehash or transfer.
- Incorrect ownership or reuse could collect the wrong directory. Run
ownership fences, current authorization, stable capture validation, and
stopped collection fallbacks are covered by tests.
- Warm reuse remains opt-in. No database migration or fleet default
changes.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, code editing, tool use, and
test execution. The exact serving model ID and context-window size are
not exposed by this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 13:33:37 -05:00
DottaandPaperclip 94e8dec56b fix(runner): preserve tool outcomes through shutdown and restart (#14734)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The runner sends authorized tool calls to the server and saves their
results.
> - A provider turn can stop while a server write is still running.
> - The old shutdown path invented a failed result that could conflict
with the real result.
> - Truncated execution input and incomplete recovery records made the
failure harder to diagnose.
> - This pull request preserves exact inputs and actual outcomes through
shutdown and restart.
> - Tests force the race and crash boundaries so safe retries do not
repeat writes.

## Linked Issues or Issue Description

**What happened?**

Stopping a turn during a server tool call could record a false failure,
then reject the actual result as a conflict. The diagnostic input
formatter could truncate instruction content before execution. A crash
during saved-result delivery could leave that delivery permanently
indeterminate. Cleanup could hide the first failure, and a retry could
overwrite earlier run logs.

**Expected behavior**

Keep dispatched tools pending until their actual result is known.
Preserve accepted input bytes. Accept identical result delivery without
failing the task. Reject conflicting results with enough evidence to
diagnose them. Recover saved-result delivery without repeating the
business operation.

**Steps to reproduce**

1. Hold an instruction update at the filesystem commit barrier.
2. Stop its provider turn before the server returns the result.
3. Release the write, deliver its result, and replay the same result.
4. Repeat with a restart before and after the delivery receipt is saved.
5. Check that there is one write and one audit row, and that the exact
result survives.

**Paperclip version or commit**

The change was developed from `44736c9c7` and rebased onto `0e5830887`.

**Deployment mode**

Self-hosted server with the native runner. Tests use local runner
processes, scripted providers, and PostgreSQL.

Related work: #12353 added durable semantic tool receipts; #12384 added
durable Codex tool recovery; #12404 bound semantic tools to ACPX
sessions. #14633 covers separate native-provider cancellation and
qualification work. This PR addresses server semantic-tool outcomes and
their durable delivery. No duplicate fix was found. AgentMail discovery
is outside this PR.

## What Changed

- Close turn admission without inventing results for dispatched tools.
Keep pending calls and accept late actual results.
- Accept identical result replay with a diagnostic warning. Include call
identity and both result hashes in real conflict errors.
- Preserve exact execution arguments. Reject prohibited or oversized
input before dispatch. Keep diagnostic previews redacted and bounded.
- Commit instruction-attempt evidence before the filesystem write. Save
completed mutation receipts so concurrent and restarted duplicates
return the first result. Recheck authorization before replay. An attempt
without a completed result stays unknown and cannot execute again.
Definite pre-write failures save and replay their original error without
another write.
- Recover an interrupted saved-result delivery only for backends with
durable result receipts. Never replay an ordinary business operation
with an unknown outcome.
- Preserve the initiating error when cleanup also fails. Record
incomplete settlement evidence. Propagate typed unknown-outcome errors
through the native tool wrapper without creating a false completed tool
result.
- Append run-log attempts and restore the durable log before appending
after local file loss. Reject incomplete restores. Publish a restored
prefix only if the destination is absent so concurrent attempts cannot
overwrite new lines.
- Add deterministic race, crash, replay, authorization, exact-content,
and log-restoration tests. Document their assertions in
`packages/paperclip-runner/docs/durable-recovery.md`.

## Verification

- Current head: `7e088f4c7fba8ebabf98ae95485a5753b013d489`. All 55
applicable checks pass; four conditional/manual checks are skipped. This
includes build, typecheck, Rust, both runner TypeScript shards, server
and workspace tests, all eight browser shards, isolated runner
compilation, and the clean-install release dry run. [CI
run](https://github.com/paperclipai/paperclip/actions/runs/36746101110).
- Greptile reviewed this exact head at 5/5 with zero new findings. All
three earlier review threads are resolved.
- Focused local verification includes 11 instruction integration tests,
23 surrounding authority/tool tests, 26 run-log tests, and 169
controller/driver tests. The post-rebase controller/transport/runtime
selection passed 415 tests. The full Rust release suite passed 617 tests
with two ignored. The real-process SIGKILL recovery test passed three
consecutive runs.
- The fault matrix in
`packages/paperclip-runner/docs/durable-recovery.md` uses explicit
barriers, real PostgreSQL rollback, durable journal reloads, and killed
runner processes. It covers late results, identical and conflicting
replay, exact long content, concurrent log restoration, lost commit
acknowledgements, and definite failure replay after the original CAS
base becomes valid again. No paid model calls are needed.
- Full local recursive typecheck and build passed during implementation.
Server typecheck and the runner TypeScript build passed after the review
fixes. The broad local repository test run was stopped after repeated
database startup timeouts. Four timing/launch failures in an earlier
broad runner run passed focused reruns without changed assertions or
timeouts. These are local verification limitations; the complete
current-head CI suite is green. An earlier CI workspace job received an
infrastructure shutdown signal; its current-head replacement passed.

## Risks

- A stopped turn can remain blocked when a dispatched operation has no
proven result. The system does not guess its outcome or rerun its
effect.
- Conflicting results still fail settlement. Existing failed or
conflicting journals are not repaired automatically.
- Accepted semantic input is limited to 480 KiB of encoded JSON to fit
the encrypted transport. Larger input fails before execution.
- Instruction filesystem writes and database receipts are not one atomic
storage operation. A separately committed attempt and audit record
survive rollback. An attempt without a completed success or definite
pre-write failure receipt remains blocked as an unknown outcome. It is
not replayed or reported as success.
- Run-log restoration now reads the durable object before appending when
the local log is missing. Failed or incomplete reads reject the append.
- No schema migration, dependency change, workflow change, or AgentMail
change is included.

## Model Used

OpenAI Codex, GPT-6, with reasoning, tool use, code execution, and test
analysis. The exact served model ID and context-window size are not
exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused suites; the broad
local run limitation is recorded above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 11:54:56 -05:00
DottaandPaperclip 3b4b270650 fix(adapters): preserve ACP terminal failure diagnostics (#14573)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The shared ACP adapter engine records agent failures for operators.
> - ACP providers can report a failure category, title, and detailed
cause.
> - Our patch kept only the category in the saved error, so an operator
could not diagnose a failure when tracing was off.
> - This pull request preserves redacted provider diagnostics in the run
error, transcript, and structured run result.
> - Operators can now inspect the provider message and any supplied
request ID or stack trace after the run ends.

## Linked Issues or Issue Description

Refs #13889 (the diagnostic gap; this PR does not update the bundled
Claude version).
Refs #14484 (related model-refusal classification; this PR retains
diagnostics for all terminal failure categories).

**What happened?**
An ACP turn failed with only `ACP agent reported a terminal service
failure.` The provider's title and details were available in memory but
absent from the saved error and transcript.

**Expected behavior**
The run retains useful provider diagnostics even when raw tracing is
disabled. Credentials remain redacted. A size limit must report
truncation instead of silently removing the cause.

**Steps to reproduce**
1. Run an ACP agent that returns an error-severity typed session
failure.
2. Include an HTTP error, request ID, and stack text in its title and
details.
3. Inspect the failed run with tracing disabled. Before this change,
only the category survives.

## What Changed

- Both pinned ACPX patches pass complete error text to the in-memory
callback, so redaction happens before truncation.
- The shared engine retains the sanitized category, title, and details
in `resultJson.terminalSessionFailure` and includes the text in the run
error and error transcript.
- Diagnostics redact configured environment values even under arbitrary
names, unknown launch-environment values, connection URL passwords, run
credentials, and common credential syntax. Known boolean settings remain
readable, while credential values are redacted even when embedded in
other text. Diagnostics remove control characters and invalid Unicode.
- Title and detail limits keep escaped transcript JSON below the
server's chunk limit. Truncated fields include an omission count. The
safe run-result projection preserves a byte-bounded diagnostic preview
when the result exceeds its byte budget, with an explicit pointer to the
full adapter-bounded run error and transcript.
- The existing UI and CLI display the error. Diagnostics do not become
assistant output. Issue continuation summaries and session-compaction
prompts receive only the generic category, preventing provider text from
becoming handoff instructions. Existing quota classification, warnings,
timeout precedence, and control-channel failure precedence remain in
place.
- Regression tests cover real ACP child processes with both pinned
versions in one-shot and persistent modes, credential redaction, request
IDs after the old 4 KiB cutoff, transcript parsing, storage bounds, and
database retrieval of oversized multibyte diagnostics.

## Verification

- Full CI on `20ad4f5f1f66c46d2c260e6ad0339cbea607b4cf`: **54 passed, 2
intentionally skipped, no pending or failing checks**. Includes
typechecking, build, all Vitest shards, Runner checks, browser E2E, and
the canary packaging/public-install dry run.
- Greptile: **5/5** on this commit. Superagent security scan passes. All
review threads are resolved.
- Local verification passed: shared ACP engine suite (395 tests); real
Claude ACP child-process and diagnostic regressions across both pinned
runtimes and both execution modes; run retrieval and model-handoff
regressions (59 tests); ACPX patch packaging (16 tests); full typecheck
and build. Affected package typechecks and focused tests were rerun
after review fixes.
- The broad local `pnpm test:run` was stopped after review edits made
its cached imports stale. Fresh targeted runs pass, including both
affected server suites. Cold-build import failures were also rerun after
dependency builds: chat integration (1,063 tests) and tool access (351
tests) pass. The final commit's complete CI matrix is green.

## Risks

- Provider diagnostic text is untrusted. This change retains more of it
in company-scoped run records. Redaction and size bounds apply before
persistence.
- Diagnostics are limited to fields the provider supplies. Old runs
cannot recover discarded error text.
- No schema migration, recovery-policy change, or new Telemetry or
OpenTelemetry export.

## Model Used

- OpenAI GPT-6 through Codex, with reasoning, repository inspection,
code editing, and test execution. The exact serving model ID and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-29 10:18:30 -05:00
DottaandPaperclip 14795136f5 fix(runtime): finalize and recover sandbox workspace exports safely (#14402)
Serialize native workspace finalization, validate streamed archives within bounded limits, and quietly recover unsafe exports from saved results. Preserve exact allocations for exhausted transient failures and provide export-only retry without rerunning the provider.

Consolidates #14314, #14315, #14329, and #14334 while preserving the already-merged finalization label changes.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-28 10:57:32 -05:00
a6c4e7a8d1 fix(heartbeat): cancel obsolete execution continuations before dispatch (#13761)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - A comment can queue execution before the task changes state.
> - The task may be completed, cancelled, deleted, or reassigned before
setup checks ownership.
> - The ownership guard must prevent that obsolete execution from
starting.
> - This pull request records that guard outcome as cancellation and
settles the wake request.
> - Missing history and authorization errors remain failures that
operators can investigate.

## Linked Issues or Issue Description

**What happened?**

A comment can queue a run just before a user marks its task Done. The
ownership guard stops setup before adapter dispatch, but the run becomes
a setup failure and can put the agent into an error state.

**Expected behavior**

Cancel obsolete work when the typed task ownership guard rejects it.
Keep genuine setup errors visible. Do not change the task's terminal
state or current owner.

**Steps to reproduce**

Queue a task continuation, then mark the task Done or Cancelled before
continuation setup. The run should settle as cancelled without invoking
the adapter. Deleting or reassigning the task must also prevent
dispatch. Missing source history or explicit user authorization must
still fail.

**Paperclip version or commit**

Refreshed against master `b2e9e82f053af14777fd68e8834fff6bb64c84c4`.
Related: #13546 covers lost issue-lock claims, and #13888 covers
persisted continuation decisions and retry budgets. This change handles
the final continuation ownership guard during setup.

Thanks to @MrBlackTongue for the original typed cancellation
implementation and regression coverage. This update preserves the
contributor's commits, resolves the master conflict, and narrows
cancellation to task ownership invalidation.

## What Changed

- Add a typed `StaleExecutionContinuationError` for
`continuation_task_ownership_changed`.
- Use existing cancellation settlement for that typed guard, including
the run, wake request, issue execution ownership, and agent state.
Suppress immediate recovery of obsolete work.
- Preserve failure classification for missing source context, missing
user authorization, and untyped errors, including an untyped error with
identical text.
- Cover Done and Cancelled tasks with real database checks. Retain
company and assignment guard coverage. Verify cancellation performs no
adapter dispatch or automatic replay.
- Document the cancellation event and its error classification in the
run-log guide.

## Verification

Current head: `b875486e81`.

- `pnpm -r typecheck`: passed.
- `pnpm build`: passed.
- `pnpm exec vitest run
server/src/__tests__/heartbeat-process-recovery.test.ts
server/src/services/execution-continuation.test.ts --maxWorkers=1
--no-file-parallelism -t 'authorized continuation context|obsolete
continuation setup|untyped continuation setup failures'`: 17 passed; 316
unrelated tests excluded by the name filter.
- `pnpm test:run`: local full run is still finishing. It is not a clean
pass: the observed failures are three company-skills cache checks
(`EACCES` while renaming read-only cache directories on macOS), one
tool-access assertion, and one comment-wake timeout. The three cache
failures reproduce in isolation in unchanged code. The tool-access check
passes in isolation, and the entire comment-wake file passes (29 tests),
including explicit feedback after completion. Final full-run totals will
be added when available.
-
[CI](https://github.com/paperclipai/paperclip/actions/runs/36282532652):
all gates green on this head. An unchanged Runner workspace-diff test
initially returned no diff; an isolated check passed, and the failed job
plus its dependent gate passed on retry with no code change. Original
failed job: 108517010989.
- Greptile: 5/5 on the full current head, with no unresolved review
threads. Branch is current with master and conflict-free.
- Diff check and local secret/PII scan passed. No customer data or
private deployment identifiers are included.

## Risks

The classification change is limited to the typed ownership guard. A
task with missing history remains a failure rather than being treated as
an expected cancellation. Existing company, task-owner, terminal-state,
and authorization checks still prevent dispatch. Cancellation retains
its specific reason in the run log and does not grant a retry. No
schema, dependency, or API changes; no migration is required.

## Model Used

OpenAI GPT-6 via Codex assisted investigation, implementation, code
review, and test execution. The exact serving model identifier and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have described the issue in-PR following the bug issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links
- [x] My branch name describes the change and contains no internal
ticket id
- [x] I have run the focused regression tests locally and they pass;
full-suite results are tracked above
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation
- [x] I have considered and documented the risks
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Devin Foley <devin@paperclip.ing>
2026-09-26 17:43:33 -07:00
Devin FoleyandPaperclip 4ca404b49a fix: record safe sandbox restore failure diagnostics (#14064)
Record a bounded diagnostic for failed workspace and staged-asset restores.
Preserve the original error, retry policy, and archive safety checks. Never
copy raw provider messages, credentials, paths, or asset names into the log.
Nested failures log once; safe fields survive throwing property getters.

Verified 129 focused restore/Claude tests, typecheck/build, and green full
PR CI. Greptile 5/5 with all review threads resolved.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-25 17:16:13 -07:00
Devin FoleyandPaperclip bd6caf51bb fix: preserve restore failure results and stop unsafe retries (#14035)
Preserve agent output and earlier execution errors when workspace restore fails. Report the restore phase and confirmed saved-plan links. Require verified repair before retrying unsafe archives, while preserving approval states and the retry budget.

Verified with full CI, 506 focused regression tests, and Greptile 5/5 with all review threads resolved.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-25 11:34:31 -07:00
Nicky LeachandPaperclip 9ed55f6931 fix: allow concurrent agent runs on one OpenAI or xAI subscription connection (#13452)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Paperclip starts local command-line sessions and stores provider
credentials through managed connections
> - One OpenAI Codex or xAI Grok subscription connection held a
credential lease for the full agent run
> - A second run then waited for the first run, and concurrent runs
could overwrite a newer credential
> - The write-back must compare fresh credentials while it holds the row
lock
> - This pull request removes the run lease, keeps the revocation guard,
and bounds Codex timestamps against the host clock
> - The benefit is safe concurrent use of one subscription connection
with newest-credential selection

## Linked Issues or Issue Description

**What happened?**

A managed OpenAI Codex or xAI Grok subscription connection held a
credential lease for the full agent run. A second run waited for the
first run to finish. The write-back gate also rejected any row change
before it compared credential freshness.

**Expected behavior**

Concurrent runs should start on one subscription connection. The server
should keep the newest valid credential and reject a credential write
after a person revokes the connection.

**Steps to reproduce**

1. Start two runs that use one OpenAI or xAI subscription connection.
2. Let both provider tools refresh the credential.
3. Finish the runs in either order.
4. Confirm that the newest valid credential remains in the connection.

**Paperclip version or commit**

`ec25bf1e4a81d1729a6d7276e6a486587cff4a0b`

**Deployment mode**

Local dev (`pnpm dev`), built from source.

**Agent adapter(s) involved**

Codex. The server path also covers xAI Grok subscription connections.

**Database mode**

Embedded PGlite for local development, and external Postgres for
deployments.

**Access context**

Both board and agent runs can use managed connections.

**Additional context**

The provider command-line tool refreshes credentials inside the sandbox.
The server copies the result back after the run. Two long runs can still
refresh one token hours apart, so the provider can reject the second
refresh. The server cannot observe that provider call.

## What Changed

- Remove the full-run credential lease for OpenAI Codex and xAI Grok
subscription connections.
- Lock and re-read the connection row before credential write-back.
- Accept only a strictly newer credential, while keeping the connection
revocation guard.
- Reject Codex freshness timestamps more than five minutes ahead of the
host clock.
- Add tests for both completion orders, xAI cleanup, revocation, the
Codex time bound, and agent hiring.
- Update the connection and run-log documentation.

## Verification

- `server/src/__tests__/ai-connections.test.ts` passes with 42 tests.
-
`packages/adapters/codex-local/src/server/codex-auth-merge-decision.test.ts`
covers the five-minute boundary and the one-millisecond overflow.
- `packages/adapters/codex-local/src/server/codex-auth-merge.test.ts`
passes.
- `server/src/__tests__/agent-hire-ai-connections.test.ts` covers OpenAI
and Anthropic.
- The project type check reports no new error in changed files.
- GitHub Actions must pass on this pull request.

## Risks

The write-back now permits concurrent runs, so the provider may reject a
later refresh when both long runs use one token. The server keeps the
revocation guard and rejects future-dated Codex timestamps. No schema
change occurs.

## Model Used

OpenAI Codex, GPT-5. Context window and exact deployment build are not
exposed in this run. The model used tool calls, code inspection, and
test verification.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` /
`Closes: #` / `Refs: #` OR (b) described the issue in-PR following the
relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-15 01:10:54 -07:00
Devin FoleyandPaperclip 667c79ded2 fix: prevent retry-exhaustion events from exhausting attention-feed memory (#13451)
## Thinking Path

> - Paperclip manages AI agents and their work.
> - Its attention feed shows failed runs whose retry budget is
exhausted.
> - Startup retention builds this feed before startup completes.
> - The feed joins every exhaustion event to the full run context, then
removes duplicate runs in JavaScript.
> - Recovery can revisit an exhausted run and append the same event
again. This multiplies the data loaded into memory.
> - This pull request selects one small row per run in PostgreSQL and
makes exhaustion writes idempotent.
> - Existing duplicate events can stay in the database without
multiplying run contexts in server memory.

## Linked Issues or Issue Description

Refs #13367. This fixes the repeated-event allocation path in the
retention feed. Other full-feed sources and sweep cadence remain
separate concerns.

**What happened?**

The attention query loaded one full run context for every matching
exhaustion event. Deduplication ran only after the driver had loaded
those rows. A run with thousands of exhaustion events therefore produced
thousands of context copies. The startup retention sweep can exhaust the
server heap while reading this result.

**Expected behavior**

The query should return one row per exhausted run and only the context
fields that the feed needs. Repeated checks of the same exhausted retry
budget should reuse the original event.

**Steps to reproduce**

1. Create a failed run with a 32 KB context and 2,500 matching
exhaustion events.
2. Build the attention feed, including dismissed items, as startup
retention does.
3. Inspect the database result before JavaScript feed processing. The
old join returns 2,500 copies of the run context.
4. Call bounded retry scheduling repeatedly for a run at its retry
limit. The old writer appends another exhaustion event on every call.

**Paperclip version or commit**

The attention query was introduced by #9380 and is present in stable
`v2026.831.1`. The retention caller is also present in that stable
release. The later recovery callback added by #13075 provides a repeated
path into the exhausted-budget writer. This change is based on
`c0fda8fac` after rebasing onto current master.

Related PR:
[#12162](https://github.com/paperclipai/paperclip/pull/12162) changes
retention cadence and newer-run suppression queries. This change
addresses the exhaustion-event join and duplicate event writes.

## What Changed

- Select the newest company-scoped exhaustion event per run with a
PostgreSQL `DISTINCT ON` subquery before joining run data.
- Project only `issueId` and `taskId` from run context. Preserve JSON
types, fallback behavior, run ordering, company filters, and run/agent
status filters.
- Reuse an exhaustion event for the same run, reason, attempt, and retry
limit under the existing run-row lock. Recognize historical events
without a migration.
- Skip sequence allocation and live publication when a receipt already
exists.
- Document the run-log behavior and add PostgreSQL regression tests.

## Verification

- Focused attention, retry-scheduling, and event-sequencing suites: 72
tests pass on the rebased head `39ce77158`.
- The large-history fixture verifies two database result rows under 4
KB, newest-message selection, task-ID fallback, and company/status
filtering. This assertion runs before feed deduplication.
- Concurrent event tests verify one receipt across two database clients,
reuse of historical receipts, distinct reason/attempt/budget keys, and
stable event sequences.
- Repeated scheduling through a new service instance produces no extra
run-log or live events.
- `pnpm --filter @paperclipai/server exec tsc --noEmit`: passes.
- `pnpm -r typecheck` and `pnpm build`: pass on the rebased head,
including native runner checks.
- Greptile: 5/5 on `39ce77158`, with no review threads.
- [GitHub
CI](https://github.com/paperclipai/paperclip/actions/runs/34929002304):
all 25 workflow jobs pass, including all server/workspace test shards,
serialized server suites, browser tests, typecheck, build, native-runner
verification, and the release dry run. Security checks also pass. The
branch has no merge conflicts.
- `pnpm test:run`: stopped with unrelated failures. Two chat integration
cases passed when rerun separately (2 passed, 993 unselected). Three
skill-cache cases failed on both this branch and the unmodified parent
commit, with `EACCES` during directory rename. The full suite is not
reported as green.

## Risks

- No schema migration or data cleanup is required. The query still scans
matching event history in PostgreSQL; its result size now scales with
exhausted runs.
- This does not bound every source in the attention feed or change
retention scheduling.
- An exhaustion receipt is emitted once per retry decision. Consumers
that observed repeated copies will now receive one event.
- Native source-event replay handling stays on its existing path.
- A deployed application boot has not been verified.

## Model Used

OpenAI GPT-6, used through Codex with reasoning, repository inspection,
code editing, and test execution. The runtime does not expose a more
specific model snapshot or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run the focused tests locally and they pass (full-suite
limitations are documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-14 21:58:34 -07:00
DottaandPaperclip f912ecaacf fix: carry AI connections through hiring and unblock task execution (#13438)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents hire other agents and assign tasks to them.
> - Managed AI connections must follow those hires across legacy and
native runners.
> - Missing accounts should pause task execution and let the user
connect from the task.
> - Subscription contention must wait without asking for new
credentials.
> - This pull request fixes these paths and the native tool and Daytona
staging failures found during live tests.
> - The result is a working hire, subtask, and connection setup flow on
local and remote runners.

## Linked Issues or Issue Description

**What happened?**

A managed Claude or Codex agent could hire a teammate without a usable
AI binding. Cross-provider hiring could fail before the user had a
chance to connect the new provider. First-time task setup did not show
the existing AI credential form inline. A busy subscription could
request a new connection. Native API replies could stop the parent after
a hire had already committed. Fresh Daytona sandboxes could fail to
extract read-only skill directories created on macOS.

**Expected behavior**

Compatible hires inherit the managed connection choice. A hire for
another provider uses the responsible user's default. If that account is
missing, the hire succeeds and the task asks for a connection.
Completing setup in the task resumes work automatically. Explicit child
auth settings and existing unmanaged login paths keep precedence.
Shared-account access checks remain in force.

**Steps to reproduce**

1. Connect a Claude or Codex parent with a managed AI account.
2. Ask it to hire one agent of each provider and create a self-assigned
subtask.
3. Assign work to both hires without connecting the second provider
first.
4. Connect the missing provider from its task card.
5. Check that all tasks finish and same-provider work uses the original
account.
6. Repeat with native runners and fresh Daytona sandboxes. The opt-in
browser suite in `tests/hiring-ai-connections/README.md` performs these
steps.

**Paperclip version or commit**

The live failures were reproduced from `f2c5e54dc`. The branch is
rebased onto `5282cabde`.

**Deployment mode**

Isolated local development instance. Legacy CLI and native runners.
Local execution and ephemeral Daytona sandboxes.

Related work: Refs #13247 for managed AI connections. Refs #13268 for
legacy credential-reference inheritance, which this branch preserves.
Refs #13432 for a concurrent managed-inheritance fix. This PR also
covers cross-provider task setup, subscription waits, native API
replies, and Daytona extraction. It permits missing responsible-user
defaults at hire time; restricted shared selections still fail.

## What Changed

- Apply managed connection defaults to both agent creation routes.
Preserve explicit auth choices and legacy credential-reference
inheritance.
- Allow hires before their responsible user connects the provider. Keep
approval gates, company boundaries, and shared-account access checks.
- Reuse the production AI credential form inside the pending task card.
Resume the task after setup.
- Retry subscription lease contention without consuming the
provider-failure allowance or creating a connection request.
- Require task execution-lock ownership when scheduling, promoting, and
dispatching subscription retries. Recheck ownership under the issue row
lock.
- Rename the HTTP operation identity at the native tool boundary so it
cannot override the runner's operation identity.
- Delay directory permission restoration during Daytona extraction.
Preserve the final read-only modes.
- Add database-backed regressions, real browser acceptance tests, and
Storybook states. Document setup and run-log behavior.

## Verification

- Six real browser scenarios passed: both parent providers on legacy
local and legacy Daytona; native Codex locally; native Claude on
Daytona. Each scenario hires both providers, completes a self-subtask
and assigned work, and connects the missing provider inline with
automatic continuation.
- Successful runs verify the account, responsible user, runner mode, and
Daytona lease. All 18 test sandboxes were deleted.
- Live authentication used API keys. Subscription inheritance, lease
contention, and retry have integration coverage. Fresh subscription
OAuth sign-in was not automated.
- Red/green tests reproduced missing bindings, missing inline forms,
subscription contention, native API reply failure, and GNU tar
permission failure.
- Seven Storybook browser checks passed. They cover both providers,
method selection, narrow layout, completion, cancellation, and invalid
credentials.
- Full local suite coverage completed before rebase. Initial timing and
fixture startup failures passed unchanged on isolated reruns. The first
full command did not exit cleanly; the remaining workspace and
serialized groups were completed separately.
- After rebase, 107 hiring/auth/retry tests and 59 native API,
task-card, and Daytona tests passed. The full workspace typecheck,
production build, and token gates passed again. Storybook build passed
before rebase.

- Review fixes: 169 hiring/retry/dispatch tests, 37 adjacent tests, and
four explicit cancellation-race cases passed. Eight cross-provider cases
cover stale auth keys on both creation routes and both runner types.
Server typecheck and build passed.

- Final CI on `ee4890837a8a4913e07453392b9a75969580dae1`: 32 checks
passed. Two optional Storybook jobs were skipped. The full server,
workspace, browser, native runner, build, typecheck, and release checks
passed.
- Three unchanged tests initially failed on a busy port, a chat row-lock
race, and preview-server readiness. Each affected job passed after one
CI rerun. Isolated local checks also passed: 41 credential tests, the
chat-concurrency case, and 25 preview-runtime tests.
- Greptile reviewed the final commit at 5/5. Both review threads are
resolved. GitHub reports no merge conflicts.

## Risks

- A missing personal account now defers authentication to the first
task. Explicit incompatible bindings and restricted shared accounts
still fail at hire time.
- An inherited personal default uses the responsible user's existing
authorization to install access for the new agent. It never copies
credentials or another user's identity.
- Subscription contention retries after a delay and rechecks task
eligibility. It does not consume the provider-failure budget.
- Native hiring uses the existing managed API-tools opt-in. Remote
native runners require a matching Linux binary and provider pack, as
documented in the acceptance README.
- No schema changes. Live tests make paid provider calls and remain
opt-in.

## Model Used

OpenAI GPT-6 (`gpt-6-astra`) in Codex, with reasoning, repository
inspection, code execution, browser automation, and API tools. The
context-window size is not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-14 16:15:48 -05:00
DottaandPaperclip 422287eecd fix: preserve runner recovery, warm sessions, and task outcomes (#13338)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The native runner connects task messages, provider execution, and
task outcomes.
> - First-time user tests exposed gaps in recovery, completion
permissions, message delivery, and Stop behavior.
> - These gaps left usable output hidden, completed work waiting for
bookkeeping, or safe work unable to continue.
> - This pull request fixes the shared lifecycle and receipt paths while
preserving process ownership and action checks.
> - Users can continue work with accurate task state and durable
messages.

## Linked Issues or Issue Description

**What happened?**

A stopped local Codex execution could remain blocked even after its
processes had stopped and its complete transcript proved that no
external action needed replay. Claude under Conservative permissions
could fail to call task completion tools. Recovery could reuse an
assistant item ID and overwrite prior output. A delivered comment could
remain marked uncertain after navigation. Stop could look like Pause or
a new recovery incident. Workspace contention could look like
cancellation. A direct reply reopening Done could enter a clarification
loop.

**Expected behavior**

Recover automatically only with verified termination and complete action
receipts. Preserve answers and messages. Keep task completion available
under Conservative permissions without broad tool access. Show crashes
as Blocked, actual human decisions as In Review, and ordinary workspace
contention as waiting. Stop the current response and allow a new
direction.

**Steps to reproduce**

1. Create ordinary response tasks with local Codex and Claude Code, then
send follow-up messages through the task composer.
2. Interrupt a disposable local Codex runner during text-only work.
Verify automatic continuation and retained output.
3. Stop a response, send a new request, answer a clarification, and
reopen completed work with another message.
4. Navigate or reload while a comment submission is pending. Confirm the
exact persisted request receipt settles it without removing newer draft
text.
5. Run two tasks in a shared Daytona workspace. Confirm waiting does not
appear as failure.

**Paperclip version or commit**

Initial acceptance baseline: `c9021c6721f91e2c74bd9fee9d3fd41c999d17b7`.
Current integration base: `6cef9743c`. Both operator-interruption and
workspace-waiting guards are preserved; native restart and legacy
permission rules remain documented.

**Deployment mode**

An isolated source-built test-drive instance, with real local Codex and
Claude Code providers and disposable Daytona environments.

Related work: #13314, #13316, #13327, #13344, #13239, #13254, #13163.
This PR addresses additional failures from ordinary task journeys,
including controller restart handoff and repeated warm sandbox setup.
Historical task status reconciliation is excluded.

## What Changed

- Persist runner ownership immediately at spawn and resume an explicitly
adopted runner even when the controller crashed before the first driver
checkpoint. Detach the controller safely across graceful restarts,
including session startup. Prevent an old finalizer from suspending or
signaling an adopted runner. Checkpoint idle warm sessions before
shutdown. Preserve the same run and queued follow-up messages.
- Scope saved legacy queue successor checks to the queue owner while
preserving ordinary task locks, operator identity, assignment gates, and
exactly-once delivery.
- Preserve managed Codex credential files when an old session is
detached for restart; normal owned cleanup still copies refreshed auth
back and removes the scoped copy.
- Reuse the bound warm shared sandbox and fully verify an existing
staged provider pack before using it. This avoids repeated uploads when
the pack is already valid.
- Add a narrow local Codex replacement path with stopped-process proof,
a closed transcript inventory, exact completion receipts, and
fresh-session lineage. Preserve no-replay holds when evidence is
incomplete. Recovery may clear only the same run's recorded Blocked
status version; manual re-blocking and dependency changes invalidate
that receipt, while queued comments do not. Later blocks stop scheduled,
queued, and final dispatch; queued/final checks re-read dependencies
even when the task status stays In Progress.
- Permit only task delivery and human-input tools through the isolated
Claude runner's exact task bridge.
- Scope assistant item identity to the provider turn and ignore only
authority-free Codex skill-change notifications during startup.
- Reconcile composer submissions by client request ID across response
loss, navigation, and reload. Retain text typed during delivery.
- Keep acknowledged run-only Stop neutral and show workspace contention
as waiting. Project exhausted native failures as Blocked.
- Restore the guarded task-page retry action for failed legacy runs,
including the server-supported explicit new-attempt path for stopped
conversation adapters. Preserve native/process recovery holds and avoid
promising Retry while a decision or execution gate hides it.
- Refresh delivered artifacts and handle direct user replies that reopen
completed work without a clarification loop.
- Check the embedded PostgreSQL PID, data directory, and actual port
before connecting or migrating.
- Document accepted behavior and add focused regressions at lifecycle,
route, transcript, and UI boundaries.

## Verification

- Final head `fece606ac2` passes the complete GitHub CI matrix: **34
green checks, two expected Storybook skips, no failures or pending
checks**, including `ci / verify`, `ci / e2e`, full runner verification,
typecheck, build, every server/workspace shard, and all browser shards.
[CI
run](https://github.com/paperclipai/paperclip/actions/runs/34727183287).
Greptile is **5/5 with no open findings**. The final two commits only
refine test fixtures; both affected suites pass 24/24 locally and in CI,
with server typecheck green.
- Complete local Vitest coverage uses the canonical groups/shards: all
635 general server suites, all 145 serialized suites, and all workspace
packages. The aggregate began on `0a8001c18` while the final queue fix
arrived: 23,903 passed, five failed, 87 skipped. The five
port/socket/timing failures passed unchanged in follow-ups (60 tests in
the exposure/file suites and 412 tests covering the serialized failures
and unrun tails). The final queue/operator-identity suites separately
passed 52/52. This is aggregate coverage plus explicit reruns, not a
pristine single-command final-head run.
- After integration with current master,
queue/operator-identity/continuation suites passed 162/162 and affected
UI suites passed 140/140. ACP Stop/continuation and legacy
task/Inbox/message browser suites passed 9/9, including both task
recovery Retry and thread Try again, automatic saved-message delivery,
exactly one new run, Done, and retained output after reload. The default
process Stop/Pause/Resume browser case passed (the native-provider case
is opt-in and skipped by default). The complete Board attachment/receipt
browser suite passed 11/11 on a disposable instance, covering both
composers, exact receipts after lost responses, no replay, bound
attachments, and newer drafts after reload.
- Blocking-intent regressions cover pre-existing Blocked, a mismatched
run/cause, an explicit manual re-block, changed dependencies, a queued
comment after failure, and a block arriving between scheduling and
provider dispatch. The negative cases reproduced before the fix. All 478
affected executor/recovery/dispatch tests passed; both database suites
ran separately after availability-probe skips in the first combined
command. The final late-dependency check passed all 143 affected
recovery/dispatch tests (zero skips) after two new negative cases
reproduced the bug.
- Focused runtime regressions cover awaited runner ownership
publication, authenticated adoption before the first checkpoint,
old-finalizer detachment, idle and busy warm-session shutdown, rejected
checkpoint propagation, provider-pack verification, and managed-Codex
credential preservation. Four managed credential detachment cases
reproduced the bug before the fix; normal owned cleanup still succeeds
exactly once.
- Live local Claude: SIGKILL 2.6 seconds into startup recovered the same
run automatically in 53 seconds, then a normal follow-up completed in 24
seconds. SIGTERM 2.5 seconds into startup preserved the same run (54
seconds) and its queued follow-up (21 seconds). Answers remained visible
and the task reached Done.
- Live Claude Daytona: a warm follow-up retained its sandbox and fell
from 121 seconds to 44 seconds. A separate cold turn took 127 seconds;
after controller shutdown and checkpointing, its follow-up completed in
33 seconds with the same sandbox, workspace, native session, and runner.
Both answers remained visible and the task was Done.
- Other live journeys covered task completion and follow-up with local
and Daytona Codex, local Codex crash recovery, Stop then new direction,
clarification response, live artifact refresh, and shared-workspace
waiting.
- Validation limits: the opt-in native composer Stop/Pause→subtree
Resume fixture exposes terminal/result ordering and subtree-cancellation
attribution bugs that can leave a child task blocked; that new finding
is assigned to a separate follow-up and is not claimed fixed here.
Default CI skips this optional native-provider fixture. Managed-Codex
credential handoff and the queue-agent integration use automated
regression evidence. Cold custom provider-pack uploads still add startup
latency.

## Risks

- Automatic replacement remains deliberately narrow: local Codex,
verified stopped identities, unchanged retained state, and a complete
text/completion-only turn. Unknown actions, partial history, or changed
ownership remain blocked.
- Claude completion permission handling changes an upstream package
patch. The exact isolated task bridge must remain pinned; unrelated
tools keep their existing permissions.
- New task failure projection changes user-visible status. No historical
status backfill or database migration is included.
- This is a broad lifecycle fix across server and UI. Live proof covers
graceful local Claude restart during startup and idle Claude Daytona
session recovery across controller shutdown. Live abrupt SIGKILL during
local Claude startup also recovered the same run. Unknown ownership or
missing action evidence still blocks reuse. Cold custom provider-pack
uploads still add startup latency; this change avoids unnecessary repeat
uploads.

## Model Used

OpenAI GPT-6 (Codex), with reasoning, code execution, browser
automation, and tool use. The exact hosted model ID and context window
are not exposed in this task.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-12 19:41:15 -05:00
DottaandPaperclip ed50a39c3f fix: preserve NUL characters in run-event payloads (#13325)
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-12 13:34:45 -05:00
DottaandPaperclip 51b0e01ead fix: resume saved user messages after execution recovery (#13270)
Preserve verified native process-stop evidence and retry saved user messages through normal continuation admission after recovery cleanup. Show the current wait reason and serialize delivery so a saved message starts one fresh turn.

Validated with 410 focused tests, typecheck, build, token gates, all PR CI checks, and Greptile 5/5.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-11 17:12:55 -05:00
DottaandPaperclip 3b550c80fa fix(codex): correct startup trust, history reads, and resume usage (#13110)
## Thinking Path

> - Paperclip runs Codex locally and in remote sandboxes.
> - The runner must preserve startup configuration and session identity.
> - Missing project trust can disable repository configuration.
> - Full-history requests use deprecated provider fields.
> - Resume usage describes old work and must not become new run usage.
> - This change corrects startup trust, state reads, and usage
classification.

## Linked Issues or Issue Description

**What happened?**

Normal Codex runs could show repository-trust and history-deprecation
warnings.
Resume could report the preceding turn's token snapshot as a late-turn
warning.
The historical last-usage value could also be attributed to the new run.

**Expected behavior**

Trust the server-selected startup root in isolated configuration. Read
lightweight
provider state and paginated evidence. Use historical cumulative usage
as a
baseline without a new charge or user-facing warning.

**Steps to reproduce**

1. Start a native Codex task in a selected repository.
2. Finish the turn and resume the provider thread.
3. Inspect provider notices, history requests, and per-run usage.
4. Repeat startup and cold resume inside a Daytona sandbox.

**Paperclip version or commit**

Codex CLI 0.153.4 is the pinned runtime and reproduced baseline.
Replayed onto master at 6abeb6733. Related authority work: Refs #13092.
This PR retains its startup cleanup and protocol-integrity checks.

**Deployment mode**

Local source checkout and disposable Daytona sandbox.

## What Changed

- Classify the exact historical resume usage event before the generic
stale-turn warning.
- Persist cumulative usage baselines across recovery of the same run.
- Use excludeTurns on resume and lightweight thread reads.
- Page turn metadata and selected turn items with cursor and identity
validation.
- Reject unsupported or incomplete history instead of guessing that
execution is idle.
- Trust the startup execution root on its host, including Git worktree
trust keys.
- Start Codex in that root and retain the selected sandbox profile on
later turns.
- Keep unrelated isolated configuration and Codex's separate hook trust
policy.
- Add Rust, TypeScript, accounting, native integration, and local
run-log documentation.

## Verification

- Codex and native-transport TypeScript: 333 passed before PR replay.
- Adjacent OpenCode/ACPX driver and accounting tests: 49 passed.
- Rust library, serialized: 226 passed. Native Codex integration: 72
passed, 1 ignored, plus two pagination regressions.
- Repository typecheck and build passed. All repository test groups have
passing coverage after fixture and resource retests; the initial
monolithic command was not clean.
- Fresh real Codex native browser tasks returned correct answers without
the three targeted notices. Answers persisted after refresh and restart.
- Real same-thread TypeScript driver tests passed locally and in
Daytona, including cold resume, configuration, skills, and an approved
harmless hook.
- Local usage summed to 64,607 tokens. Daytona usage summed to 42,737
tokens. Each sum matched its final session total exactly.
- See doc/plans/2026-09-09-codex-integration-acceptance.md for the scope
and limits of the live tests.
- After replay onto current master and review fixes: 334 Codex, backend,
and live-session tests passed, including checkpoint serialization and
real-runner process restart. TypeScript checks passed.
- The native Codex integration run passed 83 tests; the large lineage
test passed separately with the release runner (its debug build exceeded
the test deadline).
- All GitHub checks passed on the final PR head. Greptile is 5/5 with no
unresolved review threads. CI regenerates the lockfile for the added
TOML dependency, per repository policy.
- The first server shard hit a timing-dependent duplicate-key failure in
the unchanged artifact-document concurrency test. Its focused 11-test
suite passed locally. One CI retry on the same head passed all 103 files
and 1,405 tests (2 skipped): [retry
result](https://github.com/paperclipai/paperclip/actions/runs/34398832930/job/102631274667).

## Risks

- Trust applies only to the server-selected startup root and isolated
configuration. Sandbox and tool permissions remain authoritative.
- Codex still requires approval of individual hook hashes. This change
does not bypass that policy.
- Providers without the required history APIs fail explicitly.
- Daytona acceptance used the production TypeScript driver. Remote
Paperclip UI and remote Rust execution were not tested.
- No new public API, database state, recovery policy, or UI control is
included.

## Model Used

OpenAI Codex, GPT-6 (`gpt-6-astra`). Used for reasoning, code edits,
tool use,
and test execution. The exact context-window limit is not exposed in
this
session. Real-provider acceptance used Codex CLI 0.153.4 with
`gpt-5.6-sol`.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-09 15:35:18 -05:00
DottaandPaperclip 35fdc0c66b fix: make task recovery durable and preserve current requests (#13075)
Make task recovery durable and preserve the latest user request across native and legacy continuations. Keep routine recovery quiet and prevent replay when action outcomes are uncertain.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-09 09:14:25 -05:00
Dotta 7b094724e6 fix(runner): recover native sessions across restarts (#12845)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The Paperclip Runner keeps durable run and provider state outside
one server process.
> - A server restart can leave that runner alive or can interrupt it
after a provider checkpoint.
> - The old startup path used handoff intent and PID evidence, but it
did not reconstruct native ownership.
> - That gap could block the issue, create a replacement run, or start
duplicate provider work.
> - This pull request adds durable same-run recovery for coordinated and
uncoordinated restarts.
> - The benefit is exact recovery of the run, runner, session, provider,
steering, and finalization state.

## Linked Issues or Issue Description

Refs #9628. That pull request added earlier local-adapter hot-restart
work. This change adds native PRP authority reconstruction and same-run
provider resume.

Refs #10935. That pull request handles missing hot-restart snapshots.
This change also supports hard restarts with no snapshot.

Refs #11624. That pull request prevents unsafe retry after an adopted
legacy process exits. This change reconciles native terminal evidence
before provider recovery.

Refs #12070. That pull request improves process liveness checks. This
change also binds recovery to a process-start fingerprint and fails
closed on ambiguity.

**What happened?**

The server could record hot-restart intent, but startup did not rebuild
native runner ownership. A live runner could not re-register its PRP
authority. A dead runner could not resume the exact native and provider
session on the same heartbeat run. Generic recovery could then block the
issue or create replacement work.

**Expected behavior**

A live native runner must reconnect with the same PID and logical
identities. A dead runner must resume the same durable session and
heartbeat run with only a new operating-system PID. A proposed or
terminal result must finalize once before any provider turn starts.
Ambiguous process or session evidence must stay blocked without a signal
or duplicate spawn.

**Steps to reproduce**

1. Start a Paperclip Runner heartbeat and wait for an active provider
turn.
2. Restart only the Paperclip server, with or without a hot-restart
marker.
3. Observe that the old startup path does not reconstruct the native
control-plane authority.
4. Kill both the server and runner after a provider checkpoint.
5. Observe that the old path cannot resume the exact native session on
the original heartbeat run.

**Paperclip version or commit**

The defect was reproduced from commit
`1991f31fd53e7f7794d5c2e4b93be384ade2b41d`. This branch is rebased onto
the current `master`.

**Deployment mode**

Local development and self-hosted server deployments that use the local
Paperclip Runner.

## What Changed

- Added correlated hot-restart requests and version-compatible native
handoff fields.
- Added controller boot identity, process-start identity, controller
generation, recovery state, request id, and bounded history to the
native finalization ledger.
- Added transactional recovery claims for live-runner reattach,
dead-runner resume, and incomplete bootstrap.
- Added fail-closed ownership takeover rules and process identity
validation.
- Added live runner adoption to the local runner transport without a
duplicate spawn.
- Added same-run provider checkpoint resume and legacy retry-row
compatibility.
- Reconciled proposed and terminal results before runner or provider
recovery.
- Bound the HTTP and PRP listener before startup recovery and delayed
scheduling and generic reapers until classification completes.
- Added restart-aware health diagnostics, run-log recovery transitions,
durable runner diagnostics, and bounded shutdown finalizer draining.
- Moved restart-survivable diagnostics into runner-owned, pre-redacted
bounded writes; raw stdout and stderr are never persisted.
- Added process-start fencing for controller, runner, and provider PIDs;
startup classifies every candidate without an implicit cap.
- Added crash-recoverable, contention-safe development restart-request
coordination and failed-startup listener cleanup.
- Added a credential-free real-process restart suite for eight restart,
scale, and identity scenarios.
- Documented native restart operation, persistence, diagnostics, and
verification.

## Verification

- The documented native restart commands passed. They ran eight
real-process/database recovery scenarios and the live runner adoption
transport test.
- Native executor tests passed: 111 tests.
- Heartbeat recovery tests passed: 124 tests.
- Hot restart, health, and shutdown tests passed: 52 tests.
- The broader affected server suite passed: 350 tests.
- Focused native recovery and startup tests passed: 49 tests.
- Runner transport and control-plane tests passed: 63 tests.
- Runner-owned diagnostic tests passed for write-time bounding,
credential redaction, private file modes, and raw stream
non-persistence.
- Development restart coordination tests passed: 11 tests.
- Database migration checks and the partial-application/replay
regression test passed.
- Server, database, and Paperclip Runner typechecks passed.
- `git diff --check` passed.
- Full Paperclip PR CI passed, including build, canary, all five general
server shards, all five serialized server shards, all three browser E2E
shards, workspace suites, and release-registry verification.
- Greptile completed at 5/5 with no outstanding findings,
recommendations, follow-ups, or open review threads.

## Risks

- Moderate risk. This changes startup ordering and ownership transfer
for active native runs.
- The migration adds nullable columns and does not rewrite existing
rows.
- Recovery fails closed when process or durable session identity is
incomplete or contradictory.
- The first implementation supports the local Paperclip Runner. Remote
targets keep their existing behavior.
- The real-process suite covers cleanup and asserts that no runner or
provider process survives each test.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- OpenAI Codex with GPT-5. The runtime did not expose a more specific
model revision or context-window size. Repository editing, shell
execution, database tests, and real-process test execution were enabled.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-04 15:03:53 -05:00
Dotta 9964b034bb feat(runner): add hidden server PRP coordinator (#12176)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The Paperclip Runner needs a narrow server trust boundary before an
adapter can start it.
> - The package has durable runner transport, but the server does not
host or authorize that transport.
> - Native persistence exists, but no writer connects PRP events to
those records.
> - A direct adapter must not enter this path by accident.
> - This pull request adds a hidden, run-bound PRP server coordinator.
> - The benefit is a recoverable server boundary that remains
unavailable to normal execution.

## Linked Issues or Issue Description

Refs #11962

Refs #12129

Refs #12169

**Subsystem affected**

Cross-cutting. The change affects the runner package and server
orchestration.

**Problem or motivation**

The server cannot authenticate runnerd, commit PRP events before ACK,
authorize semantic tools, or enter native finalization from a durable
runner result. The application must have this hidden boundary before a
guarded adapter can use the runner.

**Proposed solution**

Add an authenticated PRP WebSocket authority and register it only for
one exact persisted native Codex run. Bind each connection and event to
the company, issue, agent, run, runner, session, turn, item, and
verified runner identity. Commit each event before its cumulative ACK.
Project only authorized same-task read tools. Rebuild the accepted
result and finalization record from durable result and terminal events.

**Alternatives considered**

The server could expose a broad runner API key or route semantic calls
through existing adapter endpoints. Those options grant too much
authority and weaken replay recovery. The server could also add the
user-facing adapter in this pull request. That option would mix rollout
selection with the transport trust boundary and make legacy
compatibility harder to review.

**Roadmap alignment**

This work supports the shipped enforced-outcomes, governed-tool, and
self-healing-run milestones. It does not add a new roadmap surface.

## What Changed

- Add the durable PRP server authority with one-use bootstrap tickets,
reconnect leases, encrypted frames, bounded state, cumulative ACKs, and
idempotent commands.
- Add `/api/runner/v1/connect/:runId`. Derive its `ws://` or `wss://`
URL from the configured Paperclip API URL.
- Register one authority only after the coordinator verifies the
complete native Codex run binding.
- Commit validated PRP events to `heartbeat_run_events` before ACK.
Reject source gaps and conflicting replays.
- Rebuild accepted results and finalization records from durable result
and terminal events. Enforce finalization owner leases and retry times.
- Project five same-task read operations. Recheck run, agent, task, and
company authority for each call.
- Keep the route hidden. No adapter selects this coordinator, and no
code starts runnerd.
- Vendor the compiled runner TypeScript runtime into the server package
while keeping the workspace package development-only for the server.
- Document the package, database writer, run-log payload, and credential
exclusions.

## Verification

- Run `pnpm --filter @paperclipai/paperclip-runner check:all`. All
TypeScript protocol checks and 69 Vitest tests pass, including
commit-before-ACK crash recovery. All 43 Rust unit tests and 13 Rust
integration tests pass. Conformance and replay parity pass.
- Run the focused server WebSocket, coordinator, package-build, and
startup-wiring suites. All 26 tests pass, including a clean-checkout
reproduction with the runner `dist` directory absent.
- Run `pnpm -r typecheck`.
- Run `pnpm test:run`.
- Run `pnpm build`.
- Confirm that the diff contains 19 files. Confirm that it contains no
workflow or `pnpm-lock.yaml` change.

## Risks

- The server installs the WebSocket route at startup. An unregistered or
malformed run path fails closed and creates no native record.
- Bootstrap tickets are one use. The private state directory uses mode
`0700`, and the state file uses mode `0600`. The file stores derived
authentication verifiers and never stores raw tickets or lease tokens.
- The journal has explicit frame, command, event-window, and file-size
bounds. A bound violation closes the runner connection or rejects the
command.
- A runner event reaches the database before its ACK. A crash between
event commit and ACK causes a byte-equivalent replay, not a second
logical effect.
- The coordinator accepts only an existing queued or running native
Codex row with exact company, task, agent, runner, session, and
completion-contract ownership.
- Existing direct adapters do not call this service. They keep their
current execution, transcript, result, and finalization paths.
- The server has no production dependency on the private runner package.
Its build copies the compiled runtime into `server/dist`; the workspace
link is development-only. This adds no external package and does not
change the lockfile.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- OpenAI Codex with GPT-5. The exact deployment ID and context-window
size are not exposed. The model used agentic reasoning, repository
tools, code execution, and test execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-08-25 14:17:14 -05:00
Nicky LeachandPaperclip d1573244b5 refactor: disambiguate the Telemetry and Observability data paths (#12128)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Paperclip records first-party events, OpenTelemetry data, and local
run-log events
> - The code and documents used one term for these three data paths
> - This naming made the required review level unclear
> - This pull request names each data path in the module names,
documents, and code comments
> - The benefit is a clear review rule without a runtime change

## Linked Issues or Issue Description

**Issue type**

Unclear or confusing.

**Where is the issue?**

`packages/shared/src/telemetry/README.md`, `doc/observability.md`,
`doc/run-log-events.md`, and the duplex instrumentation modules.

**What's wrong?**

The repository used Telemetry for first-party events, OpenTelemetry
data, and local run-log events. This usage made the data path and review
level unclear.

**Suggested fix**

Use Telemetry only for Paperclip first-party events. Use Observability
for OpenTelemetry data. Use the run log for rows in
`heartbeat_run_events`.

Related public pull requests: #8476 and #9672.

## What Changed

- Rename the duplex instrumentation modules and identifiers from
`Telemetry` to `Observability`.
- Move the Observability and run-log contracts out of the Telemetry
README.
- Add `doc/observability.md` and `doc/run-log-events.md` as the
canonical documents.
- Add a file-path review rule to `AGENTS.md`.
- Correct the remaining code comments that name the wrong data path.
- Keep all event names, payloads, database records, spans, configuration
keys, environment variables, and runtime paths unchanged.

## Verification

- `npx vitest run packages/shared/src/telemetry/readme-contract.test.ts`
passes.
- `npx vitest run packages/adapter-utils/src/published-exports.test.ts`
passes.
- `npx vitest run
packages/adapter-utils/src/acpx-engine/startup-timing.test.ts` passes
with 42 tests.
- `pnpm --filter @paperclipai/adapter-utils typecheck` passes.
- `pnpm --filter server typecheck` passes.
- The old module name does not remain in TypeScript or JSON files,
except for the intentional publication guard.
- CI and Greptile checks remain pending after PR creation.

## Risks

- The old duplex module subpath no longer has a compatibility shim. The
board accepted this intentional hard break.
- The new duplex module subpath stays blocked from package publication.
- The change has no runtime effect. The main risk is an incorrect
document or module reference.

## Model Used

OpenAI GPT-5 Codex, exact model ID `gpt-5`, with tool use and code
review support.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have described the issue in-PR with the documentation issue
fields
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-08-24 16:42:33 -07:00