Commit Graph
1883 Commits
Author SHA1 Message Date
DottaandPaperclip 29230e2000 fix(server): fence native retry cancellation through final commit
Recheck retry eligibility under the coordinator lock before creating a cancellation intent. Preserve later run and coordinator outcomes with a failed-only acknowledged-result CAS, while retaining same-intent recovery and NOWAIT conflict handling.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-30 13:54:11 -05:00
DottaandPaperclip 79db1c2a6f fix(server): admit correlated Stop for durable native retries
Check failed-run retry eligibility under the run and coordinator locks, fail closed on concurrent coordinator claims, and preserve same-caller recovery after the audited intent disables a retry. Keep terminal failures and foreign actors or intents rejected.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-30 13:35:21 -05:00
DottaandPaperclip 912dbfd54b fix(native): bind operator Stop to caller cancellation intent
Reserve an optional board request UUID under the run lock and retain it
through default Stop joins and native dispatch. Reject prior or competing
intents and require the same actor for idempotent retries.

Bind the Copilot denial fixture to its exact request and audited intent,
with versioned causal receipts and negative controls for earlier Stop.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-30 12:02:57 -05:00
DottaandPaperclip c58a8881f2 fix: validate Cursor plan waits against canonical contract policy
Reuse the production completion-contract envelope hash and exercise real contract creation and reuse in accepted-plan persistence tests.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-30 02:54:16 -05:00
DottaandPaperclip a2984a0a9d Preserve historical Cursor plan proof query bounds
Retain the original Cursor6 event filter when verifying an exact committed wait. Progress rows remain part of Cursor7 lifecycle proof without consuming the historical query budget.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-30 02:15:04 -05:00
DottaandPaperclip 770dc88a11 Bind Cursor plan waits to the native tool lifecycle
Carry the admitted native plan parent tool into canonical request identity and require its exact successful durable lifecycle before passive settlement. Preserve historical committed Cursor6 waits while versioning new admission to Cursor7.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-30 01:56:28 -05:00
DottaandPaperclip 4b94270a0b Await clean Stop continuation settlement in recovery test
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-30 01:56:28 -05:00
DottaandPaperclip ade7c83dbb Preserve committed Cursor plan waits across profile upgrades
Keep current profile qualification mandatory when first creating a wait. Recovery uses its scoped committed receipt to verify the entire original proof, so future catalog/settings changes cannot authorize task work. Cover actual persisted finalization, catalog drift, original-profile tampering, and document explicit user continuation.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-30 01:56:28 -05:00
DottaandPaperclip 72a88e124d Keep accepted Cursor plans waiting for explicit continuation
Bind normal native plan acceptance to its durable request, delivered answer, admitted run and completion contract. Commit a passive in-progress result with a visible next-message summary, preserve semantic finish priority, and suppress recovery until task-specific continuation without changing modes.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-30 01:56:28 -05:00
DottaandPaperclip f401b831f2 Stabilize route setup and attachment announcement handling
Load route modules during fixture setup so cold transforms cannot leave a timed-out request running against the next test mock state. Dismiss delayed product announcements through the normal Board UI in the attachment receipt fixture.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-30 01:32:15 -05:00
DottaandPaperclip 41503eb38f fix: avoid fresh attachment staging paths for empty wakes
Keep authorized residue scrubbing and unsafe-path checks while leaving untouched workspaces unchanged when no wake attachments were selected.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-29 23:16:23 -05:00
DottaandPaperclip 4597967f89 Handle durable cancellation before workspace failure recovery
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-29 23:15:13 -05:00
DottaandPaperclip 951b06d08a fix(runner): fence warm ACPX owners on runtime contract changes
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-29 21:40:17 -05:00
DottaandPaperclip e154e85394 test: scope Telegram lease fault to its fixture company
Fence the credential-reclamation hook used by the global recovery sweep so it cannot leave synthetic lease owners on earlier companies. Exercise a paused foreign receipt after the fault is armed and verify it settles without a leaked lease.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-29 20:12:20 -05:00
DottaandPaperclip ceb28bdba4 fix(server): bind Cursor mode at recovery identity boundaries
Reject missing, foreign, or changed observed modes before accepting suspended checkpoints or interaction-driven identity transitions.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-29 16:27:18 -05:00
DottaandPaperclip 17c6da5e32 feat(runner): carry explicit Cursor mode through configuration and transport
Preserve candidate models, reject foreign mode settings, and bind Cursor mode into session identity plumbing. Native mode admission and Rust validation follow in coordinated commits.

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-29 16:23:31 -05:00
DottaandPaperclip 4f2709a611 test(runner): check App icon parity at the server boundary
Keep both live and catalog schema assertions through public runner exports while removing the standalone suite dependency on App source.

Paperclip-Task: rich-acp-production

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-29 13:42:27 -05:00
Dotta d2734a81cb Merge ACP qualification test context and runtime readiness fixes
Paperclip-Task: rich-acp-production

* commit 'ed3a0df1c':
  test: await connection setup and persisted runtime identity
2026-09-29 13:07:17 -05:00
DottaandPaperclip ed3a0df1c1 test: await connection setup and persisted runtime identity
Complete the dialog fixture context and wait for the actual setup render. Observe the post-spawn PID update before checking the starting service identity, preserving provisioning and readiness assertions.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-29 13:04:12 -05:00
DottaandPaperclip 482471f0c1 Merge reviewed shared ACP input and continuation fixes
Co-Authored-By: Paperclip <noreply@paperclip.ing>

* codex/runner-acp-inputs:
  fix(runner): refresh settled ACPX run grants during authenticated attachment
  fix(runner): align editable draft bounds with Unicode schema lengths
  fix(ui): preserve canonical drafts in classic question cards
  Keep the legacy question fixture payload type-safe
2026-09-29 12:40:14 -05:00
DottaandPaperclip bb56af7fbd Keep the legacy question fixture payload type-safe
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-29 12:28:36 -05:00
DottaandPaperclip c5a4714add feat(runner): preserve editable question defaults through explicit response
Carry bounded text-only initialText through ACP normalization, Rust durable validation, shared contracts and the question UI. Saved edits and explicit clears survive reload. Preserve canonical submitted text while retaining legacy custom-answer normalization and existing validation.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-29 12:07:19 -05:00
DottaandPaperclip 012aa52c77 feat(runner): preserve editable question defaults through explicit response
Carry bounded text-only initialText through ACP normalization, Rust durable validation, shared contracts and the question UI. Saved edits and explicit clears survive reload. Preserve canonical submitted text while retaining legacy custom-answer normalization and existing validation.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-29 12:06:48 -05:00
DottaandPaperclip 97217c5313 Integrate reviewed rich ACP instruction and startup fixes
Co-Authored-By: Paperclip <noreply@paperclip.ing>

* codex/runner-pi-acp:
  Bind Cursor instructions to native rules and gate every ACP connection
  fix(copilot): refresh instructions under provider lifetime ownership
  fix(copilot): deliver native instructions with v4 session identity
  docs(runner): retain Pi native prompt delivery proof
  fix(runner): refresh trusted system instructions before ACPX restore
  fix(runner): refresh runtime context when restoring ACPX sessions
  test(copilot): record native system instruction delivery evidence
  fix(runner): refresh trusted system instructions before ACPX restore
  perf(runner): hash Pi runtime inventory with bounded workers
  perf(runner): bound native snapshot copy concurrency at 32
  perf(runner): bound native snapshot copy concurrency at 32
  test(copilot): preserve final question failure and offline recovery proof
  fix(runner): refresh runtime context when restoring ACPX sessions
  feat: return completed handoffs to Agent Chat (#14408)
  fix(adapters): preserve ACP terminal failure diagnostics (#14573)
  fix(runner): honor Codex effort selected in composer (#14568)
  fix: grade Codex clarification and refusal outcomes from evidence (#14570)
  fix(inbox): keep other users’ failed runs out of Mine (#14572)
2026-09-29 11:33:49 -05:00
DottaandPaperclip ad472a7755 Retry bounded company prefix collisions in chat fixtures
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-29 11:28:48 -05:00
DottaandPaperclip 1d02ab77b2 fix(runner): refresh trusted system instructions before ACPX restore
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-29 11:09:17 -05:00
DottaandPaperclip 83076d7e7c feat: return completed handoffs to Agent Chat (#14408)
Return completed Agent Chat handoffs through a durable outbox and scope each generated update to its supplied tasks. Add recovery, browser delivery, result access, and calibrated quality coverage.

Validated with two consecutive ten-case Claude/Codex campaigns, all CI checks, and a 5/5 review.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-29 10:25:52 -05:00
DottaandPaperclip 3b4b270650 fix(adapters): preserve ACP terminal failure diagnostics (#14573)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The shared ACP adapter engine records agent failures for operators.
> - ACP providers can report a failure category, title, and detailed
cause.
> - Our patch kept only the category in the saved error, so an operator
could not diagnose a failure when tracing was off.
> - This pull request preserves redacted provider diagnostics in the run
error, transcript, and structured run result.
> - Operators can now inspect the provider message and any supplied
request ID or stack trace after the run ends.

## Linked Issues or Issue Description

Refs #13889 (the diagnostic gap; this PR does not update the bundled
Claude version).
Refs #14484 (related model-refusal classification; this PR retains
diagnostics for all terminal failure categories).

**What happened?**
An ACP turn failed with only `ACP agent reported a terminal service
failure.` The provider's title and details were available in memory but
absent from the saved error and transcript.

**Expected behavior**
The run retains useful provider diagnostics even when raw tracing is
disabled. Credentials remain redacted. A size limit must report
truncation instead of silently removing the cause.

**Steps to reproduce**
1. Run an ACP agent that returns an error-severity typed session
failure.
2. Include an HTTP error, request ID, and stack text in its title and
details.
3. Inspect the failed run with tracing disabled. Before this change,
only the category survives.

## What Changed

- Both pinned ACPX patches pass complete error text to the in-memory
callback, so redaction happens before truncation.
- The shared engine retains the sanitized category, title, and details
in `resultJson.terminalSessionFailure` and includes the text in the run
error and error transcript.
- Diagnostics redact configured environment values even under arbitrary
names, unknown launch-environment values, connection URL passwords, run
credentials, and common credential syntax. Known boolean settings remain
readable, while credential values are redacted even when embedded in
other text. Diagnostics remove control characters and invalid Unicode.
- Title and detail limits keep escaped transcript JSON below the
server's chunk limit. Truncated fields include an omission count. The
safe run-result projection preserves a byte-bounded diagnostic preview
when the result exceeds its byte budget, with an explicit pointer to the
full adapter-bounded run error and transcript.
- The existing UI and CLI display the error. Diagnostics do not become
assistant output. Issue continuation summaries and session-compaction
prompts receive only the generic category, preventing provider text from
becoming handoff instructions. Existing quota classification, warnings,
timeout precedence, and control-channel failure precedence remain in
place.
- Regression tests cover real ACP child processes with both pinned
versions in one-shot and persistent modes, credential redaction, request
IDs after the old 4 KiB cutoff, transcript parsing, storage bounds, and
database retrieval of oversized multibyte diagnostics.

## Verification

- Full CI on `20ad4f5f1f66c46d2c260e6ad0339cbea607b4cf`: **54 passed, 2
intentionally skipped, no pending or failing checks**. Includes
typechecking, build, all Vitest shards, Runner checks, browser E2E, and
the canary packaging/public-install dry run.
- Greptile: **5/5** on this commit. Superagent security scan passes. All
review threads are resolved.
- Local verification passed: shared ACP engine suite (395 tests); real
Claude ACP child-process and diagnostic regressions across both pinned
runtimes and both execution modes; run retrieval and model-handoff
regressions (59 tests); ACPX patch packaging (16 tests); full typecheck
and build. Affected package typechecks and focused tests were rerun
after review fixes.
- The broad local `pnpm test:run` was stopped after review edits made
its cached imports stale. Fresh targeted runs pass, including both
affected server suites. Cold-build import failures were also rerun after
dependency builds: chat integration (1,063 tests) and tool access (351
tests) pass. The final commit's complete CI matrix is green.

## Risks

- Provider diagnostic text is untrusted. This change retains more of it
in company-scoped run records. Redaction and size bounds apply before
persistence.
- Diagnostics are limited to fields the provider supplies. Old runs
cannot recover discarded error text.
- No schema migration, recovery-policy change, or new Telemetry or
OpenTelemetry export.

## Model Used

- OpenAI GPT-6 through Codex, with reasoning, repository inspection,
code editing, and test execution. The exact serving model ID and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-29 10:18:30 -05:00
da887ea3e9 fix(runner): honor Codex effort selected in composer (#14568)
## Thinking Path

> - Paperclip manages AI agents that work on assigned tasks.
> - The task composer lets a person choose an assignee, model, and
effort for the next run.
> - A Paperclip Runner agent can use Codex as its provider.
> - The composer hid Codex effort for that agent because it checked only
the older Codex adapter.
> - The native Runner input also did not carry an effort choice to
Codex.
> - This pull request carries the chosen effort from the composer to
each Codex turn.
> - People can now select a supported effort and get the effort they
selected.

## Linked Issues or Issue Description

Refs #14322

**What happened?**

The composer showed a model but no effort slider when the assignee used
Paperclip Runner with the Codex provider. A task-level model override
also did not reach the native Runner input.

**Expected behavior**

The composer shows effort choices for a known Codex model. The next
native Codex turn uses the selected model and effort.

**Steps to reproduce**

1. Open a task composer.
2. Select an agent that uses Paperclip Runner with the Codex provider.
3. Select a known Codex model such as `gpt-6-astra`.
4. Open the assignee and model picker. The effort slider is missing
before this change.

## What Changed

- Show known Codex effort levels for Paperclip Runner Codex assignees.
- Save the task effort override in the native run input and send it to
Codex on each turn.
- Apply the task's merged model and effort overrides when the native run
starts.
- Apply a task model override for OpenCode Runner without changing the
agent's provider.
- Add Runner effort tests and desktop and mobile Storybook cases.

## Verification

- `pnpm -r typecheck` passed.
- `pnpm build` passed.
- `pnpm build-storybook` passed.
- `pnpm check:token-gates` passed.
- Focused UI, server, Runner contract, and Codex driver tests passed.
- The full CI test matrix, build, typecheck, and canary dry run passed
on the latest head.

## Risks

- Native Runner inputs add an optional Codex effort field to the current
v5 input. Older inputs keep their previous behavior.
- A known model rejects an effort that its catalog does not support.
Unknown models do not show a slider.

> This fixes an existing composer bug. I checked `ROADMAP.md`; it does
not describe this bug as planned work.

## Model Used

OpenAI Codex, GPT-6. The exact deployment ID and context window are not
exposed in this session. The model used reasoning, code execution, and
repository tools.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: OpenAI GPT-6 Astra <noreply@openai.com>
2026-09-29 09:41:05 -05:00
DottaandPaperclip 7636966452 fix(inbox): keep other users’ failed runs out of Mine (#14572)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The Mine inbox shows work that needs the current user.
> - Failed-run rows used the latest run for every agent in the company.
> - A failure from another user therefore appeared in Mine and its
badge.
> - Run list responses also omitted the responsible user needed to
filter these rows.
> - This pull request uses run ownership for personal failure routing.
> - Users see their own failures and can still inspect company failures
in All.

## Linked Issues or Issue Description

**What happened?**

An agent run started for one user failed. Its row and failure badge
appeared in another user's Mine inbox.

**Expected behavior**

Mine and its badge include failed runs for the current responsible user.
Other users' failures remain available in All and run details.

**Steps to reproduce**

1. Use a company with two human users.
2. Create a failed or timed-out run attributed to the first user.
3. Open Mine as the second user. Before this fix, the failed run appears
there and increases the badge.

**Paperclip version or commit**

Reproduced in regression tests on master at `24beb0057`.

**Deployment mode**

Authenticated deployment with multiple users. Tests also cover the local
single-user board.

Related prior work: #933 addressed inbox dismissal and badge
consistency. No duplicate ownership fix was found.

## What Changed

- Return `responsibleUserId` in normal and summary run lists.
- Share one ownership rule across both inbox versions and client/server
badges.
- Select the latest run per agent before applying the ownership filter.
This prevents old failures from resurfacing on shared agents.
- Keep unattributed historical failures in the local board's Mine view.
Hide them from authenticated users with no matching owner.
- Keep company health alerts outside the personal badge, consistent with
the client.
- Document the routing contract and add page, badge, and database
regression coverage.

## Verification

- Red: the new badge cases failed with three company failures instead of
one personal failure; eight Mine page cases failed across both inbox
versions.
- Green: 113 focused tests pass in `ui/src/lib/inbox.test.ts`,
`ui/src/pages/Inbox.test.tsx`,
`server/src/__tests__/heartbeat-list.test.ts`, and
`server/src/__tests__/inbox-dismissals.test.ts`.
- `pnpm check:token-gates` passes.
- Agent calls on behalf of a user have two additional red-to-green API
regressions.
- Full `pnpm -r typecheck` and `pnpm build` pass. Server typecheck also
passes after the agent-call fix.
- All CI test shards and browser tests pass on `243bfa681`. The
duplicate local `pnpm test:run` was stopped after the CI test lanes
completed; it did not finish locally.

## Risks

- Authenticated users no longer receive unattributed legacy failures in
Mine. Those failures remain visible in All.
- The server badge no longer counts company health alerts, matching the
existing client badge.
- No migration, run state, retry behavior, or company access rules
change.

## Model Used

- OpenAI GPT-6 through Codex, with reasoning, terminal execution, and
browser tools. The exact deployment variant and context window size are
not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-29 09:40:21 -05:00
DottaandPaperclip 24beb00575 feat(runner): add rich ACP transport and durable interaction foundation (#14430)
Add shared rich ACP transport, durable questions and permissions, verified provider packaging, and bounded activity and plan presentation. Keep Cursor, Copilot, and Pi pending their separate provider qualification.

Persist interaction settlement before publication, fence failed writes until fresh recovery, and preserve owned-process cleanup. Incorporate reviewed mainline integration with extended harness coverage.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-29 08:56:21 -05:00
DottaandPaperclip c9b93d7e8c fix: preserve terminal task owners during release (#14561)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Tasks record an assigned owner and separate checkout and execution
locks.
> - A completed task must retain its owner after execution ends.
> - The release endpoint currently clears that owner when it clears the
locks.
> - This pull request preserves the assignee of Done and Cancelled tasks
during release.
> - Unfinished tasks keep the existing relinquishment behavior.
> - The benefit is stable task attribution without retaining execution
locks.

## Linked Issues or Issue Description

**What happened?**

An agent completed an assigned task, then called the release endpoint.
The task stayed Done, but its assignee became null. The same defect
affects Cancelled tasks. It caused the legacy Claude clarification/reuse
and multiple-repository handoff E2E assertions to fail.

**Expected behavior**

Release must clear execution locks on terminal tasks and preserve their
assignee, final status, and disposition timestamps. Release of
unfinished tasks must still clear the agent assignee. Only In Progress
work returns to Todo.

**Steps to reproduce**

1. Create an assigned task with checkout and execution locks.
2. Complete or cancel the task.
3. Call `POST /api/issues/:id/release` as the assigned agent.
4. Read the saved task. Before this fix, its assignee is null.

**Paperclip version or commit**

Reproduced on master commit `d172197117a14b80a1eb2d2835a0e7cce2679656`.

**Deployment mode**

Local tests against real PostgreSQL through the production issue routes
and services. This is a core lifecycle defect, independent of the agent
adapter.

Refs: #11689, #6899, #7769. These are related open release proposals.
This is an independent fix limited to terminal task ownership. It does
not include timer scheduling changes.

## What Changed

- Preserve the current assignee when releasing Done or Cancelled tasks.
- Keep all execution-lock cleanup and existing unfinished-task behavior.
- Cover all seven task statuses through the release API and read back
saved state.
- Check disposition timestamps, activity attribution, and repeated board
cleanup.
- Update the API contract, agent reference, and CLI help.

## Verification

- Red commit `ddaabb754`: the two terminal-owner regressions failed with
`assigneeAgentId: null`; 12 other route tests passed.
- Green: all 14 route tests pass, plus the existing successor-checkout
race test (15 selected tests total).
- Command: `pnpm --filter @paperclipai/server exec vitest run
src/__tests__/issue-stale-execution-lock-routes.test.ts
src/__tests__/issues-service.test.ts -t 'stale issue execution lock
routes|does not let stale release clobber a successor checkout lock'`.
- The local host has exhausted its SysV semaphore pool. The red/green
runs used the existing test-provider hook to start disposable Docker
PostgreSQL 17 instances. Routes, services, migrations, and assertions
were unchanged. No database tests in the selected set were skipped. The
other 134 tests were excluded by the name filter.
- Capability contract and inventory drift checks pass.
- `pnpm build` and `pnpm -r typecheck` pass.
- The local `pnpm test:run` was interrupted after environment failures
while the complete sharded CI suite ran in parallel: native PostgreSQL
bootstrap fails under the host semaphore limit, and the large Git
fixture hits macOS `ENAMETOOLONG`. A focused rerun confirmed these
happen before the relevant assertions. The interrupted local run is not
counted as a full pass.
- Greptile completed on `1caeeb827e9cb658ddb71f16c2421ec20f80634e` with
**5/5**, a successful check, and no review threads.
- All CI gates pass on the current head: typecheck, build, general and
serialized tests, Runner checks, browser E2E, release packaging, and
security checks. Server shard 11 passed on one targeted retry; the first
attempt had 836 passing tests but an unhandled workspace-runtime startup
rejection caused by an existing timing window. All other successful jobs
were reused.
- No paid provider evaluations were run.

## Risks

- A caller that used release to erase ownership from terminal work will
now retain that owner. An explicit assignment update or the board
force-release option with `clearAssignee=true` can still clear it.
- No schema or migration changes. The transaction, company access,
assignee/run checks, and activity log remain in place.

## Model Used

- OpenAI GPT-6 through Codex, with tool use, code execution, and test
debugging. The session does not expose a more specific backend model
version or context-window size.
- OpenAI `gpt-6-luna` assisted with read-only test discovery and review.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-29 08:47:33 -05:00
DottaandFry 3ca196b0a6 feat(agents): persist agent files across tasks without revision history (#14420)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - An agent needs personal files across tasks and sessions.
> - AGENTS.md is one file in that directory. Supporting files need the
same persistence.
> - The Instructions Editor and agent runs must share one current
directory.
> - Concurrent runs should apply only the files they change. The last
sync of the same file wins.
> - This pull request uses existing file transport and removes temporary
copies after sync.
> - Old instruction-only sessions keep their restore contract. New saves
do not create revision history.

## Linked Issues or Issue Description

Refs #14325. This replaces its revision-oriented design with persistent
agent files. Keep #14325 unmerged.

Transport prerequisite #14416 merged first at
`d172197117a14b80a1eb2d2835a0e7cce2679656`. This PR now targets master
and remains below 100 changed files.

Related work: #4513 and #8798 cover instruction tooling. This change
handles run synchronization, cross-task personal files, browser editing,
and old-session restoration.

## What Changed

- Keep one current directory per company and agent. Point AGENT_HOME at
a temporary working copy for each active run. Keep task files and
provider HOME separate.
- Restore text, binary files, and nested folders through workspace
transport. Exclude remote agent files from task Git snapshots with a
self-ignoring file inside the reserved runtime directory; never write
through repository-controlled Git metadata.
- Collect after the provider and child processes have stopped. Keep
resumable conversation state.
- Apply changed and deleted files under the agent lock. The last sync
wins for the same file. Unrelated concurrent changes survive.
- Remove temporary copies after successful sync, rejected sync, and
staging failure. Register ownership before copying so restart recovery
can remove interrupted preparation. Retry transient synchronization up
to three times. Preserve the original remote lease reference until
deletion succeeds; restart cleanup never acquires a replacement sandbox.
Do not create captured directories or a conflict-review queue for new
runs.
- Keep browser editing, stale-draft protection, and streaming binary
downloads. Keep the instruction entry and text editor limited to 1 MiB.
- Keep historical agent-folder sync failures on their affected runs
instead of repeating them above current saved instructions. Preserve
legacy candidate review and current browser-save errors. Avoid duplicate
quota warnings while retaining separate sync failures when they describe
a different problem.
- Require target-scoped caller grants for peer instruction access, while
preserving self edits, responsible-user checks, and protected-change
consent.
- Treat full storage as a nonblocking run warning, never an agent pause
or run-admission failure. Restore already-over-quota saved folders so
ordinary agent cleanup can recover; warn on each run until cleanup. The
run detail view shows the warning.
- Allow 256 MiB per file, 2 GiB per directory, and 100,000 entries. Hash
large files as streams. Check editor-save quotas with metadata instead
of hashing unrelated files.
- Preserve old native inputs, instruction-only copies, paths, digests,
and pending legacy candidates. Adopt old revision heads once. New writes
do not append history rows.
- Add idempotent migration 0287 and verify upgrades from the preview
tables and receipts.
- Add nine interactive stories under **Agents / Persistent files**,
including automatic incoming edits, stale browser drafts, and
storage-limit diagnostics.

## Verification

- Merge candidate: `4f5390107ec6ffd80a76d1d2e85530e66f21d079`, after
merging current master and the landed transport prerequisite.
Integration required no manual conflict resolution; the feature remains
99 changed files. Full workspace typecheck, production build, token
gates, and 715 focused tests passed on this merge candidate. Fresh
Greptile review is 5/5 with no unresolved findings. All 55 checks
passed, with four conditional skips, including the build, typecheck,
browser E2E, and canary dry run. A single retry recovered four jobs
interrupted by runner shutdowns; no source changes were required.
- Historical-warning UI fix: all 6,834 UI tests across 640 files passed,
including regression coverage for three old failures, legacy preserved
edits, and warnings scoped to the affected run. Full workspace
typecheck, production build, Storybook build, and token gates passed.
Browser-verified Storybook playtests passed for Historical Failures
After Successful Save, Storage Limit, and Full Storage Run Warning.
- Review follow-ups at `4e20c9fb2`: all 18 focused tests passed,
including external Git directories, linked worktrees, symlinks,
hardlinks, and distinct I/O failures alongside storage warnings. Server
and UI typechecks, token gates, and the production build passed.
- Storage warning regressions at `0724f3012`: all 33 directory tests and
all five heartbeat-list tests passed, with no skips in their successful
runs. They cover repeated runs while full, an already-over-quota saved
folder, cleanup, warnings retained after unrelated save failures, and
bounded warnings in large result JSON. Server typecheck passed after the
final warning fixes.
- Full workspace typecheck, production build, and token gates passed
during this follow-up. Product E2E harness: 631 tests passed across 52
files; harness typecheck passed. Earlier native session/context and
directory/legacy collection suites passed 537 tests; Runner
unit/transport suites passed 329 tests.
- **Real E2E at `0724f3012` (before this follow-up):** legacy local
Codex and native Daytona Codex each passed six tasks, one server
restart, seven independent assertions, and cleanup verification. Both
prove browser-to-agent edits, agent-to-browser edits, nested/binary
restoration, per-file last-sync-wins, a successful run after an
oversized save rejection, and cleanup clearing the warning.
- Native local Codex also passed the six-task quota flow before the
final warning-retention fixes. That pass began at `918d1ed02` while the
bounded-result warning fix was being edited, so it is not claimed as
exact-final-head evidence. Its final-head rerun failed during embedded
PostgreSQL bootstrap before any provider run: the macOS host had 87,365
of 87,381 SysV semaphores occupied. No unrelated services or kernel
limits were changed.
- The final-source report intentionally records **2/3 cells passed**,
preserving the blocked native-local attempt:
`tests/runner-e2e/results/agent-files-quota-final-20260928-report/`.
Earlier failed attempts and provenance notes remain under
`tests/runner-e2e/results/agent-files-quota-final-20260928-input/` and
the original campaign directories.
- Daytona used immutable image
`ghcr.io/paperclipai/paperclip-daytona-runner@sha256:5643f0d801417cae3581833a1a3bc6715b325e028602738d2652c44cac5dc6bf`
and its exact Linux runner binary. Controller source is `0724f3012`;
image source is recorded separately.
- Legacy-session compatibility and all three ACP Stop/resume browser
regressions passed on the prior validated feature head
`169fab46d5af21caa2269b4c1b29b69c933a6951`. They assert the same
provider session is retained and interrupted writes are not replayed.
Migration upgrade tests also passed earlier.
- Nine interactive stories are under **Agents / Persistent files**,
including **Full Storage Run Warning**. Its playtest and visual browser
inspection passed; the warning states that runs continue and the editor
remains available.
- Prior-head checks on `4e20c9fb2`: 55 passed, two conditional jobs
skipped, no failures or pending checks. All eight browser E2E shards and
their aggregate passed. Fresh Greptile review is 5/5 with no findings;
all review threads are resolved, the security scan passed, and GitHub
reports no merge conflicts.
- The broad local follow-up test run was interrupted after host
semaphore exhaustion affected isolated PostgreSQL instances. It also
encountered the existing macOS long-path fixture failure and two timeout
failures. This is not a claim that the full local suite passed. Logs are
retained; focused storage/warning tests passed.

## Risks

- A later sync can overwrite an earlier edit to the same file, including
a saved browser edit. There is no text merge or retained version. This
is the intended last-sync-wins policy.
- A save that exceeds a storage limit is rejected and its temporary copy
is discarded. The run itself continues normally, and later runs restore
the last saved files with a warning until cleanup. Transient sync
failures get bounded retries. An I/O failure partway through a sync can
leave some files updated; a failed receipt does not claim whole-folder
success.
- Larger folders increase copy time, network traffic, and temporary disk
usage. Active runs still need working copies. Terminal runs do not
accumulate archives. Operators must provision disk for agents and
configured concurrency; these limits are not company-wide quotas.
- A restored old native session remains instruction-only until a fresh
session starts. Its original conflict fence and existing pending
candidates remain compatible.
- Provider processes close at the collection boundary. Conversation
resume remains available, but warm process reuse is lost.
- Backups must include the instance filesystem and database. External
bundles keep their existing behavior until explicitly moved to managed
storage.

## Model Used

OpenAI Codex, GPT-6 family. The session does not expose a more specific
model ID or context-window size. Reasoning, code execution, and browser
tools assisted this change. Real provider E2E uses `gpt-5.6-sol`.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Fry (Paperclip) <noreply@paperclip.ing>
2026-09-29 08:25:56 -05:00
Devin FoleyandPaperclip 53aad90b9e fix: retry sandbox ACP input delivery after gateway failures (#14485)
## Thinking Path

> - Paperclip coordinates agent work through execution adapters.
> - Sandbox ACP sessions send ordered input through a remote file queue.
> - A temporary provider 502 currently closes the session during an
input upload.
> - A lost response can occur after the sandbox has consumed the
message, so a blind retry can duplicate input.
> - This pull request retries gateway failures with the same sequence
and drops consumed sequences at the receiver.
> - The session can continue through a brief provider failure without
repeating a tool call.

## Linked Issues or Issue Description

**What happened?**

A sandbox ACP run can fail with `ACP agent disconnected during request
(connection_close, exit=null, signal=null)` when a provider input upload
returns HTTP 502. The bridge destroys its local socket on the first
failure and can discard the diagnostic before the proxy reads it.

**Expected behavior**

A temporary gateway failure should get a bounded retry. A lost response
after successful delivery must not duplicate input or reorder later
messages. Permanent failures must still close the session.

**Steps to reproduce**

1. Run the real sandbox process bridge with an echo child and a local
test runner.
2. Inject a provider 502 before preparation, after a chunk upload, or
after final publication and consumption.
3. Send the next input message. Before this change, the connection
closes instead of delivering it.

Searched open and closed PRs for `ACP disconnect`, `bridge retry`, and
`502 sandbox`. Related work: #13287 covers shutdown after bridge loss;
#13793 covers large launch envelopes. This change covers ordered input
delivery within a running legacy ACP session.

## What Changed

- Retry input uploads up to three times for recognized Daytona and
Cloudflare HTTP 502, 503, and 504 diagnostics, with 250 ms and 500 ms
delays.
- Give each upload separate temporary paths and discard already-consumed
input sequences, including late publication from an earlier attempt.
Clean failed attempts in the background without removing a published
message or another attempt’s files. Cleanup cannot delay retries or
shutdown.
- Keep later input behind the retry. Stop queued input on permanent
failure and flush a fixed diagnostic before closing the socket. Neither
failure-diagnostic persistence nor shutdown-warning persistence can
block teardown.
- Add real-process regression tests for lost responses, late
publication, retry exhaustion, immediate permanent failure, and
diagnostic redaction.
- Give accepted run-log file appends up to three seconds to drain before
finalization computes the size, hash, and durable copy. Close the run
handle to later appends. This waits only for file writes, independently
of later DB progress or live-event persistence. If writes remain
stalled, return null size/hash metadata and skip the final durable copy
so the run can settle. Late writes cannot restart mirroring.
- Preserve legacy comment attribution when final log size is unknown by
reading existing entries within the unchanged 2 MB scan limit. Storage
errors or a three-second read deadline return the evidence already read
instead of failing the comment listing; pagination stops at the
deadline. The deadline requests cancellation of the underlying local
stream or S3 HEAD, GET, and response stream. A separate response timeout
returns partial evidence even when filesystem I/O delays cancellation;
late reads cannot append evidence or start another page. Each listing
retains its existing batches of eight reads, without a shared admission
cap that skips readable logs under contention.
- Document the retry and log-finalization boundaries in the development
guide.

## Verification

- Final commit `347daa564b`: [Linux
CI](https://github.com/paperclipai/paperclip/actions/runs/36506995168/attempts/2)
passed. Greptile Apex review 13 scored this commit 5/5 with no new
findings; all 12 review threads are resolved.
- The final CI run initially hit a Cursor test timeout and four Discord
credential-lock contention failures. All five cases passed in isolation.
The two failed shards and their aggregate gate passed on retry without a
code change. Those intermittent failures are not claimed fixed by this
PR.
- `pnpm --filter @paperclipai/adapter-utils typecheck` passed.
- `pnpm exec vitest run
packages/adapter-utils/src/execution-target-stdin-race.test.ts
packages/adapter-utils/src/execution-target-sandbox.test.ts
packages/adapter-utils/src/sandbox-callback-bridge.test.ts`: 262 tests
passed on the final implementation, including 21 new regressions. The
original three fault-injection cases failed before the fix.
- The regressions cover failed and indefinitely stalled cleanup,
Cloudflare gateway responses and retry exhaustion, permanent errors that
must not retry, and teardown while failure logging remains indefinitely
stalled. Seven Apex regression cases failed before the review fixes.
Adapter-utils typecheck and build passed again after the final review
change.
- `pnpm exec vitest run server/src/services/run-log-store.test.ts
server/src/services/run-log-store-cancellation.test.ts`: all 25 tests
passed, including four new regressions that failed before the
finalization fix. They cover delayed and failed appends, late-write
admission, agreement between the local bytes/summary/durable copy, and a
stalled append that exhausts the three-second budget. The timeout case
verifies unknown metadata, no final upload, and no mirror restart after
late completion. New cancellation tests use the real AWS SDK against a
local HTTP server. They verify that stalled HEAD, GET, and response-body
connections close on abort and that a subsequent read succeeds. Local
range and already-aborted read cases also pass.
- `pnpm exec vitest run server/src/__tests__/issues-service.test.ts -t
'readIssueCommentRunLogText|deriveIssueCommentRunLogAttribution'`: 14
targeted tests passed. The null-size reader case, both storage-error
cases, the stalled-read case, and the cancellation/concurrent-listing
cases failed before their fixes. The new regressions verify that
timed-out reads are cancelled, subsequent listings recover, and two
concurrent listings both retain their attribution markers. A read that
ignores cancellation still returns partial evidence at three seconds and
cannot resume pagination when it finishes; this regression failed before
the response-timeout fix.
- `pnpm --filter @paperclipai/server typecheck` and `pnpm --filter
@paperclipai/server build` passed after the response-timeout change.
- Full `pnpm -r typecheck` and `pnpm build` passed earlier in this PR;
the affected packages were rechecked after review fixes.
- Full local `pnpm test:run` failed in the general-server group: 511
files passed, 40 failed, and 158 were skipped. Failures include embedded
PostgreSQL initialization, read-only cache directory renames, a macOS
long-path fixture, and a workspace exposure assertion. The PostgreSQL,
cache-permission, and long-path failures also reproduce with both
changed implementation files restored to baseline commit `24c58e479a`.
The exposure suite passes in isolation both on baseline and the fixed
branch (28 passed, 3 skipped). CI runs the full suite on Linux. Later
local test groups were not reached.
- An earlier CI run hit the Telegram retry-timing failure fixed upstream
in #14501. The branch includes that master fix. The selected recovery
test passed against a fresh, migrated PostgreSQL 16 database. The
embedded PostgreSQL runner is unavailable on this Mac; the isolated
database was stopped and removed afterward.
- No live agent turn was replayed. The tests use local child processes
and injected provider failures.

## Risks

Retries are restricted to recognized Daytona SDK and Cloudflare bridge
gateway-error messages, which survive plugin RPC serialization. Other
errors fail immediately. Temporary upload paths are now unique for all
command-managed queue writes. Receiver sequence checks prevent duplicate
input; retries do not restart an agent turn. Cleanup and failure logging
are nonblocking and best effort; session teardown remains the final
cleanup boundary. Log finalization now drains accepted local file writes
for at most three seconds and ignores later appends on the closed run
handle. A timeout leaves final size/hash unknown and skips the final
durable upload; an existing partial mirror may remain available, but it
is not claimed as a verified final snapshot. It does not wait for later
DB progress or live-event persistence. Optional attribution keeps
partial evidence when a read fails or times out. Cancellation closes S3
requests and response streams. Local filesystem I/O may finish after the
caller deadline, but a late read cannot change the returned evidence or
continue pagination. Later listings can retry after storage recovers.
There is no schema, authentication, or permission change. Revert this
commit to restore the previous behavior.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository inspection, code
editing, and local test execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally; targeted tests pass and full-suite
limitations are documented above
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 19:15:33 -07:00
Devin FoleyandPaperclip ea371b9684 fix: retain project defaults in partial workspace overrides (#14502)
## Thinking Path

> - Paperclip manages AI agents and their work.
> - Project workspace policies define how isolated worktrees are set up.
> - Tasks can override a branch without providing every setup field.
> - The resolver currently replaces the entire project strategy with
that partial override.
> - Losing an explicit setup command can run the repository fallback
script and block the task.
> - This pull request keeps enabled project defaults when the task uses
the same strategy type.

## Linked Issues or Issue Description

**What happened?**
A project uses `git_worktree` with `provisionCommand: "true"`. A task
overrides only `baseRef`. The resolver drops the command. Worktree
creation then invokes `scripts/provision-worktree.sh`, which can fail
because its required setup is absent.

**Expected behavior**
A branch override keeps the project's provision, runtime provision, and
teardown commands unless the task explicitly overrides them. A different
strategy type must not inherit those commands.

**Steps to reproduce**
Configure the project with an enabled `git_worktree` strategy and
`provisionCommand: "true"`. Give the task an isolated workspace with a
`git_worktree` strategy and a different `baseRef`. Add a failing
repository fallback provisioner. Before this change, worktree creation
invokes that script. After this change, it uses the project's explicit
command and succeeds.

Related: #4968 concerns agent strategy and working-directory fallback.
#13903 concerns gated API fields and reusable-workspace updates. #11091
concerns provision hooks on workspace reuse. None fixes partial task
overrides discarding project defaults.

## What Changed

- Merge a partial task strategy over the enabled project's strategy only
when their types match.
- Preserve explicit null values when parsing nullable strategy fields,
so they can clear project values.
- Keep explicit empty-string overrides and agent fallback behavior.
- Exclude disabled project strategies and avoid an inherited branch
template when a task pins an existing branch.
- Add policy regression coverage and a real Git worktree test with a
failing fallback script.
- Document inheritance, explicit clearing, and no-op provisioning in the
development guide.

## Verification

- Policy regression: eight failures before the fix; all 41 policy tests
pass after it.
- Real worktree regression: passes and creates a worktree using the
task's base branch without invoking the failing fallback provisioner.
- `pnpm -r typecheck`: passed.
- `pnpm build`: passed.
- `pnpm test:run`: general-server phase completed with 13,873 passed, 86
skipped, and 14 failures in unchanged macOS skills-cache and Git
long-path tests. The same failures reproduce on unmodified base code.
The command stops at that phase, so no full local pass is claimed.
- CI initially failed the existing Telegram subscription recovery test
on a 15-second timeout. The separate fix and investigation are in
#14501. A serialized job also lost its runner; GitHub reported lost
communication, and that job was rerun without source changes. All 52
final-commit checks pass, with two intentional skips. Greptile is 5/5,
with no unresolved comments or merge conflicts. The chat shard passed on
one unchanged rerun. The timeout cause remains unproven; #14501 adds
phase diagnostics for a recurrence.

## Risks

Tasks that specify a partial strategy now retain the project's omitted
fields, including setup and teardown hooks. This is the intended
behavior change. Inheritance requires an enabled project policy and
matching strategy types. Explicit task values still win. Null and empty
commands restore existing runtime defaults; they do not guarantee that
no script runs. Use `"true"` for an explicit no-op provision command. No
migration, live configuration change, or task replay is included.

## Model Used

OpenAI GPT-6 (Codex), with reasoning, terminal tools, and code
execution. The context window size is not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes:` / `Closes`
/ `Refs` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub references)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run focused tests locally and they pass; full-suite status
is recorded above
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 18:09:20 -07:00
Devin FoleyandPaperclip 90119181e7 test: control Telegram subscription retry timing (#14501)
## Thinking Path

> - Paperclip manages AI agents and their work.
> - Telegram delivery recovers subscription changes after a restart.
> - The recovery test leaves a failed action on the real one-second
retry timer.
> - A slow runner can cross that deadline before the test checks that no
retry occurred.
> - This pull request holds the fixture deadline until the explicit
restart transition.
> - The test still checks that recovery uses fresh provider options.

## Linked Issues or Issue Description

**What happened?**
The Telegram subscription recovery test expected one `setWebhook`
request but saw two. A 1.5-second delay after the first failed request
reproduces the failure.

**Expected behavior**
The test controls when the failed request becomes eligible for retry.
Host speed does not change its result.

**Steps to reproduce**
Run the test named `retries an unknown subscription mutation after
restart` with a 1.5-second delay after the first failed-action
assertion. The old fixture retries too early. The updated fixture passes
with the same delay.

Related: #13952 fixes a separate Telegram fixture cleanup problem. This
change addresses retry timing.

## What Changed

- Set the stored retry deadline to 2099 before the pre-restart
assertions.
- Keep the existing explicit epoch deadline after restart and all
provider request assertions.
- Report the current test phase only when this test fails, to diagnose
an observed intermittent CI timeout.
- Leave production retry code unchanged. Temporary delay and per-step
console tracing are not included.

## Verification

- Delayed regression: failed before the change with two requests instead
of one; passed after the change.
- Focused Telegram durable private draft Stop group: 34 tests passed.
- `pnpm -r typecheck`: passed.
- `pnpm build`: passed.
- Full local chat shard: 355 tests passed twice.
- `pnpm test:run`: general-server phase completed with 13,861 passed, 86
skipped, and 14 failures in unchanged macOS skills-cache and Git
long-path tests. The same failures reproduce on unmodified base code.
The command stops at that phase, so no full local pass is claimed.
- CI exposed a separate 15-second timeout. A diagnostic run passed all
355 shard tests, with the affected test completing in under one second.
Its cause remains unproven. Normal step logging is removed; a
failure-only phase report remains for a recurrence. The final commit
also passes the 355-test chat shard, and Greptile rates it 5/5. All 52
final-commit checks pass, with two intentional skips. There are no
unresolved review comments or merge conflicts.

## Risks

Low risk. This changes only the fixture deadline. It does not disable a
test, extend a timeout, or change production retry behavior. Existing
assertions still verify the failed action, the pending state, restart
recovery, and fresh provider options. The intermittent CI timeout is not
claimed fixed; phase diagnostics narrow the next occurrence without
changing the timeout. No documentation change is needed for a test
fixture correction.

## Model Used

OpenAI GPT-6 (Codex), with reasoning, terminal tools, and code
execution. The context window size is not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes:` / `Closes`
/ `Refs` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub references)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run focused tests locally and they pass; full-suite status
is recorded above
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes, or
explained why none is needed
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 18:08:57 -07:00
Devin FoleyandPaperclip ad1f7e98ea fix: retain diagnostic reasons for native runner failures (#14481)
Retain bounded reasons for runner identity, harness recovery, and provider-pack read failures. Preserve existing ownership and cleanup proofs and compatibility with receipt-gated chat recovery.

Verified with executor, recovery, diagnostic privacy, typecheck, build, and full CI checks.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-28 16:16:19 -07:00
Devin FoleyandPaperclip 4f6cf5b3ff fix: prevent run identity locks from blocking audit checks (#14478)
Use NO KEY UPDATE for identity locks so audit foreign-key checks can proceed while identity writers remain serialized. Preserve task-before-run ordering, company scoping, and foreign keys.

Verified with PostgreSQL concurrency regressions, focused tests, typecheck, build, and full CI.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-28 16:16:05 -07:00
0be2afcca6 feat(ui): improve task composer controls and pending input (#14322)
## Thinking Path

> - Paperclip lets operators assign tasks to AI agents and review their
work.
> - The task composer controls the next message and its assigned agent.
> - Operators needed a way to choose that agent's model and effort
without leaving the composer.
> - The old mode selector, upload button, and input cards made the
mobile composer crowded and hid normal messaging during a pending
decision.
> - Harnesses publish different model and effort capabilities, so the
picker must follow the selected agent.
> - This pull request adds one responsive composer flow, keeps pending
cards visible above it, and protects Codex ACP authentication in the
local test path.
> - Operators can choose run settings, send a message, and answer a
pending card as separate actions.

## Linked Issues or Issue Description

**Subsystem affected**

Task composer UI, issue thread interactions, Codex ACP credential
handling, and Storybook.

**Problem or motivation**

The composer did not expose model or effort for the selected agent.
Mobile actions wrapped poorly. Pending questions and confirmations
replaced the composer. A local Codex ACP test could also reuse host
authentication after the managed key was removed.

**Proposed solution**

Put assignee search, model search, exact model IDs, effort, and fast
mode in one picker. Use a mobile dialog. Replace the direct-upload plus
action and separate mode selector with an Add menu and removable Plan or
Ask chips. Place pending interaction cards above the usable composer.
Keep these cards pending after an ordinary message unless their creator
asks for comment superseding. Replace managed ACP auth files atomically
and isolate the test key from host credentials.

**Roadmap alignment**

ROADMAP.md does not list an overlapping composer milestone. This change
improves the existing task and review flows.

## What Changed

- Added the combined assignee, model, and effort picker to both task
composers. Search matches agent name, role, and harness. The server uses
a curated Codex list by default and honors instance-declared models.
Manual IDs remain available.
- Added an effort slider for known model capabilities, a conditional
Codex fast control, and reset. The picker opens in a modal on mobile.
- Added the Add menu for files, supported goals, Plan mode, and Ask
mode. Plan and Ask are exclusive removable chips. Keyboard mode cycling
remains available.
- Adjusted mobile spacing, avatars, wrapping, and Send placement.
Removed the composer divider.
- Moved pending question, confirmation, review, and related cards above
the composer. Ordinary comments now leave question and confirmation
cards pending by default. The onboarding prompt retains explicit comment
superseding.
- Updated the Storybook composer group with responsive states and the
production picker. Added UI, service, route, and browser regression
coverage.
- Isolated Codex ACP API-key authentication, skipped subscription auth
merge and shared-home copy-back for remote API-key runs, and replaced
the managed auth file atomically.

## Verification

- `pnpm -r typecheck` — passed on the final local head.
- `pnpm check:token-gates` — passed on the final local head.
- `pnpm exec vitest run server/src/__tests__/adapter-models.test.ts
ui/src/components/task-chat/ComposerRunSettingsPicker.test.tsx` — 31
tests passed, including role and harness search, declared Codex models,
and filtering general OpenAI models.
- `pnpm exec vitest run
server/src/__tests__/issue-thread-interactions-service.test.ts` — 74
tests passed.
- `pnpm exec vitest run
packages/adapters/codex-local/src/server/acp.test.ts` — 42 tests passed,
including remote API-key copy-back isolation.
- `pnpm test:run` — attempted locally; the embedded PostgreSQL test
database could not initialize on macOS. The isolated
`heartbeat-run-event-sequencing` suite reproduced that environment
failure. GitHub CI runs the full test matrix for this head.
- `pnpm build` — passed on the final head. `pnpm build-storybook` passed
after the last UI change; only server code, tests, and docs changed
afterward.
- Live local test drive — Codex ACP ran a task with a managed API key.
The test agent was restored to its default ACP configuration afterward.
- Review the interactive stories under the top-level Composer group with
`pnpm storybook`. Check a narrow desktop width and mobile Plan, Ask,
picker, and pending-question states.

## Risks

- A pending card stays open when an ordinary comment changes the
discussion. Its creator can set `supersedeOnUserComment: true` when a
new comment should replace it.
- Model and effort overrides persist on the task until reset or changed.
An unlisted manual model ID may fail when the provider runs it.
- Some harness catalogs do not report effort support. The picker hides
effort for those models.
- No database migration is required.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

OpenAI GPT-6 via Codex. This runtime does not expose the exact model ID
or context window to the task. The model used code execution and browser
tools.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: OpenAI Codex <codex@openai.com>
2026-09-28 23:02:39 +00:00
DottaandPaperclip 18e8c121d9 fix(runner): include Grok support in public installs with sandbox prerequisites (#14024)
## Thinking Path

> - Paperclip manages agents through a shared native runner.
> - Built-in harness support should ship with Paperclip's public
distribution.
> - Grok already speaks ACP; it does not require a new public bridge
package.
> - Sandbox provisioning owns the native executable and its pinned
version.
> - The runner must verify that prerequisite without downloading it
during npm installation.
> - This change separates built-in launcher identity from external
runtime identity.
> - Clean npm installation and live staging checks verify the
distribution boundary.

## Linked Issues or Issue Description

Refs #13882, #13973, #13977, #13979.

This follow-up now targets master after #13882 was squash-merged. It
replaces the private `@paperclipai/grok-acp` workspace package with
runner-owned assets. Current master is included so the branch also
contains the merged scheduler, complete-event capture, and durable
cleanup fixes.

## What Changed

- Ship Grok launcher and qualification metadata inside the runner's
compiled output and the public server's vendored runner tree.
- Remove the separate Grok npm package and all package-manager install
hooks for this runtime.
- Require the checksum-verified Grok Build 1.0.13 binary at
`/opt/paperclip/providers/grok/1.0.13/grok` in the selected execution
environment. Provision it explicitly in the Daytona image and CI setup.
- Keep native binaries outside the provider pack. Bind the built-in
launcher into the pack manifest.
- Preserve executable leases, descriptor-backed startup, credential
fences, permissions, and exact ACP model admission.
- Use `builtin:grok-acp` and `native:grok` as profile identities.
Historical package-profile sessions fail closed on resume rather than
being silently reinterpreted.
- Resolve built-in assets from the authenticated sidecar location,
including public server npm layouts. Keep the controller path out of
provider environments.
- Add clean npm tarball installation verification to the existing
trusted canary CI job and the admitted manual EC2 verification path. It
stages a unified release version and runs npm lifecycle scripts, then
verifies missing-prerequisite rejection and admission after separate
provisioning without credentials or inference.
- Include the controller-owned provider pack in stamped Cloud images.
Unstamped local images omit the pack and remain usable; remote ACPX
requires full source provenance.
- Correct CLI approval-page metadata for an already authenticated Cloud
board user; approval authorization remains unchanged.
- Honor explicit native-runner enablement in the Cloud agent picker and
direct setup page, keeping the flag disabled by default.
- Allow selecting the execution environment before connecting
credentials. Include Grok in the existing authenticated hello-probe
flow, targeting its pinned native prerequisite for runner setup.
- Recover an existing subscription sign-in conflict through an explicit
cancel-and-retry action, serialized after cancellation succeeds.
- Preserve the selected ACPX harness before normalizing config fields,
so new Grok agents use the Grok default model.
- Keep the credential-free Cloud provider pack root-owned and readable
after runtime UID remapping; verify manifest and referenced asset access
under an unrelated unprivileged UID during image builds.
- Archive prior failover backups alongside explicitly replaced harness
state, preserving evidence while preventing stale backups from blocking
a fresh replacement.
- Update Daytona image content inputs and contract tests for the
built-in assets and explicit provisioner.
- Document and regression-test the shared `approve-all` default for Grok
setup, saved configuration, and native execution. Explicitly saved
restrictions remain unchanged.

## Verification

Current merge-repair head `df09eb3e1a619430ad8419a0ee9aedd486689b05`
incorporates master `f1a394bd30cb56fb9e479f98b9f50176fe921858` after the
base PR was squash-merged. All 12 conflicts came from incoming files
identical to the tested pre-squash base. The final tree exactly matches
a three-way merge using that original base, preserving built-in Grok
distribution and removal of the obsolete private package. All 252
focused runner/UI tests, six npm-isolation tests, and token gates pass.
Fresh exact-head Greptile review is 5/5 with no outstanding findings;
security scans and EC2 native compilation pass. All current-head CI is
green: 56 successful checks/statuses and four intentional skips ([run
36468768035](https://github.com/paperclipai/paperclip/actions/runs/36468768035)).
The repository owner explicitly authorized bypassing code-owner approval
after all checks passed; no CI checks or repository protection settings
are bypassed or changed. The only remaining PR was removed from the
completed stack metadata to permit native auto-merge.

Earlier integration head `78cb306ecc41b5c96577c26c1d89153b0ef865a1`
includes master `3447609d2247e75e55d91493dda91a608364f672` (2026-09-28).
Two master advances during verification overlapped the eval catalog; the
final merge preserves Grok qualification, completion updates, and
bounded API-response reading in all 348 cells. All 77 focused
catalog/eval/workflow tests pass. Both native stack layers (#14397) are
mergeable, and both exact-head Greptile reviews are 5/5 with successful
security scans and no unresolved review threads. All current-head CI is
green: 56 successful checks/statuses and four intentional skips ([CI
attempts](https://github.com/paperclipai/paperclip/actions/runs/36447124691)).
The initial attempt lost two EC2 runners to shutdown signals and stalled
a third shard during dependency preparation; all three passed the
same-commit failed-job-only retry. Trunk code-owner requirements remain
enforced. The review summary’s non-blocking saved-asset offset
classification note concerns code already merged in #14301; those
runtime files are identical to master and outside this stack’s diff.
Historical live evidence below retains its original source revisions.
[Final public npm
verification](https://github.com/paperclipai/paperclip/actions/runs/36445542764)
passed on `76ea70cd4d13786a042af9df82f0fd7a8c85ae30`: 17 public
packages, an executed offline lifecycle sentinel, unchanged consumer
lock, built-in launcher, missing-prerequisite rejection, and verified
separately provisioned binary/command lease. Provisioning and cleanup
require no host privilege elevation; only the positive probe mounts the
temporary native binary read-only. The verifier is unchanged by the
final master merge. All six isolation tests and an offline npm smoke
test pass. The prior head had 56 green CI checks and a 5/5 review after
two unchanged tests timed out and passed a failed-job-only retry ([CI
attempts](https://github.com/paperclipai/paperclip/actions/runs/36444597313)).
All 56 recovery-display/lineage tests pass; re-review cleared the
already-covered missed-retry concern. Earlier EC2 failures remain
retained: [npm lockfile
rejection](https://github.com/paperclipai/paperclip/actions/runs/36436311203),
[missing compiler in the slim
image](https://github.com/paperclipai/paperclip/actions/runs/36440210984),
and the aggregate 15-minute test timeouts in those broad runs. Both
broad attempts passed typecheck, token gates, Product E2E type/unit
checks and build. The focused EC2 lane preserves the existing
trusted-actor and immutable-source gates.


Earlier documentation/test checkpoint
`ff244c4fd78a7ede5a3e00efe09f475f133ef33e` leaves runtime behavior
unchanged. 154 focused tests pass across configuration building, native
provider resolution, permission policy, credentials, UI configuration,
and new-agent setup (including both Grok auth modes); token gates pass.
All fresh CI is green for this head: 56 successful checks/statuses and
two intentional skips ([run
36367065119](https://github.com/paperclipai/paperclip/actions/runs/36367065119)).
Greptile is 5/5 with no new findings. Grok already inherits the shared
`approve-all` default, so unattended setup requires no manual permission
change.

Runtime head `bb5a9307991f1ac567b781970ef11b39d518e19b` fixes a final
staging continuation failure before provider startup: explicit
replacement archived the old harness but left its failover backups
active, which caused `runner_harness_state_mismatch`. The regression
fails before the fix and passes after it; all eight adjacent
recovery-safety cases also pass. Old backups remain inspectable inside
the continuity archive. All fresh CI is green at this head ([run
36360839248](https://github.com/paperclipai/paperclip/actions/runs/36360839248)),
with a 5/5 review. One unrelated Cursor test timed out in the initial
server shard; the same-commit failed-job rerun passed, and both attempts
are retained. Staging deployment is confirmed healthy on this revision.
The controller image is
`ghcr.io/paperclipai/paperclip@sha256:6ad91c487910ccd2596ff7aed0a3a3ea5233d12b51b83cd6e1402237749b9673`.
The final browser-created staging task passed on this exact revision
with API authentication: context read → structured human question →
controller restart → answer submission → same native provider session
resumed → document saved → task Done. The two turns took approximately
119s and 77s. The actual write receipt was applied, and the saved
document has exactly one revision containing the selected answer and
requested marker. Usage and cost were not reported. [Controller image
build](https://github.com/paperclipai/paperclip/actions/runs/36360889243).

- Previous integration head `a44f7dbb6b6f77cd9ed893756ca453307f281e5f`:
all CI green (53 successful checks/statuses, two intentional skips),
including repository typecheck/build/tests, native Runner tests, browser
shards, and canary installation checks. [CI run
36358672529](https://github.com/paperclipai/paperclip/actions/runs/36358672529).
Greptile is 5/5 with no unresolved findings.
- Focused checks cover Grok credentials, executable admission, launcher
assets, provider-pack paths/permissions, workflow contracts, setup
defaults, CLI authorization, and subscription conflict recovery. All 39
protocol definitions validate. Final integration checks pass 124
catalog/evidence/cache tests and nine project-form tests; token gates
pass. Some local dependency checks could not load the stale installed
dependency tree; the corresponding fresh EC2 checks pass.
- Clean public npm installation passed on EC2 at
`8b172ebcf8e02e30662d830c00f3961e3bd459ec` ([run
36164964900](https://github.com/paperclipai/paperclip/actions/runs/36164964900)):
17 unified-version packages, lifecycle scripts enabled, built-in
launcher present, no separate Grok package or npm-downloaded binary,
missing prerequisite rejected, separately provisioned native executable
and command lease verified. No credentials or inference were used.
Subsequent changes preserve this npm asset layout.
- The immutable Daytona prerequisite image is
`ghcr.io/paperclipai/paperclip-daytona-runner@sha256:98957d5be0ac774d086b6402b5849e8e6356fec70fb8c09fca6eb4ed6de918e0`,
built from `5a2db471f3ddabe77f9f80e76ed27f996cb97fba`. The previous
Cloud controller image was
`ghcr.io/paperclipai/paperclip@sha256:fd914e1ab1e45f741e8e078ff452d16f082d7ac05f9b4b3506d3a3c64150d204`,
built from `a44f7dbb6b6f77cd9ed893756ca453307f281e5f`; it is superseded
by the latest image above. Its EC2 build verified provider-pack access
under an unrelated unprivileged UID.
- Browser staging at `40f898bc4cba73c1dff4e6344a3983ba0fb247ef` passed
full Grok onboarding with the correct `grok-4.7` model, saved credential
delivery, and pinned Daytona execution. A browser-created task read
context and asked the structured human question. After a controller
restart, answering the persisted question resumed the same native
provider session, saved the requested document, and completed the task.
Actual tool outcomes and durable state agree: one question and one
document revision. The two successful turns took 42.7s and 63.1s; usage
and cost were not reported.
- Restricted policy returned the expected `approval_required` outcome.
Functional staging tests explicitly selected `approve-all`; controller
authorization and governed approvals remain enforced. Temporary board
CLI access was revoked and verified rejected (HTTP 401), and the
disposable onboarding agent was paused.

Failures remain retained: the pre-fix continuation failure (its task
remains blocked; the passing final task is fresh), the original Cloud
provider-pack permission failure, the expected restricted-policy denial,
the superseded npm staging failure, and an earlier monolithic CI
infrastructure timeout. Browser CI exposed a project alias/form race;
the final stack uses master's stronger draft-preservation fix and all
browser shards pass. Historical full subscription/API protocol and
Product rosters retain their original source revisions and do not
qualify this packaging revision. No local Docker or Rust build was used.

## Risks

The branch includes master’s draft-preservation fix for project URL
aliases. It keeps the same project’s edit form mounted and clears prior
data when the project or company changes.

Custom sandboxes and local execution hosts must provision the pinned
binary before Grok starts. Missing, changed, unsupported-platform, and
symlinked executables fail admission. The new builtin profile cannot
resume sessions created with the former private-package profile.
Existing Claude/Codex npm bridge profiles retain their package pins.
Grok restricted modes preserve the selected policy but cannot
automatically admit Paperclip calls: ACP permission metadata does not
independently bind tool authority, so those calls stop with
`approval_required`. New Grok configurations default to `approve-all`,
including API configurations that omit the mode. Existing explicitly
restricted configurations remain restricted; controller authorization
and governed approvals remain enforced.

## Model Used

OpenAI GPT-6 through Codex, with tool use and code execution. The exact
serving model identifier and context-window size are not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 14:54:44 -05:00
DottaandPaperclip 992f720262 fix: make runner task context ownership explicit (#13753)
<!-- Write all pull request text in Simplified Technical English
(ASD-STE100). -->

## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Task descriptions, comments, continuation data, skills, and
execution rules enter several agent adapters.
> - The same source can be rendered by more than one automatic input
carrier.
> - Failed resumes can also rebuild input from stale or compact context.
> - This pull request gives each Paperclip-owned source one delivery
owner and preserves the required transport boundaries.
> - It adds deterministic adapter, interaction, runner, and browser
tests for these boundaries.
> - The benefit is more predictable context delivery with explicit
evidence for later live qualification.

## Linked Issues or Issue Description

Related: #13144 removes a duplicate environment payload and bounds wake
lists. Related: #11360 addresses Hermes resume behavior. This pull
request preserves compatible active-session formats while repairing
context ownership and stale question creation.

**What happened?**

Task descriptions and comments could enter more than one automatic
context block. Native transports could wrap a complete model input in a
second task envelope. Some legacy and gateway adapters could omit the
owned assignment on ordinary tasks or rebuild a failed resume with stale
compact context. A continuation could also request a question after
newer human comments had arrived.

**Expected behavior**

Each task or comment source has one automatic model-facing owner.
Distinct comment IDs and repeated wording remain distinct. Fresh
fallback attempts rebuild the required full context. A question request
is rejected when newer queued human direction makes it stale. Harness
access policy remains owned by execution configuration.

**Steps to reproduce**

1. Build a task with a description and current comments.
2. Capture the actual adapter or runner input.
3. Compare source ownership and task-envelope nesting.
4. Queue a human comment before a continuation requests a question.
5. Trigger a failed resume and inspect the fresh retry input.
6. Run the focused adapter, interaction, runner, and browser checks.

## What Changed

- Add shared prompt-section selection at the provider-attempt boundary.
- Deliver owned assignment context through native, legacy CLI, ACP,
gateway, cloud, Pi, Kimi, Grok, Gemini, OpenCode, Cursor, OpenClaw, and
Hermes paths.
- Rebuild full or compact context after resume recovery changes the
attempt. Add native and Claude ACP tests of actual recovery requests.
- Preserve custom templates, loaded instruction files, execution
policies, and older active-session formats.
- Record continuation source metadata and reject stale question creation
under the issue-row lock.
- Add explicit Product E2E context-integrity profiles, prerequisite
gates, credential-isolation checks, and report fixtures.
- Bypass service-worker forwarding for same-origin Vite development
modules. A real Chromium test fails with resource exhaustion before the
repair and passes after it. Production asset caching keeps its existing
policy.
- Add browser diagnostics and service-worker module-loading regressions.
- Add an explicit zero-retry eval option. The default retry behavior
remains unchanged. Each campaign records its effective policy.
- Remove the model-facing working-directory sentence from four prompt
builders. Existing workspace, sandbox, permission, and custom-template
configuration remains unchanged.
- Align the everyday workflow assertion with the current 47-entry
catalog.

Compared with current upstream master, the branch carries the
context-ownership implementation and its tests, the explicit
context-integrity catalog and evidence harness, and the focused browser
regression checks.

## Verification

**Merge assessment:** focused regression evidence supports merge. This
is not full completion of the original broad qualification matrix. The
maintainer has authorized merge after fresh verification of the master
integration.

- Current head: `bbd52f82114eabf09bc7b1a7e97d54a5b43bbc00`. This
integrates current master `2f585ef26a1814fa209715242d1ca791b63e4c4e`.
All 14 conflicts are resolved. Cancellation checks, workspace
finalization, native Grok support, and both sets of tests are retained.
- Current-head Greptile: **5/5**, with no blocking findings. The review
names this exact commit. All **59 reported checks are terminal: 55
successful, 4 skipped, zero pending or failing**. This includes the full
root general and serialized suites, separate runner checks, typecheck,
build, canary, browser E2E, Docker, and security checks. The successful
legacy security status is included in that total.
- After integration: workspace typecheck and full build passed. Separate
runner checks passed: **2,160 TypeScript tests (10 skipped), 582 Rust
tests, and 39 preparation checks**. Other passing checks include 621
Product E2E harness units, 376 focused shared/adapter tests, 160
real-database/API tests, 86 Hermes tests, 18 browser-support checks, and
Product E2E typechecking. The complete root suite passed in CI. The
duplicate local monolithic root run was stopped after that CI result; it
is not counted as a completed local pass.
- New native recovery coverage retains full assignment, completion
contract, and explicit skill selection after safe replacement, for old
and prepared input formats. Full native session test file: **136/136
passed**.
- New Claude ACP coverage captures actual fresh, resumed, and
missing-session fallback requests. It verifies one assignment copy,
comment order, identical text under distinct comment IDs, and full
fallback context. Full file: **33/33 passed**. Both affected TypeScript
checks passed.
- Existing deterministic tests cover source revisions, approval and
trust boundaries, completion validation, custom templates, compatible
sessions, standalone driver wrapping, and maintained adapter transport
requests.
- Provider-free browser support: **17/17 passed** after the master
merge. Service-worker unit tests: **33/33 passed**. The module-overload
regression failed before the repair and passed after it in real
Chromium.

### Fresh live comparisons

The new batch ran exactly four Product E2E attempts. **All four passed
on the first attempt; no retries.** Each has six terminal matchers plus
the existing browser lifecycle and invariant checks.

| Exact case ID | Control | Candidate |
|---|---|---|
| `core-compatibility.runner-codex.local.plan-revise-accept` | Passed |
Passed |
|
`local-session-integrity.runner-acpx-claude.local.structured-question-restart-resume`
| Passed | Passed |

The plan case checks a revised canonical plan and revision-bound
approval before completion. The question case restarts the server before
submitting the answer, then verifies the continuation completes.

Control source is `dfa4e1bda8d50a1a01746603251a9128dbe9d0d6`. Candidate
source is `79fcdb5dece501d28064ea9da306603881b46f0c`. They use identical
frozen definitions and provider versions: Codex `0.156.0` with
`gpt-5.6-sol`; ACPX `0.13.1` / Claude ACP `0.73.0` with
`claude-sonnet-5`. The September 24 head added master browser recovery
and test-only changes. The September 28 head also integrates newer
master changes, including cancellation, workspace finalization, and
native Grok. These are frozen-source live results, not exact-head live
runs.

The candidate received one description copy where the control initially
received three. The submitted initial plan envelopes were 7,969 versus
19,097 characters. Question envelopes were 7,592 versus 18,919. These
are structural measurements, not whole-provider token or dollar savings.

### Earlier evidence and failed attempts

- The preceding fresh batch has four effective passing pairs: OpenCode
comment continuation and assigned skill, native Codex comment
continuation, and native Claude comment continuation. It retains **11
attempts: eight passed and three failed**.
- Original failures remain recorded: missing local PostgreSQL library
links before task creation; host-sleep cleanup after task/page checks
passed; and a Claude **control** session-open rejection before a model
turn. Setup was repaired identically on both worktrees. The permitted
unchanged infrastructure retries passed. The underlying Claude provider
startup error was not retained and remains unknown.
- Older R2 retains **17 passes and one failure** across 18 attempts,
including eight both-pass native/legacy Codex/Claude pairs. Its OpenCode
blank-page failure led to the service-worker repair. R2 is historical
evidence: master changed the native fixed prompt and removed duplicate
wake environment data afterward.
- The September 24 CI run initially failed one unrelated preview
readiness test (`ECONNREFUSED` on its local fixture). Its test and
production code match master. Isolated local verification passed **28
tests, 3 skipped**. One unchanged CI retry passed the full shard: **831
passed, 1 skipped**, including all **31 preview-exposure tests**. The
aggregate CI gate passed afterward. The precise startup cause remains
unknown; a port race is a hypothesis, not a proved cause.

### Limits

The original wider profile/workflow matrix, repeated trials, and remote
Daytona qualification are incomplete. These results support a focused
merge recommendation, not statistical equivalence or universal harness
qualification. Some usage receipts are missing in both variants, so no
token or dollar savings are claimed. The $500 ceiling was preserved
using conservative allowances; failed attempts and unknown charges
remain in the ledger.

Reproduce the focused additions with `pnpm exec vitest run
packages/adapters/claude-local/src/server/acp.test.ts` and `pnpm
--filter @paperclipai/paperclip-runner exec vitest run
src/native-session-runtime.test.ts`. Full checks use `pnpm -r
typecheck`, `pnpm test:run`, `pnpm build`, and the separate runner
checks. Paid evals require the frozen definitions, profiles, and
credentials; do not use `--all` as a substitute for the selected cases.

## Risks

- Context placement changes can affect model behavior. Deterministic
checks cover the selected paths, but live qualification remains
incomplete.
- The stale-question guard can reject a request when queued human
comments arrived during the run. This is intended.
- New stored inputs and model envelopes retain compatibility readers for
older active sessions.
- Custom templates may intentionally repeat content.
- Removing a model-facing working-directory sentence does not change
filesystem, command, sandbox, or permission configuration.
- The worker bypass applies only to same-origin development module
paths. Cache-policy tests preserve private-response handling and
production asset caching. Mounted HTTP fixture changes remain test-only.
- This PR does not claim measured token savings or statistical
equivalence across every harness.

## Model Used

OpenAI Codex, exact model gpt-6-astra, with repository tools and code
execution. Bounded supporting work used gpt-5.6-luna and gpt-6-luna. The
serving context-window size is not exposed in this task.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have described the issue in-PR using the required issue fields
- [x] I have not referenced internal/instance-local Paperclip issues or
links
- [x] My branch name describes the change and contains no internal
ticket id
- [x] I have run the focused local checks and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect these changes
- [x] I have considered and documented risks above
- [x] All current-head Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
for the current head
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 14:49:14 -05:00
DottaandPaperclip f1a394bd30 feat(runner): add Grok Build through native ACP (#13882)
## Thinking Path

> - Paperclip manages AI agents and governs their work.
> - Its native runner uses structured provider protocols for sessions
and tools.
> - Grok Build supports ACP over stdio, but the runner did not expose
it.
> - Native execution requires company-scoped credentials, verified
identities, and permission gates.
> - This change adds Grok through ACPX for local and Daytona execution.
> - Subscription login and explicit API-key execution have separate
credential paths.
> - Qualification grades real tool outcomes, durable state, and browser
workflows.

## Linked Issues or Issue Description

Refs #13845, #13847, #13850, #13878, #13901, #13973, #13977, #13979.

Add **Grok Build** to `paperclip_runner` with `provider: "acpx"`,
`acpxAgent: "grok"`, and model `grok-4.7`. Existing legacy Grok agents
keep their adapter. Merge the three companion fixes (#13973, #13977,
#13979) before treating the integrated Product qualification as deployed
behavior.

## What Changed

- Synchronize shared, TypeScript, Rust, server, validation, and UI
provider contracts.
- Run Grok native ACP stdio through ACPX and the authenticated Paperclip
MCP bridge. Verify the pinned executable and exact ACP model identity.
- Prefer company subscription login. Support an explicit company-secret
API key without automatic paid fallback. Fence refresh and copyback to
the same account and remove private runtime credentials after
containment.
- Preserve selected permissions, cancellation, durable session identity,
resume, and restart recovery. Keep unsupported steering and goals
unavailable. Preserve missing usage and cost as unknown.
- Package checksum-verified Grok Build 1.0.13 for Daytona with an
immutable, signed image built on EC2.
- Add deterministic admission, protocol, permissions, identity,
credential, failure, and cleanup checks. Add the maintained 39-case
protocol roster and separate subscription/API Product profiles.
- Fix live-test findings in reasoning events, reloads, idle-owner
retirement, credential-home cleanup, expired-login model discovery,
launcher pinning, and rerun evidence selection.
- Align control-plane state readers with the transport's 64 MiB bound
while retaining identity, ownership, lifecycle, and size rejection
checks.
- Stabilize two asynchronous CI assertions while retaining actual
outcome and filesystem-evidence checks.

## Verification

Current integration head `f114948376056fe0b6b34c1496ae8667b59daa63`
includes master `3447609d2247e75e55d91493dda91a608364f672` (2026-09-28).
Two master advances during verification overlapped the eval catalog; the
final merge preserves Grok qualification, completion updates, and
bounded API-response reading in all 348 cells. All 77 focused
catalog/eval/workflow tests pass. Both native stack layers (#14397) are
mergeable, and both exact-head Greptile reviews are 5/5 with successful
security scans and no unresolved review threads. All current-head CI is
green: 56 successful checks/statuses and four intentional skips ([CI
run](https://github.com/paperclipai/paperclip/actions/runs/36447097232)).
Trunk code-owner requirements remain enforced. The review summary’s
non-blocking saved-asset offset classification note concerns code
already merged in #14301; those runtime files are identical to master
and outside this stack’s diff. Historical live evidence below retains
its original source revisions.


Earlier integration checkpoint:
`24fc9b94ca0afb21ccdc8d26dbb2e4b258ad72cb`. Refreshed against master
`0f14d2612`, preserving Grok qualification alongside the new accounting
and lifecycle suites. All 124 focused catalog, evidence, and
service-worker checks pass. The current base workflow includes the
explicitly selected public-install verification lane; follow-up #14024
supplies its verifier script. CI at that earlier checkpoint was green
(56 successful checks/statuses, four intentional skips), and the review
is 5/5 with no unresolved findings. Prior feature CI at
`fd73f0a9b1ecdf4094685054028df71739ddc3e1` passed ([run
36148259902](https://github.com/paperclipai/paperclip/actions/runs/36148259902));
that is historical evidence, not a current-head result.

Paid Product measurements use frozen integrated source
`2d939a92b21dcaf5c77c88b54d96784d2ddd0699`, which combines the feature
with #13973, #13977, and #13979. That source passed all 52 CI checks and
clean 5/5 review. Later master syncs incorporate upstream changes. Their
checks remain separate from these pinned live measurements.

| Check | Result and source-pinned report |
| --- | --- |
| Subscription protocol roster | [39/39 first attempts; 206
assertions](https://d1p6rlowie26tp.cloudfront.net/runner-protocol-evals/campaigns/gha-36046839612-1/index.html),
runtime `bc6833f7`, evals `92bb4b8c` |
| API protocol roster | [39/39 first attempts; 206
assertions](https://d1p6rlowie26tp.cloudfront.net/runner-protocol-evals/campaigns/gha-35926577007-1/index.html),
runtime `4a1061c8`, evals `3213dbec` |
| Subscription full Product matrix | [16/16 first attempts; 144
assertions; cleanup
passed](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36096908572-1/index.html),
source `2d939a92` |
| Subscription core repetitions | 18/18: tool use, planning approval,
and Stop/resume each passed three times in local and Daytona profiles.
The full matrix contains repetition one; [repeat
two](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36104551060-1/index.html)
and [repeat
three](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36108100404-1/index.html)
each passed 6/6. Total: 28 unique subscription attempts at `2d939a92`. |
| API smoke and question continuation | [4/4 first attempts; cleanup
passed](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36147315401-1/index.html),
both environments at `2d939a92` |
| Historical API Product coverage | [16/16 full
matrix](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35875144860-1/index.html)
and 18/18 core repetitions at `4a1061c8`; retained as measurements of
that revision |
| Native Daytona proof | Three subscription and three API
MCP/permissions/resume runs passed at `bc6833f7`. Three expired-login
admission and fenced refresh checks passed without inference. All test
sandboxes were removed. |
| Inspectable artifacts and UI | Current-source screenshots verify
planning approval, direct Ask completion, question continuation after
controller restart, and two downloadable project revisions. The project
downloads pass 12 and 18 tests; all 40 independent artifact oracle
checks pass. |
| Provider-free checks | 116 eval-validator tests, 39 Grok definitions,
and 359 enabled/external campaign cells pass. Continuation regressions
above 2 MiB and 16 MiB failed before their fixes; 32 focused
recovery/ownership/size checks pass. |

The 32 unique current-source Product attempts have no failures, retries,
or skipped cells, and all cleanup checks pass. Whole-workflow timing,
model identity, image and provider-pack provenance, attempts, and
accounting coverage are retained in the canonical reports. The report
publisher's conservative `complete=false` flag is preserved; independent
audits verify the exact selected source catalog and immutable result
rows.

Pins: Grok Build `1.0.13 (5e9a58528b76)`, ACPX `0.13.1`, ACP model
`grok-4.7`. Linux binary SHA-256:
`edf79521581bb5e6b95abef848491a6a742e860da3e237ebe86a280d30dce4c1`.
Launcher SHA-256:
`f0b698395a3704ed2ffaf84ea19bdb20c36c8a0a70b7c629c7b6ffe144e59e55`.
Image:
`ghcr.io/paperclipai/paperclip-daytona-runner@sha256:76b24edfd850219e949418b19e4ceba690e84d51d199ade426e484953329b5e9`.
Image build source is `4196a4cd`, recorded separately from application
source `2d939a92`; each campaign verifies the image signature and
provider pack.

Original failed campaigns remain available: [continuation
bound](https://github.com/paperclipai/paperclip/actions/runs/36057718059),
[scheduler/event
capture](https://github.com/paperclipai/paperclip/actions/runs/36071063537),
and [startup cleanup plus EC2
interruption](https://github.com/paperclipai/paperclip/actions/runs/36080870743).
They retain their original grades. No Docker or Rust builds ran on the
developer laptop for these follow-ups.

## Risks

Merge packaging follow-up #14024 with this base before public release.
The follow-up replaces the private Grok bridge package with a built-in
launcher and makes the native binary an explicit sandbox prerequisite.

Three separate, reviewed fixes are part of the tested integrated
behavior: #13973 serializes task-run admission; #13977 captures complete
event evidence; #13979 durably reconciles failed Daytona creation. Each
has green CI and clean 5/5 review. Failed-create recovery has 277 plugin
tests, 92 SDK tests, host-runtime recovery tests, and a real Daytona
lost-deletion-receipt proof. The live proof uses a private file for
journal persistence; database durability is covered by host tests.
Worker death before delivery of a failure envelope remains outside that
recovery mechanism.

Subscription fixtures stage an authorized company login; interactive
browser sign-in is not qualified. Local Product profiles ran on EC2
Linux. The temporary subscription credential was removed from the
protected GitHub environment after all subscription audits, with absence
verified. Runtime homes and refresh copyback remain ownership-fenced.

Protocol results remain pinned to their original revisions; they are not
relabeled as tests of the latest feature commit. New binary/model
versions require qualification. Missing token usage and model cost
remain unknown; runtime estimates do not establish a full bill.
Automatic paid Grok scheduling remains disabled pending separate
reviewed enablement. The 64 MiB bound can increase memory use for
verbose sessions, and larger files still fail closed. No automatic
legacy-agent migration occurs.

## Model Used

OpenAI GPT-6 through Codex, with tool use and code execution. The exact
serving model identifier and context-window size are not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 13:54:36 -05:00
Devin FoleyandPaperclip e9debd3eac fix(workspaces): allow larger status output for readiness checks (#14414)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The server checks workspace contents before it allows cleanup or
branch reconciliation.
> - These checks need the full Git status output to count untracked
files.
> - Nested task worktrees can make this output exceed the scheduler's
default 1 MiB limit.
> - A scan failure then blocks an otherwise inspectable workspace.
> - This pull request raises the limit for these status checks to 32
MiB.
> - The checks retain exact counts and still protect uncommitted work
from cleanup.

## Linked Issues or Issue Description

Refs #14194 and #14253. Those changes address snapshot enumeration. This
PR addresses buffered execution-workspace status checks. Refs #13619 for
separate work on caching these checks. The related journal identity
failure is covered by #14312.

**What happened?**

Close-readiness checks failed when Git status output exceeded 1 MiB. A
workspace with thousands of long untracked paths could not report its
file count or complete readiness inspection.

**Expected behavior**

Allow up to 32 MiB of status output for execution-workspace readiness
and branch reconciliation. Preserve exact counts. Continue to block
cleanup when the workspace has uncommitted data or the scan exceeds its
bound.

**Steps to reproduce**

1. Create an execution workspace with a merged delivery.
2. Add 5,000 long untracked filenames under a nested task directory. The
status output exceeds 1 MiB.
3. Request close readiness. Before this fix the status scan fails. After
this fix it reports all 5,000 files.
4. Run the terminal-workspace sweep. Confirm that it preserves the
workspace and files.

**Paperclip version or commit**

Reproduced on master at `14795136f5` before this fix.

**Deployment mode**

Server with managed Git workspaces.

## What Changed

- Set a 32 MiB stdout bound for execution-workspace status scans.
- Add a real Git regression with 5,000 long untracked paths and a
cleanup-preservation assertion. Assert that the measured status output
exceeds 1 MiB.
- Document the bound and failure behavior.

## Verification

- The new regression failed on master before the service change: the
status result had no untracked files or count after the scan exceeded
its bound.
- `pnpm exec vitest run
server/src/__tests__/execution-workspaces-service.test.ts
server/src/services/workspace-git-operation-scheduler.test.ts` passed:
82 tests, including the new regression. The regression and server
typecheck also passed after the explicit byte-count assertion.
- `pnpm build` and `pnpm -r typecheck` passed. The full local `pnpm
test:run` was attempted and stopped after the failures listed below.
Greptile is 5/5 with zero unresolved review threads on the latest head.
All checks for head `45f93ad5ad` passed (53 successful, two intentional
skips).
- Full local validation did not pass. The attempt reproduced 13
company/runtime skill-cache permission failures, the terminal-workspace
cleanup assertion, and a heartbeat feedback timeout. It was stopped
during the general-server stage after these failures. Remaining
general-server tests, other workspace groups, and serialized-server
stages did not complete locally. Earlier clean-master checks reproduced
the cache failures and isolated cleanup retries passed. The latest-head
GitHub suite passed all of these groups.
- One GitHub browser shard initially failed because its
GitHub-connection test checked the resume-action array before the mocked
request completed. The single failed-job retry passed without a source
change. All latest-head checks are green.
- No browser suites ran locally. This change does not affect browser
behavior.

## Risks

- Each active status scan can buffer up to 32 MiB before parsing. The
existing scheduler bounds scan concurrency and queue size.
- Output above the bound still fails the scan and blocks destructive
cleanup. The change does not truncate output or change cleanup rules.
- There are no API, database, or snapshot-streaming changes.

## Model Used

OpenAI Codex, GPT-6, with repository inspection, code editing, and test
execution. The exact deployment model ID and context window are not
exposed in this session.

## Checklist


- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused suites; full
local-run failures and incomplete stages are documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 09:59:56 -07:00
Devin FoleyandPaperclip 6f40e23536 fix(logging): redact cloud authentication headers (#14413)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The server records HTTP requests to help operators diagnose
failures.
> - Cloud requests carry tenant credentials and signed assertions in
headers.
> - The HTTP logger did not redact four of these headers.
> - This pull request adds those headers to the existing redaction list.
> - Operators retain the route, method, and response status without
recording these values.

## Linked Issues or Issue Description

**What happened?**

HTTP request logs could contain cloud tenant credentials, session
identifiers, runtime identity assertions, and cloud control assertions.

**Expected behavior**

The logger must redact these header values for successful requests and
failed requests.

**Steps to reproduce**

1. Create an Express app with the production HTTP logger and redaction
configuration.
2. Send a request with the four cloud headers and distinct test values.
3. Inspect the serialized request headers for responses with status 200,
403, and 500.

**Paperclip version or commit**

Reproduced on `0f14d26123` before this fix.

**Deployment mode**

Server with cloud proxy authentication. The regression test uses an
in-process Express server.

## What Changed

- Redact `x-paperclip-cloud-tenant-token`,
`x-paperclip-cloud-session-id`, `x-paperclip-cloud-runtime-identity`,
and `x-paperclip-cloud-control` in HTTP request logs.
- Test real logger output for HTTP 200, 403, and 500 with mixed-case
request header names. Route 403 and 500 through the production error
handler. Check response bodies, log levels, and server error context.
- Check that the method, route, and status remain available.

## Verification

The `server/src/__tests__/http-log-redaction.test.ts` suite passed,
including all three new cloud-header cases.

- Rebased onto master at `14795136f5`.
- `pnpm exec vitest run server/src/__tests__/http-log-redaction.test.ts`
passed: 59 tests, including all three new response-status cases. The
suite and server typecheck also passed after the error-handler coverage
update.
- `pnpm build` and `pnpm -r typecheck` passed. The full local `pnpm
test:run` was attempted and stopped after the failures listed below.
Greptile is 5/5 with zero unresolved review threads on the latest head.
All checks for head `373d29e2f1` passed (53 successful, two intentional
skips).
- Full local validation did not pass. The attempt reproduced
company-skill cache permission failures, the terminal-workspace cleanup
assertion, a heartbeat feedback timeout, and one process-conversation
timing failure. It was stopped during the general-server stage after
these failures. Remaining general-server tests, other workspace groups,
and serialized-server stages did not complete locally. Earlier
clean-master checks reproduced the cache failures and isolated cleanup
retries passed. The latest-head GitHub suite passed all of these groups.
- No browser suites ran locally. This change does not affect browser
behavior.

## Risks

- These four values will no longer be available in HTTP logs. Route,
method, and status remain available.
- This change applies to new log entries. It does not remove old entries
or rotate credentials.
- No schema, API, or authentication behavior changes.

## Model Used

OpenAI Codex, GPT-6, with repository inspection, code editing, and test
execution. The exact deployment model ID and context window are not
exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused suites; full
local-run failures and incomplete stages are documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 09:59:15 -07:00
DottaandPaperclip 14795136f5 fix(runtime): finalize and recover sandbox workspace exports safely (#14402)
Serialize native workspace finalization, validate streamed archives within bounded limits, and quietly recover unsafe exports from saved results. Preserve exact allocations for exhausted transient failures and provide export-only retry without rerunning the provider.

Consolidates #14314, #14315, #14329, and #14334 while preserving the already-merged finalization label changes.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-28 10:57:32 -05:00
DottaandPaperclip 3447609d22 fix(runner): stream and page large API responses within capture budgets (#14301)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents use governed API tools to inspect task evidence.
> - Large API results become saved assets with short previews.
> - Reading an asset through the same tool used to create another asset,
so the agent could not reach the rest of the evidence.
> - The 10 MiB response cap also blocked useful large results. Removing
all bounds allowed excessive disk use.
> - This pull request streams responses up to 1 GiB and makes saved text
readable in bounded pages. It adds durable run budgets and capture
admission limits.
> - Agents can inspect complete evidence while tool results, memory use,
and capture work stay bounded.

## Linked Issues or Issue Description

**What happened?**
A large response became an asset. Reading that asset returned another
asset and the same preview. Responses above 10 MiB failed before the
agent could read any page.

**Expected behavior**
The agent can fetch a large response and read its saved text to EOF.
Each page stays bounded. New snapshots have a generous finite limit and
a durable run budget. Existing larger assets remain readable through
byte ranges.

**Steps to reproduce**
1. Call a GET operation that returns more than 10 MiB of text or JSON.
2. Before the fix, the tool returns `api_transport_failure`.
3. With this change, responses up to 1 GiB become streamed snapshots
with artifact references.
4. Read `GET /api/assets/{assetId}/content` with `responseText:
{offsetBytes: 0, limitBytes: 8192}`. Follow `nextOffsetBytes` until
null.

Related work: #14186 added the API fallback tools. #14218 bounded API
discovery.

## What Changed

- Add authenticated UTF-8 text windows to `call_api`, with byte offsets
and total size. Keep each page at or below 24 KiB.
- Stream new responses above 24 KiB through private temporary files into
company-owned assets. Bound each capture to 1 GiB of decoded bytes.
Reject oversized declared lengths before reading and count streamed
bytes before writing.
- Reserve capture budget in the run record before spilling. Allow 4 GiB
per run. Settle successful captures to their actual size. Failed or
interrupted captures retain their full 1 GiB reservation. Run restarts
do not reset the budget.
- Enforce a 20 GiB company snapshot quota with database reservations.
Count legacy snapshots and unfinished storage work across runs and
processes. Asset deletion frees quota.
- Limit large captures to two per company and four per server process.
Hold slots through storage upload and temporary-file cleanup. Use a
10-minute download deadline and 30-second connection/idle-read timeouts.
- Return explicit size, budget, busy, and timeout errors. Preserve
unknown outcomes for mutations whose response cannot be captured.
- Read saved assets through authenticated storage ranges, with at most
two extra bytes for UTF-8 and EOF handling. Unpaged reads return the
existing asset and digest with a bounded preview. Reads create no copies
and do not consume capture budget.
- Keep existing assets above 1 GiB readable in pages. Use safe integer
offsets and PostgreSQL `bigint` asset sizes.
- Stream large S3 uploads through ordered multipart requests. Abort
failed uploads and remove partial local files.
- Revalidate run authority during downloads. Keep company authorization,
GET-only text paging, redirect denial, and mutation replay receipts.
- Document the separate 10 MiB upload limits. This PR does not raise
memory-buffered attachment ingestion limits. Future large video uploads
need streamed ingestion and storage quotas.

## Verification

- Full workspace `pnpm -r typecheck` and `pnpm build` pass after
rebasing on master.
- Focused API and response tests: 1,761 pass. Cover declared and chunked
oversize responses, incorrect Content-Length, exact-limit success,
active-stream deadline, cancellation, cleanup, concurrency admission,
and mutation outcome handling.
- Real HTTP integration: 28 tests pass, including runnerd → PRP →
authority → HTTP, a 12 MiB snapshot, final-page/EOF reads, cross-company
denial, a persisted 3 GiB sparse asset, and large mutation receipt
replay.
- The HTTP suite verifies durable run-budget accounting, simultaneous
runs competing for company quota, legacy snapshot accounting, deletion
refunds, failed-storage reservations, cleaned-failure refunds,
metadata-rollback cleanup refunds, preservation after a lost commit
acknowledgement, and small/saved reads after capture-budget exhaustion.
- A standalone proof streams exactly 1 GiB through the production
capture helper, verifies the final bytes, and removes its temporary
file. It uses repeated 256 KiB chunks and records a peak process RSS of
191 MiB.
- Earlier storage verification covers exact S3 multipart boundaries,
cleanup/abort failures, and a 17 MiB transfer through the real AWS SDK
to a local HTTP S3 endpoint. No cloud S3 qualification was run for this
follow-up.
- The local full test run was interrupted for the company-quota changes.
A later targeted run hit exhausted macOS shared-memory slots before
tests started; two unattached PostgreSQL segments with dead owners were
reclaimed before retrying. All 55 current-head checks pass at
`aebb80ceeeee77d5a56b67bfffd835f2f846878c`, including the full CI test
suite, typecheck, build, browser suites, security scan, and Greptile
(5/5). There are no unresolved review threads. The combined rebased test
catalog also passes (48 tests).
- Earlier paging acceptance passed Daytona and separate staging at
`7739879e9`. Those runs predate the streaming and budget changes.

## Risks

- The 1 GiB response cap and 10-minute active-download deadline are
intentional product limits. Larger live results must use endpoint
pagination or a direct file workflow. Existing larger assets remain
readable through bounded ranges.
- A durable 20 GiB company snapshot quota counts stored runner-api
assets and active/orphan reservations across runs and processes. The
operator can set PAPERCLIP_RUNNER_API_COMPANY_CAPTURE_MAX_BYTES to a
finite value of at least 1 GiB. Deleting snapshots frees capacity;
possible orphan storage must be reconciled before releasing its
reservation.
- A failed capture uses its full reservation. A new large capture needs
a full 1 GiB available, even if it later completes at a smaller size.
Small reads and existing asset pages remain available.
- Concurrency limits apply per server process. The run byte budget is
shared through the database.
- The `integer` to `bigint` migration rewrites asset metadata and takes
an exclusive table lock. File bytes stay in storage.
- A live endpoint is fetched once before returning its snapshot.
Continue reading the saved artifact for stable pages. Mutations may
commit before any size or transport error; inspect state before
retrying.
- Attachment uploads and native file handoffs still default to 10 MiB.
Raising buffered ingestion paths to GiB sizes is separate work.

## Model Used

OpenAI Codex, based on GPT-6, with code execution and repository tools.
The runtime does not expose an exact serving model variant or
context-window size. The earlier paging work also used browser testing
and subagents.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 10:36:32 -05:00
DottaandPaperclip 8751e2de46 fix(ui): distinguish finalization recovery from live observation (#14326)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The task board shows which recovery actions are active.
> - A native run can stop while a person must repair its workspace.
> - The board previously called that state “Recovery in progress.”
> - The label implied that work would continue without operator action.
> - This pull request derives the label from the recovery owner and live
continuation.
> - Operators can distinguish scheduled recovery from a repair that
needs attention.

## Linked Issues or Issue Description

**What happened?**
A blocked task showed “Recovery in progress” after native finalization
stopped and no automatic continuation remained.

**Expected behavior**
Show “Recovery needed” for an idle board repair or when no live recovery
path exists. Show “Recovery in progress” while the recorded continuation
can run, including an explicitly admitted export whose exact callback is
executing even if the old recovery action remains board-owned.

**Steps to reproduce**
1. Complete a native run whose workspace export cannot be recovered
automatically.
2. Inspect the task recovery action and badge.
3. Compare the board-owned action with the old “Recovery in progress”
label.

Related work: #14314 serializes native workspace finalization and fences
stale recovery outcomes. This change reports the recovery action that
currently owns the task.

## What Changed

- Show “Recovery needed” for board-owned active-run recovery unless the
exact native export callback is positively verified as executing.
- Require the native continuation run and a live or future continuation
before showing progress.
- Project native activity from the exact company, source issue, and run.
Include active workspace export while the original heartbeat remains
failed.
- Require an executing callback before using a running export row as
evidence. Preserve activity for long exports and clear it when the
callback joins.
- Give native resume its own card explanation. Preserve ordinary
watchdog observation behavior.
- Remove the redundant ownership sentence from all six recovery-card
explanations that used it.
- Document the labels and add regression cases for stopped, scheduled,
and active recovery.

## Verification

- Copy-only follow-up (`c68aef04c`): all 167 focused recovery UI tests
and token gates pass. No UI occurrence of the removed sentence remains.
`pnpm -r typecheck`, `pnpm build`, and current-head CI pass (54
successful checks, two optional Storybook checks skipped). Greptile is
5/5 with no open review threads. The duplicate local `pnpm test:run` was
stopped after the full CI suite passed; it did not complete locally.
- Original regressions: seven failures before the change, then 52
focused cases pass.
- Review regressions: seven UI failures and nine database failures
before the follow-up. All 75 database/API recovery tests, 167 UI tests,
and 18 workspace lifecycle/finalizer tests pass. A further three RED
cases cover explicit board retry activity; one RED case rejects orphaned
running export rows after controller loss. Wrong company, issue, run,
service, phase, and completed-operation cases remain inactive.
- Final recursive typecheck, production build, token gates, and complete
local suite coverage pass. Embedded PostgreSQL startup/socket failures
passed in isolated retries with the canonical test environment; no
expected behavior was weakened. The recovery regression added during the
earlier full run passed in its final complete 75-case file.
- Before the copy-only follow-up, all 56 CI checks passed on
`c2f84cd89c905cda85c53aaf5bb83b7250900fe6`; Greptile is 5/5 with no
unresolved review threads. The final native-activity staging repeat
passed on integrated source `2bedd0f23bf4698b1f8b818f6796900647030427`.
- Verified on a separate staging instance: a real failed Daytona
workspace export retains its board-owned repair action and displays
“Recovery needed” in the task list. The repair card remains actionable
without starting another provider turn.

- Real staged export-only repair: the actual task list showed “Recovery
in progress” while the original run had a positively identified running
export operation, then Done after exact copyback of all 20,000
nonce-bound files. The accepted result and full provider
session/turn/terminal envelopes remained unchanged. All 17 independent
final checks passed; the browser downloaded the exact 19-byte result.
The separate fixture was cleaned up with independent provider-absence
verification. A control transport process restarted during repair; it
did not submit another provider turn.

- Final deployed-source repeat on
`d884e1ab046cc76004e35e6091e9e6e2c918c9eb`: explicit per-turn ephemeral
Daytona allocation, actual browser export repair, and a saved full
activity projection referencing the exact executing export operation
with no scheduled retry. The task list showed recovery in progress, then
Done; all 21 final checks passed, including 20,000 exact host files,
unchanged provider provenance, and provider deletion only after
committed copyback. The downloaded 19-byte result matched independently.
This integrates #14334; no source change was required here.

## Risks

- The label depends on the persisted recovery action. A separate runtime
defect can still stop work; this change makes that condition visible.
- Future recovery kinds must supply a valid continuation path before
they can display progress.

## Model Used

OpenAI Codex, based on GPT-6, with code execution, browser testing, and
subagent tool use. The runtime does not expose an exact serving model
variant or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 09:33:39 -05:00
DottaandPaperclip 4b38db9622 fix(runtime): stream workspace Git snapshots through disk manifests (#14253)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Managed runs copy a selected workspace to an execution environment
and restore its changes.
> - Git snapshots select the files for that copy and for later recovery.
> - A fixed output limit stops large generated trees before the run can
start.
> - Increasing the limit still keeps the complete filename lists in
memory.
> - This pull request stores those lists and merge baselines in disk
manifests.
> - Large snapshots can now complete with bounded filename buffers and
explicit failure handling.

## Linked Issues or Issue Description

Refs #14194. This is the streaming follow-up to the merged 32 MiB limit
fix.

Related: #13619 and #11621 cover workspace scan admission and demand.
This change keeps the shared scheduler and changes the snapshot data
path.

## What Changed

- Stream changed, untracked, deleted, and ignored paths through the
shared scheduler and the standalone adapter path.
- Use SQLite manifests for file selection, duplicate removal,
ignored-path lookup, baseline capture, and merge lookup.
- Set a configurable 30-minute snapshot deadline. Keep the existing
interactive scan deadlines.
- Wait for each child process and pending sink write before removing
temporary storage after failure or cancellation.
- Use NUL archive lists and bounded deletion batches. Preserve unusual
names, explicit selection, nested repositories, and source root checks.
- Store manifest references in native recovery format v2. Check their
location and digest before recovery reads. Keep v1 descriptors readable.
- Remove temporary manifests at lifecycle completion. Use fixed-size
temporary copy names for long basenames.
- Admit each manifest with a SQLite page allowance based on current disk
capacity. Keep a configurable free-space reserve and fail explicitly
when either limit is reached.
- Preserve a host file that replaces a directory deleted by the sandbox,
and continue the rest of the restore.

## Verification

- Current head `5ab622ec43cd16d35429d79dedee6a5d8e3d2df2` has 54
successful checks/statuses and two skipped Storybook jobs. No checks
failed or remain pending.
- [CI
passed](https://github.com/paperclipai/paperclip/actions/runs/36318966368):
typecheck, build, all test shards, E2E, Rust checks, and the aggregate
verify job.
- [Greptile is
5/5](https://github.com/paperclipai/paperclip/pull/14253#issuecomment-5855670338)
on the current head. All four review threads are resolved. Security
checks passed.
- 229 focused tests passed across Git sync, runtime staging, merge,
manifest integrity, native recovery, and the scheduler (214
adapter/runtime tests and 15 scheduler tests).
- A real 40,000-file fixture produces 43,428,890 filename bytes. The
original standalone and scheduled scans fail. The new test passes all
four filename paths, complete staging, exclusion of late files, unusual
names, and deletion replay.
- Recovery tests reject changed bytes, symlinks, and paths outside the
controller state directory. Adapter-utils typecheck passed.
- A test executor returned buffered output and caused two retry
integration failures. The fixture now uses the shared streaming
scheduler. All 13 tests passed with `corepack pnpm exec vitest run
server/src/__tests__/heartbeat-project-repositories.test.ts`. The same
CI shard now passes.
- Ran `pnpm -r typecheck`, `pnpm test:run`, and `pnpm build` locally.
Each full local command hit SIGKILL/exit 137 in the 4 GiB container.
These local commands did not pass. The current-head CI gates above
provide the full verification.
- A real-Git disk-capacity regression confirms a typed failure and
removal of the incomplete manifest. Repeated writer attempts cannot
exceed the permitted page count.

- Follow-up real Daytona and separate staging qualification passed with
the related archive validator (#14315) and exact-owner finalization fix
(#14314). Three successive turns copied back all 60,000 files with
39,828,890 filename bytes and five unusual names. Independent host
inventories verified every file and the pinned Git HEAD. Native,
provider, session, and process identities stayed fixed; no retry
remained. The task reached Done, and its browser-downloaded final proof
matched exactly. The reusable regression is #14316, including an
assertion of the effective environment idle policy.

## Risks

- SQLite manifests use disk space. Each receives one quarter of the
available capacity above the host reserve at creation. The reserve
defaults to 256 MiB and has a 64 MiB configuration minimum. Disk
capacity, filesystem quotas, per-path limits, Git resource use, and
execution deadlines remain limits.
- Each path and sink chunk has a 64 KiB limit. SQLite connections use a
1 MiB page cache. Invalid or incomplete records fail explicitly.
- Restore transport keeps fixed and configured archive exclusions. A
remotely created Git-ignored file can be transferred, but the host merge
excludes it through the manifest.
- Provider archive buffers, Git and tar memory, repository metadata,
legacy v1 arrays, and the separate referenced-source resolver retain
their own limits. Existing provider safety validators still buffer
textual tar listings: Daytona allows 32 MiB and Kubernetes allows 64
MiB. These separate transport limits can stop a sufficiently large
restore before merge. This change does not claim bounded total process
memory or unlimited transport size.
- New descriptors use v2. Existing v1 recovery remains supported; a
downgrade cannot read v2 descriptors.

## Model Used

OpenAI GPT-6 through Codex. The exact deployment ID and context limit
are not exposed in this run. The agent used code editing, terminal
execution, tests, and GitHub tools.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 08:40:36 -05:00
Devin FoleyandPaperclip be43c23e2d fix(recovery): preserve pending native result finalization (#14219)
Preserve native coordinator ownership while accepted results await workspace
copy-back, assessment, or arbitration. Check that ownership in the terminal
update so a coordinator recorded after the liveness read is also protected.
Keep terminal task authority and exhausted-retry cleanup unchanged.

Validation: 28 focused recovery tests, local typecheck/build, 52 passing CI
checks, and Greptile 5/5 with no open findings.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-27 00:22:49 -07:00
Devin FoleyandPaperclip 5ee9e751fb fix(runner): discover assigned tools when direct catalogs exceed limits (#14218)
Keep assigned app tools accessible when the combined Runner catalog exceeds
its operation or byte limits. Reserve task and completion tools, then expose
bounded discovery and call tools for large catalogs. Fetch oversized schemas
in reauthorized chunks without blocking later search results.

Retain task ownership, work-mode restrictions, pinned assignments, current
gateway authorization, approvals, and audit. Small catalogs stay direct.

Validation: 46 focused server tests, two Runner capacity tests, local
typecheck/build, 52 passing CI checks, and Greptile 5/5 with no open findings.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-27 00:22:16 -07:00