Commit Graph
4646 Commits
Author SHA1 Message Date
DottaandPaperclip 5aeebb20c9 Preserve privileged ACP approval projection and delivered lifecycle
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-28 12:30:19 -05:00
DottaandPaperclip 170652d2d4 Declare and validate bound ACP permission request envelopes
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-28 12:27:55 -05:00
DottaandPaperclip c8fc6047f3 Normalize optional Codex question descriptions before durable validation
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-28 12:24:56 -05:00
DottaandPaperclip 920fab4050 Record browser proof and remaining rich content gaps
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-28 12:23:19 -05:00
DottaandPaperclip 986d8153ea Verify pending Pi qualification diagnostics
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-28 12:23:19 -05:00
DottaandPaperclip 1e9ccc94e0 Retain rich ACP production UI browser verification
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-28 12:22:23 -05:00
DottaandPaperclip f52baa537e Assert rich ACP admission and notice disclosure behavior
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-28 12:22:23 -05:00
DottaandPaperclip 04b88a97f5 Preserve ACP turn ordering when negotiated controls are unchanged
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-28 12:19:29 -05:00
DottaandPaperclip 851c60a1fb Preserve complete native plan context in rich ACP interaction UI
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-28 12:17:02 -05:00
DottaandPaperclip 0adc3b1a5e Admit candidate authentication from sanitized launch bindings
Allow candidate profiles to report missing authentication before native initialization stalls.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-28 12:15:13 -05:00
DottaandPaperclip 0ece23ad7d Canonicalize native pack staging paths and format durable expiry fixtures
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-28 12:09:14 -05:00
DottaandPaperclip 2ed1540570 Merge current master and validate negotiated ACP capability envelopes
Synchronize with master 14795136f5 while retaining the recorded implementation base. Regenerate protocol contracts and preserve both input-presence and permission-scope checks.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-28 12:07:57 -05:00
DottaandPaperclip a5ff475ecb Expose negotiated ACP turn controls through durable and direct sessions
Propagate live Pi steering and native queued follow-up capabilities after session open and lazy warm prompt initialization. Preserve conservative static declarations, reject forged capabilities and malformed control modes, and admit only profile-bound candidate credentials at the Rust sidecar boundary.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-28 12:05:11 -05:00
DottaandPaperclip 51db509899 Expire durable ACP interactions when their provider turn is lost
Retain a bounded private pending-request ledger across controller and runner reconnects. Emit non-replayable expirations before terminal facts, preserve complete questions, and never restore approvals or mutations after provider loss.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-28 12:04:42 -05:00
DottaandPaperclip d4e953e443 Await provider reply delivery in the direct ACPX driver
Fail closed without reporting resolution when the pipe write is ambiguous. Preserve provider-supported Copilot session permissions.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-28 12:03:50 -05:00
DottaandPaperclip a3a71ebdb1 Bind Daytona candidate image identity to selected provider assets
Refresh the reviewed resolved dependency digest and fix ACP selector typing.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-28 11:59:18 -05:00
DottaandPaperclip 588dd00af1 Bind candidate diagnostics and recovery to explicit runner policy
Preserve Pi pricing receipt provenance, require evaluation opt-in, and record qualification boundaries.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-28 11:55:09 -05:00
DottaandPaperclip 84336a1b7d Preserve pinned Pi assistant and compaction receipt provenance
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-28 11:52:33 -05:00
DottaandPaperclip d82739164d Confirm ACP interaction replies reach the provider pipe before settlement
Bind response receipts to the active stream, request and prompt; reject cancellation, disconnect, write failure and mismatched responses before acknowledging durable delivery.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-28 11:50:35 -05:00
DottaandPaperclip 6ff8f78216 Allow explicitly selected candidate ACPX eval diagnostics
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-28 11:47:46 -05:00
DottaandPaperclip 7cf00e0072 Deliver negotiated ACP steering and queued follow-ups through durable runner controls
Bind controls to the active turn and reserve bounded delivery identities before dispatch; preserve explicit provider policy and Pi receipt estimates across both runner paths.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-28 11:43:24 -05:00
DottaandPaperclip 88f34b30f8 Bind sessionless ACP extensions to their admitted wire session
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-28 11:43:24 -05:00
DottaandPaperclip 1cee50ed48 Bind optional ACP candidate assets into local and Daytona provider packs
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-28 11:39:26 -05:00
DottaandPaperclip 411df9fa58 Bind candidate runtimes to task policy and separate Pi price estimates
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-28 11:35:14 -05:00
DottaandPaperclip a5e26ecc04 Resolve native provider assets through the runner package authority
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-28 11:34:04 -05:00
DottaandPaperclip 37e6ebbc51 Preserve rich ACP inputs and activity through runner and UI
Bind provider extensions to active turns, drain activity before settlement, preserve complete plan decisions, and expose bounded provider notice details. Keep new harness choices disabled until qualification.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-28 11:26:14 -05:00
DottaandPaperclip 2de5c04c9a Preserve scoped Pi prompt receipt provenance and cost through ACPX
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-28 11:26:11 -05:00
DottaandPaperclip 7e2e7e1c35 Fence permission requests to the admitted ACP session
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-28 11:23:59 -05:00
DottaandPaperclip 94724d1414 Preserve schema-bounded ACP display activity through durable projection
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-28 11:23:17 -05:00
DottaandPaperclip 4a278e2740 Fence ACP extension callbacks and Pi controls to active turns
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-28 11:22:09 -05:00
DottaandPaperclip ae15ff4c3f Isolate candidate ACP runtime credentials and provider homes
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-28 11:18:52 -05:00
DottaandPaperclip 384ea84546 Declare rich ACP profile capabilities and qualification gates
Keep Cursor, Copilot and Pi unavailable until qualification; require explicit new-provider models without substitution and bind new native candidates to versioned profiles.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-28 11:14:45 -05:00
DottaandPaperclip af61365ea5 Verify complete native ACP distributions under existing lifetime guardian
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-28 11:14:34 -05:00
DottaandPaperclip c810a25db2 Carry ACP permission requests through durable runner resolution
Validate provider-supported choices, fence resolutions to the active turn, retain pending state until sidecar acknowledgement, and preserve complete bounded plan documents.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-28 11:10:36 -05:00
DottaandPaperclip 8193cb2571 Add prompt-bound ACP extension requests and event hooks
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-28 11:07:38 -05:00
DottaandPaperclip 14795136f5 fix(runtime): finalize and recover sandbox workspace exports safely (#14402)
Serialize native workspace finalization, validate streamed archives within bounded limits, and quietly recover unsafe exports from saved results. Preserve exact allocations for exhausted transient failures and provide export-only retry without rerunning the provider.

Consolidates #14314, #14315, #14329, and #14334 while preserving the already-merged finalization label changes.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
canary/v2026.928.0-canary.3
2026-09-28 10:57:32 -05:00
DottaandPaperclip 3447609d22 fix(runner): stream and page large API responses within capture budgets (#14301)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents use governed API tools to inspect task evidence.
> - Large API results become saved assets with short previews.
> - Reading an asset through the same tool used to create another asset,
so the agent could not reach the rest of the evidence.
> - The 10 MiB response cap also blocked useful large results. Removing
all bounds allowed excessive disk use.
> - This pull request streams responses up to 1 GiB and makes saved text
readable in bounded pages. It adds durable run budgets and capture
admission limits.
> - Agents can inspect complete evidence while tool results, memory use,
and capture work stay bounded.

## Linked Issues or Issue Description

**What happened?**
A large response became an asset. Reading that asset returned another
asset and the same preview. Responses above 10 MiB failed before the
agent could read any page.

**Expected behavior**
The agent can fetch a large response and read its saved text to EOF.
Each page stays bounded. New snapshots have a generous finite limit and
a durable run budget. Existing larger assets remain readable through
byte ranges.

**Steps to reproduce**
1. Call a GET operation that returns more than 10 MiB of text or JSON.
2. Before the fix, the tool returns `api_transport_failure`.
3. With this change, responses up to 1 GiB become streamed snapshots
with artifact references.
4. Read `GET /api/assets/{assetId}/content` with `responseText:
{offsetBytes: 0, limitBytes: 8192}`. Follow `nextOffsetBytes` until
null.

Related work: #14186 added the API fallback tools. #14218 bounded API
discovery.

## What Changed

- Add authenticated UTF-8 text windows to `call_api`, with byte offsets
and total size. Keep each page at or below 24 KiB.
- Stream new responses above 24 KiB through private temporary files into
company-owned assets. Bound each capture to 1 GiB of decoded bytes.
Reject oversized declared lengths before reading and count streamed
bytes before writing.
- Reserve capture budget in the run record before spilling. Allow 4 GiB
per run. Settle successful captures to their actual size. Failed or
interrupted captures retain their full 1 GiB reservation. Run restarts
do not reset the budget.
- Enforce a 20 GiB company snapshot quota with database reservations.
Count legacy snapshots and unfinished storage work across runs and
processes. Asset deletion frees quota.
- Limit large captures to two per company and four per server process.
Hold slots through storage upload and temporary-file cleanup. Use a
10-minute download deadline and 30-second connection/idle-read timeouts.
- Return explicit size, budget, busy, and timeout errors. Preserve
unknown outcomes for mutations whose response cannot be captured.
- Read saved assets through authenticated storage ranges, with at most
two extra bytes for UTF-8 and EOF handling. Unpaged reads return the
existing asset and digest with a bounded preview. Reads create no copies
and do not consume capture budget.
- Keep existing assets above 1 GiB readable in pages. Use safe integer
offsets and PostgreSQL `bigint` asset sizes.
- Stream large S3 uploads through ordered multipart requests. Abort
failed uploads and remove partial local files.
- Revalidate run authority during downloads. Keep company authorization,
GET-only text paging, redirect denial, and mutation replay receipts.
- Document the separate 10 MiB upload limits. This PR does not raise
memory-buffered attachment ingestion limits. Future large video uploads
need streamed ingestion and storage quotas.

## Verification

- Full workspace `pnpm -r typecheck` and `pnpm build` pass after
rebasing on master.
- Focused API and response tests: 1,761 pass. Cover declared and chunked
oversize responses, incorrect Content-Length, exact-limit success,
active-stream deadline, cancellation, cleanup, concurrency admission,
and mutation outcome handling.
- Real HTTP integration: 28 tests pass, including runnerd → PRP →
authority → HTTP, a 12 MiB snapshot, final-page/EOF reads, cross-company
denial, a persisted 3 GiB sparse asset, and large mutation receipt
replay.
- The HTTP suite verifies durable run-budget accounting, simultaneous
runs competing for company quota, legacy snapshot accounting, deletion
refunds, failed-storage reservations, cleaned-failure refunds,
metadata-rollback cleanup refunds, preservation after a lost commit
acknowledgement, and small/saved reads after capture-budget exhaustion.
- A standalone proof streams exactly 1 GiB through the production
capture helper, verifies the final bytes, and removes its temporary
file. It uses repeated 256 KiB chunks and records a peak process RSS of
191 MiB.
- Earlier storage verification covers exact S3 multipart boundaries,
cleanup/abort failures, and a 17 MiB transfer through the real AWS SDK
to a local HTTP S3 endpoint. No cloud S3 qualification was run for this
follow-up.
- The local full test run was interrupted for the company-quota changes.
A later targeted run hit exhausted macOS shared-memory slots before
tests started; two unattached PostgreSQL segments with dead owners were
reclaimed before retrying. All 55 current-head checks pass at
`aebb80ceeeee77d5a56b67bfffd835f2f846878c`, including the full CI test
suite, typecheck, build, browser suites, security scan, and Greptile
(5/5). There are no unresolved review threads. The combined rebased test
catalog also passes (48 tests).
- Earlier paging acceptance passed Daytona and separate staging at
`7739879e9`. Those runs predate the streaming and budget changes.

## Risks

- The 1 GiB response cap and 10-minute active-download deadline are
intentional product limits. Larger live results must use endpoint
pagination or a direct file workflow. Existing larger assets remain
readable through bounded ranges.
- A durable 20 GiB company snapshot quota counts stored runner-api
assets and active/orphan reservations across runs and processes. The
operator can set PAPERCLIP_RUNNER_API_COMPANY_CAPTURE_MAX_BYTES to a
finite value of at least 1 GiB. Deleting snapshots frees capacity;
possible orphan storage must be reconciled before releasing its
reservation.
- A failed capture uses its full reservation. A new large capture needs
a full 1 GiB available, even if it later completes at a smaller size.
Small reads and existing asset pages remain available.
- Concurrency limits apply per server process. The run byte budget is
shared through the database.
- The `integer` to `bigint` migration rewrites asset metadata and takes
an exclusive table lock. File bytes stay in storage.
- A live endpoint is fetched once before returning its snapshot.
Continue reading the saved artifact for stable pages. Mutations may
commit before any size or transport error; inspect state before
retrying.
- Attachment uploads and native file handoffs still default to 10 MiB.
Raising buffered ingestion paths to GiB sizes is separate work.

## Model Used

OpenAI Codex, based on GPT-6, with code execution and repository tools.
The runtime does not expose an exact serving model variant or
context-window size. The earlier paging work also used browser testing
and subagents.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.928.0-canary.2
2026-09-28 10:36:32 -05:00
DottaandPaperclip cea8dda472 test: evaluate completion updates after native task handoffs (#13969)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Users can delegate work through onboarding and Agent Chat.
> - A completed task does not prove that its result reached the original
conversation.
> - Existing tests do not isolate completion after the source chat
becomes idle.
> - This pull request adds four explicit native-runner probes across
Claude and Codex.
> - The probes preserve the result and reply so we can separate delivery
failures from inaccurate answers.

## Linked Issues or Issue Description

Refs #13775. Refs #13813.

These evals extend native-runner qualification. They measure completion
updates before we choose a product change.

## What Changed

- Add the opt-in `completion-updates` suite with two stories for each
native provider.
- Test completion in the existing onboarding task flow and after an
Agent Chat handoff becomes idle.
- Gate the chat worker on a brief inside its managed project workspace.
Prove the source is idle before releasing the worker.
- Check durable task completion, saved output, a subsequent source
reply, and rendered access to the result.
- Preserve replies, task state, screenshots, run events, and a separate
semantic review rubric.
- Add grader regression tests and update the documented eval contract.
- Preserve the suites added on master and include four completion cases
in the 306-cell catalog.

Production behavior and prompts are unchanged.

## Verification

- Passed all 565 eval support tests across 45 files after merging
current master: `node node_modules/vitest/vitest.mjs run --config
tests/runner-e2e/vitest.config.ts`.
- Passed eval TypeScript: `node node_modules/typescript/bin/tsc -p
tests/runner-e2e/tsconfig.json`.
- Confirmed four selected cells: `node cli/node_modules/tsx/dist/cli.mjs
tests/runner-e2e/launch.ts --list --suite completion-updates`.
- Four-cell behavior campaign on source
`ad47cf1da2b1e36f19f4227cfeb53998720b0b5b`:
https://github.com/paperclipai/paperclip/actions/runs/36072337485.
- A screenshot-only follow-up waits for the restored source reply to
render after result-link navigation. Its one-cell Claude onboarding
verification passed on final head:
https://github.com/paperclipai/paperclip/actions/runs/36075716141.
Corrected report:
https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36075716141-1/.
The original four-cell onboarding screenshots caught navigation loading;
its saved reply evidence remains valid. The follow-up again found stale
wording: "That work will run next" was posted 38 seconds after the child
was Done. The four-cell campaign keeps its original source and
measurements.
- Published evidence:
https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36072337485-1/.
- Suite definition:
`afba4d85d6c53d9f64c08b37a2e9cc20481b78f5bd7e2fa012045e2c69444d9d`,
version 6. Models: native `gpt-5.6-sol` and `claude-sonnet-5`, local
execution, one attempt per cell. All four cleanup checks passed.
Onboarding billing coverage is partial; reported zero cost must not be
read as a free run.

| Story | Automated delivery/access | Separate semantic review |
| --- | --- | --- |
| Codex onboarding | Pass | Pass: accurate completion reply with an
accessible result |
| Claude onboarding | Pass | Fail: reply says it will save the note once
the task runs, after the note is already saved and the task is Done |
| Codex idle chat handoff | Fail | Worker completed and saved the note;
no completion reply during the full observation window |
| Claude idle chat handoff | Fail | Worker completed and saved the note;
no completion reply during the full observation window |

Both chat cases positively recorded the source waiting and the worker at
the brief gate before release. Both saved outputs include the brief-only
start time. The opt-in campaign is red because it exposes current
behavior. It is not a required merge gate. The PR does not fix that
product behavior. Semantic review is a recorded human/agent assessment
of retained evidence; it is not an automated prose-quality judge.
- Second campaign:
https://github.com/paperclipai/paperclip/actions/runs/36071065098. Codex
chat reached the idle boundary and completed its task, then received no
completion reply during the full window. Claude onboarding again
returned a stale handoff answer. Claude chat exceeded the prior
110-second handoff setup budget; this revision raises that bounded setup
window to 180 seconds.
- Retained baseline:
https://github.com/paperclipai/paperclip/actions/runs/36069427676.
Onboarding passed delivery/access for both providers, but Claude gave a
stale handoff answer. Chat cases stopped at fixture problems; they do
not establish a completion-delivery failure. This revision fixes the
workspace path and competing reference requirements.
- On the previous head `4023a2a3c28d45c9eb2c42d452ce99ffba5c7b73`, 54 PR
checks passed and two were skipped, including typecheck, tests, and
build. Broad checks ran in CI, not locally. That head received Greptile
5/5 with no unresolved findings. The unchanged mobile
repository-settings browser test passed on one targeted retry after a
detached/disabled Save-button timeout.

- Merged current master in `9b4491e1f` and resolved the catalog-count
conflict. Eval support tests and eval TypeScript pass locally. All
individual CI jobs passed on this merge commit, including build,
typecheck, server tests, runner checks, and browser shards. The final
aggregate check also passed: 54 checks passed and two were skipped.
Greptile reviewed this exact commit at 5/5 with no unresolved findings.

## Risks

- These explicit probes can expose current product failures. They do not
change the default paid test selection.
- Mechanical delivery and result access do not establish answer
accuracy. The preserved reply still requires semantic review.
- A fixture failure before the idle boundary or worker completion cannot
establish a completion-update failure.
- The handoff setup window lasts three minutes. The worker brief wait is
bounded at four minutes. The observation window lasts two minutes after
worker completion. It retains later replies without erasing earlier
accessible delivery.

## Model Used

OpenAI Codex, GPT-6 (`gpt-6-astra`), with reasoning, repository
inspection, code execution, and GitHub tool use. The runtime does not
expose the context window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 09:58:14 -05:00
DottaandPaperclip 8751e2de46 fix(ui): distinguish finalization recovery from live observation (#14326)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The task board shows which recovery actions are active.
> - A native run can stop while a person must repair its workspace.
> - The board previously called that state “Recovery in progress.”
> - The label implied that work would continue without operator action.
> - This pull request derives the label from the recovery owner and live
continuation.
> - Operators can distinguish scheduled recovery from a repair that
needs attention.

## Linked Issues or Issue Description

**What happened?**
A blocked task showed “Recovery in progress” after native finalization
stopped and no automatic continuation remained.

**Expected behavior**
Show “Recovery needed” for an idle board repair or when no live recovery
path exists. Show “Recovery in progress” while the recorded continuation
can run, including an explicitly admitted export whose exact callback is
executing even if the old recovery action remains board-owned.

**Steps to reproduce**
1. Complete a native run whose workspace export cannot be recovered
automatically.
2. Inspect the task recovery action and badge.
3. Compare the board-owned action with the old “Recovery in progress”
label.

Related work: #14314 serializes native workspace finalization and fences
stale recovery outcomes. This change reports the recovery action that
currently owns the task.

## What Changed

- Show “Recovery needed” for board-owned active-run recovery unless the
exact native export callback is positively verified as executing.
- Require the native continuation run and a live or future continuation
before showing progress.
- Project native activity from the exact company, source issue, and run.
Include active workspace export while the original heartbeat remains
failed.
- Require an executing callback before using a running export row as
evidence. Preserve activity for long exports and clear it when the
callback joins.
- Give native resume its own card explanation. Preserve ordinary
watchdog observation behavior.
- Remove the redundant ownership sentence from all six recovery-card
explanations that used it.
- Document the labels and add regression cases for stopped, scheduled,
and active recovery.

## Verification

- Copy-only follow-up (`c68aef04c`): all 167 focused recovery UI tests
and token gates pass. No UI occurrence of the removed sentence remains.
`pnpm -r typecheck`, `pnpm build`, and current-head CI pass (54
successful checks, two optional Storybook checks skipped). Greptile is
5/5 with no open review threads. The duplicate local `pnpm test:run` was
stopped after the full CI suite passed; it did not complete locally.
- Original regressions: seven failures before the change, then 52
focused cases pass.
- Review regressions: seven UI failures and nine database failures
before the follow-up. All 75 database/API recovery tests, 167 UI tests,
and 18 workspace lifecycle/finalizer tests pass. A further three RED
cases cover explicit board retry activity; one RED case rejects orphaned
running export rows after controller loss. Wrong company, issue, run,
service, phase, and completed-operation cases remain inactive.
- Final recursive typecheck, production build, token gates, and complete
local suite coverage pass. Embedded PostgreSQL startup/socket failures
passed in isolated retries with the canonical test environment; no
expected behavior was weakened. The recovery regression added during the
earlier full run passed in its final complete 75-case file.
- Before the copy-only follow-up, all 56 CI checks passed on
`c2f84cd89c905cda85c53aaf5bb83b7250900fe6`; Greptile is 5/5 with no
unresolved review threads. The final native-activity staging repeat
passed on integrated source `2bedd0f23bf4698b1f8b818f6796900647030427`.
- Verified on a separate staging instance: a real failed Daytona
workspace export retains its board-owned repair action and displays
“Recovery needed” in the task list. The repair card remains actionable
without starting another provider turn.

- Real staged export-only repair: the actual task list showed “Recovery
in progress” while the original run had a positively identified running
export operation, then Done after exact copyback of all 20,000
nonce-bound files. The accepted result and full provider
session/turn/terminal envelopes remained unchanged. All 17 independent
final checks passed; the browser downloaded the exact 19-byte result.
The separate fixture was cleaned up with independent provider-absence
verification. A control transport process restarted during repair; it
did not submit another provider turn.

- Final deployed-source repeat on
`d884e1ab046cc76004e35e6091e9e6e2c918c9eb`: explicit per-turn ephemeral
Daytona allocation, actual browser export repair, and a saved full
activity projection referencing the exact executing export operation
with no scheduled retry. The task list showed recovery in progress, then
Done; all 21 final checks passed, including 20,000 exact host files,
unchanged provider provenance, and provider deletion only after
committed copyback. The downloaded 19-byte result matched independently.
This integrates #14334; no source change was required here.

## Risks

- The label depends on the persisted recovery action. A separate runtime
defect can still stop work; this change makes that condition visible.
- Future recovery kinds must supply a valid continuation path before
they can display progress.

## Model Used

OpenAI Codex, based on GPT-6, with code execution, browser testing, and
subagent tool use. The runtime does not expose an exact serving model
variant or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.928.0-canary.1
2026-09-28 09:33:39 -05:00
DottaandPaperclip cbc5132e6c fix(adapters): expose a verified provider stop before workspace restoration (#14311)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agent adapters own local or remote provider processes.
> - Run cleanup must know when the final provider process has stopped.
> - A remote timeout or a lost transport does not prove that the
provider stopped.
> - Retried provider invocations also make an earlier stop signal stale.
> - This pull request adds a verified final-invocation stop callback
before workspace restoration.
> - The dependent instruction revision change uses that boundary to
preserve private instruction edits safely.

## Linked Issues or Issue Description

**What happened?**

Adapter completion did not expose a reliable point between provider
shutdown and workspace restoration. Cleanup could lose provider-written
files, or treat a remote timeout as proof that a process stopped.

**Expected behavior**

Cleanup runs once after the final provider invocation has a verified
stop receipt and before workspace restoration. An incomplete remote
command keeps collection pending.

**Steps to reproduce**

1. Run a remote provider that returns a timeout without a numeric exit
code.
2. Let the adapter return or retry the provider.
3. Attempt to collect provider-written files during cleanup. The adapter
has no verified final-process boundary to use.

This is the prerequisite for the stacked canonical instruction revision
pull request. It has no database or UI dependency. Related #13291
verifies remote termination for later recovery; this change exposes the
earlier adapter-owned stop boundary before workspace restoration. The
scopes do not duplicate each other.

## What Changed

- Add a stop callback to adapter execution context and a
final-invocation fence.
- Confirm local child closure and complete remote exit receipts. Reject
remote timeouts, missing exits, transport failures, and SSH exit 255 as
stop proof.
- Invoke collection once before workspace restoration in eight CLI
adapters and at the confirmed ACP stop boundary.
- Preserve the stop observation when a later log flush fails.
- Run bridge and workspace cleanup in `finally` even when collection
rejects. ACP records a safe error without exposing a raw filesystem
path. Grok keeps collection errors separate from workspace restore
failures and preserves completed provider results when both cleanup
steps fail.
- Correct the existing Cursor test shell fixture so bounded remote file
reads run against real fixture files.

## Verification

- Three stop-boundary regressions failed before the callback
implementation and passed after it.
- Independent prerequisite branch: 377 tests passed across 24 adapter,
process-target, and ACP suites. Two added collector-rejection tests
failed before the cleanup fix and passed after it.
- A third regression reproduced Grok misclassifying a collection failure
as failed workspace restoration. Two additional cases covered completed
and failed provider turns when collection and restore both fail. The
Grok and restore-classifier suites passed 48 tests.
- All nine affected package typechecks and affected package builds
passed; Grok checks passed again after its classification fix.
- Integrated instruction branch: native local, legacy local, and native
Daytona each passed three browser tasks with exact persisted bytes,
fresh-task readback, history/restore, and explicit conflict resolution.
Unchanged warm Daytona passed three turns. Legacy Codex passed all five
checks again after the exception-safe cleanup fix.
- Alternate staging passed the same three-task native Daytona flow:
exact stopped-run save, independent downloaded readback, browser
history/restore, and explicit resolution of a real concurrent edit. The
deployed source was `14c3d810c9e05625121b3d27767aea9317b03125`, which
covers the initial adapter callback. Later Grok cleanup failures are
qualified by the adapter tests above.

- The final combined native Daytona flow passed again on deployed
`2bedd0f23bf4698b1f8b818f6796900647030427`: three fresh tasks proved
ordinary instruction edits, independent readback, History/Restore,
concurrent board conflict, and explicit candidate resolution. This
native staging flow does not claim to exercise the Grok adapter.

## Risks

- An unverified remote stop intentionally does not trigger collection. A
later controller with verified stop evidence must recover it or report
the copy unavailable.
- The callback is optional. Callers that do not register it retain their
existing behavior.
- The callback runs before workspace restoration and can delay cleanup
if its caller does not bound its own work. The dependent instruction
collector uses bounded reads and retries. A rejected callback still
permits bridge and workspace cleanup; it cannot claim an instruction
save.

## Model Used

OpenAI Codex, GPT-6, with tool use and code execution. The runtime does
not expose the exact deployment variant or context-window size. Multiple
Codex agents implemented and verified the change.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 08:45:53 -05:00
DottaandPaperclip 4b38db9622 fix(runtime): stream workspace Git snapshots through disk manifests (#14253)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Managed runs copy a selected workspace to an execution environment
and restore its changes.
> - Git snapshots select the files for that copy and for later recovery.
> - A fixed output limit stops large generated trees before the run can
start.
> - Increasing the limit still keeps the complete filename lists in
memory.
> - This pull request stores those lists and merge baselines in disk
manifests.
> - Large snapshots can now complete with bounded filename buffers and
explicit failure handling.

## Linked Issues or Issue Description

Refs #14194. This is the streaming follow-up to the merged 32 MiB limit
fix.

Related: #13619 and #11621 cover workspace scan admission and demand.
This change keeps the shared scheduler and changes the snapshot data
path.

## What Changed

- Stream changed, untracked, deleted, and ignored paths through the
shared scheduler and the standalone adapter path.
- Use SQLite manifests for file selection, duplicate removal,
ignored-path lookup, baseline capture, and merge lookup.
- Set a configurable 30-minute snapshot deadline. Keep the existing
interactive scan deadlines.
- Wait for each child process and pending sink write before removing
temporary storage after failure or cancellation.
- Use NUL archive lists and bounded deletion batches. Preserve unusual
names, explicit selection, nested repositories, and source root checks.
- Store manifest references in native recovery format v2. Check their
location and digest before recovery reads. Keep v1 descriptors readable.
- Remove temporary manifests at lifecycle completion. Use fixed-size
temporary copy names for long basenames.
- Admit each manifest with a SQLite page allowance based on current disk
capacity. Keep a configurable free-space reserve and fail explicitly
when either limit is reached.
- Preserve a host file that replaces a directory deleted by the sandbox,
and continue the rest of the restore.

## Verification

- Current head `5ab622ec43cd16d35429d79dedee6a5d8e3d2df2` has 54
successful checks/statuses and two skipped Storybook jobs. No checks
failed or remain pending.
- [CI
passed](https://github.com/paperclipai/paperclip/actions/runs/36318966368):
typecheck, build, all test shards, E2E, Rust checks, and the aggregate
verify job.
- [Greptile is
5/5](https://github.com/paperclipai/paperclip/pull/14253#issuecomment-5855670338)
on the current head. All four review threads are resolved. Security
checks passed.
- 229 focused tests passed across Git sync, runtime staging, merge,
manifest integrity, native recovery, and the scheduler (214
adapter/runtime tests and 15 scheduler tests).
- A real 40,000-file fixture produces 43,428,890 filename bytes. The
original standalone and scheduled scans fail. The new test passes all
four filename paths, complete staging, exclusion of late files, unusual
names, and deletion replay.
- Recovery tests reject changed bytes, symlinks, and paths outside the
controller state directory. Adapter-utils typecheck passed.
- A test executor returned buffered output and caused two retry
integration failures. The fixture now uses the shared streaming
scheduler. All 13 tests passed with `corepack pnpm exec vitest run
server/src/__tests__/heartbeat-project-repositories.test.ts`. The same
CI shard now passes.
- Ran `pnpm -r typecheck`, `pnpm test:run`, and `pnpm build` locally.
Each full local command hit SIGKILL/exit 137 in the 4 GiB container.
These local commands did not pass. The current-head CI gates above
provide the full verification.
- A real-Git disk-capacity regression confirms a typed failure and
removal of the incomplete manifest. Repeated writer attempts cannot
exceed the permitted page count.

- Follow-up real Daytona and separate staging qualification passed with
the related archive validator (#14315) and exact-owner finalization fix
(#14314). Three successive turns copied back all 60,000 files with
39,828,890 filename bytes and five unusual names. Independent host
inventories verified every file and the pinned Git HEAD. Native,
provider, session, and process identities stayed fixed; no retry
remained. The task reached Done, and its browser-downloaded final proof
matched exactly. The reusable regression is #14316, including an
assertion of the effective environment idle policy.

## Risks

- SQLite manifests use disk space. Each receives one quarter of the
available capacity above the host reserve at creation. The reserve
defaults to 256 MiB and has a 64 MiB configuration minimum. Disk
capacity, filesystem quotas, per-path limits, Git resource use, and
execution deadlines remain limits.
- Each path and sink chunk has a 64 KiB limit. SQLite connections use a
1 MiB page cache. Invalid or incomplete records fail explicitly.
- Restore transport keeps fixed and configured archive exclusions. A
remotely created Git-ignored file can be transferred, but the host merge
excludes it through the manifest.
- Provider archive buffers, Git and tar memory, repository metadata,
legacy v1 arrays, and the separate referenced-source resolver retain
their own limits. Existing provider safety validators still buffer
textual tar listings: Daytona allows 32 MiB and Kubernetes allows 64
MiB. These separate transport limits can stop a sufficiently large
restore before merge. This change does not claim bounded total process
memory or unlimited transport size.
- New descriptors use v2. Existing v1 recovery remains supported; a
downgrade cannot read v2 descriptors.

## Model Used

OpenAI GPT-6 through Codex. The exact deployment ID and context limit
are not exposed in this run. The agent used code editing, terminal
execution, tests, and GitHub tools.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 08:40:36 -05:00
DottaandPaperclip 890d11137f fix(ui): register artifact tabs without opening the panel (#14193)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Tasks keep agent outputs in the Artifacts tab.
> - An output can arrive while the user writes a message or reads a
document.
> - Opening the side panel on arrival interrupts that work, especially
on mobile.
> - This pull request adds the Artifacts tab without opening the panel
or changing the selected tab.
> - Users can open their outputs when they choose.

## Linked Issues or Issue Description

**What happened?**
New agent outputs opened the task side panel or mobile drawer. An
arrival could also replace the selected document or workspace file.
Existing outputs did not always register an Artifacts tab.

**Expected behavior**
Register one Artifacts tab for existing and new outputs. Keep a closed
panel closed. Preserve composer focus, the selected tab, and document or
file links.

**Steps to reproduce**
Open a task from the inbox. Close its side panel. Enter a message draft.
Create an agent output in that task. The panel must stay closed and the
draft must keep focus. Open the panel to see the Artifacts tab. Repeat
on a mobile viewport.

**Paperclip version or commit**
Base commit: 0f14d2612.

Related work: #11226 and #11551.

## What Changed

- Register existing outputs and later arrivals without opening the panel
or selecting Artifacts.
- Keep open documents, workspace-file links, and the tab launcher
unchanged.
- Handle each document deep-link request once so query refreshes
preserve later manual selection.
- Deduplicate attachment and work-product arrivals. Preserve dismissed
tabs across repeated refreshes.
- Add desktop and mobile browser regression tests. Update artifact
presentation documentation.

## Verification

- The closed-panel regression failed before the fix in unit and
real-browser tests.
- All 158 focused UI tests pass, including the original deep-link cases.
- UI typecheck and token gates pass on this branch. Full typecheck and
production build passed on the passive-arrival candidate before the
existing PR integration.
- The two local desktop/mobile browser cases pass against real
API-created artifacts.
- Desktop and mobile staging checks pass on the combined staging
candidate. Artifact arrival preserved a closed pane, draft text, and
composer focus. Explicitly opening the pane showed the Artifacts tab.
- All 56 current-head check contexts are successful or intentionally
skipped at `17ad904455b9378552f07a6f6e51402c6d164688`, including full
typecheck, test shards, build, and browser suites. Greptile is 5/5 on
that commit with no unresolved review threads.
- The interrupted local broad validation was resumed; the remaining
serialized 64 files and 1,035 tests pass. Local database startup
failures passed after stale test resources were released.

## Risks

- Outputs no longer reveal the panel automatically. Users open the panel
to view them.
- The document request guard must still allow a new explicit deep link.
The regression tests cover this case.
- There are no API or database changes.

## Model Used

- OpenAI Codex, GPT-6, with code editing, shell tools, GitHub tools, and
browser verification. The exact serving model ID and context-window size
are not exposed in this session. Earlier implementation model metadata
is not available.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 08:38:00 -05:00
DottaandPaperclip bacc0e6a98 docs: remove automatic agent escalation from coordination skill (#14188)
## Thinking Path

> - Paperclip manages work for AI agents.
> - The coordination skill tells agents how to handle blocked work.
> - The skill directs blocked work to managers and other agents.
> - Those agents may lack the same access or authority.
> - These extra assignments delay the required human action.
> - This pull request removes automatic escalation advice.
> - Agents must identify the missing capability and use the correct
approval or human-input path.

## Linked Issues or Issue Description

**Issue type**

Incorrect information.

**Where is the issue?**

The critical rules in `skills/paperclip/SKILL.md` and blocker guidance
in `skills/paperclip/references/api-reference.md`.

**What's wrong?**

The skill tells agents to escalate through `chainOfCommand`, ask another
agent for help, and avoid human help. A manager title does not grant
permission to fix a connection or complete an administrator action.

**Suggested fix**

Remove blanket escalation and agent-first rules. Keep normal delegation
when the recipient has a concrete capability for a bounded task. Use
saved human-input interactions or existing connection and approval flows
for human-only actions.

Searches for open PRs with “escalation”, “chainOfCommand”, and “ask
another agent” found no duplicate skill change. Recovery-routing PRs
change server behavior, which is outside this change.

## What Changed

- Remove the chain-of-command escalation rule and both repeated
agent-first directives.
- Replace manager handoffs in the API reference with direct blocker
handling.
- Keep reporting fields, normal delegation, approval gates, and the ban
on bypassing permission denials.
- Keep the ban on cancelling cross-team tasks. Request a decision
instead of automatically assigning the task to a manager.

## Verification

- Final head `1483e82cc3fe10a7c910230b3779635d6d54c162`: Greptile 5/5,
no outstanding findings, all CI checks pass (optional Storybook jobs
skipped).

- Created a fresh worktree at `.worktrees/skill-blocker-guidance` from
the current master commit.
- Ran `git diff --check`: passed.
- Ran Node assertions against both documents: passed. Removed directives
are absent. Human-input and capability-based delegation guidance is
present.
- Checked that checkout conflict, approval, dependency, and normal
delegation instructions remain.
- Attempted `pnpm -r typecheck`, `pnpm test:run`, and `pnpm build`. None
could run because this environment has no `pnpm` executable.
- Refreshed both generated capability inventories and their derived
contract after skill heading positions changed. Both generator and
live-inventory checks pass. The eval and MCP baselines remain unchanged.
- Ran `node --test
packages/paperclip-runner/scripts/check-capability-inventory.test.mjs`:
all four tests pass.
- Ran `node
packages/paperclip-runner/scripts/check-capability-inventory.mjs` and
`node packages/paperclip-runner/scripts/generate-capability-contract.mjs
--check`: both pass.
- Added explicit requester routing for agent and human scope questions.
Focused Node assertions pass.

## Risks

- Agents can request human input earlier for actions that require human
authority.
- Existing runs or installed copies can keep old skill text until
refreshed.
- This change does not alter server recovery routing or permission
checks.

## Model Used

OpenAI Codex, with reasoning, tool use, and shell execution. The runtime
does not expose a verifiable exact model ID or context window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.928.0-canary.0
2026-09-28 08:37:26 -05:00
DottaandPaperclip 0f14d26123 fix(runtime): allow bounded large untracked workspace snapshots (#14194)
## Thinking Path

> - Paperclip manages AI agents and their work.
> - Remote runs need a snapshot of the task workspace before the agent
starts.
> - The snapshot lists untracked filenames through the shared Git scan
scheduler.
> - A generated directory with a few thousand long filenames can exceed
the 1 MiB output limit.
> - This stops setup and prevents the agent from continuing its task.
> - This change gives that listing a 32 MiB bound and keeps the explicit
file snapshot.
> - Normal generated trees can now pass setup, while larger snapshots
still fail at a finite limit.

## Linked Issues or Issue Description

Refs #11572 and #12214 for the existing bounded scan and ignore-scan
protections.

**What happened?**

An agent continuation failed during workspace setup with `Workspace Git
scan exceeded its output limit`. A real Git fixture reproduces the
untracked-file path: 5,000 long filenames in one generated directory
exceed its 1 MiB output limit.

**Expected behavior**

The workspace snapshot must support ordinary generated trees with
thousands of files. It must keep a finite output bound and select
explicit files before staging.

**Steps to reproduce**

1. Commit a base file in a Git repository.
2. Add 5,000 untracked files with long names in one new directory.
3. Call `readGitWorkspaceSnapshot` with the normal scan limits.
4. Observe the output-limit error before this change.

**Paperclip version or commit**

Base commit: `640dee1`.

**Deployment mode**

Remote sandbox execution from source.

## What Changed

- Increase the untracked-file snapshot output bound from 1 MiB to 32
MiB.
- Keep explicit file selection, the shared scheduler, the timeout, and
the other scan bounds.
- Test that 40,000 long filenames above the old 8 MiB bound reach the
snapshot. Reuse those files with deeper paths to prove that output above
32 MiB still fails.
- Test that files created after the snapshot, including an ignored
secret, stay out of the overlay archive.
- Check the workspace root identity and reject selected paths with
symlinked parent directories before upload.
- Test root replacement after path resolution, including root-level
selected files.
- Preserve existing workspace root aliases by capturing the resolved
root before snapshot selection. Test that later alias retargeting cannot
change the archive contents.
- Stop staging on permission and I/O errors; continue to allow missing
files.
- Test these failures and preserve selected symlink entries.
- Accept valid case-renamed directories by checking ancestor file types.
A modeled case-insensitive regression failed before this correction and
now passes.
- Give the 40,000-file fixture enough time to remove its files.
- Document the larger bound and staging behavior.

## Verification

- Red: the 5,000-file regression failed with `stdout maxBuffer length
exceeded` before the fix.
- A separate check through the real server scheduler reproduced
`workspace_git_scan_output_limit` on the original code. The revised code
selected all 5,000 files.
- The late-file regression failed against the first PR revision because
the archive contained `drafts/late.secret`. It passes with the final
explicit-file approach.
- The staging regressions failed before the review fix: a substituted
parent directory and permission/I/O errors were accepted. All three
cases now stop before upload.
- Green: 134 tests passed across `git-workspace-sync.test.ts` and
`sandbox-managed-runtime.test.ts` with Vitest 4.1.11. This includes
complete selection above 8 MiB and rejection above 32 MiB.
- `pnpm --filter @paperclipai/adapter-utils... typecheck` passed after
the revision.
- The module-boundary check and `git diff --check` passed.
- `pnpm -r typecheck` and `pnpm build` stopped in the Rust runner steps
because this environment has no `cargo` executable.
- The full `pnpm test:run` attempt ended with `SIGKILL` during the
general server suite. It did not finish. That full-suite result belongs
to the earlier revision. Fresh checks are required for this revision.

- The new regression fails at the old 8 MiB bound with `stdout maxBuffer
length exceeded`. All 134 focused tests pass with the 32 MiB change.
- The affected typechecks and module-boundary check pass. The full local
typecheck requires Cargo, which is absent in this environment.
- Final verification for `b43e9bc95e55382c6a9bfe200487c164770c8be8`: 54
successful checks/statuses and two skipped Storybook checks. No checks
remain pending or failed.
- The [CI
run](https://github.com/paperclipai/paperclip/actions/runs/36315569674)
passes on attempt 2. The first attempt had one unrelated preview-fixture
readiness timeout. That exact test passed locally; its CI shard passed
on the single rerun.
- [Greptile reports
5/5](https://github.com/paperclipai/paperclip/pull/14194#issuecomment-5852045501)
on this revision. All review threads are resolved.
- [Security review accepts the documented memory
tradeoff](https://github.com/paperclipai/paperclip/pull/14194#discussion_r4115149465)
for this finite mitigation. The separate streaming follow-up will remove
full-list buffering.
- This revision also passes affected local typechecks and the
module-boundary check. Full local typecheck/build stop because Cargo is
absent. The full local Vitest attempt was stopped after about 18 minutes
once all remote gates passed; it did not complete locally.

## Risks

- Each untracked-file scan can buffer up to 32 MiB instead of 1 MiB. The
scheduler still limits concurrent scans and execution time.
- Snapshots above 32 MiB still fail with the existing error. Tracked and
ignored-file scan bounds stay unchanged.
- A workspace that replaces a selected path’s parent with a symlink now
fails staging.
- Detailed logs from the reported host were unavailable. The exact
command that exceeded its limit on that host is unconfirmed.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, tool use, and code
execution. The runtime does not expose a more specific model build or
context window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub references)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-27 06:42:19 -05:00
DottaandPaperclip 7e0f051d3b Align task property icons with assignee avatars (#13319)
## Thinking Path

> - Paperclip is the open source app that people use to manage AI agents
for work.
> - The task properties panel helps operators inspect and update task
data.
> - The status, assignee, and project rows used different leading icon
sizes.
> - The mixed sizes made the rows look uneven.
> - The assignee avatar already provides the correct 24px visual size.
> - This pull request sets the status and project visuals to that same
size.
> - The benefit is a tidy and consistent properties panel.

## Linked Issues or Issue Description

**What happened?**

The task properties panel showed a 12px status icon, a 24px assignee
avatar, and a 16px project tile.

**Expected behavior**

The status, assignee, and project rows must use the same 24px leading
visual size.

**Steps to reproduce**

1. Open a task.
2. Open the Properties panel.
3. Compare the Status, Assignee, and Project rows.

## What Changed

- Keep the Status glyph at 16px and use spacing to give it the 24px
avatar footprint.
- Set the Project row tile to the 24px small tile size.
- Add a regression check for the Status row size.

## Verification

- `pnpm exec vitest run ui/src/components/IssueProperties.test.tsx` (72
tests passed)
- `pnpm check:token-gates` (all gates clean)
- `pnpm --filter @paperclipai/ui typecheck` (passed)
- Full repository typecheck and build reached the Rust runner step, but
this environment does not contain `cargo`.
- The full test command was stopped after unrelated chat integration
tests did not finish. The focused UI suite passed.

## Risks

- Low risk. This change only changes visual sizes in three property
rows.
- The wider status footprint can use slightly more horizontal space in a
narrow panel.

> This fix does not overlap with planned core work in `ROADMAP.md`.

## Model Used

- OpenAI Codex, GPT-5.6, tool use and code execution enabled.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.927.0-canary.5 nightly/v2026.928.0-nightly.0
2026-09-27 11:35:39 +00:00
DottaandPaperclip d423fc4411 fix(ui): use large mobile selector modals (#14250)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The new task dialog lets an operator assign work on a phone.
> - The mobile picker used fixed coordinates inside the Radix wrapper.
> - iOS Safari already makes fixed coordinates relative to the visual
viewport.
> - The old code added the visual viewport offset a second time.
> - This pull request removes the second offset and gives the mobile
wrapper full viewport geometry.
> - The benefit is a visible picker and search field when the iOS
keyboard opens.

## Linked Issues or Issue Description

**What happened?**

The assignee and project pickers could move outside the visible screen
in the new task dialog on iOS.

**What did you expect to happen?**

The picker and its search field must stay visible above the bottom edge
and the software keyboard.

**Steps to reproduce**

Open the new task dialog on an iPhone. Select Assignee or Project. The
picker can move outside the visual viewport when Safari pans the page.

**Paperclip version or commit**

The problem exists after the change in PR #13343.

**Deployment mode**

The problem affects the board UI in local and authenticated modes.

Refs #13343

## What Changed

- Use visual viewport local coordinates for the new task dialog.
- Give the mobile Radix popper wrapper full viewport geometry.
- Render both shared selector components as large, titled mobile modals.
- Keep desktop selectors as anchored popovers.
- Cover the new-task assignee, new-task project, composer assignee, and
generic searchable selector in Storybook.
- Add iPhone-sized Storybook states for the assignee and project
pickers.
- Update viewport tests for the corrected coordinate model.

## Verification

- `pnpm --filter @paperclipai/ui exec vitest run
src/components/NewIssueDialog.test.tsx
src/components/InlineEntitySelector.test.tsx`
- `pnpm check:token-gates`
- Playwright used the iPhone 14 device profile. All mobile modals were
inside the 390 by 664 CSS viewport.
- Each mobile modal was `x=16, y=16, width=358, height=632`.
- A desktop check kept the anchored popover at `width=320, height=211`.

## Risks

- Low risk. The CSS only changes narrow mobile viewports.
- The full-screen popper wrapper does not receive pointer events. The
picker still receives pointer events.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. See `CONTRIBUTING.md`.

## Model Used

- OpenAI GPT-5.6, Codex agent with tool use and code execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal task
ID
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant Storybook documentation to reflect my
changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-27 11:33:24 +00:00
DottaandPaperclip 74508357fa fix(ui): keep mobile task pickers in view (#14249)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The new task dialog lets an operator assign work on a phone.
> - The mobile picker used fixed coordinates inside the Radix wrapper.
> - iOS Safari already makes fixed coordinates relative to the visual
viewport.
> - The old code added the visual viewport offset a second time.
> - This pull request removes the second offset and gives the mobile
wrapper full viewport geometry.
> - The benefit is a visible picker and search field when the iOS
keyboard opens.

## Linked Issues or Issue Description

**What happened?**

The assignee and project pickers could move outside the visible screen
in the new task dialog on iOS.

**What did you expect to happen?**

The picker and its search field must stay visible above the bottom edge
and the software keyboard.

**Steps to reproduce**

Open the new task dialog on an iPhone. Select Assignee or Project. The
picker can move outside the visual viewport when Safari pans the page.

**Paperclip version or commit**

The problem exists after the change in PR #13343.

**Deployment mode**

The problem affects the board UI in local and authenticated modes.

Refs #13343

## What Changed

- Use visual viewport local coordinates for the new task dialog.
- Give the mobile Radix popper wrapper full viewport geometry.
- Keep the mobile picker at the visible bottom edge.
- Add iPhone-sized Storybook states for the assignee and project
pickers.
- Update viewport tests for the corrected coordinate model.

## Verification

- `pnpm --filter @paperclipai/ui exec vitest run
src/components/NewIssueDialog.test.tsx
src/components/InlineEntitySelector.test.tsx`
- `pnpm check:token-gates`
- Playwright used the iPhone 14 device profile. Both picker boxes were
inside the 390 by 664 CSS viewport.
- The assignee picker box was `x=16, y=366, width=358, height=282`.
- The project picker box was `x=16, y=366, width=358, height=282`.

## Risks

- Low risk. The CSS only changes narrow mobile viewports.
- The full-screen popper wrapper does not receive pointer events. The
picker still receives pointer events.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. See `CONTRIBUTING.md`.

## Model Used

- OpenAI GPT-5.6, Codex agent with tool use and code execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal task
ID
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant Storybook documentation to reflect my
changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-27 11:02:15 +00:00
Devin FoleyandPaperclip be43c23e2d fix(recovery): preserve pending native result finalization (#14219)
Preserve native coordinator ownership while accepted results await workspace
copy-back, assessment, or arbitration. Check that ownership in the terminal
update so a coordinator recorded after the liveness read is also protected.
Keep terminal task authority and exhausted-retry cleanup unchanged.

Validation: 28 focused recovery tests, local typecheck/build, 52 passing CI
checks, and Greptile 5/5 with no open findings.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
canary/v2026.927.0-canary.4 nightly/v2026.927.0-nightly.0
2026-09-27 00:22:49 -07:00
Devin FoleyandPaperclip 5ee9e751fb fix(runner): discover assigned tools when direct catalogs exceed limits (#14218)
Keep assigned app tools accessible when the combined Runner catalog exceeds
its operation or byte limits. Reserve task and completion tools, then expose
bounded discovery and call tools for large catalogs. Fetch oversized schemas
in reauthorized chunks without blocking later search results.

Retain task ownership, work-mode restrictions, pinned assignments, current
gateway authorization, approvals, and audit. Small catalogs stay direct.

Validation: 46 focused server tests, two Runner capacity tests, local
typecheck/build, 52 passing CI checks, and Greptile 5/5 with no open findings.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
canary/v2026.927.0-canary.3
2026-09-27 00:22:16 -07:00
DottaandPaperclip 640dee1802 fix(runner): acquire reusable leases when switching to warm sessions (#14187)
## Thinking Path
> - Paperclip coordinates agent work across task turns.
> - Warm runners need reusable sandbox leases.
> - Lease acquisition read only the environment reuse setting.
> - Startup then read the agent's new warm setting and rejected the
ephemeral lease.
> - This fix requests reuse before acquisition, with provider checks
intact.

## Linked Issues or Issue Description
**What happened?**
Switching an existing native runner task from per-turn to warm fails
with `runner_warm_environment_requires_reusable_lease`. Environment
`reuseLease` defaults to false. Acquisition therefore creates an
ephemeral lease, and capability narrowing correctly denies reuse.

**Expected behavior**
Changing lifecycle between turns requests a compatible lease without
replacing the task workspace or changing the shared environment.

**Steps to reproduce**
1. Start a native runner task with inherited lifecycle and environment
reuse disabled.
2. Change the agent from per-turn to warm.
3. Continue the task on a reusable-capable sandbox provider.

**Paperclip version or commit**
Reproduced on `a6c4e7a`. Related: #12904 established warm workspace
continuity; this fixes the earlier lease-selection mismatch.

## What Changed
- Pass agent settings into acquisition and derive a run-scoped reuse
request.
- Respect explicit environment lifecycle overrides; leave other adapters
unchanged.
- Keep provider capability, ownership, cleanup and restore gates intact.
- Add orchestration and database-backed transition tests; document the
contract.

## Verification
- Red: three new assertions failed before the fix.
- Green: 156 tests pass in environment-run-orchestrator,
environment-runtime, native-sandbox-lifecycle, and
environment-execution-target-capabilities.
- Database-backed regression proves ephemeral → reusable → resumed
lease, same workspace, and unchanged stored environment. Provider RPCs
are mocked, not live Daytona.
- `git diff --check` passes.
- Latest-head CI passes typecheck, build, tests, runner checks, E2E, and
canary dry run. Greptile: 5/5; no open threads.
- Recovery uses the persisted lifecycle. Direct regression: red before,
six cases green after. CI runs the added Vitest cases.
- Full local checks were unavailable (missing dependencies and earlier
memory limits).

## Risks
Warm mode now requests retained resources despite an environment's
default `reuseLease:false`. Existing idle timeout and cleanup still
apply. Unsupported providers remain denied. No migration, credential
[REDACTED], or shared-environment mutation.

## Model Used
OpenAI Codex agent; reasoning, code editing and test execution. Exact
model ID and context size were not exposed to this run.

## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details available)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links
- [x] My branch name describes the change and contains no internal
Paperclip ticket [REDACTED] or instance-derived details
- [x] I have run targeted tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-26 21:23:33 -05:00