Commit Graph
4623 Commits
Author SHA1 Message Date
Devin FoleyandPaperclip 4f6cf5b3ff fix: prevent run identity locks from blocking audit checks (#14478)
Use NO KEY UPDATE for identity locks so audit foreign-key checks can proceed while identity writers remain serialized. Preserve task-before-run ordering, company scoping, and foreign keys.

Verified with PostgreSQL concurrency regressions, focused tests, typecheck, build, and full CI.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-28 16:16:05 -07:00
Devin FoleyandPaperclip faf72cb1d5 fix: preserve company context across browser hot reload (#14482)
## Thinking Path

> - Paperclip uses a shared company context for browser providers and
consumers.
> - Vite can load a new consumer module while an older provider is
mounted.
> - Recreating the context disconnects that consumer from the mounted
provider.
> - The consumer then reports a missing provider even though one is in
its React ancestry.
> - This change preserves the context object across development module
refreshes.
> - Bounded global error diagnostics distinguish development bundles and
otherwise context-free promise rejections.

## Linked Issues or Issue Description

**What happened?**
A refreshed company consumer can throw `useCompany must be used within a
CompanyProvider`. A real Vite and Chromium reproduction confirms that a
retained provider and a refreshed consumer can hold different context
objects. Global promise rejections also lack the bounded document state
already attached to React boundary errors.

**Expected behavior**
A refreshed consumer should read the mounted provider. Error reports
should identify the loaded bundle mode and bounded browser state while
preserving monitoring opt-in, sign-out, and privacy behavior.

**Steps to reproduce**
Run `pnpm test:e2e:browser-context`. The isolated Vite fixture renders
the real CompanyProvider, imports a new timestamped consumer module, and
renders that consumer below the retained provider. The test fails before
the context change and passes after it. The SDK regression invokes its
real unhandled-rejection handler with an undefined reason.

## What Changed

- Keep the React context object in Vite's per-module `hot.data`. Account
values stay in React.
- Add an isolated browser regression with mocked API responses and no
live instance, discovered by the existing Chrome CI shards.
- Add document-state diagnostics to global errors while preserving
earlier boundary snapshots.
- Tag events with development or production bundle mode and the type of
an unhandled rejected value.
- Document the new test command and diagnostic fields.

## Verification

- Company context, browser context, and Sentry suites: 59 passed.
- `pnpm test:e2e:browser-context`: passed in Chromium. The original
context code fails the reproduction.
- Real SDK tests preserve DSN/sign-out behavior and omit request
context, breadcrumbs, and private DOM data.
- `pnpm check:token-gates`: passed.
- `pnpm -r typecheck`: passed.
- `pnpm build`: passed.
- Local `pnpm test:run` exited in the general-server phase: 13,819
passed, 86 skipped, 21 failures in unchanged filesystem and host-port
suites. Cache permission and long-path failures also reproduce on the
unmodified base; seven other failures involve local runtime port
ownership. This is not a full local pass.
- Full Linux CI and review passed at a0b5270662 (53 successful checks;
two conditional checks skipped). Three jobs interrupted by a runner
shutdown passed on rerun. Greptile is 5/5 with no unresolved comments.
The real browser regression also passed in the normal Chrome CI shard
(3.2 seconds).
- Code-owner approval remains required because the dedicated test
command changes `package.json`.
- Scanned the diff and PR text for credentials and private data before
pushing.

## Risks

The context cache applies only to development hot reload. It retains the
context object, not account state. Production continues to create an
ordinary React context. The diagnostic hook adds only bounded state and
type values, preserves boundary snapshots, and returns the original
event if enrichment fails. It does not suppress errors or restore raw
breadcrumbs. The global rejection diagnostics do not identify the
promise's originating operation by themselves.

## Model Used

OpenAI GPT-6 through Codex, with repository inspection, code editing,
and command execution. The runtime does not expose a more specific model
revision or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused suites and
Chromium regression; full local limitations documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 16:15:14 -07:00
0be2afcca6 feat(ui): improve task composer controls and pending input (#14322)
## Thinking Path

> - Paperclip lets operators assign tasks to AI agents and review their
work.
> - The task composer controls the next message and its assigned agent.
> - Operators needed a way to choose that agent's model and effort
without leaving the composer.
> - The old mode selector, upload button, and input cards made the
mobile composer crowded and hid normal messaging during a pending
decision.
> - Harnesses publish different model and effort capabilities, so the
picker must follow the selected agent.
> - This pull request adds one responsive composer flow, keeps pending
cards visible above it, and protects Codex ACP authentication in the
local test path.
> - Operators can choose run settings, send a message, and answer a
pending card as separate actions.

## Linked Issues or Issue Description

**Subsystem affected**

Task composer UI, issue thread interactions, Codex ACP credential
handling, and Storybook.

**Problem or motivation**

The composer did not expose model or effort for the selected agent.
Mobile actions wrapped poorly. Pending questions and confirmations
replaced the composer. A local Codex ACP test could also reuse host
authentication after the managed key was removed.

**Proposed solution**

Put assignee search, model search, exact model IDs, effort, and fast
mode in one picker. Use a mobile dialog. Replace the direct-upload plus
action and separate mode selector with an Add menu and removable Plan or
Ask chips. Place pending interaction cards above the usable composer.
Keep these cards pending after an ordinary message unless their creator
asks for comment superseding. Replace managed ACP auth files atomically
and isolate the test key from host credentials.

**Roadmap alignment**

ROADMAP.md does not list an overlapping composer milestone. This change
improves the existing task and review flows.

## What Changed

- Added the combined assignee, model, and effort picker to both task
composers. Search matches agent name, role, and harness. The server uses
a curated Codex list by default and honors instance-declared models.
Manual IDs remain available.
- Added an effort slider for known model capabilities, a conditional
Codex fast control, and reset. The picker opens in a modal on mobile.
- Added the Add menu for files, supported goals, Plan mode, and Ask
mode. Plan and Ask are exclusive removable chips. Keyboard mode cycling
remains available.
- Adjusted mobile spacing, avatars, wrapping, and Send placement.
Removed the composer divider.
- Moved pending question, confirmation, review, and related cards above
the composer. Ordinary comments now leave question and confirmation
cards pending by default. The onboarding prompt retains explicit comment
superseding.
- Updated the Storybook composer group with responsive states and the
production picker. Added UI, service, route, and browser regression
coverage.
- Isolated Codex ACP API-key authentication, skipped subscription auth
merge and shared-home copy-back for remote API-key runs, and replaced
the managed auth file atomically.

## Verification

- `pnpm -r typecheck` — passed on the final local head.
- `pnpm check:token-gates` — passed on the final local head.
- `pnpm exec vitest run server/src/__tests__/adapter-models.test.ts
ui/src/components/task-chat/ComposerRunSettingsPicker.test.tsx` — 31
tests passed, including role and harness search, declared Codex models,
and filtering general OpenAI models.
- `pnpm exec vitest run
server/src/__tests__/issue-thread-interactions-service.test.ts` — 74
tests passed.
- `pnpm exec vitest run
packages/adapters/codex-local/src/server/acp.test.ts` — 42 tests passed,
including remote API-key copy-back isolation.
- `pnpm test:run` — attempted locally; the embedded PostgreSQL test
database could not initialize on macOS. The isolated
`heartbeat-run-event-sequencing` suite reproduced that environment
failure. GitHub CI runs the full test matrix for this head.
- `pnpm build` — passed on the final head. `pnpm build-storybook` passed
after the last UI change; only server code, tests, and docs changed
afterward.
- Live local test drive — Codex ACP ran a task with a managed API key.
The test agent was restored to its default ACP configuration afterward.
- Review the interactive stories under the top-level Composer group with
`pnpm storybook`. Check a narrow desktop width and mobile Plan, Ask,
picker, and pending-question states.

## Risks

- A pending card stays open when an ordinary comment changes the
discussion. Its creator can set `supersedeOnUserComment: true` when a
new comment should replace it.
- Model and effort overrides persist on the task until reset or changed.
An unlisted manual model ID may fail when the provider runs it.
- Some harness catalogs do not report effort support. The picker hides
effort for those models.
- No database migration is required.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

OpenAI GPT-6 via Codex. This runtime does not expose the exact model ID
or context window to the task. The model used code execution and browser
tools.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: OpenAI Codex <codex@openai.com>
canary/v2026.928.0-canary.10
2026-09-28 23:02:39 +00:00
DottaandPaperclip 24c58e479a Improve task artifacts with rich cards and editable stories (#14469)
Render eight artifact card types from real task records and share them with editable Storybook stories. Preserve document review and media/file actions, and load bounded CSV previews on request.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
canary/v2026.928.0-canary.9
2026-09-28 15:38:06 -05:00
DottaandPaperclip 18e8c121d9 fix(runner): include Grok support in public installs with sandbox prerequisites (#14024)
## Thinking Path

> - Paperclip manages agents through a shared native runner.
> - Built-in harness support should ship with Paperclip's public
distribution.
> - Grok already speaks ACP; it does not require a new public bridge
package.
> - Sandbox provisioning owns the native executable and its pinned
version.
> - The runner must verify that prerequisite without downloading it
during npm installation.
> - This change separates built-in launcher identity from external
runtime identity.
> - Clean npm installation and live staging checks verify the
distribution boundary.

## Linked Issues or Issue Description

Refs #13882, #13973, #13977, #13979.

This follow-up now targets master after #13882 was squash-merged. It
replaces the private `@paperclipai/grok-acp` workspace package with
runner-owned assets. Current master is included so the branch also
contains the merged scheduler, complete-event capture, and durable
cleanup fixes.

## What Changed

- Ship Grok launcher and qualification metadata inside the runner's
compiled output and the public server's vendored runner tree.
- Remove the separate Grok npm package and all package-manager install
hooks for this runtime.
- Require the checksum-verified Grok Build 1.0.13 binary at
`/opt/paperclip/providers/grok/1.0.13/grok` in the selected execution
environment. Provision it explicitly in the Daytona image and CI setup.
- Keep native binaries outside the provider pack. Bind the built-in
launcher into the pack manifest.
- Preserve executable leases, descriptor-backed startup, credential
fences, permissions, and exact ACP model admission.
- Use `builtin:grok-acp` and `native:grok` as profile identities.
Historical package-profile sessions fail closed on resume rather than
being silently reinterpreted.
- Resolve built-in assets from the authenticated sidecar location,
including public server npm layouts. Keep the controller path out of
provider environments.
- Add clean npm tarball installation verification to the existing
trusted canary CI job and the admitted manual EC2 verification path. It
stages a unified release version and runs npm lifecycle scripts, then
verifies missing-prerequisite rejection and admission after separate
provisioning without credentials or inference.
- Include the controller-owned provider pack in stamped Cloud images.
Unstamped local images omit the pack and remain usable; remote ACPX
requires full source provenance.
- Correct CLI approval-page metadata for an already authenticated Cloud
board user; approval authorization remains unchanged.
- Honor explicit native-runner enablement in the Cloud agent picker and
direct setup page, keeping the flag disabled by default.
- Allow selecting the execution environment before connecting
credentials. Include Grok in the existing authenticated hello-probe
flow, targeting its pinned native prerequisite for runner setup.
- Recover an existing subscription sign-in conflict through an explicit
cancel-and-retry action, serialized after cancellation succeeds.
- Preserve the selected ACPX harness before normalizing config fields,
so new Grok agents use the Grok default model.
- Keep the credential-free Cloud provider pack root-owned and readable
after runtime UID remapping; verify manifest and referenced asset access
under an unrelated unprivileged UID during image builds.
- Archive prior failover backups alongside explicitly replaced harness
state, preserving evidence while preventing stale backups from blocking
a fresh replacement.
- Update Daytona image content inputs and contract tests for the
built-in assets and explicit provisioner.
- Document and regression-test the shared `approve-all` default for Grok
setup, saved configuration, and native execution. Explicitly saved
restrictions remain unchanged.

## Verification

Current merge-repair head `df09eb3e1a619430ad8419a0ee9aedd486689b05`
incorporates master `f1a394bd30cb56fb9e479f98b9f50176fe921858` after the
base PR was squash-merged. All 12 conflicts came from incoming files
identical to the tested pre-squash base. The final tree exactly matches
a three-way merge using that original base, preserving built-in Grok
distribution and removal of the obsolete private package. All 252
focused runner/UI tests, six npm-isolation tests, and token gates pass.
Fresh exact-head Greptile review is 5/5 with no outstanding findings;
security scans and EC2 native compilation pass. All current-head CI is
green: 56 successful checks/statuses and four intentional skips ([run
36468768035](https://github.com/paperclipai/paperclip/actions/runs/36468768035)).
The repository owner explicitly authorized bypassing code-owner approval
after all checks passed; no CI checks or repository protection settings
are bypassed or changed. The only remaining PR was removed from the
completed stack metadata to permit native auto-merge.

Earlier integration head `78cb306ecc41b5c96577c26c1d89153b0ef865a1`
includes master `3447609d2247e75e55d91493dda91a608364f672` (2026-09-28).
Two master advances during verification overlapped the eval catalog; the
final merge preserves Grok qualification, completion updates, and
bounded API-response reading in all 348 cells. All 77 focused
catalog/eval/workflow tests pass. Both native stack layers (#14397) are
mergeable, and both exact-head Greptile reviews are 5/5 with successful
security scans and no unresolved review threads. All current-head CI is
green: 56 successful checks/statuses and four intentional skips ([CI
attempts](https://github.com/paperclipai/paperclip/actions/runs/36447124691)).
The initial attempt lost two EC2 runners to shutdown signals and stalled
a third shard during dependency preparation; all three passed the
same-commit failed-job-only retry. Trunk code-owner requirements remain
enforced. The review summary’s non-blocking saved-asset offset
classification note concerns code already merged in #14301; those
runtime files are identical to master and outside this stack’s diff.
Historical live evidence below retains its original source revisions.
[Final public npm
verification](https://github.com/paperclipai/paperclip/actions/runs/36445542764)
passed on `76ea70cd4d13786a042af9df82f0fd7a8c85ae30`: 17 public
packages, an executed offline lifecycle sentinel, unchanged consumer
lock, built-in launcher, missing-prerequisite rejection, and verified
separately provisioned binary/command lease. Provisioning and cleanup
require no host privilege elevation; only the positive probe mounts the
temporary native binary read-only. The verifier is unchanged by the
final master merge. All six isolation tests and an offline npm smoke
test pass. The prior head had 56 green CI checks and a 5/5 review after
two unchanged tests timed out and passed a failed-job-only retry ([CI
attempts](https://github.com/paperclipai/paperclip/actions/runs/36444597313)).
All 56 recovery-display/lineage tests pass; re-review cleared the
already-covered missed-retry concern. Earlier EC2 failures remain
retained: [npm lockfile
rejection](https://github.com/paperclipai/paperclip/actions/runs/36436311203),
[missing compiler in the slim
image](https://github.com/paperclipai/paperclip/actions/runs/36440210984),
and the aggregate 15-minute test timeouts in those broad runs. Both
broad attempts passed typecheck, token gates, Product E2E type/unit
checks and build. The focused EC2 lane preserves the existing
trusted-actor and immutable-source gates.


Earlier documentation/test checkpoint
`ff244c4fd78a7ede5a3e00efe09f475f133ef33e` leaves runtime behavior
unchanged. 154 focused tests pass across configuration building, native
provider resolution, permission policy, credentials, UI configuration,
and new-agent setup (including both Grok auth modes); token gates pass.
All fresh CI is green for this head: 56 successful checks/statuses and
two intentional skips ([run
36367065119](https://github.com/paperclipai/paperclip/actions/runs/36367065119)).
Greptile is 5/5 with no new findings. Grok already inherits the shared
`approve-all` default, so unattended setup requires no manual permission
change.

Runtime head `bb5a9307991f1ac567b781970ef11b39d518e19b` fixes a final
staging continuation failure before provider startup: explicit
replacement archived the old harness but left its failover backups
active, which caused `runner_harness_state_mismatch`. The regression
fails before the fix and passes after it; all eight adjacent
recovery-safety cases also pass. Old backups remain inspectable inside
the continuity archive. All fresh CI is green at this head ([run
36360839248](https://github.com/paperclipai/paperclip/actions/runs/36360839248)),
with a 5/5 review. One unrelated Cursor test timed out in the initial
server shard; the same-commit failed-job rerun passed, and both attempts
are retained. Staging deployment is confirmed healthy on this revision.
The controller image is
`ghcr.io/paperclipai/paperclip@sha256:6ad91c487910ccd2596ff7aed0a3a3ea5233d12b51b83cd6e1402237749b9673`.
The final browser-created staging task passed on this exact revision
with API authentication: context read → structured human question →
controller restart → answer submission → same native provider session
resumed → document saved → task Done. The two turns took approximately
119s and 77s. The actual write receipt was applied, and the saved
document has exactly one revision containing the selected answer and
requested marker. Usage and cost were not reported. [Controller image
build](https://github.com/paperclipai/paperclip/actions/runs/36360889243).

- Previous integration head `a44f7dbb6b6f77cd9ed893756ca453307f281e5f`:
all CI green (53 successful checks/statuses, two intentional skips),
including repository typecheck/build/tests, native Runner tests, browser
shards, and canary installation checks. [CI run
36358672529](https://github.com/paperclipai/paperclip/actions/runs/36358672529).
Greptile is 5/5 with no unresolved findings.
- Focused checks cover Grok credentials, executable admission, launcher
assets, provider-pack paths/permissions, workflow contracts, setup
defaults, CLI authorization, and subscription conflict recovery. All 39
protocol definitions validate. Final integration checks pass 124
catalog/evidence/cache tests and nine project-form tests; token gates
pass. Some local dependency checks could not load the stale installed
dependency tree; the corresponding fresh EC2 checks pass.
- Clean public npm installation passed on EC2 at
`8b172ebcf8e02e30662d830c00f3961e3bd459ec` ([run
36164964900](https://github.com/paperclipai/paperclip/actions/runs/36164964900)):
17 unified-version packages, lifecycle scripts enabled, built-in
launcher present, no separate Grok package or npm-downloaded binary,
missing prerequisite rejected, separately provisioned native executable
and command lease verified. No credentials or inference were used.
Subsequent changes preserve this npm asset layout.
- The immutable Daytona prerequisite image is
`ghcr.io/paperclipai/paperclip-daytona-runner@sha256:98957d5be0ac774d086b6402b5849e8e6356fec70fb8c09fca6eb4ed6de918e0`,
built from `5a2db471f3ddabe77f9f80e76ed27f996cb97fba`. The previous
Cloud controller image was
`ghcr.io/paperclipai/paperclip@sha256:fd914e1ab1e45f741e8e078ff452d16f082d7ac05f9b4b3506d3a3c64150d204`,
built from `a44f7dbb6b6f77cd9ed893756ca453307f281e5f`; it is superseded
by the latest image above. Its EC2 build verified provider-pack access
under an unrelated unprivileged UID.
- Browser staging at `40f898bc4cba73c1dff4e6344a3983ba0fb247ef` passed
full Grok onboarding with the correct `grok-4.7` model, saved credential
delivery, and pinned Daytona execution. A browser-created task read
context and asked the structured human question. After a controller
restart, answering the persisted question resumed the same native
provider session, saved the requested document, and completed the task.
Actual tool outcomes and durable state agree: one question and one
document revision. The two successful turns took 42.7s and 63.1s; usage
and cost were not reported.
- Restricted policy returned the expected `approval_required` outcome.
Functional staging tests explicitly selected `approve-all`; controller
authorization and governed approvals remain enforced. Temporary board
CLI access was revoked and verified rejected (HTTP 401), and the
disposable onboarding agent was paused.

Failures remain retained: the pre-fix continuation failure (its task
remains blocked; the passing final task is fresh), the original Cloud
provider-pack permission failure, the expected restricted-policy denial,
the superseded npm staging failure, and an earlier monolithic CI
infrastructure timeout. Browser CI exposed a project alias/form race;
the final stack uses master's stronger draft-preservation fix and all
browser shards pass. Historical full subscription/API protocol and
Product rosters retain their original source revisions and do not
qualify this packaging revision. No local Docker or Rust build was used.

## Risks

The branch includes master’s draft-preservation fix for project URL
aliases. It keeps the same project’s edit form mounted and clears prior
data when the project or company changes.

Custom sandboxes and local execution hosts must provision the pinned
binary before Grok starts. Missing, changed, unsupported-platform, and
symlinked executables fail admission. The new builtin profile cannot
resume sessions created with the former private-package profile.
Existing Claude/Codex npm bridge profiles retain their package pins.
Grok restricted modes preserve the selected policy but cannot
automatically admit Paperclip calls: ACP permission metadata does not
independently bind tool authority, so those calls stop with
`approval_required`. New Grok configurations default to `approve-all`,
including API configurations that omit the mode. Existing explicitly
restricted configurations remain restricted; controller authorization
and governed approvals remain enforced.

## Model Used

OpenAI GPT-6 through Codex, with tool use and code execution. The exact
serving model identifier and context-window size are not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.928.0-canary.8
2026-09-28 14:54:44 -05:00
DottaandPaperclip 992f720262 fix: make runner task context ownership explicit (#13753)
<!-- Write all pull request text in Simplified Technical English
(ASD-STE100). -->

## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Task descriptions, comments, continuation data, skills, and
execution rules enter several agent adapters.
> - The same source can be rendered by more than one automatic input
carrier.
> - Failed resumes can also rebuild input from stale or compact context.
> - This pull request gives each Paperclip-owned source one delivery
owner and preserves the required transport boundaries.
> - It adds deterministic adapter, interaction, runner, and browser
tests for these boundaries.
> - The benefit is more predictable context delivery with explicit
evidence for later live qualification.

## Linked Issues or Issue Description

Related: #13144 removes a duplicate environment payload and bounds wake
lists. Related: #11360 addresses Hermes resume behavior. This pull
request preserves compatible active-session formats while repairing
context ownership and stale question creation.

**What happened?**

Task descriptions and comments could enter more than one automatic
context block. Native transports could wrap a complete model input in a
second task envelope. Some legacy and gateway adapters could omit the
owned assignment on ordinary tasks or rebuild a failed resume with stale
compact context. A continuation could also request a question after
newer human comments had arrived.

**Expected behavior**

Each task or comment source has one automatic model-facing owner.
Distinct comment IDs and repeated wording remain distinct. Fresh
fallback attempts rebuild the required full context. A question request
is rejected when newer queued human direction makes it stale. Harness
access policy remains owned by execution configuration.

**Steps to reproduce**

1. Build a task with a description and current comments.
2. Capture the actual adapter or runner input.
3. Compare source ownership and task-envelope nesting.
4. Queue a human comment before a continuation requests a question.
5. Trigger a failed resume and inspect the fresh retry input.
6. Run the focused adapter, interaction, runner, and browser checks.

## What Changed

- Add shared prompt-section selection at the provider-attempt boundary.
- Deliver owned assignment context through native, legacy CLI, ACP,
gateway, cloud, Pi, Kimi, Grok, Gemini, OpenCode, Cursor, OpenClaw, and
Hermes paths.
- Rebuild full or compact context after resume recovery changes the
attempt. Add native and Claude ACP tests of actual recovery requests.
- Preserve custom templates, loaded instruction files, execution
policies, and older active-session formats.
- Record continuation source metadata and reject stale question creation
under the issue-row lock.
- Add explicit Product E2E context-integrity profiles, prerequisite
gates, credential-isolation checks, and report fixtures.
- Bypass service-worker forwarding for same-origin Vite development
modules. A real Chromium test fails with resource exhaustion before the
repair and passes after it. Production asset caching keeps its existing
policy.
- Add browser diagnostics and service-worker module-loading regressions.
- Add an explicit zero-retry eval option. The default retry behavior
remains unchanged. Each campaign records its effective policy.
- Remove the model-facing working-directory sentence from four prompt
builders. Existing workspace, sandbox, permission, and custom-template
configuration remains unchanged.
- Align the everyday workflow assertion with the current 47-entry
catalog.

Compared with current upstream master, the branch carries the
context-ownership implementation and its tests, the explicit
context-integrity catalog and evidence harness, and the focused browser
regression checks.

## Verification

**Merge assessment:** focused regression evidence supports merge. This
is not full completion of the original broad qualification matrix. The
maintainer has authorized merge after fresh verification of the master
integration.

- Current head: `bbd52f82114eabf09bc7b1a7e97d54a5b43bbc00`. This
integrates current master `2f585ef26a1814fa209715242d1ca791b63e4c4e`.
All 14 conflicts are resolved. Cancellation checks, workspace
finalization, native Grok support, and both sets of tests are retained.
- Current-head Greptile: **5/5**, with no blocking findings. The review
names this exact commit. All **59 reported checks are terminal: 55
successful, 4 skipped, zero pending or failing**. This includes the full
root general and serialized suites, separate runner checks, typecheck,
build, canary, browser E2E, Docker, and security checks. The successful
legacy security status is included in that total.
- After integration: workspace typecheck and full build passed. Separate
runner checks passed: **2,160 TypeScript tests (10 skipped), 582 Rust
tests, and 39 preparation checks**. Other passing checks include 621
Product E2E harness units, 376 focused shared/adapter tests, 160
real-database/API tests, 86 Hermes tests, 18 browser-support checks, and
Product E2E typechecking. The complete root suite passed in CI. The
duplicate local monolithic root run was stopped after that CI result; it
is not counted as a completed local pass.
- New native recovery coverage retains full assignment, completion
contract, and explicit skill selection after safe replacement, for old
and prepared input formats. Full native session test file: **136/136
passed**.
- New Claude ACP coverage captures actual fresh, resumed, and
missing-session fallback requests. It verifies one assignment copy,
comment order, identical text under distinct comment IDs, and full
fallback context. Full file: **33/33 passed**. Both affected TypeScript
checks passed.
- Existing deterministic tests cover source revisions, approval and
trust boundaries, completion validation, custom templates, compatible
sessions, standalone driver wrapping, and maintained adapter transport
requests.
- Provider-free browser support: **17/17 passed** after the master
merge. Service-worker unit tests: **33/33 passed**. The module-overload
regression failed before the repair and passed after it in real
Chromium.

### Fresh live comparisons

The new batch ran exactly four Product E2E attempts. **All four passed
on the first attempt; no retries.** Each has six terminal matchers plus
the existing browser lifecycle and invariant checks.

| Exact case ID | Control | Candidate |
|---|---|---|
| `core-compatibility.runner-codex.local.plan-revise-accept` | Passed |
Passed |
|
`local-session-integrity.runner-acpx-claude.local.structured-question-restart-resume`
| Passed | Passed |

The plan case checks a revised canonical plan and revision-bound
approval before completion. The question case restarts the server before
submitting the answer, then verifies the continuation completes.

Control source is `dfa4e1bda8d50a1a01746603251a9128dbe9d0d6`. Candidate
source is `79fcdb5dece501d28064ea9da306603881b46f0c`. They use identical
frozen definitions and provider versions: Codex `0.156.0` with
`gpt-5.6-sol`; ACPX `0.13.1` / Claude ACP `0.73.0` with
`claude-sonnet-5`. The September 24 head added master browser recovery
and test-only changes. The September 28 head also integrates newer
master changes, including cancellation, workspace finalization, and
native Grok. These are frozen-source live results, not exact-head live
runs.

The candidate received one description copy where the control initially
received three. The submitted initial plan envelopes were 7,969 versus
19,097 characters. Question envelopes were 7,592 versus 18,919. These
are structural measurements, not whole-provider token or dollar savings.

### Earlier evidence and failed attempts

- The preceding fresh batch has four effective passing pairs: OpenCode
comment continuation and assigned skill, native Codex comment
continuation, and native Claude comment continuation. It retains **11
attempts: eight passed and three failed**.
- Original failures remain recorded: missing local PostgreSQL library
links before task creation; host-sleep cleanup after task/page checks
passed; and a Claude **control** session-open rejection before a model
turn. Setup was repaired identically on both worktrees. The permitted
unchanged infrastructure retries passed. The underlying Claude provider
startup error was not retained and remains unknown.
- Older R2 retains **17 passes and one failure** across 18 attempts,
including eight both-pass native/legacy Codex/Claude pairs. Its OpenCode
blank-page failure led to the service-worker repair. R2 is historical
evidence: master changed the native fixed prompt and removed duplicate
wake environment data afterward.
- The September 24 CI run initially failed one unrelated preview
readiness test (`ECONNREFUSED` on its local fixture). Its test and
production code match master. Isolated local verification passed **28
tests, 3 skipped**. One unchanged CI retry passed the full shard: **831
passed, 1 skipped**, including all **31 preview-exposure tests**. The
aggregate CI gate passed afterward. The precise startup cause remains
unknown; a port race is a hypothesis, not a proved cause.

### Limits

The original wider profile/workflow matrix, repeated trials, and remote
Daytona qualification are incomplete. These results support a focused
merge recommendation, not statistical equivalence or universal harness
qualification. Some usage receipts are missing in both variants, so no
token or dollar savings are claimed. The $500 ceiling was preserved
using conservative allowances; failed attempts and unknown charges
remain in the ledger.

Reproduce the focused additions with `pnpm exec vitest run
packages/adapters/claude-local/src/server/acp.test.ts` and `pnpm
--filter @paperclipai/paperclip-runner exec vitest run
src/native-session-runtime.test.ts`. Full checks use `pnpm -r
typecheck`, `pnpm test:run`, `pnpm build`, and the separate runner
checks. Paid evals require the frozen definitions, profiles, and
credentials; do not use `--all` as a substitute for the selected cases.

## Risks

- Context placement changes can affect model behavior. Deterministic
checks cover the selected paths, but live qualification remains
incomplete.
- The stale-question guard can reject a request when queued human
comments arrived during the run. This is intended.
- New stored inputs and model envelopes retain compatibility readers for
older active sessions.
- Custom templates may intentionally repeat content.
- Removing a model-facing working-directory sentence does not change
filesystem, command, sandbox, or permission configuration.
- The worker bypass applies only to same-origin development module
paths. Cache-policy tests preserve private-response handling and
production asset caching. Mounted HTTP fixture changes remain test-only.
- This PR does not claim measured token savings or statistical
equivalence across every harness.

## Model Used

OpenAI Codex, exact model gpt-6-astra, with repository tools and code
execution. Bounded supporting work used gpt-5.6-luna and gpt-6-luna. The
serving context-window size is not exposed in this task.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have described the issue in-PR using the required issue fields
- [x] I have not referenced internal/instance-local Paperclip issues or
links
- [x] My branch name describes the change and contains no internal
ticket id
- [x] I have run the focused local checks and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect these changes
- [x] I have considered and documented risks above
- [x] All current-head Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
for the current head
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.928.0-canary.7
2026-09-28 14:49:14 -05:00
DottaandPaperclip 2f585ef26a fix(ui): preserve newer drafts after repeated receipt cleanup (#14332)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The task composer keeps unsent text across page reloads.
> - A server receipt confirms a submitted message after a reload.
> - Replayed cleanup can apply the old text offset to a newer draft
twice.
> - This pull request reconciles each confirmed attempt once and
preserves both tabs’ unsent intent.
> - The full newer draft stays available for the next send.

## Linked Issues or Issue Description

**What happened?**

Reloading the classic composer while a save was pending could truncate a
newer draft. Replayed receipt cleanup changed `A newer draft written
while delivery was pending.` into `nding.`. The storage helper rejected
the duplicate settlement, but the effect still changed the editor and
its body reference.

**Expected behavior**

A receipt settles its retained submission once. Later cleanup must
preserve the newer draft in both the editor and browser storage.

**Steps to reproduce**

1. Send a comment and hold its HTTP response after the server accepts
it.
2. Type a newer draft, then reload the page.
3. Restore the matching receipt under React StrictMode.
4. Inspect the newer draft after effect replay and unmount.

The deterministic regression reproduces this on current master. The
existing browser test exposed the issue during #14329 verification.
Related draft persistence code came from #13338. A search found no
separate fix for this duplicate settlement.

## What Changed

- Reconcile each confirmed attempt once in memory so duplicate effects
cannot trim its newer draft again.
- Unlock confirmed drafts when storage writes fail. Never acquire
another tab’s pending receipt or assume its text is newer. Keep a
conflicting local draft and its attachments in tab-scoped session
storage while preserving the shared draft unchanged.
- Add component regressions for StrictMode replay, exact editor and
stored bytes, foreign receipts, attachment-only differences, reload
recovery, unavailable storage, and task navigation before autosave. A
brief notice explains when this tab has a separate draft.

## Verification

- RED: the new test received `nding.` instead of the full newer draft
before the fix.
- Review RED: three cases reproduced a locked composer after failed
storage writes or another tab's settlement. A further negative case
preserves a newer retained attempt and its attachments.
- Cross-tab review RED: four deterministic cases reproduced lost local
text, foreign receipt takeover, attachment loss when text matched, and
overwriting newer stored text.
- GREEN: 118 tests across `IssueChatThread`, `composer-draft`, and
`comment-submit-draft`, including same-mounted A → B → A navigation and
a full recovery-storage failure. Two further lifecycle RED tests verify
finishing a recovery returns to the shared draft on a later visit while
continued typing and attachments retain recovery.
- Both existing browser reload cases passed against a fresh server and
database on exact head `ceb80aca77fc8cc0f813c328ba87025b3e1a2222` (42.5
seconds), with classic mode enabled and disabled. The browser flow
verifies one original comment, a preserved draft, and a successful
second send.
- UI typecheck, token gates, and the shipped static UI build passed. The
initial typecheck required the fresh worktree's plugin SDK build; the
retry passed after that dependency built.
- The session-only failure mock also preserves localStorage on platforms
where both share the Storage prototype; this fixes the Linux workspace
test failure.
- All 56 checks passed on `4c30dcccc4fe5b90dcd08dd1d90475ea89e4a376`;
Greptile is 5/5 with zero open review threads. An unrelated runner
baseline scan hit its existing 100 ms deadline once; its isolated test
and the single failed CI job passed on retry without source changes.
This four-file UI fix does not change server or database code.


- Final alternate-staging acceptance passed on deployed source
`d884e1ab046cc76004e35e6091e9e6e2c918c9eb`, with exact health checked
before and after. Three real browser cases used shipped static UI and
real API saves: reload during an accepted-but-unacknowledged save in
both composer variants, plus two classic tabs with different drafts. The
test deliberately held only its own accepted POST acknowledgement and
temporarily withheld its own GET receipt from one tab to reproduce
settlement ordering. Both drafts survived independent reloads, the
shared draft remained intact after the other tab sent, and finishing
recovery returned to the shared draft on a later visit. Exact request
IDs/counts and server comment bytes passed; no provider was invoked.
Screenshots preserve the live drafts before fixture cleanup.
- An ordinary authenticated Chrome/CUA visit independently showed the
exact original and newer comments; an additional unsent draft survived
reload and was then cleared. All three isolated fixtures are complete,
browser contexts are closed, and original user tasks were untouched.

## Risks

- Reconciliation uses the exact request ID and draft key. Conflicting
drafts stay separate; the tab recovery survives same-tab reload through
session storage and does not outlive the browser tab. If recovery
storage is unavailable, the editor keeps its in-memory text and
explicitly warns the user to copy it before leaving.
- Attachment selections remain bounded to 20 per draft; if combining
equal-text snapshots would exceed that limit, each original selection
stays in its respective draft.
- No API, schema, or migration changes. The status notice uses existing
design tokens.

## Model Used

OpenAI Codex, GPT-6, with tool use and code execution. The runtime does
not expose the exact deployment variant or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 14:03:36 -05:00
DottaandPaperclip f1a394bd30 feat(runner): add Grok Build through native ACP (#13882)
## Thinking Path

> - Paperclip manages AI agents and governs their work.
> - Its native runner uses structured provider protocols for sessions
and tools.
> - Grok Build supports ACP over stdio, but the runner did not expose
it.
> - Native execution requires company-scoped credentials, verified
identities, and permission gates.
> - This change adds Grok through ACPX for local and Daytona execution.
> - Subscription login and explicit API-key execution have separate
credential paths.
> - Qualification grades real tool outcomes, durable state, and browser
workflows.

## Linked Issues or Issue Description

Refs #13845, #13847, #13850, #13878, #13901, #13973, #13977, #13979.

Add **Grok Build** to `paperclip_runner` with `provider: "acpx"`,
`acpxAgent: "grok"`, and model `grok-4.7`. Existing legacy Grok agents
keep their adapter. Merge the three companion fixes (#13973, #13977,
#13979) before treating the integrated Product qualification as deployed
behavior.

## What Changed

- Synchronize shared, TypeScript, Rust, server, validation, and UI
provider contracts.
- Run Grok native ACP stdio through ACPX and the authenticated Paperclip
MCP bridge. Verify the pinned executable and exact ACP model identity.
- Prefer company subscription login. Support an explicit company-secret
API key without automatic paid fallback. Fence refresh and copyback to
the same account and remove private runtime credentials after
containment.
- Preserve selected permissions, cancellation, durable session identity,
resume, and restart recovery. Keep unsupported steering and goals
unavailable. Preserve missing usage and cost as unknown.
- Package checksum-verified Grok Build 1.0.13 for Daytona with an
immutable, signed image built on EC2.
- Add deterministic admission, protocol, permissions, identity,
credential, failure, and cleanup checks. Add the maintained 39-case
protocol roster and separate subscription/API Product profiles.
- Fix live-test findings in reasoning events, reloads, idle-owner
retirement, credential-home cleanup, expired-login model discovery,
launcher pinning, and rerun evidence selection.
- Align control-plane state readers with the transport's 64 MiB bound
while retaining identity, ownership, lifecycle, and size rejection
checks.
- Stabilize two asynchronous CI assertions while retaining actual
outcome and filesystem-evidence checks.

## Verification

Current integration head `f114948376056fe0b6b34c1496ae8667b59daa63`
includes master `3447609d2247e75e55d91493dda91a608364f672` (2026-09-28).
Two master advances during verification overlapped the eval catalog; the
final merge preserves Grok qualification, completion updates, and
bounded API-response reading in all 348 cells. All 77 focused
catalog/eval/workflow tests pass. Both native stack layers (#14397) are
mergeable, and both exact-head Greptile reviews are 5/5 with successful
security scans and no unresolved review threads. All current-head CI is
green: 56 successful checks/statuses and four intentional skips ([CI
run](https://github.com/paperclipai/paperclip/actions/runs/36447097232)).
Trunk code-owner requirements remain enforced. The review summary’s
non-blocking saved-asset offset classification note concerns code
already merged in #14301; those runtime files are identical to master
and outside this stack’s diff. Historical live evidence below retains
its original source revisions.


Earlier integration checkpoint:
`24fc9b94ca0afb21ccdc8d26dbb2e4b258ad72cb`. Refreshed against master
`0f14d2612`, preserving Grok qualification alongside the new accounting
and lifecycle suites. All 124 focused catalog, evidence, and
service-worker checks pass. The current base workflow includes the
explicitly selected public-install verification lane; follow-up #14024
supplies its verifier script. CI at that earlier checkpoint was green
(56 successful checks/statuses, four intentional skips), and the review
is 5/5 with no unresolved findings. Prior feature CI at
`fd73f0a9b1ecdf4094685054028df71739ddc3e1` passed ([run
36148259902](https://github.com/paperclipai/paperclip/actions/runs/36148259902));
that is historical evidence, not a current-head result.

Paid Product measurements use frozen integrated source
`2d939a92b21dcaf5c77c88b54d96784d2ddd0699`, which combines the feature
with #13973, #13977, and #13979. That source passed all 52 CI checks and
clean 5/5 review. Later master syncs incorporate upstream changes. Their
checks remain separate from these pinned live measurements.

| Check | Result and source-pinned report |
| --- | --- |
| Subscription protocol roster | [39/39 first attempts; 206
assertions](https://d1p6rlowie26tp.cloudfront.net/runner-protocol-evals/campaigns/gha-36046839612-1/index.html),
runtime `bc6833f7`, evals `92bb4b8c` |
| API protocol roster | [39/39 first attempts; 206
assertions](https://d1p6rlowie26tp.cloudfront.net/runner-protocol-evals/campaigns/gha-35926577007-1/index.html),
runtime `4a1061c8`, evals `3213dbec` |
| Subscription full Product matrix | [16/16 first attempts; 144
assertions; cleanup
passed](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36096908572-1/index.html),
source `2d939a92` |
| Subscription core repetitions | 18/18: tool use, planning approval,
and Stop/resume each passed three times in local and Daytona profiles.
The full matrix contains repetition one; [repeat
two](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36104551060-1/index.html)
and [repeat
three](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36108100404-1/index.html)
each passed 6/6. Total: 28 unique subscription attempts at `2d939a92`. |
| API smoke and question continuation | [4/4 first attempts; cleanup
passed](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36147315401-1/index.html),
both environments at `2d939a92` |
| Historical API Product coverage | [16/16 full
matrix](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35875144860-1/index.html)
and 18/18 core repetitions at `4a1061c8`; retained as measurements of
that revision |
| Native Daytona proof | Three subscription and three API
MCP/permissions/resume runs passed at `bc6833f7`. Three expired-login
admission and fenced refresh checks passed without inference. All test
sandboxes were removed. |
| Inspectable artifacts and UI | Current-source screenshots verify
planning approval, direct Ask completion, question continuation after
controller restart, and two downloadable project revisions. The project
downloads pass 12 and 18 tests; all 40 independent artifact oracle
checks pass. |
| Provider-free checks | 116 eval-validator tests, 39 Grok definitions,
and 359 enabled/external campaign cells pass. Continuation regressions
above 2 MiB and 16 MiB failed before their fixes; 32 focused
recovery/ownership/size checks pass. |

The 32 unique current-source Product attempts have no failures, retries,
or skipped cells, and all cleanup checks pass. Whole-workflow timing,
model identity, image and provider-pack provenance, attempts, and
accounting coverage are retained in the canonical reports. The report
publisher's conservative `complete=false` flag is preserved; independent
audits verify the exact selected source catalog and immutable result
rows.

Pins: Grok Build `1.0.13 (5e9a58528b76)`, ACPX `0.13.1`, ACP model
`grok-4.7`. Linux binary SHA-256:
`edf79521581bb5e6b95abef848491a6a742e860da3e237ebe86a280d30dce4c1`.
Launcher SHA-256:
`f0b698395a3704ed2ffaf84ea19bdb20c36c8a0a70b7c629c7b6ffe144e59e55`.
Image:
`ghcr.io/paperclipai/paperclip-daytona-runner@sha256:76b24edfd850219e949418b19e4ceba690e84d51d199ade426e484953329b5e9`.
Image build source is `4196a4cd`, recorded separately from application
source `2d939a92`; each campaign verifies the image signature and
provider pack.

Original failed campaigns remain available: [continuation
bound](https://github.com/paperclipai/paperclip/actions/runs/36057718059),
[scheduler/event
capture](https://github.com/paperclipai/paperclip/actions/runs/36071063537),
and [startup cleanup plus EC2
interruption](https://github.com/paperclipai/paperclip/actions/runs/36080870743).
They retain their original grades. No Docker or Rust builds ran on the
developer laptop for these follow-ups.

## Risks

Merge packaging follow-up #14024 with this base before public release.
The follow-up replaces the private Grok bridge package with a built-in
launcher and makes the native binary an explicit sandbox prerequisite.

Three separate, reviewed fixes are part of the tested integrated
behavior: #13973 serializes task-run admission; #13977 captures complete
event evidence; #13979 durably reconciles failed Daytona creation. Each
has green CI and clean 5/5 review. Failed-create recovery has 277 plugin
tests, 92 SDK tests, host-runtime recovery tests, and a real Daytona
lost-deletion-receipt proof. The live proof uses a private file for
journal persistence; database durability is covered by host tests.
Worker death before delivery of a failure envelope remains outside that
recovery mechanism.

Subscription fixtures stage an authorized company login; interactive
browser sign-in is not qualified. Local Product profiles ran on EC2
Linux. The temporary subscription credential was removed from the
protected GitHub environment after all subscription audits, with absence
verified. Runtime homes and refresh copyback remain ownership-fenced.

Protocol results remain pinned to their original revisions; they are not
relabeled as tests of the latest feature commit. New binary/model
versions require qualification. Missing token usage and model cost
remain unknown; runtime estimates do not establish a full bill.
Automatic paid Grok scheduling remains disabled pending separate
reviewed enablement. The 64 MiB bound can increase memory use for
verbose sessions, and larger files still fail closed. No automatic
legacy-agent migration occurs.

## Model Used

OpenAI GPT-6 through Codex, with tool use and code execution. The exact
serving model identifier and context-window size are not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 13:54:36 -05:00
DottaandPaperclip d9d2147171 fix(auth): keep Cloud tenants on the Cloud sign-in flow (#14407)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Cloud owns human identity and passes a verified identity to each
tenant.
> - The tenant can report no session while the Cloud session is still
valid.
> - The access gate and direct `/auth` route then show the instance
password form.
> - This pull request sends those users through the configured Cloud
entry endpoint.
> - Cloud can renew the tenant session or show its login page, then
return to the original task.

## Linked Issues or Issue Description

**What happened?**

A Cloud tenant can display the self-hosted email/password form after an
instance session check returns no session. This gives Cloud users the
wrong login method.

**Expected behavior**

An active Cloud session renews tenant access automatically. A signed-out
user signs in through Cloud. Staging and production use their own
configured Cloud origins. Self-hosted instances keep their instance
login form.

**Steps to reproduce**

1. Open a Cloud tenant task or an `/auth?next=...` link.
2. Keep the Cloud session active but make the instance session check
return 401.
3. Observe the instance password form instead of Cloud session recovery.

**Deployment mode**

Cloud-managed authenticated instances. No database or server API
changes.

Searched related authentication PRs. Native self-hosted OIDC support in
#10411 is a separate feature; this change uses the existing Cloud entry
contract.

## What Changed

- Wait for deployment metadata before showing an instance login form.
- Use the health response's Cloud origin and stack slug for session
recovery.
- Preserve the tenant path, query, and fragment. Reject external and
recursive login return targets.
- Limit automatic recovery per tab. Show a manual Cloud retry after
failed recovery. Show service failures as errors.
- Keep self-hosted login and local trusted access. Add focused tests,
browser regressions, deployment documentation, and an unavailable-state
design example.

## Verification

- `pnpm -r typecheck` and `pnpm build` passed.
- UI typecheck and `pnpm check:token-gates` passed after the final UI
edits.
- All UI tests passed: 639 files, 6,777 tests.
- 63 focused Vitest tests passed across Auth, CloudAccessGate, Cloud
links, and recovery coordination.
- `pnpm exec playwright test --config tests/e2e/playwright.config.ts
tests/e2e/cloud-auth.spec.ts`: 5 passed. The tests use the real tenant
UI and database with a simulated Cloud HTTP endpoint. They cover both
Cloud origins, direct auth/task links, no password-form flash, preserved
URLs, reload, and self-hosted login.
- Hands-on browser test used the real Cloud gateway and a fresh tenant
build with disposable local data. Active Cloud session plus a forced
missing instance session returned to the task. An expired tenant cookie
also renewed automatically and returned to the task. Removing both
sessions reached the real Cloud email/social login UI. Persistent
failure stopped at the retry screen; retry succeeded after removing the
injected fault. The fixture used a loopback transport adapter and a
simulated signed-out OIDC issuer. No production session or deployment
was changed.
- The default local browser startup hit the host's embedded PostgreSQL
resource limit. The passing run used a separate disposable database on
the test PostgreSQL process.
- The full local `pnpm test:run` sweep was stopped after about 31
minutes once CI completed the full suite. It had reported 33 failures in
the unchanged runner API unit/integration files; both files pass in
isolation (1,749 + 28 tests). The local sweep did not reach the later
workspace/serialized groups. CI completed all of those groups
successfully.
- Greptile reviewed commit `d603fd4e39455de44da9dae81b72197096c0e1e8` at
5/5 with its only thread resolved. All CI gates are green on this
commit, including all general/serialized server groups, workspace tests,
Runner checks, typecheck, build, and all eight browser shards
([run](https://github.com/paperclipai/paperclip/actions/runs/36446230697)).

## Risks

- Recovery depends on valid Cloud origin and stack metadata. Incomplete
metadata shows an unavailable message instead of a password form.
- Browsers with session storage disabled use the manual Cloud link,
since automatic retries cannot be bounded across documents.
- The external identity provider's email/social login was not completed
in this local test. Existing Cloud authentication owns that flow.
- No migration, credential format, membership rule, or production
deployment changes.

## Model Used

OpenAI GPT-6 through Codex. The exact served variant and context-window
limit are not exposed in this session. Used reasoning, repository tools,
code execution, and browser testing.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.928.0-canary.6
2026-09-28 13:45:27 -05:00
DottaandPaperclip 270afd2fb8 feat(ui): show running commit in staging account menu (#14410)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The account menu shows the signed-in user's identity.
> - Staging users need to know which server commit is running after a
deploy.
> - The health endpoint already returns that commit, but the menu does
not show it.
> - This pull request adds the short commit below the email on staging
hosts.
> - Users can open the menu to check a deploy without opening deployment
tools.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

The account menu on staging instances.

**Current behavior**

The menu shows the user's name and email. It does not show the running
server commit.

**Proposed behavior**

On `*.staging.paperclip.app`, show `SHA 8751e2d` below the email. Use
the current `/api/health` commit. Show the full SHA on hover and link to
the commit on GitHub. Refresh the health query when the menu opens. Hide
the label on other hosts and when commit metadata is unavailable.

**Reason and benefit**

A user can confirm which commit a staging instance runs after an
automatic deploy.

**Breaking changes**

None. The server already returns the commit field.

Refs #14060 for related account-menu work. This change adds deployment
information only.

## What Changed

- Add the existing health response commit field to the UI type.
- Share the staging host check and a separate health query across both
account-menu variants. A menu refresh failure leaves the access gate
health state unchanged.
- Show a short SHA below the email. Link to the full commit on GitHub
and include the full SHA in its accessible name and hover title.
- Document the staging label and cover staging hosts, other hosts,
missing metadata, refresh on reopen, and request failure isolation.

## Verification

- `pnpm exec vitest run ui/src/components/SidebarAccountMenu.test.tsx` —
24 tests passed.
- `pnpm check:token-gates` — passed.
- `pnpm -r typecheck` — passed.
- `pnpm build` — passed. The UI build and typecheck also passed again
after review fixes.
- Full test matrix — passed on the latest commit in
[CI](https://github.com/paperclipai/paperclip/actions/runs/36447810742).
The local `pnpm test:run` was stopped before completion after CI
finished the same suites. The focused local tests, full local typecheck,
and full local build passed.
- Browser shard 6 passed on one rerun. Its first attempt lost part of
the draft text in the existing attachment-receipt reload test. No code
changed for the rerun.
- Rendered the real account menu in a local browser fixture with a
staging hostname condition and mocked health response. Confirmed the SHA
fits below the email in the dark menu.
- Greptile review — 5/5, all review threads resolved.
- Manual check after deployment: open the menu on a staging host and
compare the SHA with `/api/health`. Open the menu again after a deploy
to refresh it. Confirm the label is absent on production and localhost.

## Risks

- Low risk. Each menu opening on staging can make one additional health
request.
- The host check applies to `*.staging.paperclip.app`. Other staging
domains will need an explicit update.
- The label identifies the running server commit. It can briefly show
cached data while the request completes.

## Model Used

OpenAI Codex, GPT-6. The exact model variant and context window are not
exposed in this session. Used code execution, repository inspection, and
browser inspection.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#123` / `Refs #123` OR (b) described the issue in-PR following the
relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 13:43:30 -05:00
Devin FoleyandPaperclip e9debd3eac fix(workspaces): allow larger status output for readiness checks (#14414)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The server checks workspace contents before it allows cleanup or
branch reconciliation.
> - These checks need the full Git status output to count untracked
files.
> - Nested task worktrees can make this output exceed the scheduler's
default 1 MiB limit.
> - A scan failure then blocks an otherwise inspectable workspace.
> - This pull request raises the limit for these status checks to 32
MiB.
> - The checks retain exact counts and still protect uncommitted work
from cleanup.

## Linked Issues or Issue Description

Refs #14194 and #14253. Those changes address snapshot enumeration. This
PR addresses buffered execution-workspace status checks. Refs #13619 for
separate work on caching these checks. The related journal identity
failure is covered by #14312.

**What happened?**

Close-readiness checks failed when Git status output exceeded 1 MiB. A
workspace with thousands of long untracked paths could not report its
file count or complete readiness inspection.

**Expected behavior**

Allow up to 32 MiB of status output for execution-workspace readiness
and branch reconciliation. Preserve exact counts. Continue to block
cleanup when the workspace has uncommitted data or the scan exceeds its
bound.

**Steps to reproduce**

1. Create an execution workspace with a merged delivery.
2. Add 5,000 long untracked filenames under a nested task directory. The
status output exceeds 1 MiB.
3. Request close readiness. Before this fix the status scan fails. After
this fix it reports all 5,000 files.
4. Run the terminal-workspace sweep. Confirm that it preserves the
workspace and files.

**Paperclip version or commit**

Reproduced on master at `14795136f5` before this fix.

**Deployment mode**

Server with managed Git workspaces.

## What Changed

- Set a 32 MiB stdout bound for execution-workspace status scans.
- Add a real Git regression with 5,000 long untracked paths and a
cleanup-preservation assertion. Assert that the measured status output
exceeds 1 MiB.
- Document the bound and failure behavior.

## Verification

- The new regression failed on master before the service change: the
status result had no untracked files or count after the scan exceeded
its bound.
- `pnpm exec vitest run
server/src/__tests__/execution-workspaces-service.test.ts
server/src/services/workspace-git-operation-scheduler.test.ts` passed:
82 tests, including the new regression. The regression and server
typecheck also passed after the explicit byte-count assertion.
- `pnpm build` and `pnpm -r typecheck` passed. The full local `pnpm
test:run` was attempted and stopped after the failures listed below.
Greptile is 5/5 with zero unresolved review threads on the latest head.
All checks for head `45f93ad5ad` passed (53 successful, two intentional
skips).
- Full local validation did not pass. The attempt reproduced 13
company/runtime skill-cache permission failures, the terminal-workspace
cleanup assertion, and a heartbeat feedback timeout. It was stopped
during the general-server stage after these failures. Remaining
general-server tests, other workspace groups, and serialized-server
stages did not complete locally. Earlier clean-master checks reproduced
the cache failures and isolated cleanup retries passed. The latest-head
GitHub suite passed all of these groups.
- One GitHub browser shard initially failed because its
GitHub-connection test checked the resume-action array before the mocked
request completed. The single failed-job retry passed without a source
change. All latest-head checks are green.
- No browser suites ran locally. This change does not affect browser
behavior.

## Risks

- Each active status scan can buffer up to 32 MiB before parsing. The
existing scheduler bounds scan concurrency and queue size.
- Output above the bound still fails the scan and blocks destructive
cleanup. The change does not truncate output or change cleanup rules.
- There are no API, database, or snapshot-streaming changes.

## Model Used

OpenAI Codex, GPT-6, with repository inspection, code editing, and test
execution. The exact deployment model ID and context window are not
exposed in this session.

## Checklist


- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused suites; full
local-run failures and incomplete stages are documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.928.0-canary.5
2026-09-28 09:59:56 -07:00
Devin FoleyandPaperclip 6f40e23536 fix(logging): redact cloud authentication headers (#14413)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The server records HTTP requests to help operators diagnose
failures.
> - Cloud requests carry tenant credentials and signed assertions in
headers.
> - The HTTP logger did not redact four of these headers.
> - This pull request adds those headers to the existing redaction list.
> - Operators retain the route, method, and response status without
recording these values.

## Linked Issues or Issue Description

**What happened?**

HTTP request logs could contain cloud tenant credentials, session
identifiers, runtime identity assertions, and cloud control assertions.

**Expected behavior**

The logger must redact these header values for successful requests and
failed requests.

**Steps to reproduce**

1. Create an Express app with the production HTTP logger and redaction
configuration.
2. Send a request with the four cloud headers and distinct test values.
3. Inspect the serialized request headers for responses with status 200,
403, and 500.

**Paperclip version or commit**

Reproduced on `0f14d26123` before this fix.

**Deployment mode**

Server with cloud proxy authentication. The regression test uses an
in-process Express server.

## What Changed

- Redact `x-paperclip-cloud-tenant-token`,
`x-paperclip-cloud-session-id`, `x-paperclip-cloud-runtime-identity`,
and `x-paperclip-cloud-control` in HTTP request logs.
- Test real logger output for HTTP 200, 403, and 500 with mixed-case
request header names. Route 403 and 500 through the production error
handler. Check response bodies, log levels, and server error context.
- Check that the method, route, and status remain available.

## Verification

The `server/src/__tests__/http-log-redaction.test.ts` suite passed,
including all three new cloud-header cases.

- Rebased onto master at `14795136f5`.
- `pnpm exec vitest run server/src/__tests__/http-log-redaction.test.ts`
passed: 59 tests, including all three new response-status cases. The
suite and server typecheck also passed after the error-handler coverage
update.
- `pnpm build` and `pnpm -r typecheck` passed. The full local `pnpm
test:run` was attempted and stopped after the failures listed below.
Greptile is 5/5 with zero unresolved review threads on the latest head.
All checks for head `373d29e2f1` passed (53 successful, two intentional
skips).
- Full local validation did not pass. The attempt reproduced
company-skill cache permission failures, the terminal-workspace cleanup
assertion, a heartbeat feedback timeout, and one process-conversation
timing failure. It was stopped during the general-server stage after
these failures. Remaining general-server tests, other workspace groups,
and serialized-server stages did not complete locally. Earlier
clean-master checks reproduced the cache failures and isolated cleanup
retries passed. The latest-head GitHub suite passed all of these groups.
- No browser suites ran locally. This change does not affect browser
behavior.

## Risks

- These four values will no longer be available in HTTP logs. Route,
method, and status remain available.
- This change applies to new log entries. It does not remove old entries
or rotate credentials.
- No schema, API, or authentication behavior changes.

## Model Used

OpenAI Codex, GPT-6, with repository inspection, code editing, and test
execution. The exact deployment model ID and context window are not
exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused suites; full
local-run failures and incomplete stages are documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.928.0-canary.4
2026-09-28 09:59:15 -07:00
DottaandPaperclip 14795136f5 fix(runtime): finalize and recover sandbox workspace exports safely (#14402)
Serialize native workspace finalization, validate streamed archives within bounded limits, and quietly recover unsafe exports from saved results. Preserve exact allocations for exhausted transient failures and provide export-only retry without rerunning the provider.

Consolidates #14314, #14315, #14329, and #14334 while preserving the already-merged finalization label changes.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
canary/v2026.928.0-canary.3
2026-09-28 10:57:32 -05:00
DottaandPaperclip 3447609d22 fix(runner): stream and page large API responses within capture budgets (#14301)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents use governed API tools to inspect task evidence.
> - Large API results become saved assets with short previews.
> - Reading an asset through the same tool used to create another asset,
so the agent could not reach the rest of the evidence.
> - The 10 MiB response cap also blocked useful large results. Removing
all bounds allowed excessive disk use.
> - This pull request streams responses up to 1 GiB and makes saved text
readable in bounded pages. It adds durable run budgets and capture
admission limits.
> - Agents can inspect complete evidence while tool results, memory use,
and capture work stay bounded.

## Linked Issues or Issue Description

**What happened?**
A large response became an asset. Reading that asset returned another
asset and the same preview. Responses above 10 MiB failed before the
agent could read any page.

**Expected behavior**
The agent can fetch a large response and read its saved text to EOF.
Each page stays bounded. New snapshots have a generous finite limit and
a durable run budget. Existing larger assets remain readable through
byte ranges.

**Steps to reproduce**
1. Call a GET operation that returns more than 10 MiB of text or JSON.
2. Before the fix, the tool returns `api_transport_failure`.
3. With this change, responses up to 1 GiB become streamed snapshots
with artifact references.
4. Read `GET /api/assets/{assetId}/content` with `responseText:
{offsetBytes: 0, limitBytes: 8192}`. Follow `nextOffsetBytes` until
null.

Related work: #14186 added the API fallback tools. #14218 bounded API
discovery.

## What Changed

- Add authenticated UTF-8 text windows to `call_api`, with byte offsets
and total size. Keep each page at or below 24 KiB.
- Stream new responses above 24 KiB through private temporary files into
company-owned assets. Bound each capture to 1 GiB of decoded bytes.
Reject oversized declared lengths before reading and count streamed
bytes before writing.
- Reserve capture budget in the run record before spilling. Allow 4 GiB
per run. Settle successful captures to their actual size. Failed or
interrupted captures retain their full 1 GiB reservation. Run restarts
do not reset the budget.
- Enforce a 20 GiB company snapshot quota with database reservations.
Count legacy snapshots and unfinished storage work across runs and
processes. Asset deletion frees quota.
- Limit large captures to two per company and four per server process.
Hold slots through storage upload and temporary-file cleanup. Use a
10-minute download deadline and 30-second connection/idle-read timeouts.
- Return explicit size, budget, busy, and timeout errors. Preserve
unknown outcomes for mutations whose response cannot be captured.
- Read saved assets through authenticated storage ranges, with at most
two extra bytes for UTF-8 and EOF handling. Unpaged reads return the
existing asset and digest with a bounded preview. Reads create no copies
and do not consume capture budget.
- Keep existing assets above 1 GiB readable in pages. Use safe integer
offsets and PostgreSQL `bigint` asset sizes.
- Stream large S3 uploads through ordered multipart requests. Abort
failed uploads and remove partial local files.
- Revalidate run authority during downloads. Keep company authorization,
GET-only text paging, redirect denial, and mutation replay receipts.
- Document the separate 10 MiB upload limits. This PR does not raise
memory-buffered attachment ingestion limits. Future large video uploads
need streamed ingestion and storage quotas.

## Verification

- Full workspace `pnpm -r typecheck` and `pnpm build` pass after
rebasing on master.
- Focused API and response tests: 1,761 pass. Cover declared and chunked
oversize responses, incorrect Content-Length, exact-limit success,
active-stream deadline, cancellation, cleanup, concurrency admission,
and mutation outcome handling.
- Real HTTP integration: 28 tests pass, including runnerd → PRP →
authority → HTTP, a 12 MiB snapshot, final-page/EOF reads, cross-company
denial, a persisted 3 GiB sparse asset, and large mutation receipt
replay.
- The HTTP suite verifies durable run-budget accounting, simultaneous
runs competing for company quota, legacy snapshot accounting, deletion
refunds, failed-storage reservations, cleaned-failure refunds,
metadata-rollback cleanup refunds, preservation after a lost commit
acknowledgement, and small/saved reads after capture-budget exhaustion.
- A standalone proof streams exactly 1 GiB through the production
capture helper, verifies the final bytes, and removes its temporary
file. It uses repeated 256 KiB chunks and records a peak process RSS of
191 MiB.
- Earlier storage verification covers exact S3 multipart boundaries,
cleanup/abort failures, and a 17 MiB transfer through the real AWS SDK
to a local HTTP S3 endpoint. No cloud S3 qualification was run for this
follow-up.
- The local full test run was interrupted for the company-quota changes.
A later targeted run hit exhausted macOS shared-memory slots before
tests started; two unattached PostgreSQL segments with dead owners were
reclaimed before retrying. All 55 current-head checks pass at
`aebb80ceeeee77d5a56b67bfffd835f2f846878c`, including the full CI test
suite, typecheck, build, browser suites, security scan, and Greptile
(5/5). There are no unresolved review threads. The combined rebased test
catalog also passes (48 tests).
- Earlier paging acceptance passed Daytona and separate staging at
`7739879e9`. Those runs predate the streaming and budget changes.

## Risks

- The 1 GiB response cap and 10-minute active-download deadline are
intentional product limits. Larger live results must use endpoint
pagination or a direct file workflow. Existing larger assets remain
readable through bounded ranges.
- A durable 20 GiB company snapshot quota counts stored runner-api
assets and active/orphan reservations across runs and processes. The
operator can set PAPERCLIP_RUNNER_API_COMPANY_CAPTURE_MAX_BYTES to a
finite value of at least 1 GiB. Deleting snapshots frees capacity;
possible orphan storage must be reconciled before releasing its
reservation.
- A failed capture uses its full reservation. A new large capture needs
a full 1 GiB available, even if it later completes at a smaller size.
Small reads and existing asset pages remain available.
- Concurrency limits apply per server process. The run byte budget is
shared through the database.
- The `integer` to `bigint` migration rewrites asset metadata and takes
an exclusive table lock. File bytes stay in storage.
- A live endpoint is fetched once before returning its snapshot.
Continue reading the saved artifact for stable pages. Mutations may
commit before any size or transport error; inspect state before
retrying.
- Attachment uploads and native file handoffs still default to 10 MiB.
Raising buffered ingestion paths to GiB sizes is separate work.

## Model Used

OpenAI Codex, based on GPT-6, with code execution and repository tools.
The runtime does not expose an exact serving model variant or
context-window size. The earlier paging work also used browser testing
and subagents.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.928.0-canary.2
2026-09-28 10:36:32 -05:00
DottaandPaperclip cea8dda472 test: evaluate completion updates after native task handoffs (#13969)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Users can delegate work through onboarding and Agent Chat.
> - A completed task does not prove that its result reached the original
conversation.
> - Existing tests do not isolate completion after the source chat
becomes idle.
> - This pull request adds four explicit native-runner probes across
Claude and Codex.
> - The probes preserve the result and reply so we can separate delivery
failures from inaccurate answers.

## Linked Issues or Issue Description

Refs #13775. Refs #13813.

These evals extend native-runner qualification. They measure completion
updates before we choose a product change.

## What Changed

- Add the opt-in `completion-updates` suite with two stories for each
native provider.
- Test completion in the existing onboarding task flow and after an
Agent Chat handoff becomes idle.
- Gate the chat worker on a brief inside its managed project workspace.
Prove the source is idle before releasing the worker.
- Check durable task completion, saved output, a subsequent source
reply, and rendered access to the result.
- Preserve replies, task state, screenshots, run events, and a separate
semantic review rubric.
- Add grader regression tests and update the documented eval contract.
- Preserve the suites added on master and include four completion cases
in the 306-cell catalog.

Production behavior and prompts are unchanged.

## Verification

- Passed all 565 eval support tests across 45 files after merging
current master: `node node_modules/vitest/vitest.mjs run --config
tests/runner-e2e/vitest.config.ts`.
- Passed eval TypeScript: `node node_modules/typescript/bin/tsc -p
tests/runner-e2e/tsconfig.json`.
- Confirmed four selected cells: `node cli/node_modules/tsx/dist/cli.mjs
tests/runner-e2e/launch.ts --list --suite completion-updates`.
- Four-cell behavior campaign on source
`ad47cf1da2b1e36f19f4227cfeb53998720b0b5b`:
https://github.com/paperclipai/paperclip/actions/runs/36072337485.
- A screenshot-only follow-up waits for the restored source reply to
render after result-link navigation. Its one-cell Claude onboarding
verification passed on final head:
https://github.com/paperclipai/paperclip/actions/runs/36075716141.
Corrected report:
https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36075716141-1/.
The original four-cell onboarding screenshots caught navigation loading;
its saved reply evidence remains valid. The follow-up again found stale
wording: "That work will run next" was posted 38 seconds after the child
was Done. The four-cell campaign keeps its original source and
measurements.
- Published evidence:
https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36072337485-1/.
- Suite definition:
`afba4d85d6c53d9f64c08b37a2e9cc20481b78f5bd7e2fa012045e2c69444d9d`,
version 6. Models: native `gpt-5.6-sol` and `claude-sonnet-5`, local
execution, one attempt per cell. All four cleanup checks passed.
Onboarding billing coverage is partial; reported zero cost must not be
read as a free run.

| Story | Automated delivery/access | Separate semantic review |
| --- | --- | --- |
| Codex onboarding | Pass | Pass: accurate completion reply with an
accessible result |
| Claude onboarding | Pass | Fail: reply says it will save the note once
the task runs, after the note is already saved and the task is Done |
| Codex idle chat handoff | Fail | Worker completed and saved the note;
no completion reply during the full observation window |
| Claude idle chat handoff | Fail | Worker completed and saved the note;
no completion reply during the full observation window |

Both chat cases positively recorded the source waiting and the worker at
the brief gate before release. Both saved outputs include the brief-only
start time. The opt-in campaign is red because it exposes current
behavior. It is not a required merge gate. The PR does not fix that
product behavior. Semantic review is a recorded human/agent assessment
of retained evidence; it is not an automated prose-quality judge.
- Second campaign:
https://github.com/paperclipai/paperclip/actions/runs/36071065098. Codex
chat reached the idle boundary and completed its task, then received no
completion reply during the full window. Claude onboarding again
returned a stale handoff answer. Claude chat exceeded the prior
110-second handoff setup budget; this revision raises that bounded setup
window to 180 seconds.
- Retained baseline:
https://github.com/paperclipai/paperclip/actions/runs/36069427676.
Onboarding passed delivery/access for both providers, but Claude gave a
stale handoff answer. Chat cases stopped at fixture problems; they do
not establish a completion-delivery failure. This revision fixes the
workspace path and competing reference requirements.
- On the previous head `4023a2a3c28d45c9eb2c42d452ce99ffba5c7b73`, 54 PR
checks passed and two were skipped, including typecheck, tests, and
build. Broad checks ran in CI, not locally. That head received Greptile
5/5 with no unresolved findings. The unchanged mobile
repository-settings browser test passed on one targeted retry after a
detached/disabled Save-button timeout.

- Merged current master in `9b4491e1f` and resolved the catalog-count
conflict. Eval support tests and eval TypeScript pass locally. All
individual CI jobs passed on this merge commit, including build,
typecheck, server tests, runner checks, and browser shards. The final
aggregate check also passed: 54 checks passed and two were skipped.
Greptile reviewed this exact commit at 5/5 with no unresolved findings.

## Risks

- These explicit probes can expose current product failures. They do not
change the default paid test selection.
- Mechanical delivery and result access do not establish answer
accuracy. The preserved reply still requires semantic review.
- A fixture failure before the idle boundary or worker completion cannot
establish a completion-update failure.
- The handoff setup window lasts three minutes. The worker brief wait is
bounded at four minutes. The observation window lasts two minutes after
worker completion. It retains later replies without erasing earlier
accessible delivery.

## Model Used

OpenAI Codex, GPT-6 (`gpt-6-astra`), with reasoning, repository
inspection, code execution, and GitHub tool use. The runtime does not
expose the context window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 09:58:14 -05:00
DottaandPaperclip 8751e2de46 fix(ui): distinguish finalization recovery from live observation (#14326)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The task board shows which recovery actions are active.
> - A native run can stop while a person must repair its workspace.
> - The board previously called that state “Recovery in progress.”
> - The label implied that work would continue without operator action.
> - This pull request derives the label from the recovery owner and live
continuation.
> - Operators can distinguish scheduled recovery from a repair that
needs attention.

## Linked Issues or Issue Description

**What happened?**
A blocked task showed “Recovery in progress” after native finalization
stopped and no automatic continuation remained.

**Expected behavior**
Show “Recovery needed” for an idle board repair or when no live recovery
path exists. Show “Recovery in progress” while the recorded continuation
can run, including an explicitly admitted export whose exact callback is
executing even if the old recovery action remains board-owned.

**Steps to reproduce**
1. Complete a native run whose workspace export cannot be recovered
automatically.
2. Inspect the task recovery action and badge.
3. Compare the board-owned action with the old “Recovery in progress”
label.

Related work: #14314 serializes native workspace finalization and fences
stale recovery outcomes. This change reports the recovery action that
currently owns the task.

## What Changed

- Show “Recovery needed” for board-owned active-run recovery unless the
exact native export callback is positively verified as executing.
- Require the native continuation run and a live or future continuation
before showing progress.
- Project native activity from the exact company, source issue, and run.
Include active workspace export while the original heartbeat remains
failed.
- Require an executing callback before using a running export row as
evidence. Preserve activity for long exports and clear it when the
callback joins.
- Give native resume its own card explanation. Preserve ordinary
watchdog observation behavior.
- Remove the redundant ownership sentence from all six recovery-card
explanations that used it.
- Document the labels and add regression cases for stopped, scheduled,
and active recovery.

## Verification

- Copy-only follow-up (`c68aef04c`): all 167 focused recovery UI tests
and token gates pass. No UI occurrence of the removed sentence remains.
`pnpm -r typecheck`, `pnpm build`, and current-head CI pass (54
successful checks, two optional Storybook checks skipped). Greptile is
5/5 with no open review threads. The duplicate local `pnpm test:run` was
stopped after the full CI suite passed; it did not complete locally.
- Original regressions: seven failures before the change, then 52
focused cases pass.
- Review regressions: seven UI failures and nine database failures
before the follow-up. All 75 database/API recovery tests, 167 UI tests,
and 18 workspace lifecycle/finalizer tests pass. A further three RED
cases cover explicit board retry activity; one RED case rejects orphaned
running export rows after controller loss. Wrong company, issue, run,
service, phase, and completed-operation cases remain inactive.
- Final recursive typecheck, production build, token gates, and complete
local suite coverage pass. Embedded PostgreSQL startup/socket failures
passed in isolated retries with the canonical test environment; no
expected behavior was weakened. The recovery regression added during the
earlier full run passed in its final complete 75-case file.
- Before the copy-only follow-up, all 56 CI checks passed on
`c2f84cd89c905cda85c53aaf5bb83b7250900fe6`; Greptile is 5/5 with no
unresolved review threads. The final native-activity staging repeat
passed on integrated source `2bedd0f23bf4698b1f8b818f6796900647030427`.
- Verified on a separate staging instance: a real failed Daytona
workspace export retains its board-owned repair action and displays
“Recovery needed” in the task list. The repair card remains actionable
without starting another provider turn.

- Real staged export-only repair: the actual task list showed “Recovery
in progress” while the original run had a positively identified running
export operation, then Done after exact copyback of all 20,000
nonce-bound files. The accepted result and full provider
session/turn/terminal envelopes remained unchanged. All 17 independent
final checks passed; the browser downloaded the exact 19-byte result.
The separate fixture was cleaned up with independent provider-absence
verification. A control transport process restarted during repair; it
did not submit another provider turn.

- Final deployed-source repeat on
`d884e1ab046cc76004e35e6091e9e6e2c918c9eb`: explicit per-turn ephemeral
Daytona allocation, actual browser export repair, and a saved full
activity projection referencing the exact executing export operation
with no scheduled retry. The task list showed recovery in progress, then
Done; all 21 final checks passed, including 20,000 exact host files,
unchanged provider provenance, and provider deletion only after
committed copyback. The downloaded 19-byte result matched independently.
This integrates #14334; no source change was required here.

## Risks

- The label depends on the persisted recovery action. A separate runtime
defect can still stop work; this change makes that condition visible.
- Future recovery kinds must supply a valid continuation path before
they can display progress.

## Model Used

OpenAI Codex, based on GPT-6, with code execution, browser testing, and
subagent tool use. The runtime does not expose an exact serving model
variant or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.928.0-canary.1
2026-09-28 09:33:39 -05:00
DottaandPaperclip cbc5132e6c fix(adapters): expose a verified provider stop before workspace restoration (#14311)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agent adapters own local or remote provider processes.
> - Run cleanup must know when the final provider process has stopped.
> - A remote timeout or a lost transport does not prove that the
provider stopped.
> - Retried provider invocations also make an earlier stop signal stale.
> - This pull request adds a verified final-invocation stop callback
before workspace restoration.
> - The dependent instruction revision change uses that boundary to
preserve private instruction edits safely.

## Linked Issues or Issue Description

**What happened?**

Adapter completion did not expose a reliable point between provider
shutdown and workspace restoration. Cleanup could lose provider-written
files, or treat a remote timeout as proof that a process stopped.

**Expected behavior**

Cleanup runs once after the final provider invocation has a verified
stop receipt and before workspace restoration. An incomplete remote
command keeps collection pending.

**Steps to reproduce**

1. Run a remote provider that returns a timeout without a numeric exit
code.
2. Let the adapter return or retry the provider.
3. Attempt to collect provider-written files during cleanup. The adapter
has no verified final-process boundary to use.

This is the prerequisite for the stacked canonical instruction revision
pull request. It has no database or UI dependency. Related #13291
verifies remote termination for later recovery; this change exposes the
earlier adapter-owned stop boundary before workspace restoration. The
scopes do not duplicate each other.

## What Changed

- Add a stop callback to adapter execution context and a
final-invocation fence.
- Confirm local child closure and complete remote exit receipts. Reject
remote timeouts, missing exits, transport failures, and SSH exit 255 as
stop proof.
- Invoke collection once before workspace restoration in eight CLI
adapters and at the confirmed ACP stop boundary.
- Preserve the stop observation when a later log flush fails.
- Run bridge and workspace cleanup in `finally` even when collection
rejects. ACP records a safe error without exposing a raw filesystem
path. Grok keeps collection errors separate from workspace restore
failures and preserves completed provider results when both cleanup
steps fail.
- Correct the existing Cursor test shell fixture so bounded remote file
reads run against real fixture files.

## Verification

- Three stop-boundary regressions failed before the callback
implementation and passed after it.
- Independent prerequisite branch: 377 tests passed across 24 adapter,
process-target, and ACP suites. Two added collector-rejection tests
failed before the cleanup fix and passed after it.
- A third regression reproduced Grok misclassifying a collection failure
as failed workspace restoration. Two additional cases covered completed
and failed provider turns when collection and restore both fail. The
Grok and restore-classifier suites passed 48 tests.
- All nine affected package typechecks and affected package builds
passed; Grok checks passed again after its classification fix.
- Integrated instruction branch: native local, legacy local, and native
Daytona each passed three browser tasks with exact persisted bytes,
fresh-task readback, history/restore, and explicit conflict resolution.
Unchanged warm Daytona passed three turns. Legacy Codex passed all five
checks again after the exception-safe cleanup fix.
- Alternate staging passed the same three-task native Daytona flow:
exact stopped-run save, independent downloaded readback, browser
history/restore, and explicit resolution of a real concurrent edit. The
deployed source was `14c3d810c9e05625121b3d27767aea9317b03125`, which
covers the initial adapter callback. Later Grok cleanup failures are
qualified by the adapter tests above.

- The final combined native Daytona flow passed again on deployed
`2bedd0f23bf4698b1f8b818f6796900647030427`: three fresh tasks proved
ordinary instruction edits, independent readback, History/Restore,
concurrent board conflict, and explicit candidate resolution. This
native staging flow does not claim to exercise the Grok adapter.

## Risks

- An unverified remote stop intentionally does not trigger collection. A
later controller with verified stop evidence must recover it or report
the copy unavailable.
- The callback is optional. Callers that do not register it retain their
existing behavior.
- The callback runs before workspace restoration and can delay cleanup
if its caller does not bound its own work. The dependent instruction
collector uses bounded reads and retries. A rejected callback still
permits bridge and workspace cleanup; it cannot claim an instruction
save.

## Model Used

OpenAI Codex, GPT-6, with tool use and code execution. The runtime does
not expose the exact deployment variant or context-window size. Multiple
Codex agents implemented and verified the change.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 08:45:53 -05:00
DottaandPaperclip 4b38db9622 fix(runtime): stream workspace Git snapshots through disk manifests (#14253)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Managed runs copy a selected workspace to an execution environment
and restore its changes.
> - Git snapshots select the files for that copy and for later recovery.
> - A fixed output limit stops large generated trees before the run can
start.
> - Increasing the limit still keeps the complete filename lists in
memory.
> - This pull request stores those lists and merge baselines in disk
manifests.
> - Large snapshots can now complete with bounded filename buffers and
explicit failure handling.

## Linked Issues or Issue Description

Refs #14194. This is the streaming follow-up to the merged 32 MiB limit
fix.

Related: #13619 and #11621 cover workspace scan admission and demand.
This change keeps the shared scheduler and changes the snapshot data
path.

## What Changed

- Stream changed, untracked, deleted, and ignored paths through the
shared scheduler and the standalone adapter path.
- Use SQLite manifests for file selection, duplicate removal,
ignored-path lookup, baseline capture, and merge lookup.
- Set a configurable 30-minute snapshot deadline. Keep the existing
interactive scan deadlines.
- Wait for each child process and pending sink write before removing
temporary storage after failure or cancellation.
- Use NUL archive lists and bounded deletion batches. Preserve unusual
names, explicit selection, nested repositories, and source root checks.
- Store manifest references in native recovery format v2. Check their
location and digest before recovery reads. Keep v1 descriptors readable.
- Remove temporary manifests at lifecycle completion. Use fixed-size
temporary copy names for long basenames.
- Admit each manifest with a SQLite page allowance based on current disk
capacity. Keep a configurable free-space reserve and fail explicitly
when either limit is reached.
- Preserve a host file that replaces a directory deleted by the sandbox,
and continue the rest of the restore.

## Verification

- Current head `5ab622ec43cd16d35429d79dedee6a5d8e3d2df2` has 54
successful checks/statuses and two skipped Storybook jobs. No checks
failed or remain pending.
- [CI
passed](https://github.com/paperclipai/paperclip/actions/runs/36318966368):
typecheck, build, all test shards, E2E, Rust checks, and the aggregate
verify job.
- [Greptile is
5/5](https://github.com/paperclipai/paperclip/pull/14253#issuecomment-5855670338)
on the current head. All four review threads are resolved. Security
checks passed.
- 229 focused tests passed across Git sync, runtime staging, merge,
manifest integrity, native recovery, and the scheduler (214
adapter/runtime tests and 15 scheduler tests).
- A real 40,000-file fixture produces 43,428,890 filename bytes. The
original standalone and scheduled scans fail. The new test passes all
four filename paths, complete staging, exclusion of late files, unusual
names, and deletion replay.
- Recovery tests reject changed bytes, symlinks, and paths outside the
controller state directory. Adapter-utils typecheck passed.
- A test executor returned buffered output and caused two retry
integration failures. The fixture now uses the shared streaming
scheduler. All 13 tests passed with `corepack pnpm exec vitest run
server/src/__tests__/heartbeat-project-repositories.test.ts`. The same
CI shard now passes.
- Ran `pnpm -r typecheck`, `pnpm test:run`, and `pnpm build` locally.
Each full local command hit SIGKILL/exit 137 in the 4 GiB container.
These local commands did not pass. The current-head CI gates above
provide the full verification.
- A real-Git disk-capacity regression confirms a typed failure and
removal of the incomplete manifest. Repeated writer attempts cannot
exceed the permitted page count.

- Follow-up real Daytona and separate staging qualification passed with
the related archive validator (#14315) and exact-owner finalization fix
(#14314). Three successive turns copied back all 60,000 files with
39,828,890 filename bytes and five unusual names. Independent host
inventories verified every file and the pinned Git HEAD. Native,
provider, session, and process identities stayed fixed; no retry
remained. The task reached Done, and its browser-downloaded final proof
matched exactly. The reusable regression is #14316, including an
assertion of the effective environment idle policy.

## Risks

- SQLite manifests use disk space. Each receives one quarter of the
available capacity above the host reserve at creation. The reserve
defaults to 256 MiB and has a 64 MiB configuration minimum. Disk
capacity, filesystem quotas, per-path limits, Git resource use, and
execution deadlines remain limits.
- Each path and sink chunk has a 64 KiB limit. SQLite connections use a
1 MiB page cache. Invalid or incomplete records fail explicitly.
- Restore transport keeps fixed and configured archive exclusions. A
remotely created Git-ignored file can be transferred, but the host merge
excludes it through the manifest.
- Provider archive buffers, Git and tar memory, repository metadata,
legacy v1 arrays, and the separate referenced-source resolver retain
their own limits. Existing provider safety validators still buffer
textual tar listings: Daytona allows 32 MiB and Kubernetes allows 64
MiB. These separate transport limits can stop a sufficiently large
restore before merge. This change does not claim bounded total process
memory or unlimited transport size.
- New descriptors use v2. Existing v1 recovery remains supported; a
downgrade cannot read v2 descriptors.

## Model Used

OpenAI GPT-6 through Codex. The exact deployment ID and context limit
are not exposed in this run. The agent used code editing, terminal
execution, tests, and GitHub tools.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 08:40:36 -05:00
DottaandPaperclip 890d11137f fix(ui): register artifact tabs without opening the panel (#14193)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Tasks keep agent outputs in the Artifacts tab.
> - An output can arrive while the user writes a message or reads a
document.
> - Opening the side panel on arrival interrupts that work, especially
on mobile.
> - This pull request adds the Artifacts tab without opening the panel
or changing the selected tab.
> - Users can open their outputs when they choose.

## Linked Issues or Issue Description

**What happened?**
New agent outputs opened the task side panel or mobile drawer. An
arrival could also replace the selected document or workspace file.
Existing outputs did not always register an Artifacts tab.

**Expected behavior**
Register one Artifacts tab for existing and new outputs. Keep a closed
panel closed. Preserve composer focus, the selected tab, and document or
file links.

**Steps to reproduce**
Open a task from the inbox. Close its side panel. Enter a message draft.
Create an agent output in that task. The panel must stay closed and the
draft must keep focus. Open the panel to see the Artifacts tab. Repeat
on a mobile viewport.

**Paperclip version or commit**
Base commit: 0f14d2612.

Related work: #11226 and #11551.

## What Changed

- Register existing outputs and later arrivals without opening the panel
or selecting Artifacts.
- Keep open documents, workspace-file links, and the tab launcher
unchanged.
- Handle each document deep-link request once so query refreshes
preserve later manual selection.
- Deduplicate attachment and work-product arrivals. Preserve dismissed
tabs across repeated refreshes.
- Add desktop and mobile browser regression tests. Update artifact
presentation documentation.

## Verification

- The closed-panel regression failed before the fix in unit and
real-browser tests.
- All 158 focused UI tests pass, including the original deep-link cases.
- UI typecheck and token gates pass on this branch. Full typecheck and
production build passed on the passive-arrival candidate before the
existing PR integration.
- The two local desktop/mobile browser cases pass against real
API-created artifacts.
- Desktop and mobile staging checks pass on the combined staging
candidate. Artifact arrival preserved a closed pane, draft text, and
composer focus. Explicitly opening the pane showed the Artifacts tab.
- All 56 current-head check contexts are successful or intentionally
skipped at `17ad904455b9378552f07a6f6e51402c6d164688`, including full
typecheck, test shards, build, and browser suites. Greptile is 5/5 on
that commit with no unresolved review threads.
- The interrupted local broad validation was resumed; the remaining
serialized 64 files and 1,035 tests pass. Local database startup
failures passed after stale test resources were released.

## Risks

- Outputs no longer reveal the panel automatically. Users open the panel
to view them.
- The document request guard must still allow a new explicit deep link.
The regression tests cover this case.
- There are no API or database changes.

## Model Used

- OpenAI Codex, GPT-6, with code editing, shell tools, GitHub tools, and
browser verification. The exact serving model ID and context-window size
are not exposed in this session. Earlier implementation model metadata
is not available.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 08:38:00 -05:00
DottaandPaperclip bacc0e6a98 docs: remove automatic agent escalation from coordination skill (#14188)
## Thinking Path

> - Paperclip manages work for AI agents.
> - The coordination skill tells agents how to handle blocked work.
> - The skill directs blocked work to managers and other agents.
> - Those agents may lack the same access or authority.
> - These extra assignments delay the required human action.
> - This pull request removes automatic escalation advice.
> - Agents must identify the missing capability and use the correct
approval or human-input path.

## Linked Issues or Issue Description

**Issue type**

Incorrect information.

**Where is the issue?**

The critical rules in `skills/paperclip/SKILL.md` and blocker guidance
in `skills/paperclip/references/api-reference.md`.

**What's wrong?**

The skill tells agents to escalate through `chainOfCommand`, ask another
agent for help, and avoid human help. A manager title does not grant
permission to fix a connection or complete an administrator action.

**Suggested fix**

Remove blanket escalation and agent-first rules. Keep normal delegation
when the recipient has a concrete capability for a bounded task. Use
saved human-input interactions or existing connection and approval flows
for human-only actions.

Searches for open PRs with “escalation”, “chainOfCommand”, and “ask
another agent” found no duplicate skill change. Recovery-routing PRs
change server behavior, which is outside this change.

## What Changed

- Remove the chain-of-command escalation rule and both repeated
agent-first directives.
- Replace manager handoffs in the API reference with direct blocker
handling.
- Keep reporting fields, normal delegation, approval gates, and the ban
on bypassing permission denials.
- Keep the ban on cancelling cross-team tasks. Request a decision
instead of automatically assigning the task to a manager.

## Verification

- Final head `1483e82cc3fe10a7c910230b3779635d6d54c162`: Greptile 5/5,
no outstanding findings, all CI checks pass (optional Storybook jobs
skipped).

- Created a fresh worktree at `.worktrees/skill-blocker-guidance` from
the current master commit.
- Ran `git diff --check`: passed.
- Ran Node assertions against both documents: passed. Removed directives
are absent. Human-input and capability-based delegation guidance is
present.
- Checked that checkout conflict, approval, dependency, and normal
delegation instructions remain.
- Attempted `pnpm -r typecheck`, `pnpm test:run`, and `pnpm build`. None
could run because this environment has no `pnpm` executable.
- Refreshed both generated capability inventories and their derived
contract after skill heading positions changed. Both generator and
live-inventory checks pass. The eval and MCP baselines remain unchanged.
- Ran `node --test
packages/paperclip-runner/scripts/check-capability-inventory.test.mjs`:
all four tests pass.
- Ran `node
packages/paperclip-runner/scripts/check-capability-inventory.mjs` and
`node packages/paperclip-runner/scripts/generate-capability-contract.mjs
--check`: both pass.
- Added explicit requester routing for agent and human scope questions.
Focused Node assertions pass.

## Risks

- Agents can request human input earlier for actions that require human
authority.
- Existing runs or installed copies can keep old skill text until
refreshed.
- This change does not alter server recovery routing or permission
checks.

## Model Used

OpenAI Codex, with reasoning, tool use, and shell execution. The runtime
does not expose a verifiable exact model ID or context window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.928.0-canary.0
2026-09-28 08:37:26 -05:00
DottaandPaperclip 0f14d26123 fix(runtime): allow bounded large untracked workspace snapshots (#14194)
## Thinking Path

> - Paperclip manages AI agents and their work.
> - Remote runs need a snapshot of the task workspace before the agent
starts.
> - The snapshot lists untracked filenames through the shared Git scan
scheduler.
> - A generated directory with a few thousand long filenames can exceed
the 1 MiB output limit.
> - This stops setup and prevents the agent from continuing its task.
> - This change gives that listing a 32 MiB bound and keeps the explicit
file snapshot.
> - Normal generated trees can now pass setup, while larger snapshots
still fail at a finite limit.

## Linked Issues or Issue Description

Refs #11572 and #12214 for the existing bounded scan and ignore-scan
protections.

**What happened?**

An agent continuation failed during workspace setup with `Workspace Git
scan exceeded its output limit`. A real Git fixture reproduces the
untracked-file path: 5,000 long filenames in one generated directory
exceed its 1 MiB output limit.

**Expected behavior**

The workspace snapshot must support ordinary generated trees with
thousands of files. It must keep a finite output bound and select
explicit files before staging.

**Steps to reproduce**

1. Commit a base file in a Git repository.
2. Add 5,000 untracked files with long names in one new directory.
3. Call `readGitWorkspaceSnapshot` with the normal scan limits.
4. Observe the output-limit error before this change.

**Paperclip version or commit**

Base commit: `640dee1`.

**Deployment mode**

Remote sandbox execution from source.

## What Changed

- Increase the untracked-file snapshot output bound from 1 MiB to 32
MiB.
- Keep explicit file selection, the shared scheduler, the timeout, and
the other scan bounds.
- Test that 40,000 long filenames above the old 8 MiB bound reach the
snapshot. Reuse those files with deeper paths to prove that output above
32 MiB still fails.
- Test that files created after the snapshot, including an ignored
secret, stay out of the overlay archive.
- Check the workspace root identity and reject selected paths with
symlinked parent directories before upload.
- Test root replacement after path resolution, including root-level
selected files.
- Preserve existing workspace root aliases by capturing the resolved
root before snapshot selection. Test that later alias retargeting cannot
change the archive contents.
- Stop staging on permission and I/O errors; continue to allow missing
files.
- Test these failures and preserve selected symlink entries.
- Accept valid case-renamed directories by checking ancestor file types.
A modeled case-insensitive regression failed before this correction and
now passes.
- Give the 40,000-file fixture enough time to remove its files.
- Document the larger bound and staging behavior.

## Verification

- Red: the 5,000-file regression failed with `stdout maxBuffer length
exceeded` before the fix.
- A separate check through the real server scheduler reproduced
`workspace_git_scan_output_limit` on the original code. The revised code
selected all 5,000 files.
- The late-file regression failed against the first PR revision because
the archive contained `drafts/late.secret`. It passes with the final
explicit-file approach.
- The staging regressions failed before the review fix: a substituted
parent directory and permission/I/O errors were accepted. All three
cases now stop before upload.
- Green: 134 tests passed across `git-workspace-sync.test.ts` and
`sandbox-managed-runtime.test.ts` with Vitest 4.1.11. This includes
complete selection above 8 MiB and rejection above 32 MiB.
- `pnpm --filter @paperclipai/adapter-utils... typecheck` passed after
the revision.
- The module-boundary check and `git diff --check` passed.
- `pnpm -r typecheck` and `pnpm build` stopped in the Rust runner steps
because this environment has no `cargo` executable.
- The full `pnpm test:run` attempt ended with `SIGKILL` during the
general server suite. It did not finish. That full-suite result belongs
to the earlier revision. Fresh checks are required for this revision.

- The new regression fails at the old 8 MiB bound with `stdout maxBuffer
length exceeded`. All 134 focused tests pass with the 32 MiB change.
- The affected typechecks and module-boundary check pass. The full local
typecheck requires Cargo, which is absent in this environment.
- Final verification for `b43e9bc95e55382c6a9bfe200487c164770c8be8`: 54
successful checks/statuses and two skipped Storybook checks. No checks
remain pending or failed.
- The [CI
run](https://github.com/paperclipai/paperclip/actions/runs/36315569674)
passes on attempt 2. The first attempt had one unrelated preview-fixture
readiness timeout. That exact test passed locally; its CI shard passed
on the single rerun.
- [Greptile reports
5/5](https://github.com/paperclipai/paperclip/pull/14194#issuecomment-5852045501)
on this revision. All review threads are resolved.
- [Security review accepts the documented memory
tradeoff](https://github.com/paperclipai/paperclip/pull/14194#discussion_r4115149465)
for this finite mitigation. The separate streaming follow-up will remove
full-list buffering.
- This revision also passes affected local typechecks and the
module-boundary check. Full local typecheck/build stop because Cargo is
absent. The full local Vitest attempt was stopped after about 18 minutes
once all remote gates passed; it did not complete locally.

## Risks

- Each untracked-file scan can buffer up to 32 MiB instead of 1 MiB. The
scheduler still limits concurrent scans and execution time.
- Snapshots above 32 MiB still fail with the existing error. Tracked and
ignored-file scan bounds stay unchanged.
- A workspace that replaces a selected path’s parent with a symlink now
fails staging.
- Detailed logs from the reported host were unavailable. The exact
command that exceeded its limit on that host is unconfirmed.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, tool use, and code
execution. The runtime does not expose a more specific model build or
context window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub references)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-27 06:42:19 -05:00
DottaandPaperclip 7e0f051d3b Align task property icons with assignee avatars (#13319)
## Thinking Path

> - Paperclip is the open source app that people use to manage AI agents
for work.
> - The task properties panel helps operators inspect and update task
data.
> - The status, assignee, and project rows used different leading icon
sizes.
> - The mixed sizes made the rows look uneven.
> - The assignee avatar already provides the correct 24px visual size.
> - This pull request sets the status and project visuals to that same
size.
> - The benefit is a tidy and consistent properties panel.

## Linked Issues or Issue Description

**What happened?**

The task properties panel showed a 12px status icon, a 24px assignee
avatar, and a 16px project tile.

**Expected behavior**

The status, assignee, and project rows must use the same 24px leading
visual size.

**Steps to reproduce**

1. Open a task.
2. Open the Properties panel.
3. Compare the Status, Assignee, and Project rows.

## What Changed

- Keep the Status glyph at 16px and use spacing to give it the 24px
avatar footprint.
- Set the Project row tile to the 24px small tile size.
- Add a regression check for the Status row size.

## Verification

- `pnpm exec vitest run ui/src/components/IssueProperties.test.tsx` (72
tests passed)
- `pnpm check:token-gates` (all gates clean)
- `pnpm --filter @paperclipai/ui typecheck` (passed)
- Full repository typecheck and build reached the Rust runner step, but
this environment does not contain `cargo`.
- The full test command was stopped after unrelated chat integration
tests did not finish. The focused UI suite passed.

## Risks

- Low risk. This change only changes visual sizes in three property
rows.
- The wider status footprint can use slightly more horizontal space in a
narrow panel.

> This fix does not overlap with planned core work in `ROADMAP.md`.

## Model Used

- OpenAI Codex, GPT-5.6, tool use and code execution enabled.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.927.0-canary.5 nightly/v2026.928.0-nightly.0
2026-09-27 11:35:39 +00:00
DottaandPaperclip d423fc4411 fix(ui): use large mobile selector modals (#14250)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The new task dialog lets an operator assign work on a phone.
> - The mobile picker used fixed coordinates inside the Radix wrapper.
> - iOS Safari already makes fixed coordinates relative to the visual
viewport.
> - The old code added the visual viewport offset a second time.
> - This pull request removes the second offset and gives the mobile
wrapper full viewport geometry.
> - The benefit is a visible picker and search field when the iOS
keyboard opens.

## Linked Issues or Issue Description

**What happened?**

The assignee and project pickers could move outside the visible screen
in the new task dialog on iOS.

**What did you expect to happen?**

The picker and its search field must stay visible above the bottom edge
and the software keyboard.

**Steps to reproduce**

Open the new task dialog on an iPhone. Select Assignee or Project. The
picker can move outside the visual viewport when Safari pans the page.

**Paperclip version or commit**

The problem exists after the change in PR #13343.

**Deployment mode**

The problem affects the board UI in local and authenticated modes.

Refs #13343

## What Changed

- Use visual viewport local coordinates for the new task dialog.
- Give the mobile Radix popper wrapper full viewport geometry.
- Render both shared selector components as large, titled mobile modals.
- Keep desktop selectors as anchored popovers.
- Cover the new-task assignee, new-task project, composer assignee, and
generic searchable selector in Storybook.
- Add iPhone-sized Storybook states for the assignee and project
pickers.
- Update viewport tests for the corrected coordinate model.

## Verification

- `pnpm --filter @paperclipai/ui exec vitest run
src/components/NewIssueDialog.test.tsx
src/components/InlineEntitySelector.test.tsx`
- `pnpm check:token-gates`
- Playwright used the iPhone 14 device profile. All mobile modals were
inside the 390 by 664 CSS viewport.
- Each mobile modal was `x=16, y=16, width=358, height=632`.
- A desktop check kept the anchored popover at `width=320, height=211`.

## Risks

- Low risk. The CSS only changes narrow mobile viewports.
- The full-screen popper wrapper does not receive pointer events. The
picker still receives pointer events.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. See `CONTRIBUTING.md`.

## Model Used

- OpenAI GPT-5.6, Codex agent with tool use and code execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal task
ID
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant Storybook documentation to reflect my
changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-27 11:33:24 +00:00
DottaandPaperclip 74508357fa fix(ui): keep mobile task pickers in view (#14249)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The new task dialog lets an operator assign work on a phone.
> - The mobile picker used fixed coordinates inside the Radix wrapper.
> - iOS Safari already makes fixed coordinates relative to the visual
viewport.
> - The old code added the visual viewport offset a second time.
> - This pull request removes the second offset and gives the mobile
wrapper full viewport geometry.
> - The benefit is a visible picker and search field when the iOS
keyboard opens.

## Linked Issues or Issue Description

**What happened?**

The assignee and project pickers could move outside the visible screen
in the new task dialog on iOS.

**What did you expect to happen?**

The picker and its search field must stay visible above the bottom edge
and the software keyboard.

**Steps to reproduce**

Open the new task dialog on an iPhone. Select Assignee or Project. The
picker can move outside the visual viewport when Safari pans the page.

**Paperclip version or commit**

The problem exists after the change in PR #13343.

**Deployment mode**

The problem affects the board UI in local and authenticated modes.

Refs #13343

## What Changed

- Use visual viewport local coordinates for the new task dialog.
- Give the mobile Radix popper wrapper full viewport geometry.
- Keep the mobile picker at the visible bottom edge.
- Add iPhone-sized Storybook states for the assignee and project
pickers.
- Update viewport tests for the corrected coordinate model.

## Verification

- `pnpm --filter @paperclipai/ui exec vitest run
src/components/NewIssueDialog.test.tsx
src/components/InlineEntitySelector.test.tsx`
- `pnpm check:token-gates`
- Playwright used the iPhone 14 device profile. Both picker boxes were
inside the 390 by 664 CSS viewport.
- The assignee picker box was `x=16, y=366, width=358, height=282`.
- The project picker box was `x=16, y=366, width=358, height=282`.

## Risks

- Low risk. The CSS only changes narrow mobile viewports.
- The full-screen popper wrapper does not receive pointer events. The
picker still receives pointer events.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. See `CONTRIBUTING.md`.

## Model Used

- OpenAI GPT-5.6, Codex agent with tool use and code execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal task
ID
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant Storybook documentation to reflect my
changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-27 11:02:15 +00:00
Devin FoleyandPaperclip be43c23e2d fix(recovery): preserve pending native result finalization (#14219)
Preserve native coordinator ownership while accepted results await workspace
copy-back, assessment, or arbitration. Check that ownership in the terminal
update so a coordinator recorded after the liveness read is also protected.
Keep terminal task authority and exhausted-retry cleanup unchanged.

Validation: 28 focused recovery tests, local typecheck/build, 52 passing CI
checks, and Greptile 5/5 with no open findings.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
canary/v2026.927.0-canary.4 nightly/v2026.927.0-nightly.0
2026-09-27 00:22:49 -07:00
Devin FoleyandPaperclip 5ee9e751fb fix(runner): discover assigned tools when direct catalogs exceed limits (#14218)
Keep assigned app tools accessible when the combined Runner catalog exceeds
its operation or byte limits. Reserve task and completion tools, then expose
bounded discovery and call tools for large catalogs. Fetch oversized schemas
in reauthorized chunks without blocking later search results.

Retain task ownership, work-mode restrictions, pinned assignments, current
gateway authorization, approvals, and audit. Small catalogs stay direct.

Validation: 46 focused server tests, two Runner capacity tests, local
typecheck/build, 52 passing CI checks, and Greptile 5/5 with no open findings.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
canary/v2026.927.0-canary.3
2026-09-27 00:22:16 -07:00
DottaandPaperclip 640dee1802 fix(runner): acquire reusable leases when switching to warm sessions (#14187)
## Thinking Path
> - Paperclip coordinates agent work across task turns.
> - Warm runners need reusable sandbox leases.
> - Lease acquisition read only the environment reuse setting.
> - Startup then read the agent's new warm setting and rejected the
ephemeral lease.
> - This fix requests reuse before acquisition, with provider checks
intact.

## Linked Issues or Issue Description
**What happened?**
Switching an existing native runner task from per-turn to warm fails
with `runner_warm_environment_requires_reusable_lease`. Environment
`reuseLease` defaults to false. Acquisition therefore creates an
ephemeral lease, and capability narrowing correctly denies reuse.

**Expected behavior**
Changing lifecycle between turns requests a compatible lease without
replacing the task workspace or changing the shared environment.

**Steps to reproduce**
1. Start a native runner task with inherited lifecycle and environment
reuse disabled.
2. Change the agent from per-turn to warm.
3. Continue the task on a reusable-capable sandbox provider.

**Paperclip version or commit**
Reproduced on `a6c4e7a`. Related: #12904 established warm workspace
continuity; this fixes the earlier lease-selection mismatch.

## What Changed
- Pass agent settings into acquisition and derive a run-scoped reuse
request.
- Respect explicit environment lifecycle overrides; leave other adapters
unchanged.
- Keep provider capability, ownership, cleanup and restore gates intact.
- Add orchestration and database-backed transition tests; document the
contract.

## Verification
- Red: three new assertions failed before the fix.
- Green: 156 tests pass in environment-run-orchestrator,
environment-runtime, native-sandbox-lifecycle, and
environment-execution-target-capabilities.
- Database-backed regression proves ephemeral → reusable → resumed
lease, same workspace, and unchanged stored environment. Provider RPCs
are mocked, not live Daytona.
- `git diff --check` passes.
- Latest-head CI passes typecheck, build, tests, runner checks, E2E, and
canary dry run. Greptile: 5/5; no open threads.
- Recovery uses the persisted lifecycle. Direct regression: red before,
six cases green after. CI runs the added Vitest cases.
- Full local checks were unavailable (missing dependencies and earlier
memory limits).

## Risks
Warm mode now requests retained resources despite an environment's
default `reuseLease:false`. Existing idle timeout and cleanup still
apply. Unsupported providers remain denied. No migration, credential
[REDACTED], or shared-environment mutation.

## Model Used
OpenAI Codex agent; reasoning, code editing and test execution. Exact
model ID and context size were not exposed to this run.

## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details available)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links
- [x] My branch name describes the change and contains no internal
Paperclip ticket [REDACTED] or instance-derived details
- [x] I have run targeted tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-26 21:23:33 -05:00
DottaandPaperclip f2ed0b65c4 fix(runner): enable API tools by default (#14186)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The native runner exposes tools for company tasks.
> - API search and call tools cover operations without a dedicated tool.
> - The current default hides these tools unless an operator sets an
environment variable.
> - This pull request enables the tools when that variable is absent.
> - Operators can still disable the tools or restrict them to selected
companies.

## Linked Issues or Issue Description

Refs #13003, which added the guarded API tools.

**What happened?**
The native runner does not advertise `search_api` or `call_api` with the
default server configuration.

**Expected behavior**
The tools are available without a special environment variable. Existing
authorization checks still apply.

**Steps to reproduce**
Remove `PAPERCLIP_RUNNER_API_TOOLS_ENABLED` and
`PAPERCLIP_RUNNER_API_TOOLS_COMPANY_IDS`. Create a normal runner
authority. Inspect its tool definitions.

**Paperclip version or commit**
a6c4e7a.

## What Changed

- Enable API tools when the server flag is absent.
- Keep explicit disable, invalid-value rejection, company restrictions,
and binding restrictions.
- Test default tool definitions and run HTTP integration tests without
the enabling flag.
- Update operator and hiring documentation. The existing shared gate
also controls `hire_agent`.
- Use the same policy in the E2E evidence summary so an unset flag is
not reported as disabled.
- Isolate default-availability tests from operator environment
variables.

## Verification

- Red test: two new rollout policy assertions failed before the fix.
- Focused policy, authority, and HTTP tests: 3 files passed; 41 tests
passed, 2 skipped. The two runnerd transport cases require a Rust-built
binary absent from this workspace.
- Command: `pnpm exec vitest run
server/src/services/native-runtime/runner-api-rollout.test.ts
server/src/services/native-runtime/paperclip-runner-tool-authority.test.ts
server/src/services/native-runtime/runner-api.integration.test.ts`.
- Attempted `pnpm -r typecheck`: blocked by missing `cargo` in this
workspace.
- Attempted `pnpm build`: terminated at the 4 GiB memory limit.
- Attempted `pnpm test:run`: stopped after memory pressure to run
focused tests alone.
- Repeated the 41 passing focused tests with an inherited disabled flag
and a foreign-company restriction; test isolation passed.
- `pnpm test:e2e:runner:unit`: 44 files and 543 tests passed.
- `pnpm test:e2e:runner:typecheck` exceeded the workspace memory limit,
including a retry with bounded Go memory settings.
- CI results will be recorded before handoff.

## Risks

- More native runs can discover API tools and the existing `hire_agent`
tool by default.
- The change does not remove company, run, mode, credential, lifecycle,
or approval checks.
- Explicit operator restrictions still take precedence. No database
migration is required.

## Model Used

OpenAI Codex agent. The runtime does not expose the exact model ID or
context-window size. Used reasoning, repository tools, shell execution,
and tests.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.927.0-canary.2
2026-09-26 21:16:33 -05:00
a6c4e7a8d1 fix(heartbeat): cancel obsolete execution continuations before dispatch (#13761)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - A comment can queue execution before the task changes state.
> - The task may be completed, cancelled, deleted, or reassigned before
setup checks ownership.
> - The ownership guard must prevent that obsolete execution from
starting.
> - This pull request records that guard outcome as cancellation and
settles the wake request.
> - Missing history and authorization errors remain failures that
operators can investigate.

## Linked Issues or Issue Description

**What happened?**

A comment can queue a run just before a user marks its task Done. The
ownership guard stops setup before adapter dispatch, but the run becomes
a setup failure and can put the agent into an error state.

**Expected behavior**

Cancel obsolete work when the typed task ownership guard rejects it.
Keep genuine setup errors visible. Do not change the task's terminal
state or current owner.

**Steps to reproduce**

Queue a task continuation, then mark the task Done or Cancelled before
continuation setup. The run should settle as cancelled without invoking
the adapter. Deleting or reassigning the task must also prevent
dispatch. Missing source history or explicit user authorization must
still fail.

**Paperclip version or commit**

Refreshed against master `b2e9e82f053af14777fd68e8834fff6bb64c84c4`.
Related: #13546 covers lost issue-lock claims, and #13888 covers
persisted continuation decisions and retry budgets. This change handles
the final continuation ownership guard during setup.

Thanks to @MrBlackTongue for the original typed cancellation
implementation and regression coverage. This update preserves the
contributor's commits, resolves the master conflict, and narrows
cancellation to task ownership invalidation.

## What Changed

- Add a typed `StaleExecutionContinuationError` for
`continuation_task_ownership_changed`.
- Use existing cancellation settlement for that typed guard, including
the run, wake request, issue execution ownership, and agent state.
Suppress immediate recovery of obsolete work.
- Preserve failure classification for missing source context, missing
user authorization, and untyped errors, including an untyped error with
identical text.
- Cover Done and Cancelled tasks with real database checks. Retain
company and assignment guard coverage. Verify cancellation performs no
adapter dispatch or automatic replay.
- Document the cancellation event and its error classification in the
run-log guide.

## Verification

Current head: `b875486e81`.

- `pnpm -r typecheck`: passed.
- `pnpm build`: passed.
- `pnpm exec vitest run
server/src/__tests__/heartbeat-process-recovery.test.ts
server/src/services/execution-continuation.test.ts --maxWorkers=1
--no-file-parallelism -t 'authorized continuation context|obsolete
continuation setup|untyped continuation setup failures'`: 17 passed; 316
unrelated tests excluded by the name filter.
- `pnpm test:run`: local full run is still finishing. It is not a clean
pass: the observed failures are three company-skills cache checks
(`EACCES` while renaming read-only cache directories on macOS), one
tool-access assertion, and one comment-wake timeout. The three cache
failures reproduce in isolation in unchanged code. The tool-access check
passes in isolation, and the entire comment-wake file passes (29 tests),
including explicit feedback after completion. Final full-run totals will
be added when available.
-
[CI](https://github.com/paperclipai/paperclip/actions/runs/36282532652):
all gates green on this head. An unchanged Runner workspace-diff test
initially returned no diff; an isolated check passed, and the failed job
plus its dependent gate passed on retry with no code change. Original
failed job: 108517010989.
- Greptile: 5/5 on the full current head, with no unresolved review
threads. Branch is current with master and conflict-free.
- Diff check and local secret/PII scan passed. No customer data or
private deployment identifiers are included.

## Risks

The classification change is limited to the typed ownership guard. A
task with missing history remains a failure rather than being treated as
an expected cancellation. Existing company, task-owner, terminal-state,
and authorization checks still prevent dispatch. Cancellation retains
its specific reason in the run log and does not grant a retry. No
schema, dependency, or API changes; no migration is required.

## Model Used

OpenAI GPT-6 via Codex assisted investigation, implementation, code
review, and test execution. The exact serving model identifier and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have described the issue in-PR following the bug issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links
- [x] My branch name describes the change and contains no internal
ticket id
- [x] I have run the focused regression tests locally and they pass;
full-suite results are tracked above
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation
- [x] I have considered and documented the risks
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Devin Foley <devin@paperclip.ing>
canary/v2026.927.0-canary.1
2026-09-26 17:43:33 -07:00
Devin FoleyandPaperclip b2e9e82f05 fix: stop remote Grok runs before continuing queued messages (#14100)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The execution service owns each run and saves messages sent while it
runs.
> - Interrupt must stop the current executor before it delivers those
messages.
> - Remote Grok commands did not register the host cancellation control.
> - A cancelled task run could still write Done and prevent queue
recovery.
> - This pull request connects remote cancellation and revokes cancelled
run writes.
> - Saved input can use the existing queue admission rules after
verified cleanup.

## Linked Issues or Issue Description

**What happened?**

Interrupting a queued message marked a remote Grok run cancelled before
its sandbox stopped. The old run could still post a reply and mark the
task Done. Its saved follow-up remained deferred behind execution
recovery.

**Expected behavior**

Stop revokes run write authority and waits for verified termination.
Saved messages remain durable and enter one successor through normal
admission after cleanup.

**Steps to reproduce**

1. Run a task with `grok_local` in a remote sandbox.
2. Send a follow-up and use Interrupt while the command runs.
3. Let the old command attempt a task status update after cancellation.
4. Observe the task disposition and the saved message queue.

**Paperclip version or commit**

The gap is present in master at `d3e0f0a238`.

**Deployment mode**

Authenticated server with a Daytona sandbox.

Related work: #14028 and #14046 handle bounded continuation. #13291
covers infrastructure interruption and verified remote cleanup. #13332
addresses atomic recovery holds. This change handles direct Grok
operator cancellation and stale task writes.

## What Changed

- Register remote Grok cancellation before preparation. Keep command
ownership until the host confirms sandbox termination.
- Reuse the sandbox cancellation boundary for the direct CLI invocation.
Reject fresh attempts after cancellation and preserve workspace restore
failure evidence.
- Reject writes from cancelled task JWTs and runs with a pending stop.
Preserve diagnostic reads and existing conversation error codes.
- Recheck run authority under a database lock before task updates and
interaction responses commit.
- Preserve authorized handoffs that stop their own run. Only the
server-issued stop receipt for that request permits the final task
update.
- Add tests for hung commands, unverified stops, early cancellation,
copy-back failures, late Done, late interaction responses, authorized
handoffs, exact lease receipts, and one queue successor across
concurrent restart sweeps.
- Document the cancellation and write-authority contract.

## Verification

- Targeted adapter, cancellation-boundary, authentication,
queued-message, interaction-service, and activity-route tests passed.
The expanded run passed 214 tests; one new test had an incomplete
fixture. After correcting the fixture, all 8 selected follow-up cases
passed.
- `pnpm -r typecheck`: passed on
`179c86caf1bf0d89914a503d46e24af7e4b8c557`.
- `pnpm build`: passed on the same commit.
- `pnpm test:run`: the general-server group completed with 13,521
passed, 99 skipped, and 18 failed tests. It then stopped, so the
remaining local groups did not run. Five Slack, email, and wake-batching
failures passed on focused reruns after correcting the local
environment. The remaining 13 failures reproduce as `EACCES` on rename
in unchanged skill-cache code on macOS. Two custom-image suite setup
hooks also failed to start embedded PostgreSQL after the machine
exhausted shared-memory slots; all 31 tests in that file passed on rerun
after the local resource issue was resolved. CI covers all test groups.
- CI: 53 checks passed and 2 were skipped on the latest commit,
including the aggregate verification gate. The last server shard passed
on its single rerun after a preview-server startup timeout. The affected
file also passed locally with 28 passed and 3 skipped.
- Greptile: 5/5 on the latest commit. Both review threads are resolved.
- No live deployment or staging task mutation has been performed.

## Risks

- Stopping the sandbox can prevent file copy-back. The result preserves
workspace restore failure evidence; termination does not imply restored
files.
- If provider termination fails, the adapter keeps ownership of its
outstanding command and does not acknowledge Stop.
- The write restriction now applies to ordinary cancelled tasks. Reads
remain allowed. Task and interaction checks add a shared run-row lock to
agent mutations. An exact server-issued receipt permits the task request
that stopped its own run to complete its handoff.
- Existing terminal tasks are not reopened automatically. An operator
must correct a historical late Done before its saved queue can continue.
- No schema migration or UI change.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository inspection, code
execution, and test tools. The precise backend revision and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.927.0-canary.0
2026-09-26 17:07:07 -07:00
DottaandCodie 01d9a12185 fix: make keyboard shortcut enablement a personal preference (#14141)
Store keyboard shortcut enablement per user and expose it in Profile settings.

Co-Authored-By: Codie <Codie@users.noreply.github.com>
2026-09-26 11:50:04 -05:00
Devin FoleyandPaperclip d3e0f0a238 fix(server): validate CLI auth challenge IDs (#14095)
Reject malformed CLI challenge UUIDs before database access while preserving secret and authentication precedence. Document the HTTP responses and pin the supervisor test fixture to the CI-selected Node executable.

Validated by 21 focused tests, root typecheck/build, and full CI. Mac broad-suite baseline limitations are documented in the PR. Greptile 5/5.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
canary/v2026.926.0-canary.4 nightly/v2026.926.0-nightly.0
2026-09-26 00:48:14 -07:00
Devin FoleyandPaperclip 1e3a148f46 fix(connections): retain safe broker rejection diagnostics (#14098)
Preserve fixed broker rejection reason codes within strict size/time limits while retaining public error codes and status. Unknown bodies remain generic; no raw response or credential material enters the error.

Validated by 33 consumer tests, a synthetic producer HTTP contract fixture, root typecheck/build, and full CI. Greptile 5/5.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-26 00:47:47 -07:00
Devin FoleyandPaperclip ffa32373bc fix(runtime): honor provider acquisition timeout defaults (#14097)
Declare provider acquisition budgets so slow Daytona creation does not hit the host’s 30-second fallback. Bound creation and setup to one deadline and preserve scoped cleanup ownership after timeout. Legacy drivers retain their original call shape.

Validated by 238 provider/manifest tests, focused database and heartbeat regressions, root and standalone provider typecheck/build, and full CI. Local broad tests also expose recorded Mac baseline limitations. Greptile 5/5.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
canary/v2026.926.0-canary.3
2026-09-26 00:47:21 -07:00
Devin FoleyandPaperclip 7f3c06dac4 refactor(ui): remove the legacy Cloud organization switcher (#14061)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Its sidebar provides navigation between organizations.
> - The generic organization switcher slot lets installed plugins own
that navigation.
> - Both Core UI shells still contain the old Cloud portfolio menu.
> - Cloud now uses its private Account plugin for this menu.
> - This pull request removes the duplicate Cloud menu and keeps the
built-in company menu.
> - This reduces Cloud-specific code without changing the plugin host
contract.

## Linked Issues or Issue Description

Refs #13832 and #13854. Related #14060 changes the account popup, not
this organization switcher.

**What existing behavior does this improve?**

Organization navigation in both sidebar shells.

**Current behavior**

Core retains Cloud portfolio fetching, stack rows and stack-entry links
behind the plugin replacement slot.

**Proposed behavior**

The installed switcher plugin owns Cloud navigation. Core lists local
companies when no usable replacement exists. The Members-page Cloud
invitation action keeps its existing portfolio API; it is a live caller,
not an old-image fallback.

## What Changed

- Remove Cloud portfolio queries, stack rendering and Cloud
creation/entry branches from both built-in menus.
- Remove the unused stack-entry URL helper.
- Keep company selection, ordering, invitations, logout and plugin error
handling. Hide local company creation on managed hosts, where the server
forbids it.
- Update switcher tests and the navigation contract.

## Verification

- `pnpm -r typecheck` passed, including Rust checks with the installed
Cargo toolchain on PATH.
- `pnpm exec vitest run --project @paperclipai/ui`: 637 files and 6,727
tests passed.
- Focused switcher, plugin host and Cloud link tests: 26 passed after
the managed-host creation guard.
- `pnpm check:token-gates` passed.
- Full `pnpm test:run` was attempted, then stopped after failures in
unchanged server tests. Targeted reproduction found an ancestor
skills-directory collision for Slack and macOS EACCES errors renaming
the company skills cache. Other local failures appeared in email
connector skill setup and a process-turn test. This is not a local
full-suite pass; clean Linux CI covers the complete suite.
- Full `pnpm build`, UI production build and Storybook build passed. All
[latest-head CI
checks](https://github.com/paperclipai/paperclip/actions/runs/36201814444)
passed, including the full Linux test matrix, browser tests, typecheck,
build and release checks. Greptile is 5/5 with no unresolved comments.

## Risks

- A managed host without a usable switcher plugin now gets the ordinary
company menu. It no longer gets the old Cloud portfolio menu, and local
company creation remains unavailable there.
- The Cloud portfolio endpoint remains required by the Members-page
invitation action. This PR does not remove that live endpoint or change
its authorization.
- No database, authentication, plugin protocol or deployment changes.

## Model Used

OpenAI Codex, GPT-6, with reasoning, repository inspection and code
execution. The exact deployment identifier and context-window size are
not exposed by this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.926.0-canary.2
2026-09-25 22:58:52 -07:00
Devin FoleyandPaperclip 4ca404b49a fix: record safe sandbox restore failure diagnostics (#14064)
Record a bounded diagnostic for failed workspace and staged-asset restores.
Preserve the original error, retry policy, and archive safety checks. Never
copy raw provider messages, credentials, paths, or asset names into the log.
Nested failures log once; safe fields survive throwing property getters.

Verified 129 focused restore/Claude tests, typecheck/build, and green full
PR CI. Greptile 5/5 with all review threads resolved.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
canary/v2026.926.0-canary.1
2026-09-25 17:16:13 -07:00
Devin FoleyandPaperclip c3ddb288b1 fix: validate work product execution workspace references (#14063)
Reject invalid and cross-company execution workspace references with a useful
422 before changing the work product. Hold the validated reference through
the transaction so concurrent deletion cannot turn validation into a 500.

Verified 13 focused tests, full typecheck/build, and green PR CI. Greptile
5/5 with no unresolved comments. The known UI copy-toast flake passed on an
unchanged-commit retry and in a focused local run.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
canary/v2026.926.0-canary.0
2026-09-25 17:06:30 -07:00
DottaandPaperclip 96bf004a79 fix: use persisted state for lifecycle continuation and retry budgets (#13888)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Its control plane decides when a task can continue, wait, stop, or
complete.
> - Legacy continuation could change when an agent changed its wording
without changing task state.
> - Shared attempt counts also let repair and infrastructure retries
affect each other's limits.
> - This pull request uses persisted state and separate, bounded
allowances for these decisions.
> - If automatic repair stops, the task explains what happened and
offers a guarded retry.
> - Paired tests and real-provider evaluations verify that Stop,
approvals, ownership, and spending limits remain authoritative.

## Linked Issues or Issue Description

Related work: Refs #13761, Refs #11126, Refs #13610. These cover
obsolete continuation dispatch and retry storms. Open and closed issues
and PRs were searched for related lifecycle, continuation, and retry
work.

**What happened?**
Legacy continuation depended on English wording and progress heuristics.
Repair, failure retry, and productive continuation could consume shared
counts. When bounded repair stopped, the task showed a technical
recovery message without a clear next action.

**Expected behavior**
Persisted disposition and owned execution paths determine the next
action. Missing disposition prompts bounded agent repair. Explicit work
mode determines planning mode. Narrative changes and raw activity counts
cannot replenish allowances. An exhausted repair shows a readable
notice. An explicit retry checks current controls and preserves the
assigned agent.

**Steps to reproduce**
Run `pnpm test:lifecycle-baseline`. The paired probes keep structured
state constant while varying completion, planning, blocker, and progress
prose. Run the explicit `lifecycle-baseline` and
`continuation-accounting` Product E2E suites for real-provider coverage.
In Storybook, open **Design previews / Recovery notice** to inspect the
production component's normal, pending, acknowledged, unavailable,
failure, and mobile states.

## What Changed

- Hide the image attachment button, icon, and drop/paste hint in answer
composers. Image paste and drop support remains available.
- Merge current master and retain both browser regression sets. Use a
production-stamped service worker in the offline recovery browser
fixture.
- Share one state-based legacy continuation decision across immediate,
delayed, and recovered dispatch. Bind bounded repairs to their source
run and episode.
- Remove title and description wording from work-mode authority. Agents
can still write requested plans in execution mode.
- Persist separate failure-retry and productive-continuation counters.
Disposition repair and resource waits cannot consume or reset those
allowances.
- Validate delayed repair identity, then recheck current gates before
provider dispatch. Fence native startup cancellation.
- Show **Agent needs attention**, a plain-language explanation, **Retry
agent**, and expandable details in both task interfaces. Report request
progress, acknowledgement, and errors inline.
- Store typed recovery notice metadata. Recognize older active notices
only through exact stored action and run IDs. Notice text never grants
retry authority.
- Use the existing recovery-action endpoint for retry. Recheck current
action, status, owner, agent availability, dependencies, active runs,
pending questions and confirmations, approvals, pause controls, and
budget. Duplicate requests do not wake twice.
- Add component, page, route, database, contract, and Storybook
coverage. Keep the scenario inventory and executable evals here.
Historical reports and snapshots live in the [commit-pinned
paperclip-evals
archive](https://github.com/paperclipai/paperclip-evals/blob/ce3e5afcd4a1184650f586a2b5b8be5874c66c8b/experiments/2026-09-lifecycle-authority/README.md).
- Preserve unsaved project fields while the same project URL changes to
its canonical alias. Do not reuse data across projects or companies.
This separate fix addresses the repeated repository-editor browser
failure without changing the browser test.
- Keep the development service worker from intercepting Vite module
reloads. Update the connection-intent browser fixture to record progress
and completion through the agent API.

## Verification

Merge preparation on September 25, commit
`c1e8e4b7ddd9fbc4913ed55ce21b8e12906c2f97`:

- Merged master `bd2030932` and resolved the browser test-list conflict
by keeping both sets of regressions.
- Deterministic lifecycle baseline: 1,090/1,090 assertions passed; no
failures, skips, or missing selected evidence. Unit 423, runner 184,
database integration 397, grading 86.
- Browser support: 17/17 passed. The offline recovery test first failed
with an unstamped development worker, then passed with the production
stamp. Its assertions are unchanged.
- Focused interaction UI and offline fallback tests: 19/19 passed.
Verified the custom-answer composer in Storybook: no attachment controls
or hint; entering an answer enables Next.
- Recursive typecheck, production build, token gates, and diff checks
passed. The worktree is clean. No new real-provider campaign was run.
- Current CI and review: [Current PR CI
passed](https://github.com/paperclipai/paperclip/actions/runs/36166011243):
55 successful checks and two optional Storybook skips. Greptile scored
this exact commit 5/5. Hiding the question attachment controls is an
intentional UI change; paste/drop remains available.

Earlier recovery UI verification, commit
`21be0fec0e90e86b6d662b8ee4831847cd041cdb`:

- Recursive typecheck, production build, token gates, and diff checks
passed.
- Focused UI coverage: 338 tests passed across six suites (336 before
the interaction guard, with the two affected suites rerun at 149 passed
after it). Covers both task interfaces, the real page mutation,
pending/error acknowledgement, stale state, and unavailable controls.
- Recovery database integration: 352 tests passed before the interaction
guard. The complete recovery-action and mutation-route suites passed 181
tests after it. The two new pending question/confirmation regressions
failed before the fix and passed afterward, including
resolved-interaction controls. Shared validator suite: 31 passed. E2E
catalog suites: 34 passed.
- Browser inspection passed for light/dark themes, mobile layout,
expandable details, pending retry, acknowledgement, failure, and
disabled retry. Storybook renders the production component; its request
is simulated.
- The broad local run hit two chat callback-order wait failures and was
stopped after all CI unit/database/runner shards passed. Both local
failures passed when rerun without the competing full-suite process.
- CI exposed a repeated project-repository draft-loss race during
canonical redirects. A new unit regression failed before the fix; all
nine project-page tests now pass, including controls for other projects
and companies. Both unchanged repository browser tests passed against a
fresh local server. UI typecheck, production UI build, and token gates
passed after this fix.
- [Earlier PR CI
passed](https://github.com/paperclipai/paperclip/actions/runs/36072486798)
on `21be0fec0e90e86b6d662b8ee4831847cd041cdb`: 55 successful checks, two
optional Storybook skips, and no failed or pending checks. The
repository browser shard passed with the production fix. Greptile is 5/5
on this exact commit with no unresolved review threads. The PR is
mergeable.

Historical, source-qualified lifecycle evidence:

- Lifecycle baseline: 1,074 assertions. Native session coverage: 447
tests. Product E2E support: 515 tests. Browser support: 11 tests. Full
earlier verification is retained in the archive.
- [Real-provider campaign: 8/8 passed, zero
retries](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35881382080-1/index.html),
source `e88d210417280140b44a36449027290adcb1aeaa`. Evidence and cleanup
checks passed. This includes deliberately exhausted repair cases that
correctly remain blocked; it does not mean every task finished Done.
This campaign predates the recovery UI change.
- Archive migration verified all 16 original JSON files byte-for-byte
and all 24 checksum entries. App tests do not need private archive
access. [Archive PR
#27](https://github.com/paperclipai/paperclip-evals/pull/27) is merged.

## Risks

- Agents that omit durable disposition receive at most two repair
attempts by default. Prose-only completion exposes missing state rather
than silently changing scheduling.
- A retry is an explicit board action. The server rechecks current
controls. A successful response confirms the task returned to To do; it
does not claim that the provider has already started.
- Existing notice metadata remains valid. Only older active notices with
matching structured evidence receive the new UI. Historical notices
without that evidence keep their existing rendering. No schema migration
is required.
- Old run records require conservative retry accounting. Tests cover old
counters, alternating retry lanes, restarts, and exhausted repairs.
- Historical snapshots require private `paperclip-evals` access. The app
index retains public campaign links. Live campaigns qualify specific
sources and scenarios; no new real-provider campaign has run for the
recovery UI commit.

> This fixes existing lifecycle and recovery behavior and does not
duplicate planned core work.

## Model Used

OpenAI GPT-6 through Codex assisted implementation, reasoning, code
execution, and review. The exact serving model ID and context window are
not exposed in this task. Historical real-provider evaluations used
Codex model `gpt-5.6-sol`, separately from the implementation assistant.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.925.0-canary.6
2026-09-25 15:28:11 -07:00
Devin FoleyandPaperclip 6bc830b62b fix(recovery): escalate an issue whose interrupted run has spent its retry budget (#14046)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The stranded-issue sweeper is what puts an assigned issue back on a
live path when its run dies.
> - A failed or interrupted run gets a bounded transient retry; when
that budget is spent, the retry scheduler queues nothing and reports the
exhaustion.
> - The sweeper treated that "nothing queued" like every other "nothing
queued" and skipped the issue, on every tick, forever: `in_progress`, no
run, no path, no notice.
> - Three server restarts in a row (a deploy storm) are enough to spend
the budget, so this is reachable in ordinary operation.
> - This pull request makes the sweeper escalate a spent budget to a
board-owned recovery action, the same visible `blocked` state its other
dead ends use.
> - The benefit is that no issue can sit assigned and silent after its
retries run out.

## Linked Issues or Issue Description

No existing issue found. Related work: #14028 (this branch's
predecessor: queue drain-time wakes, retry interrupted corrective runs)
fixed two neighbouring gaps but not this one.

**What happened?**

A routine-created issue's run was interrupted by a graceful server
shutdown, and both bounded transient retries were interrupted by the
next two shutdowns. The retry scheduler logged "Bounded retry exhausted
after 2 scheduled attempts; no further automatic retry will be queued"
and stopped. On every tick after that, `reconcileStrandedAssignedIssues`
reached the generic continuation lane, `enqueueStrandedIssueRecovery`
delegated to `scheduleRecoveryRetry`, which returned null (exhausted),
and the sweeper counted the issue as `skipped`. The issue stayed
`in_progress` with no run, no recovery action, and no comment for hours
until a person woke the agent by hand. The same hole exists in the
assigned-`todo` dispatch lane.

**Expected behavior**

When the transient retry budget for an interrupted or failed run is
spent, the sweeper should treat it as the dead end it is: escalate the
issue to `blocked` with a board-owned recovery action and a notice,
exactly as it does when a continuation retry chain or an assignment
retry chain is exhausted. Leaving the budget spent across restarts is
correct and unchanged; leaving the issue silent is not.

**Steps to reproduce**

1. Assign an agent using a conversation adapter (its interrupted runs
carry `conversationContinuation`, so legacy reconciliation does not
terminalize them) an issue and let a run start.
2. Interrupt the run with a graceful shutdown, let the transient retry
start, interrupt it, and repeat once more so `scheduledRetryAttempt`
reaches 2.
3. Run the stranded-issue sweep. Before this change: `skipped` every
tick, issue `in_progress`, no run, no recovery action. After:
`escalated`, issue `blocked`, one active board-owned
`issue_recovery_actions` row, a "No live execution path" notice.

**Paperclip version or commit**

master at bd6caf51bb (2026-09-25).

**Deployment mode**

Managed cloud instance restarted by fleet deploys; the code path is the
same for any operator whose server restarts more often than the retry
budget allows.

## What Changed

- `recoveryService` gains an optional `transientRetryBudgetSpent(run)`
dependency; `heartbeatService` wires it as
`executionFailureRetryCount(run) >=
BOUNDED_TRANSIENT_HEARTBEAT_RETRY_MAX_ATTEMPTS`, the same check
`scheduleBoundedRetryForRun` applies.
- `enqueueStrandedIssueRecovery` takes an optional `outcome`
out-parameter and sets `retryExhausted` when the failed predecessor's
retry returned nothing **because** the budget is spent. A null return
without it still means another authority owns the run (native runtime,
legacy reconciliation) and the caller leaves it alone; the deliberate
"failure recovery cannot fall through into the continuation queue" rule
is unchanged.
- The generic `in_progress` continuation lane and the assigned-`todo`
dispatch lane escalate on `retryExhausted` via
`escalateStrandedAssignedIssue` with a "No live execution path" notice
(danger tone), which creates the board-owned source-scoped recovery
action and moves the issue to `blocked`. Every other
`enqueueStrandedIssueRecovery` caller is unchanged.
- Tests (`heartbeat-process-recovery.test.ts`): an `in_progress` issue
with a spent budget escalates (blocked, one active board-owned action,
notice, no successor run, idempotent on the next sweep); an interrupted
run with budget remaining still gets its transient retry; an assigned
`todo` issue with a spent dispatch budget escalates. The existing guard
"does not reset an exhausted incident budget on server restart" keeps
its no-successor-run assertion and now expects the board escalation
instead of nothing, with a comment on why.

## Verification

```
pnpm -r --filter './packages/**' build
cd server
npx vitest run src/__tests__/heartbeat-process-recovery.test.ts \
  src/__tests__/issue-recovery-actions.test.ts \
  src/services/recovery/successful-run-handoff.test.ts \
  src/__tests__/heartbeat-task-drain-admission-release.test.ts \
  src/__tests__/attention-service.test.ts \
  src/__tests__/heartbeat-comment-wake-batching.test.ts
npx tsc --noEmit -p tsconfig.json
```

Live reproduction: the exact stuck state (issue `in_progress`, latest
run `interrupted` with `scheduledRetryAttempt` 2, the "Bounded retry
exhausted" lifecycle event, sweeper `skipped` every tick) was observed
on a managed instance running current master before this change was
written.

## Risks

- Behavior change is limited to runs whose transient budget is already
spent, which previously produced no action at all. Nothing new is
retried; the change only adds the escalation, so no retry loop can be
introduced.
- The board-owned action spawns no run. Resolving it (restore to the
owner) re-dispatches through the existing recovery-action routes, the
same flow as every other stranded escalation.
- Native-runtime and legacy-reconciliation predecessors are untouched:
they return before the exhaustion check.

## Model Used

Claude (Anthropic) — `claude-fable-5-1`, extended thinking, tool use
(Claude Code CLI). The incident diagnosis and the choice to escalate
rather than re-dispatch were steered by the maintainer.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.925.0-canary.5
2026-09-25 13:17:48 -07:00
Devin FoleyandPaperclip bd6caf51bb fix: preserve restore failure results and stop unsafe retries (#14035)
Preserve agent output and earlier execution errors when workspace restore fails. Report the restore phase and confirmed saved-plan links. Require verified repair before retrying unsafe archives, while preserving approval states and the retry budget.

Verified with full CI, 506 focused regression tests, and Greptile 5/5 with all review threads resolved.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
canary/v2026.925.0-canary.4
2026-09-25 11:34:31 -07:00
Devin FoleyandPaperclip 5b09d66183 fix(heartbeat): keep drain-time wakes queued and retry interrupted disposition handoffs (#14028)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The server's heartbeat scheduler queues, admits, and recovers agent
runs.
> - A control plane arms an instance task drain before it restarts the
process, so no new run starts mid-restart.
> - Two recovery paths treated that transient hold as a permanent
verdict: a wake that arrived during a drain was written as `skipped` and
never replayed, and a corrective disposition run that the restart
interrupted counted as an exhausted attempt, so the issue went to
`blocked`.
> - Both outcomes leave an assigned issue with no run and no path until
a person notices.
> - This pull request keeps drain-time wakes in the durable queue and
gives an interrupted corrective run the same bounded transient retry any
interrupted run gets.
> - The benefit is that a graceful restart never strands or blocks an
issue by itself.

## Linked Issues or Issue Description

No existing issue found. I searched open PRs that touch the task drain
(#13527 adds opt-in termination of active runs; #13965 handles deferred
wakes after a stale cancellation). Neither replays drain-time wakes or
changes the corrective-handoff escalation.

**What happened?**

1. An operator accepted a plan confirmation while the instance was in a
task drain (a control plane armed it before a deploy restart).
`enqueueWakeup` wrote the assignee's continuation wake with `status:
"skipped"` and `heartbeatSkip.reason: "task_drain"`. The drain is
process-local and cleared on the restart seconds later, but nothing
replays skipped wakes. The issue stayed in `todo` with no run for 20
hours until a person retried it by hand. The stranded-issue sweeper did
not help: a `todo` issue whose latest run succeeded is treated as
deliberately handed back.
2. On another issue, the watchdog correctly raised
`successful_run_missing_state` and queued the single corrective handoff
run. A graceful server shutdown (the same deploy restart) interrupted
that run with `server_shutdown_interrupted`. On the next boot,
`reconcileStrandedAssignedIssues` saw a corrective run at
`handoffAttempt >= maxHandoffAttempts` and escalated the issue to
`blocked` with a board-owned `missing_disposition` recovery action. The
agent never got to finish one corrective attempt.

**Expected behavior**

A task drain holds admission, not the request. A wake that arrives
during a drain must run once the drain lifts or the process restarts. An
interrupted corrective run is not evidence that the agent could not
choose a disposition; it should be retried like any interrupted run, and
escalate only when that retry budget is spent or a finished attempt
still leaves no disposition.

**Steps to reproduce**

1. Assign an agent an issue and `POST /api/instance/task-drain` with a
Cloud control assertion (or call `startTaskDrain({})` in a test).
2. Trigger any wake for that agent (assignment, comment, or an accepted
`request_confirmation`).
3. Observe the `agent_wakeup_requests` row: `status = skipped`, `reason
= heartbeat.scheduling_suppressed`, `payload.heartbeatSkip.reason =
task_drain`. Stop the drain or restart: no run is ever created for that
wake.
4. For the second case: let a `finish_successful_run_handoff` run be
interrupted by SIGTERM during a graceful shutdown, then boot. The issue
moves to `blocked` on the startup sweep without any retry.

**Paperclip version or commit**

master at e2f1a66aa7 (2026-09-25).

**Deployment mode**

Managed cloud instance behind a control plane that arms task drains
before deploy restarts; the code paths are the same for any operator who
uses the drain endpoint.

## What Changed

- `enqueueWakeup` no longer writes a wake as `skipped` when the only
suppression is `task_drain`. The wake and its queued run land in the
durable queue; `startNextQueuedRunForAgent` and `executeRun` already
refuse to admit work while the drain is active, and the boot-time
`resumeQueuedRuns` pass picks it up after the restart.
`worktree_instance` and `database_restore_in_progress` suppression still
write `skipped`: those holds are not transient restarts.
- `reconcileStrandedAssignedIssues`: when the latest run is a corrective
successful-run handoff at its attempt cap **and** that run is
`interrupted`, the sweeper schedules the bounded transient retry (via
`enqueueStrandedIssueRecovery` → `scheduleRecoveryRetry`) instead of
escalating. The retry keeps the handoff context, so it is still the
corrective run. If the retry budget is spent, or a finished attempt
still leaves no disposition, escalation proceeds exactly as before.
Failed corrective runs (adapter or provider failures) are unchanged:
they still escalate immediately.
- New sweep counter `successfulRunHandoffRetried`, included in the
startup and periodic recovery log lines.
- Tests: `heartbeat-task-drain-admission-release.test.ts` gains a case
that a wake during a drain is queued, held while the drain is active
(quiescence still reports true), and runs to completion once the drain
lifts. `heartbeat-process-recovery.test.ts` gains two cases: an
interrupted corrective run is retried with its handoff context and the
issue stays `in_progress`; an interrupted corrective run with a spent
transient budget escalates to `blocked` with the usual recovery action
evidence.

## Verification

```
pnpm -r --filter './packages/**' build
cd server
npx vitest run src/__tests__/heartbeat-process-recovery.test.ts            # 296 passed
npx vitest run src/__tests__/heartbeat-task-drain-admission-release.test.ts \
  src/__tests__/heartbeat-worktree-suppression.test.ts \
  src/__tests__/heartbeat-scheduling-suppression.test.ts \
  src/__tests__/heartbeat-task-drain.test.ts \
  src/__tests__/instance-settings-routes.test.ts \
  src/services/recovery/successful-run-handoff.test.ts \
  src/__tests__/issue-recovery-actions.test.ts \
  src/__tests__/heartbeat-comment-wake-batching.test.ts \
  src/__tests__/attention-service.test.ts                                    # 226 passed
npx tsc --noEmit -p tsconfig.json                                           # clean
```

Live reproduction of the first defect: after this change was written, a
board wake issued against a draining managed instance (running current
master) was again recorded as `skipped` with `heartbeatSkip.reason =
task_drain`, which is the exact row the new test asserts no longer
appears.

## Risks

- A wake that lands during a drain now waits in the queue instead of
being dropped. `getTaskDrainStatus().pendingWakes` counts only in-flight
enqueue promises, so quiescence is unchanged and a control plane waiting
for the drain is not held longer. If a drain is stopped without a
restart, the queued run starts on the next scheduler tick.
- The interrupted-handoff retry reuses the existing transient retry
budget (two attempts) and the existing `hasActiveExecutionPath` skip, so
a sweep cannot double-schedule. Native-runtime corrective runs still
return `null` from the recovery enqueue and escalate as before.
- Self-hosted instances that never call the drain endpoint see no change
on the first path; the second path only changes behavior for corrective
runs interrupted by a graceful shutdown.

## Model Used

Claude (Anthropic) — `claude-fable-5-1`, extended thinking, tool use
(Claude Code CLI). Human-directed: diagnosis of the two production
incidents, the fix design, and review were steered by the maintainer.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.925.0-canary.3
2026-09-25 11:28:54 -07:00
DottaandPaperclip bd20309323 fix(ui): keep mobile task pickers above the keyboard (#14022)
## Thinking Path

> - Paperclip helps people manage AI agents and their tasks.
> - The New Task dialog lets a person select an assignee and a project.
> - A mobile browser reduces the visible viewport when the software
keyboard opens.
> - The dialog accepted invalid viewport data, and the pickers stayed
inside a transformed container.
> - This behavior could collapse the dialog or move a picker search
field above the visible area.
> - This pull request validates viewport data and puts mobile pickers in
the visible viewport.
> - The benefit is that a mobile user can see and use each picker while
the keyboard is open.

## Linked Issues or Issue Description

**What happened?**

On a mobile device, the Assignee and Project pickers in the New Task
dialog could move above the visible viewport. A short invalid viewport
value could also collapse the dialog to a line.

**Expected behavior**

The dialog and each open picker must stay in the visible viewport while
the software keyboard is open.

**Steps to reproduce**

1. Open the New Task dialog in a mobile browser.
2. Open the Assignee picker or the Project picker.
3. Focus the picker search field so that the software keyboard opens.
4. Observe that the picker can move above the visible viewport.

**Paperclip version or commit**

`efce9356b5`

**Deployment mode**

Local dev and hosted browser UI.

## What Changed

- Ignore zero, negative, and non-finite Visual Viewport measurements.
- Keep the last safe dialog geometry until the browser gives a valid
measurement.
- Put entity pickers outside the transformed dialog container.
- Size and position the mobile picker from the dialog Visual Viewport
values.
- Add unit tests for invalid viewport recovery.
- Add Chromium tests for the Assignee and Project picker states.

## Verification

- `pnpm exec vitest run ui/src/components/NewIssueDialog.test.tsx`
passed with 34 tests.
- The Chromium viewport test passed 35 of 35 runs with five repeats and
no retries.
- `pnpm --filter @paperclipai/ui typecheck` passed.
- `pnpm check:token-gates` passed.
- `pnpm build-storybook` passed.
- `pnpm -r typecheck` passed.
- `pnpm build` passed.
- A full local Vitest run reached restricted workspace-runtime tests
that require sibling worktree and runtime writes. Remote CI will run the
supported test environment.

## Risks

- Risk is low because the new picker layout applies only to mobile
widths.
- The layout depends on Visual Viewport data when the browser supplies
valid values.
- Unit and browser tests cover invalid data, mobile pickers, tablet
layout, and desktop layout.

> This change fixes a focused UI bug. It does not duplicate a core
feature in `ROADMAP.md`.

## Model Used

- OpenAI Codex with GPT-5. The agent used reasoning, repository tools,
code execution, and browser automation. The runtime did not expose the
context window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.925.0-canary.2
2026-09-25 08:37:26 -07:00
Devin FoleyandPaperclip e2f1a66aa7 feat: support default-hidden experimental settings (#13980)
## Thinking Path

> - Paperclip is the open source control plane for AI-agent companies.
> - Operators can hide settings that their users must not change.
> - The server shares the effective restrictions with the UI and
settings API.
> - An explicit list must change whenever a new experimental flag is
added.
> - This pull request adds a wildcard with named exceptions to the
existing setting.
> - Core applies the policy to its current catalog, so new flags stay
hidden automatically.

## Linked Issues or Issue Description

Related: #11823 introduced settings visibility. #13907 added workspace
isolation visibility. I searched related PRs and issues and found no
duplicate wildcard implementation.

**What existing behavior does this improve?**

Operator control of experimental setting visibility through
`PAPERCLIP_HIDDEN_SETTINGS`.

**Subsystem affected**

Shared settings policy and its existing server health and mutation
consumers.

**Current behavior**

Operators must name every hidden experimental toggle. A new Core flag
can become visible until the operator updates that list.

**Proposed behavior**

`instance.experimental.*` hides current and future experimental toggles.
Entries such as `!instance.experimental.enableEnvironments` leave named
controls available. Explicit hidden keys and the hidden parent page take
precedence over exceptions.

**Reason and benefit**

Operators can maintain a short list of allowed controls instead of a
second copy of Core's full feature catalog.

**Breaking changes**

Existing explicit lists and unset configuration keep their behavior. The
new syntax is opt-in. Older images ignore it, so operators must retain
explicit restrictions until those images are upgraded. Visibility does
not change feature values.

## What Changed

- Expand the wildcard into concrete catalog keys in the shared parser.
- Limit exceptions to known experimental controls and preserve explicit
restrictions in either input order.
- Test a synthetic future catalog addition, duplicate and invalid
entries, API rejection, same-value echoes, and the effective health
payload.
- Document the syntax and the transition for deployments with mixed
image versions.

## Verification

- Targeted parser, future-catalog, health, and settings-route tests
pass: 95 tests across four files.
- `pnpm -r typecheck` passes, including Rust checks, with the installed
Cargo directory on PATH.
- `pnpm build` passes.
- All current-head CI gates pass, including the full test shards, Rust,
build, browser E2E, and canary dry run:
https://github.com/paperclipai/paperclip/actions/runs/36083292578. One
unchanged runtime-exposure cold-start test passed on its first retry.
- The full local `pnpm test:run` did not pass on macOS/Node 25: the
first server group reported 13,356 passed, 18 failed, and 99 skipped,
with six failed files (including two failed suite setups). Failures were
in unchanged runtime/company skill cache, chat/email connector fixtures,
embedded-Postgres setup, and workspace cleanup tests. A standalone
filesystem probe reproduced the read-only-directory rename permission
failure. Missing connector fixture paths, database startup failures, and
two integration assertions also occurred; the remaining local groups
were not reached after this group failed. The corresponding CI lanes all
pass. These local failures are not claimed as fixed by this PR.
- Browser suites were not run because this changes the shared policy,
not UI rendering or browser workflows. Health payload and route tests
cover the shared UI/API contract.

## Risks

A malformed exception remains hidden and is reported as unknown.
Exceptions cannot override an explicit hidden toggle or parent page.
Older images ignore wildcard syntax; keep their explicit list during a
mixed-version rollout. No schema or feature-value changes are included.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, tool use, and code
execution. The exact runtime variant and context-window size are not
exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-25 08:07:50 -07:00
DottaandPaperclip 1dbffb4c04 fix(daytona): clean up sandbox allocation after create failure (#13979)
## Thinking Path

> - Paperclip manages AI agents and their execution environments.
> - The Daytona provider creates sandboxes before it returns a lease.
> - Daytona can allocate a sandbox, then time out while waiting for it
to start.
> - The SDK throws without returning the allocated sandbox handle.
> - Paperclip previously lost the resource identity and could leave a
running sandbox behind.
> - This pull request keeps an exact creation identity and deletes a
matching failed allocation.

## Linked Issues or Issue Description

Related: #13882. Searches for Daytona create cleanup and orphan fixes
found no duplicate for this path.

**What happened?**

In [Grok qualification campaign
36080870743](https://github.com/paperclipai/paperclip/actions/runs/36080870743),
Daytona sandbox creation exceeded 300 seconds. No native runner or
provider session started. The SDK threw, but an ownership-filtered
provider lookup found the allocated sandbox still running. The
disposable resource was then deleted separately. The original attempt
and its failure remain retained.

**Expected behavior**

A failed create must retain enough identity to clean up an allocated
sandbox. Cleanup must never delete another attempt's resource. If the
provider cannot confirm cleanup during sandbox lease acquisition, the
host must retain a durable cleanup record and retry the exact owned
allocation after restart.

**Steps to reproduce**

1. Have Daytona allocate a sandbox for an environment acquisition.
2. Make the SDK throw while waiting for startup, before it returns the
sandbox handle.
3. Before this fix, acquisition fails without deleting the allocated
sandbox.
4. The regression tests reproduce this failure and verify exact-owner
deletion.

**Paperclip version or commit**

Observed at `1b0551bb7c8de3c54f4bee64dbe2c88328b3645e`. The same
creation path exists on master; this fix starts from `efce9356b`.

## What Changed

- Assign each cold create a unique provider name and creation-attempt
label.
- On create failure, look up that exact name and verify every ownership
label before deletion.
- Wait for provider deletion with bounded lookup and deletion calls.
- Preserve the original create error after successful cleanup. Report
both errors and the provider name when cleanup cannot be confirmed.
- Treat a missing lookup after an uncertain create as unconfirmed
cleanup: a delayed provider request might still create the sandbox.
- Transfer only validated ownership fields through the acquire-lease RPC
error; provider exceptions and arbitrary error data are not serialized.
- Persist a pending-cleanup lease before the host retries deletion.
Reuse the existing cleanup sweep and durable spool, retaining the
original environment scope after its row is deleted.
- Fence retries by company, environment, run, provider account, unique
creation name, and every ownership label. Missing provisional-name
lookups remain unresolved.
- Journal an ownership-verified provider ID before host-driven deletion.
Its absence then confirms cleanup after a lost deletion reply or final
database update; failed journal writes block deletion.
- Keep ownership labels separate from the mutable input Daytona modifies
during create.
- Report confirmed host-side deletion accurately while still rejecting
the failed acquisition.
- Add protocol, malformed-evidence, cross-scope denial,
controller-restart, and deleted-environment regressions alongside the
original bounded cleanup tests.

## Verification

- SDK suite: 92 passed. Daytona plugin suite: 277 passed; six opt-in
live tests skipped. Environment-runtime suite: 103 passed, then all four
focused observation/recovery variants passed after adding a
foreign-observation case. SDK and provider TypeScript checks and `git
diff --check` pass.
- Regressions failed before the respective fixes: missing durable
ownership, misleading successful-cleanup message, and SDK mutation of
the ownership labels.
- Controlled real-Daytona proof at
`d2ac1c95b32ca64daf039c0427137657f351c0e0`: create a real sandbox;
inject an SDK failure and a failed immediate lookup; persist the
validated ownership envelope; have a child journal the observed provider
ID before deleting; deliberately drop its successful deletion receipt;
reconcile absence from another fresh process and confirm it with a
separate provider read. Passed, with no manual cleanup and zero model
calls. The live proof uses an owner-only envelope file; actual host
database/spool recovery is tested separately in the environment-runtime
suite.
- The preceding controlled attempt failed because the real SDK mutated
the labels object. That failure is retained. The harness deleted its
exact owned sandbox, and a fresh lookup confirmed absence. The fix
copies the SDK input labels separately from the ownership snapshot.
- [Repository
CI](https://github.com/paperclipai/paperclip/actions/runs/36095925399)
passes on this exact head, with a clean 5/5 review. The unchanged Cursor
remote-command test passed its one bounded rerun after a 10-second
timeout; the same test had already passed on the combined tree. The
original failure is retained. Earlier `199f0852` CI and
immediate-deletion proof are retained, without being promoted to the new
source.
- Combined-source verification uses temporary PR #13990 against the Grok
feature branch. No Docker or Rust build runs on the developer laptop.

## Risks

Creation failures now add up to ten seconds for lookup and fifteen
seconds for deletion confirmation. The SDK still owns the initial
creation timeout. If a request materializes only after the failed
lookup, its durable pending record remains eligible for later cleanup; a
never-observed creation name is never reported deleted from a 404. A
previously observed provider ID can be reconciled as deleted. The host
and SDK must be deployed together. Durable handoff applies to
plugin-backed sandbox lease acquisition used by heartbeat and runner
login. Probe and custom-image interactive setup retain immediate cleanup
and explicit failure reporting, but do not use this lease-recovery path.
A worker crash before delivering the failure envelope still cannot be
recovered through this mechanism. A cleanup error never returns a lease.
Successful creation behavior is unchanged except for the provider name
and ownership label.

## Model Used

OpenAI GPT-6 through Codex, with repository tools and code execution.
The exact serving identifier and context-window size are not exposed in
this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-25 10:06:57 -05:00
DottaandPaperclip f56aee5423 fix(evals): capture complete durable run event streams (#13977)
## Thinking Path

> - Paperclip manages AI agents and records their durable work outcomes.
> - Product E2E checks those outcomes through the browser and public
API.
> - A successful long run can emit more than 1,000 durable events.
> - The harness read one page and missed the later completion evidence.
> - This pull request reads every page before it checks runtime
invariants.
> - Invalid or incomplete capture still fails. A missing page cannot
produce a pass.

## Linked Issues or Issue Description

Related: #13882 (Grok qualification). A duplicate search found no
existing event-pagination fix.

**What happened?**

The local structured-question case in [campaign
36071063537](https://github.com/paperclipai/paperclip/actions/runs/36071063537)
completed the task and passed its six outcome matchers. It failed native
runtime invariants because the capture contained exactly 1,000 events.
The last captured event preceded the run's completion by more than a
minute. The API caps each response at 1,000 rows. The harness did not
request the next page.

**Expected behavior**

Read the complete durable event stream through the public API before
checking semantic-result and terminal-event counts. Reject incomplete or
malformed evidence.

**Steps to reproduce**

1. Complete a native task that emits more than 1,000 durable events.
2. Place the semantic-result and terminal events after row 1,000.
3. Capture the run with the Product E2E harness.
4. Before this fix, the invariant checker sees only the first page.

**Paperclip version or commit**

Observed at `4196a4cd76db434854b679035e4146c7f69689ce`. The same
single-page capture exists on master. The original failed result remains
unchanged; missing historical tail evidence is not reconstructed or
graded as a pass.

## What Changed

- Add a bounded event collector that advances through the public
`afterSeq` cursor.
- Use it for task success/failure evidence and shared chat run evidence.
- Reject invalid pages, missing or non-increasing sequence numbers,
repeated cursors, failed later requests, and an exhausted page limit.
- Test completion events beyond the first page, exact page boundaries,
and malformed evidence.
- Correct the existing Everyday catalog test from 38 to the maintained
47 cells. The suite stays explicit-only.
- Document the complete-capture requirement and its bound.

## Verification

- Product harness typecheck passes.
- The full credential-free harness suite passed 477 tests. After adding
the chat integration regression, all 34 chat evidence tests pass.
- Seventeen pagination tests cover the valid tail and malformed-evidence
cases.
- `git diff --check` passes.
- [Full repository
CI](https://github.com/paperclipai/paperclip/actions/runs/36080683422)
passes at `aed6f79c089226be79e75dcf390969561ec4f787`: 52 successful
checks and two intentional skips, including Rust, typecheck, build,
server tests, and browser shards. Greptile gives this exact head 5/5
with no findings. No local Docker, Rust build, browser suite, or model
invocation was used for this change.
- [Live Grok
requalification](https://github.com/paperclipai/paperclip/actions/runs/36080870743)
is running on combined source
`1b0551bb7c8de3c54f4bee64dbe2c88328b3645e`, with the runner, scheduler,
and evidence fixes. Its full credential-free harness passes all 493
tests and typecheck. Live results are pending; the original campaign
remains a failed measurement.

## Risks

Long runs need more read-only API requests and larger private evidence
files. Capture stops with an explicit error after 100 full pages. This
changes neither production APIs nor provider behavior. It does not relax
an invariant or change a historical grade.

## Model Used

OpenAI GPT-6 through Codex, with repository tools and code execution.
The exact serving identifier and context-window size are not exposed in
this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-25 10:06:33 -05:00
DottaandPaperclip c96c751f39 fix(heartbeat): claim task ownership with queued runs (#13973)
## Thinking Path

> - Paperclip manages agents and their tasks.
> - The scheduler can start several runs for one agent.
> - Each task still needs one execution owner.
> - Task creation and recovery can queue assignments for the same task.
> - The scheduler previously marked both runs running before starting
either executor.
> - This change claims the run and task ownership in one transaction.
> - Tasks with different IDs retain concurrent execution.

## Linked Issues or Issue Description

Refs #13882. Related queue work: #11830 and #13965; those address
different admission and cancellation paths.

**What happened?**

The Grok subscription Product campaign exposed two assignment runs for
one task. Task creation and periodic recovery started runs within 200
ms. Both acquired Daytona sandboxes. One failed before provider work
with `paperclip_runner_attachment_staging_not_authorized` because the
other owned the task. Fixture cleanup then found an active session. The
original failed result is retained in [campaign
36071063537](https://github.com/paperclipai/paperclip/actions/runs/36071063537).

**Expected behavior**

Only the task execution owner may start provider setup. Competing queued
work must wait. Different tasks may use the agent's available slots.

**Steps to reproduce**

1. Set an agent's concurrent run limit to two.
2. Queue task-creation and recovery assignment runs for the same task.
3. Resume the queue while holding adapter execution open.
4. Before this fix, both runs become running. Only one has the task
execution lock.

**Paperclip version or commit**

Reproduced on master `8781f06a8` and in the Grok campaign at `4196a4cd`.

## What Changed

- Use the same company-scoped task ownership gate for assignments,
direct comments, and queued comments.
- Leave competing work queued while another live run owns the task.
Transfer a terminal pointer only after the tracked executor and durable
environment leases/finalization settle, including cleanup on another
controller.
- Commit the running state and task execution owner together.
- Preserve the dedicated review path and the native replacement checkout
guard.
- Add sixteen database regressions for assignment/comment orderings,
unrelated concurrent tasks, and cross-controller cleanup fences.

## Verification

- Live subscription Product E2E planning passed on its first attempt in
both Daytona (351,197 ms) and local execution (268,883 ms), with 6/6
matchers and cleanup passing in each, on combined source
`1b0551bb7c8de3c54f4bee64dbe2c88328b3645e`
([campaign](https://github.com/paperclipai/paperclip/actions/runs/36080870743)).
The local screenshots and persisted state show the same plan revised,
the first approval rejected, the revised approval accepted, and one
completion. The full campaign remains unqualified because of a separate
startup retry and EC2 Spot interruption.
- The new same-task regression failed before the fix: two running rows
instead of one.
- All seven focused concurrency cases passed. Three comment cases failed
before the shared gate was added. Five cross-controller cleanup cases
failed before the durable gate was added. A warm-retention regression
also failed before its successful release receipt was admitted;
missing/failed receipts and a different retention policy stay blocked.
- `node ../node_modules/vitest/vitest.mjs run
src/__tests__/heartbeat-stale-queue-invalidation.test.ts` from `server`:
48 passed. Native cleanup admission adds one passing test; task-drain
release adds two. Total focused checks: 51 passed.
- `git diff --check`: passed.
- Current-head
[CI](https://github.com/paperclipai/paperclip/actions/runs/36078495577)
passes repository typecheck, tests, Rust checks, build, and browser
suites. The unrelated repository-form browser shard passed one bounded
rerun without assertion or code changes; its original navigation failure
remains retained. No local Docker or Rust build was used.
- The failed test's exact two sandboxes were checked: one was already
absent; the other was deleted after its company, environment, and run
labels matched.

## Risks

The claim transaction adds a task row lock. It follows task-before-run
lock order. A competing run stays queued until the current owner
releases the task. The change does not alter attachment authorization,
provider permissions, schemas, or public APIs. Review is 5/5 with all
threads resolved on `23c0d13b6`. All CI checks pass on that exact head.
Live Grok requalification will use the combined runner and scheduler
source.

## Model Used

OpenAI GPT-6 through Codex, with repository tools and code execution.
The exact serving identifier and context-window size are not exposed in
this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

### Live regression verification

The Daytona plan/revise/accept lifecycle in [Grok campaign
36080870743](https://github.com/paperclipai/paperclip/actions/runs/36080870743)
passed on its first attempt at combined source
`1b0551bb7c8de3c54f4bee64dbe2c88328b3645e`, in 351,197 ms, with all six
outcome matchers and cleanup passing. The earlier competing-run failure
remains retained. This is one live planning result; the complete Grok
matrix and repeated qualification are separate gates, and this campaign
also has separately retained infrastructure failures.

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.925.0-canary.1
2026-09-25 10:06:05 -05:00
DottaandPaperclip efce9356b5 fix(ui): offer recovery when the app fails before React starts (#13970)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The browser must load its JavaScript before React can render a task.
> - A failed import can stop that process before the React error
boundary exists.
> - The HTML entry then leaves an empty page with no recovery action.
> - This PR adds a small recovery screen that works without React.
> - The user can retry the same page and return to saved task content.

## Linked Issues or Issue Description

Refs #13824 and #13895. This is a follow-up to their browser startup
investigation.

**What happened?**

Interrupting the app bundle or a required import leaves an empty React
root. A startup exception has the same effect. A React error boundary
cannot handle these failures because React has not started.

**Expected behavior**

The page must explain the startup failure and offer a manual retry. A
late successful load must dismiss the recovery message without a reload.

**Steps to reproduce**

1. Open a saved task in the browser.
2. Abort the application bundle request, or make a required module
return HTTP 503.
3. Observe the empty page before this change. With this change, use
Reload page after the fault clears and verify the saved task and
comment.

**Paperclip version or commit**

The failing regression baseline used master at `8781f06a8`.

**Deployment mode**

Local source build and compiled UI. Tests cover both initial navigation
and a page controlled by the production service worker.

The exact cause of the older intermittent Vite stall remains
unconfirmed. Forty app loads and thirty replays of retained responses
did not reproduce it. This PR fixes the missing recovery path; it does
not claim to remove that historical cause. A normal HTTP 304 response is
not a failure.

## What Changed

- Add an inline startup guard and recovery screen in the HTML entry. It
does not depend on the app module graph.
- Show a manual reload action after a startup error or after 30 seconds
without rendered root content.
- Remove the notice, timer, observer, and error listeners when the app
starts. Never reload automatically.
- Keep the recovery screen outside the React root so it cannot satisfy
app-readiness checks.
- Add browser tests for interrupted imports, a stalled import, an
evaluation error, service-worker-controlled retry, repeated offline
retry, and cleanup after successful startup.
- Return a static, uncached HTML retry screen when a
service-worker-controlled navigation fails offline. It contains no task
content.
- Add a full-app test that retries an interrupted compiled bundle and
checks the saved task, comment, composer, route, and absence of agent
runs.
- Document the coverage and the limits of the historical diagnosis.

## Verification

- Red baseline: four recovery cases failed; the normal-startup case
passed. After the change, all five recovery cases passed. The review
found an offline retry gap; that additional case failed before the
worker fix and passed afterward.
- Full provider-free browser-support suite: 16 passed.
- Compiled-app browser tests: four passed, including saved-task reload,
interrupted-bundle recovery, slow-CPU service-worker reload, and sidebar
navigation.
- Expanded service-worker, offline response, PWA, and worker build-ID
unit tests: 37 passed. The two old plain-text offline expectations were
reproduced as failures and updated for the HTML retry contract.
- UI production build, full local repository typecheck (`pnpm -r
typecheck`), runner-E2E typecheck, and design token checks passed.
- Manual browser check: a temporary server failed the compiled bundle
once. The recovery screen appeared. Clicking Reload page restored the
same saved task, comment, and composer.
- Full local `pnpm build` passed.
- Full local `pnpm test:run` was attempted with a bounded deadline and
stopped after it timed out. Workspace runtime/cleanup tests reported
timeouts on this host. The monolithic local run is not a pass. The
focused tests above and the complete Linux CI run provide the successful
verification.
- Final-head [CI
run](https://github.com/paperclipai/paperclip/actions/runs/36072201966)
passed. All 53 check runs succeeded; the two Storybook jobs were
intentionally skipped. The legacy security status also passed.
- Greptile reviewed `f83e0f51fb760541d83353f2c1df4e182f3948f9`: 5/5.
Both review findings are fixed and resolved.

## Risks

- The guard only handles startup before React renders root content.
Existing React boundaries handle later rendering errors.
- A slow startup can show the message after 30 seconds. A later
successful render removes it; the page does not reload by itself.
- The fallback uses native HTML when the app stylesheet is unavailable.
- The worker changes only its offline navigation response. It returns
static HTML with a reload button and `Cache-Control: no-store`. Its
cache allowlist, private-response protections, task state, provider
prompts, and grading rules stay unchanged.
- This does not establish or fix the unknown cause of the historical
intermittent Vite stall.

## Model Used

OpenAI GPT-6 through Codex. The session exposes the GPT-6 family but not
an exact served model ID or context window size. Used reasoning, code
editing, shell tools, and browser testing. No subagents were used.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.925.0-canary.0 nightly/v2026.925.0-nightly.0
2026-09-24 19:36:45 -05:00
DottaandPaperclip 8781f06a87 feat(connections): enable MCP aggregators by default (#13964)
## Thinking Path

> - Paperclip helps people manage AI agents for work.
> - Connections let those agents use external services with explicit
access rules.
> - Zapier, Arcade, Composio Connect, and Executor already have setup
and runtime support.
> - Their experimental switch still blocks discovery and setup by
default.
> - This pull request removes those gates and the Settings toggle.
> - Users can connect these providers without enabling an experiment.

## Linked Issues or Issue Description

Refs #13755. Refs #13941.

**What existing behavior does this improve?**
Apps browsing, inline setup, and agent connection search for the four
MCP aggregators.

**Current behavior**
An instance must enable the MCP aggregators experiment before users or
agents can start setup.

**Proposed behavior**
All four providers are available by default on local and managed
instances. Old stored and managed values still parse but cannot disable
them.

## What Changed

- Remove the aggregator gates from Apps, inline setup, server setup, and
agent search.
- Remove the Settings toggle and its UI hook.
- Retain the old setting key only for upgrade compatibility. Normalize
it to true and ignore managed overrides, as Apps already does.
- Replace opt-in fixtures with default-on coverage. Test old false
values, all four setup flows, provider choice, and the removed toggle.
- Update current connector guidance and remove the opt-in from the
runner acceptance fixture.

## Verification

- 306 focused tests passed across eight files: shared remote MCP
contracts; server remote MCP lifecycle, aggregator fallback, settings
normalization, and managed overlay; UI Apps browsing, setup, and
experimental settings.
- Server and UI TypeScript checks passed.
- UI token gates and `git diff --check` passed.
- The full local suite was not run, per the maintainer's instruction.
All 54 CI checks passed; two checks were skipped. One unrelated
workspace-preview readiness timeout passed on one failed-shard retry.
- The setup fixtures use simulated MCP responses. This change does not
claim new live provider acceptance.

## Risks

- Existing instances now show all four providers, even if the old flag
was false. This is intentional.
- External provider choice, credentials, company isolation, agent
grants, and tool policies still apply. Showing a connector does not
authorize an external account.
- No data migration is required. The compatibility key keeps old managed
configuration documents valid.
- Historical Zapier live acceptance remains incomplete in the existing
evidence report. The maintainer explicitly requested the default-on
rollout for all four existing providers; the report records that scoped
exception.

## Model Used

OpenAI GPT-6 (`gpt-6-astra`) through Codex, with reasoning, repository
tools, and test execution. The context window size is not exposed in
this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.924.0-canary.5
2026-09-24 17:31:32 -05:00
DottaandPaperclip aa8fc86331 feat(connections): prefer native apps and ask users to choose external providers (#13941)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Connections give agents governed access to external services.
> - Native connections should remain the first choice for a supported
app.
> - Other apps may be available through an external MCP provider.
> - The user must know which external provider handles the connection
and choose it before setup.
> - This pull request adds ranked alternatives and server-authored
instructions to connection search.
> - Agents can follow the returned instructions while Paperclip
validates saved choices and access.

## Linked Issues or Issue Description

Related: #13879, which fixed inline MCP provider setup. This PR adds
discovery and provider selection on top of that work.

**Subsystem affected**

Cross-cutting: shared connection contracts, server search and intent
services, native runtime, CLI, inline setup UI, and evals.

**Problem or motivation**

An agent cannot offer a clear external-provider choice when Paperclip
has no native connection for an app. Adding provider-specific branches
to the core prompt would make those instructions harder to maintain.

**Proposed solution**

Prefer a native connection. Otherwise return verified alternatives in
Composio, Arcade, Executor, Zapier order. Include an external-service
disclosure, a question with None, and the next instruction in the search
result. Validate the saved human choice before creating a selected
fallback setup card. Reuse existing provider accounts and verify
underlying app access separately.

**Alternatives considered**

Do not silently choose a provider. Do not claim that broad execution
tools prove support for every app. Reuse existing questions and
connection intents rather than add another connection model.

**Roadmap alignment**

Extends the existing MCP Tool Gateway & Apps and Agent evals & feedback
capabilities. The MCP aggregators experiment remains the gate. No
duplicate provider-routing PR was found in the public search.

## What Changed

- Add a dated support index and authorized cached-tool evidence for
external routes.
- Return provider questions and next-step instructions from
`connections_search`.
- Preserve pending choices and declines across continuation. Validate
company, task, agent, human, app, and current route eligibility.
- Carry the selected app into new setup and account reuse, validate
explicit provider requests against persisted human messages, and
distinguish provider readiness from app authorization.
- Sync native, MCP, REST, and CLI contracts. Keep core agent
instructions provider-neutral.
- Add production-component Storybooks, focused database tests, and three
real-agent browser eval cases.
- Record the plan, observed failures, fixes, passing evidence, and
acceptance limits.

## Verification

- Latest head `586f0e6cd`: 54 checks passed, 2 skipped; Greptile 5/5 and
all review threads resolved.

- After rebasing on master `18dac1e1e`: 64 focused shared, validator,
route-contract, and database tests passed; server typecheck passed.
- Embedded-browser test drive on the rebased head: native Jira card,
HubSpot external-provider question, Arcade account reuse, one actual MCP
read against a local synthetic fixture, reload persistence, and None
preventing further calls. A real OpenAI-backed agent performed discovery
and continuation.
- UX observation: the agent initially combined mutually exclusive
request fields; the server rejected it and the agent recovered without
changing access. This extra retry remains visible in the transcript.
- After rebase: 23 focused eval grader/catalog tests, affected
TypeScript checks, token gates, production UI build, and Storybook build
passed. Full local tests are intentionally excluded at the maintainer's
request.
- Before rebase: four browser/real-agent attempts passed: native Jira,
None, and reuse of the second provider on two Codex profiles.
- Browser evals used an isolated deterministic MCP fixture through the
real Paperclip gateway. They do not prove production compatibility with
all four providers.
- Review `Apps / Connections / Provider choice` in Storybook. Choose
Arcade, continue through Access, and verify the app name,
external-service disclosure, and URL configuration.
- The detailed verification report is
`doc/connections/2026-09-23-aggregator-routing-verification.md`.

## Risks

- The public support index is finite and can age. Account capability and
app authorization still require verification after selection.
- Existing installed-tool permissions remain in effect. Provider choice
is not a new execution permission boundary. An early Mini attempt
skipped search; clearer provider-neutral instructions made the targeted
rerun pass. This is not a measured reliability rate.
- Explicit requests skip provider confirmation only when a clear
persisted human message or saved provider choice supports them. Other
phrasing falls back to confirmation; the agent query alone is not
consent. These routes do not add tool permissions.
- No database migration or legacy Composio broker is added.
Real-provider acceptance remains separate from fixture proof.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, code
execution, and browser tools. The exact deployment variant and
context-window size are not exposed in this session. Product evals
separately used the repository's primary Codex and Codex Mini profiles;
those agents supplied test behavior, not independent provider
compatibility proof.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.924.0-canary.4
2026-09-24 16:36:49 -05:00
DottaandPaperclip 18dac1e1ef feat(connections): add experimental memory providers and remote MCP access (#13942)
## Thinking Path

> - Paperclip manages AI agents and their work.
> - Connections give agents governed access to external tools.
> - Agents need durable memory across tasks and execution environments.
> - Mem0, Zep, Supermemory, Cognee, and Honcho provide hosted memory
tools.
> - This pull request adds their setup flows behind an experimental
toggle.
> - It also delivers assigned MCP tools through the native remote Codex
runner.
> - Operators can connect a provider once and use the same governed
tools locally or in Daytona.

## Linked Issues or Issue Description

**Problem or motivation**

The Apps catalog lacks a complete set of memory providers. Remote native
Codex agents also need access to assigned managed MCP tools without
receiving provider credentials.

**Proposed solution**

Add five memory connectors behind the disabled-by-default Experimental
memory connectors setting. Use the existing connection setup and
permissions UI. Default their tools to Allowed. Preserve company
boundaries, operator permission changes, provider scopes, and audit
attribution.

**Alternatives considered**

Direct provider credentials in each sandbox would duplicate setup and
bypass the managed gateway. The remote runner instead uses its existing
protocol channel to call the gateway on the server.

**Roadmap alignment**

This maintainer-requested experiment supports the Memory / Knowledge and
Connected Apps roadmap areas. It adds provider connections without
introducing a separate memory UI. Related connector authoring
documentation is tracked in #13692; no duplicate memory-provider
implementation was found.

## What Changed

- Add provider definitions, official branding, and the experimental
setting for all five providers.
- Use OAuth for Zep and Supermemory, API credentials for Mem0 and
Honcho, and a bundled Cloud API bridge for Cognee with no runtime
downloads or subprocesses.
- Default memory tools to Allowed and classify destructive actions
explicitly.
- Fix personal remote credential resolution and propagate provider tool
errors.
- Relay assigned managed MCP tools to remote native Codex through the
runner protocol. Recheck current authority for each call and rotate
stale tool contracts.
- Add 21 Storybook states and complete OAuth walkthrough fixtures.
- Document provider research, sanitized tool inventory, and live local
and Daytona proof.

## Verification

- Passed workspace typecheck: `pnpm -r typecheck`.
- Passed production build: `pnpm build`.
- Full local `pnpm test:run`: 13,273 passed, with failures from
process/readiness timeouts under parallel load. Reran all 16 affected
suites with one worker: 456 passed, leaving two macOS `/var` versus
`/private/var` path assertions. Both passed with `TMPDIR=/private/tmp`.
No test failures remain unverified. Latest-head remote CI passes all 54
checks (two optional Storybook jobs skipped). Greptile is 5/5 with no
unresolved findings.
- Latest Cognee gateway regression: 71 passed, including public
deployment without a runtime host and immediate recovery after a
provider error. Bundled bridge tests: 25 passed.
- Browser setup and real agent tasks exercised all five providers. Mem0,
Cognee, Zep, and Honcho have successful store/retrieve proof.
- Supermemory now has scoped read/write consent. Local storage and real
Daytona write, document read, and semantic recall passed; indexing
completion was verified before claiming success.
- All five providers were exercised through a real Daytona sandbox and
its native runner MCP relay. Provider credentials remained on the
server. The final bundled Cognee bridge also passed a fresh Daytona
store/recall run. All disposable sandboxes were removed and verified
absent after testing.
- Storybook is rebuilt and contains the experimental toggle, catalog,
setup, permissions, and error states. The Zep and Supermemory access
steps advance correctly.

## Risks

- Provider OAuth scopes and plan limits remain independent of Paperclip
tool permissions. An Allowed tool can still be rejected by the provider.
- Providers can queue memory indexing; save acceptance does not prove
that semantic recall is ready.
- Remote tool contracts must stay synchronized with current connection
authority. Regression tests cover revocation and stale contracts.
- This change has no database migration. Existing connections remain
usable when the experimental catalog toggle is disabled.

## Model Used

OpenAI Codex, GPT-6 Astra (`gpt-6-astra`), with reasoning, repository
edits, code execution, and browser automation. The exact context window
size is not exposed in this session. Live acceptance agents used
`gpt-5.6-sol`.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-24 15:54:38 -05:00