Commit Graph
1915 Commits
Author SHA1 Message Date
Devin FoleyandPaperclip e9debd3eac fix(workspaces): allow larger status output for readiness checks (#14414)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The server checks workspace contents before it allows cleanup or
branch reconciliation.
> - These checks need the full Git status output to count untracked
files.
> - Nested task worktrees can make this output exceed the scheduler's
default 1 MiB limit.
> - A scan failure then blocks an otherwise inspectable workspace.
> - This pull request raises the limit for these status checks to 32
MiB.
> - The checks retain exact counts and still protect uncommitted work
from cleanup.

## Linked Issues or Issue Description

Refs #14194 and #14253. Those changes address snapshot enumeration. This
PR addresses buffered execution-workspace status checks. Refs #13619 for
separate work on caching these checks. The related journal identity
failure is covered by #14312.

**What happened?**

Close-readiness checks failed when Git status output exceeded 1 MiB. A
workspace with thousands of long untracked paths could not report its
file count or complete readiness inspection.

**Expected behavior**

Allow up to 32 MiB of status output for execution-workspace readiness
and branch reconciliation. Preserve exact counts. Continue to block
cleanup when the workspace has uncommitted data or the scan exceeds its
bound.

**Steps to reproduce**

1. Create an execution workspace with a merged delivery.
2. Add 5,000 long untracked filenames under a nested task directory. The
status output exceeds 1 MiB.
3. Request close readiness. Before this fix the status scan fails. After
this fix it reports all 5,000 files.
4. Run the terminal-workspace sweep. Confirm that it preserves the
workspace and files.

**Paperclip version or commit**

Reproduced on master at `14795136f5` before this fix.

**Deployment mode**

Server with managed Git workspaces.

## What Changed

- Set a 32 MiB stdout bound for execution-workspace status scans.
- Add a real Git regression with 5,000 long untracked paths and a
cleanup-preservation assertion. Assert that the measured status output
exceeds 1 MiB.
- Document the bound and failure behavior.

## Verification

- The new regression failed on master before the service change: the
status result had no untracked files or count after the scan exceeded
its bound.
- `pnpm exec vitest run
server/src/__tests__/execution-workspaces-service.test.ts
server/src/services/workspace-git-operation-scheduler.test.ts` passed:
82 tests, including the new regression. The regression and server
typecheck also passed after the explicit byte-count assertion.
- `pnpm build` and `pnpm -r typecheck` passed. The full local `pnpm
test:run` was attempted and stopped after the failures listed below.
Greptile is 5/5 with zero unresolved review threads on the latest head.
All checks for head `45f93ad5ad` passed (53 successful, two intentional
skips).
- Full local validation did not pass. The attempt reproduced 13
company/runtime skill-cache permission failures, the terminal-workspace
cleanup assertion, and a heartbeat feedback timeout. It was stopped
during the general-server stage after these failures. Remaining
general-server tests, other workspace groups, and serialized-server
stages did not complete locally. Earlier clean-master checks reproduced
the cache failures and isolated cleanup retries passed. The latest-head
GitHub suite passed all of these groups.
- One GitHub browser shard initially failed because its
GitHub-connection test checked the resume-action array before the mocked
request completed. The single failed-job retry passed without a source
change. All latest-head checks are green.
- No browser suites ran locally. This change does not affect browser
behavior.

## Risks

- Each active status scan can buffer up to 32 MiB before parsing. The
existing scheduler bounds scan concurrency and queue size.
- Output above the bound still fails the scan and blocks destructive
cleanup. The change does not truncate output or change cleanup rules.
- There are no API, database, or snapshot-streaming changes.

## Model Used

OpenAI Codex, GPT-6, with repository inspection, code editing, and test
execution. The exact deployment model ID and context window are not
exposed in this session.

## Checklist


- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused suites; full
local-run failures and incomplete stages are documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 09:59:56 -07:00
Devin FoleyandPaperclip 6f40e23536 fix(logging): redact cloud authentication headers (#14413)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The server records HTTP requests to help operators diagnose
failures.
> - Cloud requests carry tenant credentials and signed assertions in
headers.
> - The HTTP logger did not redact four of these headers.
> - This pull request adds those headers to the existing redaction list.
> - Operators retain the route, method, and response status without
recording these values.

## Linked Issues or Issue Description

**What happened?**

HTTP request logs could contain cloud tenant credentials, session
identifiers, runtime identity assertions, and cloud control assertions.

**Expected behavior**

The logger must redact these header values for successful requests and
failed requests.

**Steps to reproduce**

1. Create an Express app with the production HTTP logger and redaction
configuration.
2. Send a request with the four cloud headers and distinct test values.
3. Inspect the serialized request headers for responses with status 200,
403, and 500.

**Paperclip version or commit**

Reproduced on `0f14d26123` before this fix.

**Deployment mode**

Server with cloud proxy authentication. The regression test uses an
in-process Express server.

## What Changed

- Redact `x-paperclip-cloud-tenant-token`,
`x-paperclip-cloud-session-id`, `x-paperclip-cloud-runtime-identity`,
and `x-paperclip-cloud-control` in HTTP request logs.
- Test real logger output for HTTP 200, 403, and 500 with mixed-case
request header names. Route 403 and 500 through the production error
handler. Check response bodies, log levels, and server error context.
- Check that the method, route, and status remain available.

## Verification

The `server/src/__tests__/http-log-redaction.test.ts` suite passed,
including all three new cloud-header cases.

- Rebased onto master at `14795136f5`.
- `pnpm exec vitest run server/src/__tests__/http-log-redaction.test.ts`
passed: 59 tests, including all three new response-status cases. The
suite and server typecheck also passed after the error-handler coverage
update.
- `pnpm build` and `pnpm -r typecheck` passed. The full local `pnpm
test:run` was attempted and stopped after the failures listed below.
Greptile is 5/5 with zero unresolved review threads on the latest head.
All checks for head `373d29e2f1` passed (53 successful, two intentional
skips).
- Full local validation did not pass. The attempt reproduced
company-skill cache permission failures, the terminal-workspace cleanup
assertion, a heartbeat feedback timeout, and one process-conversation
timing failure. It was stopped during the general-server stage after
these failures. Remaining general-server tests, other workspace groups,
and serialized-server stages did not complete locally. Earlier
clean-master checks reproduced the cache failures and isolated cleanup
retries passed. The latest-head GitHub suite passed all of these groups.
- No browser suites ran locally. This change does not affect browser
behavior.

## Risks

- These four values will no longer be available in HTTP logs. Route,
method, and status remain available.
- This change applies to new log entries. It does not remove old entries
or rotate credentials.
- No schema, API, or authentication behavior changes.

## Model Used

OpenAI Codex, GPT-6, with repository inspection, code editing, and test
execution. The exact deployment model ID and context window are not
exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused suites; full
local-run failures and incomplete stages are documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 09:59:15 -07:00
DottaandPaperclip 14795136f5 fix(runtime): finalize and recover sandbox workspace exports safely (#14402)
Serialize native workspace finalization, validate streamed archives within bounded limits, and quietly recover unsafe exports from saved results. Preserve exact allocations for exhausted transient failures and provide export-only retry without rerunning the provider.

Consolidates #14314, #14315, #14329, and #14334 while preserving the already-merged finalization label changes.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-28 10:57:32 -05:00
DottaandPaperclip 3447609d22 fix(runner): stream and page large API responses within capture budgets (#14301)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents use governed API tools to inspect task evidence.
> - Large API results become saved assets with short previews.
> - Reading an asset through the same tool used to create another asset,
so the agent could not reach the rest of the evidence.
> - The 10 MiB response cap also blocked useful large results. Removing
all bounds allowed excessive disk use.
> - This pull request streams responses up to 1 GiB and makes saved text
readable in bounded pages. It adds durable run budgets and capture
admission limits.
> - Agents can inspect complete evidence while tool results, memory use,
and capture work stay bounded.

## Linked Issues or Issue Description

**What happened?**
A large response became an asset. Reading that asset returned another
asset and the same preview. Responses above 10 MiB failed before the
agent could read any page.

**Expected behavior**
The agent can fetch a large response and read its saved text to EOF.
Each page stays bounded. New snapshots have a generous finite limit and
a durable run budget. Existing larger assets remain readable through
byte ranges.

**Steps to reproduce**
1. Call a GET operation that returns more than 10 MiB of text or JSON.
2. Before the fix, the tool returns `api_transport_failure`.
3. With this change, responses up to 1 GiB become streamed snapshots
with artifact references.
4. Read `GET /api/assets/{assetId}/content` with `responseText:
{offsetBytes: 0, limitBytes: 8192}`. Follow `nextOffsetBytes` until
null.

Related work: #14186 added the API fallback tools. #14218 bounded API
discovery.

## What Changed

- Add authenticated UTF-8 text windows to `call_api`, with byte offsets
and total size. Keep each page at or below 24 KiB.
- Stream new responses above 24 KiB through private temporary files into
company-owned assets. Bound each capture to 1 GiB of decoded bytes.
Reject oversized declared lengths before reading and count streamed
bytes before writing.
- Reserve capture budget in the run record before spilling. Allow 4 GiB
per run. Settle successful captures to their actual size. Failed or
interrupted captures retain their full 1 GiB reservation. Run restarts
do not reset the budget.
- Enforce a 20 GiB company snapshot quota with database reservations.
Count legacy snapshots and unfinished storage work across runs and
processes. Asset deletion frees quota.
- Limit large captures to two per company and four per server process.
Hold slots through storage upload and temporary-file cleanup. Use a
10-minute download deadline and 30-second connection/idle-read timeouts.
- Return explicit size, budget, busy, and timeout errors. Preserve
unknown outcomes for mutations whose response cannot be captured.
- Read saved assets through authenticated storage ranges, with at most
two extra bytes for UTF-8 and EOF handling. Unpaged reads return the
existing asset and digest with a bounded preview. Reads create no copies
and do not consume capture budget.
- Keep existing assets above 1 GiB readable in pages. Use safe integer
offsets and PostgreSQL `bigint` asset sizes.
- Stream large S3 uploads through ordered multipart requests. Abort
failed uploads and remove partial local files.
- Revalidate run authority during downloads. Keep company authorization,
GET-only text paging, redirect denial, and mutation replay receipts.
- Document the separate 10 MiB upload limits. This PR does not raise
memory-buffered attachment ingestion limits. Future large video uploads
need streamed ingestion and storage quotas.

## Verification

- Full workspace `pnpm -r typecheck` and `pnpm build` pass after
rebasing on master.
- Focused API and response tests: 1,761 pass. Cover declared and chunked
oversize responses, incorrect Content-Length, exact-limit success,
active-stream deadline, cancellation, cleanup, concurrency admission,
and mutation outcome handling.
- Real HTTP integration: 28 tests pass, including runnerd → PRP →
authority → HTTP, a 12 MiB snapshot, final-page/EOF reads, cross-company
denial, a persisted 3 GiB sparse asset, and large mutation receipt
replay.
- The HTTP suite verifies durable run-budget accounting, simultaneous
runs competing for company quota, legacy snapshot accounting, deletion
refunds, failed-storage reservations, cleaned-failure refunds,
metadata-rollback cleanup refunds, preservation after a lost commit
acknowledgement, and small/saved reads after capture-budget exhaustion.
- A standalone proof streams exactly 1 GiB through the production
capture helper, verifies the final bytes, and removes its temporary
file. It uses repeated 256 KiB chunks and records a peak process RSS of
191 MiB.
- Earlier storage verification covers exact S3 multipart boundaries,
cleanup/abort failures, and a 17 MiB transfer through the real AWS SDK
to a local HTTP S3 endpoint. No cloud S3 qualification was run for this
follow-up.
- The local full test run was interrupted for the company-quota changes.
A later targeted run hit exhausted macOS shared-memory slots before
tests started; two unattached PostgreSQL segments with dead owners were
reclaimed before retrying. All 55 current-head checks pass at
`aebb80ceeeee77d5a56b67bfffd835f2f846878c`, including the full CI test
suite, typecheck, build, browser suites, security scan, and Greptile
(5/5). There are no unresolved review threads. The combined rebased test
catalog also passes (48 tests).
- Earlier paging acceptance passed Daytona and separate staging at
`7739879e9`. Those runs predate the streaming and budget changes.

## Risks

- The 1 GiB response cap and 10-minute active-download deadline are
intentional product limits. Larger live results must use endpoint
pagination or a direct file workflow. Existing larger assets remain
readable through bounded ranges.
- A durable 20 GiB company snapshot quota counts stored runner-api
assets and active/orphan reservations across runs and processes. The
operator can set PAPERCLIP_RUNNER_API_COMPANY_CAPTURE_MAX_BYTES to a
finite value of at least 1 GiB. Deleting snapshots frees capacity;
possible orphan storage must be reconciled before releasing its
reservation.
- A failed capture uses its full reservation. A new large capture needs
a full 1 GiB available, even if it later completes at a smaller size.
Small reads and existing asset pages remain available.
- Concurrency limits apply per server process. The run byte budget is
shared through the database.
- The `integer` to `bigint` migration rewrites asset metadata and takes
an exclusive table lock. File bytes stay in storage.
- A live endpoint is fetched once before returning its snapshot.
Continue reading the saved artifact for stable pages. Mutations may
commit before any size or transport error; inspect state before
retrying.
- Attachment uploads and native file handoffs still default to 10 MiB.
Raising buffered ingestion paths to GiB sizes is separate work.

## Model Used

OpenAI Codex, based on GPT-6, with code execution and repository tools.
The runtime does not expose an exact serving model variant or
context-window size. The earlier paging work also used browser testing
and subagents.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 10:36:32 -05:00
DottaandPaperclip 8751e2de46 fix(ui): distinguish finalization recovery from live observation (#14326)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The task board shows which recovery actions are active.
> - A native run can stop while a person must repair its workspace.
> - The board previously called that state “Recovery in progress.”
> - The label implied that work would continue without operator action.
> - This pull request derives the label from the recovery owner and live
continuation.
> - Operators can distinguish scheduled recovery from a repair that
needs attention.

## Linked Issues or Issue Description

**What happened?**
A blocked task showed “Recovery in progress” after native finalization
stopped and no automatic continuation remained.

**Expected behavior**
Show “Recovery needed” for an idle board repair or when no live recovery
path exists. Show “Recovery in progress” while the recorded continuation
can run, including an explicitly admitted export whose exact callback is
executing even if the old recovery action remains board-owned.

**Steps to reproduce**
1. Complete a native run whose workspace export cannot be recovered
automatically.
2. Inspect the task recovery action and badge.
3. Compare the board-owned action with the old “Recovery in progress”
label.

Related work: #14314 serializes native workspace finalization and fences
stale recovery outcomes. This change reports the recovery action that
currently owns the task.

## What Changed

- Show “Recovery needed” for board-owned active-run recovery unless the
exact native export callback is positively verified as executing.
- Require the native continuation run and a live or future continuation
before showing progress.
- Project native activity from the exact company, source issue, and run.
Include active workspace export while the original heartbeat remains
failed.
- Require an executing callback before using a running export row as
evidence. Preserve activity for long exports and clear it when the
callback joins.
- Give native resume its own card explanation. Preserve ordinary
watchdog observation behavior.
- Remove the redundant ownership sentence from all six recovery-card
explanations that used it.
- Document the labels and add regression cases for stopped, scheduled,
and active recovery.

## Verification

- Copy-only follow-up (`c68aef04c`): all 167 focused recovery UI tests
and token gates pass. No UI occurrence of the removed sentence remains.
`pnpm -r typecheck`, `pnpm build`, and current-head CI pass (54
successful checks, two optional Storybook checks skipped). Greptile is
5/5 with no open review threads. The duplicate local `pnpm test:run` was
stopped after the full CI suite passed; it did not complete locally.
- Original regressions: seven failures before the change, then 52
focused cases pass.
- Review regressions: seven UI failures and nine database failures
before the follow-up. All 75 database/API recovery tests, 167 UI tests,
and 18 workspace lifecycle/finalizer tests pass. A further three RED
cases cover explicit board retry activity; one RED case rejects orphaned
running export rows after controller loss. Wrong company, issue, run,
service, phase, and completed-operation cases remain inactive.
- Final recursive typecheck, production build, token gates, and complete
local suite coverage pass. Embedded PostgreSQL startup/socket failures
passed in isolated retries with the canonical test environment; no
expected behavior was weakened. The recovery regression added during the
earlier full run passed in its final complete 75-case file.
- Before the copy-only follow-up, all 56 CI checks passed on
`c2f84cd89c905cda85c53aaf5bb83b7250900fe6`; Greptile is 5/5 with no
unresolved review threads. The final native-activity staging repeat
passed on integrated source `2bedd0f23bf4698b1f8b818f6796900647030427`.
- Verified on a separate staging instance: a real failed Daytona
workspace export retains its board-owned repair action and displays
“Recovery needed” in the task list. The repair card remains actionable
without starting another provider turn.

- Real staged export-only repair: the actual task list showed “Recovery
in progress” while the original run had a positively identified running
export operation, then Done after exact copyback of all 20,000
nonce-bound files. The accepted result and full provider
session/turn/terminal envelopes remained unchanged. All 17 independent
final checks passed; the browser downloaded the exact 19-byte result.
The separate fixture was cleaned up with independent provider-absence
verification. A control transport process restarted during repair; it
did not submit another provider turn.

- Final deployed-source repeat on
`d884e1ab046cc76004e35e6091e9e6e2c918c9eb`: explicit per-turn ephemeral
Daytona allocation, actual browser export repair, and a saved full
activity projection referencing the exact executing export operation
with no scheduled retry. The task list showed recovery in progress, then
Done; all 21 final checks passed, including 20,000 exact host files,
unchanged provider provenance, and provider deletion only after
committed copyback. The downloaded 19-byte result matched independently.
This integrates #14334; no source change was required here.

## Risks

- The label depends on the persisted recovery action. A separate runtime
defect can still stop work; this change makes that condition visible.
- Future recovery kinds must supply a valid continuation path before
they can display progress.

## Model Used

OpenAI Codex, based on GPT-6, with code execution, browser testing, and
subagent tool use. The runtime does not expose an exact serving model
variant or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 09:33:39 -05:00
DottaandPaperclip 4b38db9622 fix(runtime): stream workspace Git snapshots through disk manifests (#14253)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Managed runs copy a selected workspace to an execution environment
and restore its changes.
> - Git snapshots select the files for that copy and for later recovery.
> - A fixed output limit stops large generated trees before the run can
start.
> - Increasing the limit still keeps the complete filename lists in
memory.
> - This pull request stores those lists and merge baselines in disk
manifests.
> - Large snapshots can now complete with bounded filename buffers and
explicit failure handling.

## Linked Issues or Issue Description

Refs #14194. This is the streaming follow-up to the merged 32 MiB limit
fix.

Related: #13619 and #11621 cover workspace scan admission and demand.
This change keeps the shared scheduler and changes the snapshot data
path.

## What Changed

- Stream changed, untracked, deleted, and ignored paths through the
shared scheduler and the standalone adapter path.
- Use SQLite manifests for file selection, duplicate removal,
ignored-path lookup, baseline capture, and merge lookup.
- Set a configurable 30-minute snapshot deadline. Keep the existing
interactive scan deadlines.
- Wait for each child process and pending sink write before removing
temporary storage after failure or cancellation.
- Use NUL archive lists and bounded deletion batches. Preserve unusual
names, explicit selection, nested repositories, and source root checks.
- Store manifest references in native recovery format v2. Check their
location and digest before recovery reads. Keep v1 descriptors readable.
- Remove temporary manifests at lifecycle completion. Use fixed-size
temporary copy names for long basenames.
- Admit each manifest with a SQLite page allowance based on current disk
capacity. Keep a configurable free-space reserve and fail explicitly
when either limit is reached.
- Preserve a host file that replaces a directory deleted by the sandbox,
and continue the rest of the restore.

## Verification

- Current head `5ab622ec43cd16d35429d79dedee6a5d8e3d2df2` has 54
successful checks/statuses and two skipped Storybook jobs. No checks
failed or remain pending.
- [CI
passed](https://github.com/paperclipai/paperclip/actions/runs/36318966368):
typecheck, build, all test shards, E2E, Rust checks, and the aggregate
verify job.
- [Greptile is
5/5](https://github.com/paperclipai/paperclip/pull/14253#issuecomment-5855670338)
on the current head. All four review threads are resolved. Security
checks passed.
- 229 focused tests passed across Git sync, runtime staging, merge,
manifest integrity, native recovery, and the scheduler (214
adapter/runtime tests and 15 scheduler tests).
- A real 40,000-file fixture produces 43,428,890 filename bytes. The
original standalone and scheduled scans fail. The new test passes all
four filename paths, complete staging, exclusion of late files, unusual
names, and deletion replay.
- Recovery tests reject changed bytes, symlinks, and paths outside the
controller state directory. Adapter-utils typecheck passed.
- A test executor returned buffered output and caused two retry
integration failures. The fixture now uses the shared streaming
scheduler. All 13 tests passed with `corepack pnpm exec vitest run
server/src/__tests__/heartbeat-project-repositories.test.ts`. The same
CI shard now passes.
- Ran `pnpm -r typecheck`, `pnpm test:run`, and `pnpm build` locally.
Each full local command hit SIGKILL/exit 137 in the 4 GiB container.
These local commands did not pass. The current-head CI gates above
provide the full verification.
- A real-Git disk-capacity regression confirms a typed failure and
removal of the incomplete manifest. Repeated writer attempts cannot
exceed the permitted page count.

- Follow-up real Daytona and separate staging qualification passed with
the related archive validator (#14315) and exact-owner finalization fix
(#14314). Three successive turns copied back all 60,000 files with
39,828,890 filename bytes and five unusual names. Independent host
inventories verified every file and the pinned Git HEAD. Native,
provider, session, and process identities stayed fixed; no retry
remained. The task reached Done, and its browser-downloaded final proof
matched exactly. The reusable regression is #14316, including an
assertion of the effective environment idle policy.

## Risks

- SQLite manifests use disk space. Each receives one quarter of the
available capacity above the host reserve at creation. The reserve
defaults to 256 MiB and has a 64 MiB configuration minimum. Disk
capacity, filesystem quotas, per-path limits, Git resource use, and
execution deadlines remain limits.
- Each path and sink chunk has a 64 KiB limit. SQLite connections use a
1 MiB page cache. Invalid or incomplete records fail explicitly.
- Restore transport keeps fixed and configured archive exclusions. A
remotely created Git-ignored file can be transferred, but the host merge
excludes it through the manifest.
- Provider archive buffers, Git and tar memory, repository metadata,
legacy v1 arrays, and the separate referenced-source resolver retain
their own limits. Existing provider safety validators still buffer
textual tar listings: Daytona allows 32 MiB and Kubernetes allows 64
MiB. These separate transport limits can stop a sufficiently large
restore before merge. This change does not claim bounded total process
memory or unlimited transport size.
- New descriptors use v2. Existing v1 recovery remains supported; a
downgrade cannot read v2 descriptors.

## Model Used

OpenAI GPT-6 through Codex. The exact deployment ID and context limit
are not exposed in this run. The agent used code editing, terminal
execution, tests, and GitHub tools.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 08:40:36 -05:00
Devin FoleyandPaperclip be43c23e2d fix(recovery): preserve pending native result finalization (#14219)
Preserve native coordinator ownership while accepted results await workspace
copy-back, assessment, or arbitration. Check that ownership in the terminal
update so a coordinator recorded after the liveness read is also protected.
Keep terminal task authority and exhausted-retry cleanup unchanged.

Validation: 28 focused recovery tests, local typecheck/build, 52 passing CI
checks, and Greptile 5/5 with no open findings.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-27 00:22:49 -07:00
Devin FoleyandPaperclip 5ee9e751fb fix(runner): discover assigned tools when direct catalogs exceed limits (#14218)
Keep assigned app tools accessible when the combined Runner catalog exceeds
its operation or byte limits. Reserve task and completion tools, then expose
bounded discovery and call tools for large catalogs. Fetch oversized schemas
in reauthorized chunks without blocking later search results.

Retain task ownership, work-mode restrictions, pinned assignments, current
gateway authorization, approvals, and audit. Small catalogs stay direct.

Validation: 46 focused server tests, two Runner capacity tests, local
typecheck/build, 52 passing CI checks, and Greptile 5/5 with no open findings.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-27 00:22:16 -07:00
DottaandPaperclip 640dee1802 fix(runner): acquire reusable leases when switching to warm sessions (#14187)
## Thinking Path
> - Paperclip coordinates agent work across task turns.
> - Warm runners need reusable sandbox leases.
> - Lease acquisition read only the environment reuse setting.
> - Startup then read the agent's new warm setting and rejected the
ephemeral lease.
> - This fix requests reuse before acquisition, with provider checks
intact.

## Linked Issues or Issue Description
**What happened?**
Switching an existing native runner task from per-turn to warm fails
with `runner_warm_environment_requires_reusable_lease`. Environment
`reuseLease` defaults to false. Acquisition therefore creates an
ephemeral lease, and capability narrowing correctly denies reuse.

**Expected behavior**
Changing lifecycle between turns requests a compatible lease without
replacing the task workspace or changing the shared environment.

**Steps to reproduce**
1. Start a native runner task with inherited lifecycle and environment
reuse disabled.
2. Change the agent from per-turn to warm.
3. Continue the task on a reusable-capable sandbox provider.

**Paperclip version or commit**
Reproduced on `a6c4e7a`. Related: #12904 established warm workspace
continuity; this fixes the earlier lease-selection mismatch.

## What Changed
- Pass agent settings into acquisition and derive a run-scoped reuse
request.
- Respect explicit environment lifecycle overrides; leave other adapters
unchanged.
- Keep provider capability, ownership, cleanup and restore gates intact.
- Add orchestration and database-backed transition tests; document the
contract.

## Verification
- Red: three new assertions failed before the fix.
- Green: 156 tests pass in environment-run-orchestrator,
environment-runtime, native-sandbox-lifecycle, and
environment-execution-target-capabilities.
- Database-backed regression proves ephemeral → reusable → resumed
lease, same workspace, and unchanged stored environment. Provider RPCs
are mocked, not live Daytona.
- `git diff --check` passes.
- Latest-head CI passes typecheck, build, tests, runner checks, E2E, and
canary dry run. Greptile: 5/5; no open threads.
- Recovery uses the persisted lifecycle. Direct regression: red before,
six cases green after. CI runs the added Vitest cases.
- Full local checks were unavailable (missing dependencies and earlier
memory limits).

## Risks
Warm mode now requests retained resources despite an environment's
default `reuseLease:false`. Existing idle timeout and cleanup still
apply. Unsupported providers remain denied. No migration, credential
[REDACTED], or shared-environment mutation.

## Model Used
OpenAI Codex agent; reasoning, code editing and test execution. Exact
model ID and context size were not exposed to this run.

## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details available)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links
- [x] My branch name describes the change and contains no internal
Paperclip ticket [REDACTED] or instance-derived details
- [x] I have run targeted tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-26 21:23:33 -05:00
DottaandPaperclip f2ed0b65c4 fix(runner): enable API tools by default (#14186)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The native runner exposes tools for company tasks.
> - API search and call tools cover operations without a dedicated tool.
> - The current default hides these tools unless an operator sets an
environment variable.
> - This pull request enables the tools when that variable is absent.
> - Operators can still disable the tools or restrict them to selected
companies.

## Linked Issues or Issue Description

Refs #13003, which added the guarded API tools.

**What happened?**
The native runner does not advertise `search_api` or `call_api` with the
default server configuration.

**Expected behavior**
The tools are available without a special environment variable. Existing
authorization checks still apply.

**Steps to reproduce**
Remove `PAPERCLIP_RUNNER_API_TOOLS_ENABLED` and
`PAPERCLIP_RUNNER_API_TOOLS_COMPANY_IDS`. Create a normal runner
authority. Inspect its tool definitions.

**Paperclip version or commit**
a6c4e7a.

## What Changed

- Enable API tools when the server flag is absent.
- Keep explicit disable, invalid-value rejection, company restrictions,
and binding restrictions.
- Test default tool definitions and run HTTP integration tests without
the enabling flag.
- Update operator and hiring documentation. The existing shared gate
also controls `hire_agent`.
- Use the same policy in the E2E evidence summary so an unset flag is
not reported as disabled.
- Isolate default-availability tests from operator environment
variables.

## Verification

- Red test: two new rollout policy assertions failed before the fix.
- Focused policy, authority, and HTTP tests: 3 files passed; 41 tests
passed, 2 skipped. The two runnerd transport cases require a Rust-built
binary absent from this workspace.
- Command: `pnpm exec vitest run
server/src/services/native-runtime/runner-api-rollout.test.ts
server/src/services/native-runtime/paperclip-runner-tool-authority.test.ts
server/src/services/native-runtime/runner-api.integration.test.ts`.
- Attempted `pnpm -r typecheck`: blocked by missing `cargo` in this
workspace.
- Attempted `pnpm build`: terminated at the 4 GiB memory limit.
- Attempted `pnpm test:run`: stopped after memory pressure to run
focused tests alone.
- Repeated the 41 passing focused tests with an inherited disabled flag
and a foreign-company restriction; test isolation passed.
- `pnpm test:e2e:runner:unit`: 44 files and 543 tests passed.
- `pnpm test:e2e:runner:typecheck` exceeded the workspace memory limit,
including a retry with bounded Go memory settings.
- CI results will be recorded before handoff.

## Risks

- More native runs can discover API tools and the existing `hire_agent`
tool by default.
- The change does not remove company, run, mode, credential, lifecycle,
or approval checks.
- Explicit operator restrictions still take precedence. No database
migration is required.

## Model Used

OpenAI Codex agent. The runtime does not expose the exact model ID or
context-window size. Used reasoning, repository tools, shell execution,
and tests.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-26 21:16:33 -05:00
a6c4e7a8d1 fix(heartbeat): cancel obsolete execution continuations before dispatch (#13761)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - A comment can queue execution before the task changes state.
> - The task may be completed, cancelled, deleted, or reassigned before
setup checks ownership.
> - The ownership guard must prevent that obsolete execution from
starting.
> - This pull request records that guard outcome as cancellation and
settles the wake request.
> - Missing history and authorization errors remain failures that
operators can investigate.

## Linked Issues or Issue Description

**What happened?**

A comment can queue a run just before a user marks its task Done. The
ownership guard stops setup before adapter dispatch, but the run becomes
a setup failure and can put the agent into an error state.

**Expected behavior**

Cancel obsolete work when the typed task ownership guard rejects it.
Keep genuine setup errors visible. Do not change the task's terminal
state or current owner.

**Steps to reproduce**

Queue a task continuation, then mark the task Done or Cancelled before
continuation setup. The run should settle as cancelled without invoking
the adapter. Deleting or reassigning the task must also prevent
dispatch. Missing source history or explicit user authorization must
still fail.

**Paperclip version or commit**

Refreshed against master `b2e9e82f053af14777fd68e8834fff6bb64c84c4`.
Related: #13546 covers lost issue-lock claims, and #13888 covers
persisted continuation decisions and retry budgets. This change handles
the final continuation ownership guard during setup.

Thanks to @MrBlackTongue for the original typed cancellation
implementation and regression coverage. This update preserves the
contributor's commits, resolves the master conflict, and narrows
cancellation to task ownership invalidation.

## What Changed

- Add a typed `StaleExecutionContinuationError` for
`continuation_task_ownership_changed`.
- Use existing cancellation settlement for that typed guard, including
the run, wake request, issue execution ownership, and agent state.
Suppress immediate recovery of obsolete work.
- Preserve failure classification for missing source context, missing
user authorization, and untyped errors, including an untyped error with
identical text.
- Cover Done and Cancelled tasks with real database checks. Retain
company and assignment guard coverage. Verify cancellation performs no
adapter dispatch or automatic replay.
- Document the cancellation event and its error classification in the
run-log guide.

## Verification

Current head: `b875486e81`.

- `pnpm -r typecheck`: passed.
- `pnpm build`: passed.
- `pnpm exec vitest run
server/src/__tests__/heartbeat-process-recovery.test.ts
server/src/services/execution-continuation.test.ts --maxWorkers=1
--no-file-parallelism -t 'authorized continuation context|obsolete
continuation setup|untyped continuation setup failures'`: 17 passed; 316
unrelated tests excluded by the name filter.
- `pnpm test:run`: local full run is still finishing. It is not a clean
pass: the observed failures are three company-skills cache checks
(`EACCES` while renaming read-only cache directories on macOS), one
tool-access assertion, and one comment-wake timeout. The three cache
failures reproduce in isolation in unchanged code. The tool-access check
passes in isolation, and the entire comment-wake file passes (29 tests),
including explicit feedback after completion. Final full-run totals will
be added when available.
-
[CI](https://github.com/paperclipai/paperclip/actions/runs/36282532652):
all gates green on this head. An unchanged Runner workspace-diff test
initially returned no diff; an isolated check passed, and the failed job
plus its dependent gate passed on retry with no code change. Original
failed job: 108517010989.
- Greptile: 5/5 on the full current head, with no unresolved review
threads. Branch is current with master and conflict-free.
- Diff check and local secret/PII scan passed. No customer data or
private deployment identifiers are included.

## Risks

The classification change is limited to the typed ownership guard. A
task with missing history remains a failure rather than being treated as
an expected cancellation. Existing company, task-owner, terminal-state,
and authorization checks still prevent dispatch. Cancellation retains
its specific reason in the run log and does not grant a retry. No
schema, dependency, or API changes; no migration is required.

## Model Used

OpenAI GPT-6 via Codex assisted investigation, implementation, code
review, and test execution. The exact serving model identifier and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have described the issue in-PR following the bug issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links
- [x] My branch name describes the change and contains no internal
ticket id
- [x] I have run the focused regression tests locally and they pass;
full-suite results are tracked above
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation
- [x] I have considered and documented the risks
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Devin Foley <devin@paperclip.ing>
2026-09-26 17:43:33 -07:00
Devin FoleyandPaperclip b2e9e82f05 fix: stop remote Grok runs before continuing queued messages (#14100)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The execution service owns each run and saves messages sent while it
runs.
> - Interrupt must stop the current executor before it delivers those
messages.
> - Remote Grok commands did not register the host cancellation control.
> - A cancelled task run could still write Done and prevent queue
recovery.
> - This pull request connects remote cancellation and revokes cancelled
run writes.
> - Saved input can use the existing queue admission rules after
verified cleanup.

## Linked Issues or Issue Description

**What happened?**

Interrupting a queued message marked a remote Grok run cancelled before
its sandbox stopped. The old run could still post a reply and mark the
task Done. Its saved follow-up remained deferred behind execution
recovery.

**Expected behavior**

Stop revokes run write authority and waits for verified termination.
Saved messages remain durable and enter one successor through normal
admission after cleanup.

**Steps to reproduce**

1. Run a task with `grok_local` in a remote sandbox.
2. Send a follow-up and use Interrupt while the command runs.
3. Let the old command attempt a task status update after cancellation.
4. Observe the task disposition and the saved message queue.

**Paperclip version or commit**

The gap is present in master at `d3e0f0a238`.

**Deployment mode**

Authenticated server with a Daytona sandbox.

Related work: #14028 and #14046 handle bounded continuation. #13291
covers infrastructure interruption and verified remote cleanup. #13332
addresses atomic recovery holds. This change handles direct Grok
operator cancellation and stale task writes.

## What Changed

- Register remote Grok cancellation before preparation. Keep command
ownership until the host confirms sandbox termination.
- Reuse the sandbox cancellation boundary for the direct CLI invocation.
Reject fresh attempts after cancellation and preserve workspace restore
failure evidence.
- Reject writes from cancelled task JWTs and runs with a pending stop.
Preserve diagnostic reads and existing conversation error codes.
- Recheck run authority under a database lock before task updates and
interaction responses commit.
- Preserve authorized handoffs that stop their own run. Only the
server-issued stop receipt for that request permits the final task
update.
- Add tests for hung commands, unverified stops, early cancellation,
copy-back failures, late Done, late interaction responses, authorized
handoffs, exact lease receipts, and one queue successor across
concurrent restart sweeps.
- Document the cancellation and write-authority contract.

## Verification

- Targeted adapter, cancellation-boundary, authentication,
queued-message, interaction-service, and activity-route tests passed.
The expanded run passed 214 tests; one new test had an incomplete
fixture. After correcting the fixture, all 8 selected follow-up cases
passed.
- `pnpm -r typecheck`: passed on
`179c86caf1bf0d89914a503d46e24af7e4b8c557`.
- `pnpm build`: passed on the same commit.
- `pnpm test:run`: the general-server group completed with 13,521
passed, 99 skipped, and 18 failed tests. It then stopped, so the
remaining local groups did not run. Five Slack, email, and wake-batching
failures passed on focused reruns after correcting the local
environment. The remaining 13 failures reproduce as `EACCES` on rename
in unchanged skill-cache code on macOS. Two custom-image suite setup
hooks also failed to start embedded PostgreSQL after the machine
exhausted shared-memory slots; all 31 tests in that file passed on rerun
after the local resource issue was resolved. CI covers all test groups.
- CI: 53 checks passed and 2 were skipped on the latest commit,
including the aggregate verification gate. The last server shard passed
on its single rerun after a preview-server startup timeout. The affected
file also passed locally with 28 passed and 3 skipped.
- Greptile: 5/5 on the latest commit. Both review threads are resolved.
- No live deployment or staging task mutation has been performed.

## Risks

- Stopping the sandbox can prevent file copy-back. The result preserves
workspace restore failure evidence; termination does not imply restored
files.
- If provider termination fails, the adapter keeps ownership of its
outstanding command and does not acknowledge Stop.
- The write restriction now applies to ordinary cancelled tasks. Reads
remain allowed. Task and interaction checks add a shared run-row lock to
agent mutations. An exact server-issued receipt permits the task request
that stopped its own run to complete its handoff.
- Existing terminal tasks are not reopened automatically. An operator
must correct a historical late Done before its saved queue can continue.
- No schema migration or UI change.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository inspection, code
execution, and test tools. The precise backend revision and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-26 17:07:07 -07:00
DottaandCodie 01d9a12185 fix: make keyboard shortcut enablement a personal preference (#14141)
Store keyboard shortcut enablement per user and expose it in Profile settings.

Co-Authored-By: Codie <Codie@users.noreply.github.com>
2026-09-26 11:50:04 -05:00
Devin FoleyandPaperclip d3e0f0a238 fix(server): validate CLI auth challenge IDs (#14095)
Reject malformed CLI challenge UUIDs before database access while preserving secret and authentication precedence. Document the HTTP responses and pin the supervisor test fixture to the CI-selected Node executable.

Validated by 21 focused tests, root typecheck/build, and full CI. Mac broad-suite baseline limitations are documented in the PR. Greptile 5/5.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-26 00:48:14 -07:00
Devin FoleyandPaperclip 1e3a148f46 fix(connections): retain safe broker rejection diagnostics (#14098)
Preserve fixed broker rejection reason codes within strict size/time limits while retaining public error codes and status. Unknown bodies remain generic; no raw response or credential material enters the error.

Validated by 33 consumer tests, a synthetic producer HTTP contract fixture, root typecheck/build, and full CI. Greptile 5/5.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-26 00:47:47 -07:00
Devin FoleyandPaperclip ffa32373bc fix(runtime): honor provider acquisition timeout defaults (#14097)
Declare provider acquisition budgets so slow Daytona creation does not hit the host’s 30-second fallback. Bound creation and setup to one deadline and preserve scoped cleanup ownership after timeout. Legacy drivers retain their original call shape.

Validated by 238 provider/manifest tests, focused database and heartbeat regressions, root and standalone provider typecheck/build, and full CI. Local broad tests also expose recorded Mac baseline limitations. Greptile 5/5.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-26 00:47:21 -07:00
Devin FoleyandPaperclip c3ddb288b1 fix: validate work product execution workspace references (#14063)
Reject invalid and cross-company execution workspace references with a useful
422 before changing the work product. Hold the validated reference through
the transaction so concurrent deletion cannot turn validation into a 500.

Verified 13 focused tests, full typecheck/build, and green PR CI. Greptile
5/5 with no unresolved comments. The known UI copy-toast flake passed on an
unchanged-commit retry and in a focused local run.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-25 17:06:30 -07:00
DottaandPaperclip 96bf004a79 fix: use persisted state for lifecycle continuation and retry budgets (#13888)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Its control plane decides when a task can continue, wait, stop, or
complete.
> - Legacy continuation could change when an agent changed its wording
without changing task state.
> - Shared attempt counts also let repair and infrastructure retries
affect each other's limits.
> - This pull request uses persisted state and separate, bounded
allowances for these decisions.
> - If automatic repair stops, the task explains what happened and
offers a guarded retry.
> - Paired tests and real-provider evaluations verify that Stop,
approvals, ownership, and spending limits remain authoritative.

## Linked Issues or Issue Description

Related work: Refs #13761, Refs #11126, Refs #13610. These cover
obsolete continuation dispatch and retry storms. Open and closed issues
and PRs were searched for related lifecycle, continuation, and retry
work.

**What happened?**
Legacy continuation depended on English wording and progress heuristics.
Repair, failure retry, and productive continuation could consume shared
counts. When bounded repair stopped, the task showed a technical
recovery message without a clear next action.

**Expected behavior**
Persisted disposition and owned execution paths determine the next
action. Missing disposition prompts bounded agent repair. Explicit work
mode determines planning mode. Narrative changes and raw activity counts
cannot replenish allowances. An exhausted repair shows a readable
notice. An explicit retry checks current controls and preserves the
assigned agent.

**Steps to reproduce**
Run `pnpm test:lifecycle-baseline`. The paired probes keep structured
state constant while varying completion, planning, blocker, and progress
prose. Run the explicit `lifecycle-baseline` and
`continuation-accounting` Product E2E suites for real-provider coverage.
In Storybook, open **Design previews / Recovery notice** to inspect the
production component's normal, pending, acknowledged, unavailable,
failure, and mobile states.

## What Changed

- Hide the image attachment button, icon, and drop/paste hint in answer
composers. Image paste and drop support remains available.
- Merge current master and retain both browser regression sets. Use a
production-stamped service worker in the offline recovery browser
fixture.
- Share one state-based legacy continuation decision across immediate,
delayed, and recovered dispatch. Bind bounded repairs to their source
run and episode.
- Remove title and description wording from work-mode authority. Agents
can still write requested plans in execution mode.
- Persist separate failure-retry and productive-continuation counters.
Disposition repair and resource waits cannot consume or reset those
allowances.
- Validate delayed repair identity, then recheck current gates before
provider dispatch. Fence native startup cancellation.
- Show **Agent needs attention**, a plain-language explanation, **Retry
agent**, and expandable details in both task interfaces. Report request
progress, acknowledgement, and errors inline.
- Store typed recovery notice metadata. Recognize older active notices
only through exact stored action and run IDs. Notice text never grants
retry authority.
- Use the existing recovery-action endpoint for retry. Recheck current
action, status, owner, agent availability, dependencies, active runs,
pending questions and confirmations, approvals, pause controls, and
budget. Duplicate requests do not wake twice.
- Add component, page, route, database, contract, and Storybook
coverage. Keep the scenario inventory and executable evals here.
Historical reports and snapshots live in the [commit-pinned
paperclip-evals
archive](https://github.com/paperclipai/paperclip-evals/blob/ce3e5afcd4a1184650f586a2b5b8be5874c66c8b/experiments/2026-09-lifecycle-authority/README.md).
- Preserve unsaved project fields while the same project URL changes to
its canonical alias. Do not reuse data across projects or companies.
This separate fix addresses the repeated repository-editor browser
failure without changing the browser test.
- Keep the development service worker from intercepting Vite module
reloads. Update the connection-intent browser fixture to record progress
and completion through the agent API.

## Verification

Merge preparation on September 25, commit
`c1e8e4b7ddd9fbc4913ed55ce21b8e12906c2f97`:

- Merged master `bd2030932` and resolved the browser test-list conflict
by keeping both sets of regressions.
- Deterministic lifecycle baseline: 1,090/1,090 assertions passed; no
failures, skips, or missing selected evidence. Unit 423, runner 184,
database integration 397, grading 86.
- Browser support: 17/17 passed. The offline recovery test first failed
with an unstamped development worker, then passed with the production
stamp. Its assertions are unchanged.
- Focused interaction UI and offline fallback tests: 19/19 passed.
Verified the custom-answer composer in Storybook: no attachment controls
or hint; entering an answer enables Next.
- Recursive typecheck, production build, token gates, and diff checks
passed. The worktree is clean. No new real-provider campaign was run.
- Current CI and review: [Current PR CI
passed](https://github.com/paperclipai/paperclip/actions/runs/36166011243):
55 successful checks and two optional Storybook skips. Greptile scored
this exact commit 5/5. Hiding the question attachment controls is an
intentional UI change; paste/drop remains available.

Earlier recovery UI verification, commit
`21be0fec0e90e86b6d662b8ee4831847cd041cdb`:

- Recursive typecheck, production build, token gates, and diff checks
passed.
- Focused UI coverage: 338 tests passed across six suites (336 before
the interaction guard, with the two affected suites rerun at 149 passed
after it). Covers both task interfaces, the real page mutation,
pending/error acknowledgement, stale state, and unavailable controls.
- Recovery database integration: 352 tests passed before the interaction
guard. The complete recovery-action and mutation-route suites passed 181
tests after it. The two new pending question/confirmation regressions
failed before the fix and passed afterward, including
resolved-interaction controls. Shared validator suite: 31 passed. E2E
catalog suites: 34 passed.
- Browser inspection passed for light/dark themes, mobile layout,
expandable details, pending retry, acknowledgement, failure, and
disabled retry. Storybook renders the production component; its request
is simulated.
- The broad local run hit two chat callback-order wait failures and was
stopped after all CI unit/database/runner shards passed. Both local
failures passed when rerun without the competing full-suite process.
- CI exposed a repeated project-repository draft-loss race during
canonical redirects. A new unit regression failed before the fix; all
nine project-page tests now pass, including controls for other projects
and companies. Both unchanged repository browser tests passed against a
fresh local server. UI typecheck, production UI build, and token gates
passed after this fix.
- [Earlier PR CI
passed](https://github.com/paperclipai/paperclip/actions/runs/36072486798)
on `21be0fec0e90e86b6d662b8ee4831847cd041cdb`: 55 successful checks, two
optional Storybook skips, and no failed or pending checks. The
repository browser shard passed with the production fix. Greptile is 5/5
on this exact commit with no unresolved review threads. The PR is
mergeable.

Historical, source-qualified lifecycle evidence:

- Lifecycle baseline: 1,074 assertions. Native session coverage: 447
tests. Product E2E support: 515 tests. Browser support: 11 tests. Full
earlier verification is retained in the archive.
- [Real-provider campaign: 8/8 passed, zero
retries](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35881382080-1/index.html),
source `e88d210417280140b44a36449027290adcb1aeaa`. Evidence and cleanup
checks passed. This includes deliberately exhausted repair cases that
correctly remain blocked; it does not mean every task finished Done.
This campaign predates the recovery UI change.
- Archive migration verified all 16 original JSON files byte-for-byte
and all 24 checksum entries. App tests do not need private archive
access. [Archive PR
#27](https://github.com/paperclipai/paperclip-evals/pull/27) is merged.

## Risks

- Agents that omit durable disposition receive at most two repair
attempts by default. Prose-only completion exposes missing state rather
than silently changing scheduling.
- A retry is an explicit board action. The server rechecks current
controls. A successful response confirms the task returned to To do; it
does not claim that the provider has already started.
- Existing notice metadata remains valid. Only older active notices with
matching structured evidence receive the new UI. Historical notices
without that evidence keep their existing rendering. No schema migration
is required.
- Old run records require conservative retry accounting. Tests cover old
counters, alternating retry lanes, restarts, and exhausted repairs.
- Historical snapshots require private `paperclip-evals` access. The app
index retains public campaign links. Live campaigns qualify specific
sources and scenarios; no new real-provider campaign has run for the
recovery UI commit.

> This fixes existing lifecycle and recovery behavior and does not
duplicate planned core work.

## Model Used

OpenAI GPT-6 through Codex assisted implementation, reasoning, code
execution, and review. The exact serving model ID and context window are
not exposed in this task. Historical real-provider evaluations used
Codex model `gpt-5.6-sol`, separately from the implementation assistant.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-25 15:28:11 -07:00
Devin FoleyandPaperclip 6bc830b62b fix(recovery): escalate an issue whose interrupted run has spent its retry budget (#14046)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The stranded-issue sweeper is what puts an assigned issue back on a
live path when its run dies.
> - A failed or interrupted run gets a bounded transient retry; when
that budget is spent, the retry scheduler queues nothing and reports the
exhaustion.
> - The sweeper treated that "nothing queued" like every other "nothing
queued" and skipped the issue, on every tick, forever: `in_progress`, no
run, no path, no notice.
> - Three server restarts in a row (a deploy storm) are enough to spend
the budget, so this is reachable in ordinary operation.
> - This pull request makes the sweeper escalate a spent budget to a
board-owned recovery action, the same visible `blocked` state its other
dead ends use.
> - The benefit is that no issue can sit assigned and silent after its
retries run out.

## Linked Issues or Issue Description

No existing issue found. Related work: #14028 (this branch's
predecessor: queue drain-time wakes, retry interrupted corrective runs)
fixed two neighbouring gaps but not this one.

**What happened?**

A routine-created issue's run was interrupted by a graceful server
shutdown, and both bounded transient retries were interrupted by the
next two shutdowns. The retry scheduler logged "Bounded retry exhausted
after 2 scheduled attempts; no further automatic retry will be queued"
and stopped. On every tick after that, `reconcileStrandedAssignedIssues`
reached the generic continuation lane, `enqueueStrandedIssueRecovery`
delegated to `scheduleRecoveryRetry`, which returned null (exhausted),
and the sweeper counted the issue as `skipped`. The issue stayed
`in_progress` with no run, no recovery action, and no comment for hours
until a person woke the agent by hand. The same hole exists in the
assigned-`todo` dispatch lane.

**Expected behavior**

When the transient retry budget for an interrupted or failed run is
spent, the sweeper should treat it as the dead end it is: escalate the
issue to `blocked` with a board-owned recovery action and a notice,
exactly as it does when a continuation retry chain or an assignment
retry chain is exhausted. Leaving the budget spent across restarts is
correct and unchanged; leaving the issue silent is not.

**Steps to reproduce**

1. Assign an agent using a conversation adapter (its interrupted runs
carry `conversationContinuation`, so legacy reconciliation does not
terminalize them) an issue and let a run start.
2. Interrupt the run with a graceful shutdown, let the transient retry
start, interrupt it, and repeat once more so `scheduledRetryAttempt`
reaches 2.
3. Run the stranded-issue sweep. Before this change: `skipped` every
tick, issue `in_progress`, no run, no recovery action. After:
`escalated`, issue `blocked`, one active board-owned
`issue_recovery_actions` row, a "No live execution path" notice.

**Paperclip version or commit**

master at bd6caf51bb (2026-09-25).

**Deployment mode**

Managed cloud instance restarted by fleet deploys; the code path is the
same for any operator whose server restarts more often than the retry
budget allows.

## What Changed

- `recoveryService` gains an optional `transientRetryBudgetSpent(run)`
dependency; `heartbeatService` wires it as
`executionFailureRetryCount(run) >=
BOUNDED_TRANSIENT_HEARTBEAT_RETRY_MAX_ATTEMPTS`, the same check
`scheduleBoundedRetryForRun` applies.
- `enqueueStrandedIssueRecovery` takes an optional `outcome`
out-parameter and sets `retryExhausted` when the failed predecessor's
retry returned nothing **because** the budget is spent. A null return
without it still means another authority owns the run (native runtime,
legacy reconciliation) and the caller leaves it alone; the deliberate
"failure recovery cannot fall through into the continuation queue" rule
is unchanged.
- The generic `in_progress` continuation lane and the assigned-`todo`
dispatch lane escalate on `retryExhausted` via
`escalateStrandedAssignedIssue` with a "No live execution path" notice
(danger tone), which creates the board-owned source-scoped recovery
action and moves the issue to `blocked`. Every other
`enqueueStrandedIssueRecovery` caller is unchanged.
- Tests (`heartbeat-process-recovery.test.ts`): an `in_progress` issue
with a spent budget escalates (blocked, one active board-owned action,
notice, no successor run, idempotent on the next sweep); an interrupted
run with budget remaining still gets its transient retry; an assigned
`todo` issue with a spent dispatch budget escalates. The existing guard
"does not reset an exhausted incident budget on server restart" keeps
its no-successor-run assertion and now expects the board escalation
instead of nothing, with a comment on why.

## Verification

```
pnpm -r --filter './packages/**' build
cd server
npx vitest run src/__tests__/heartbeat-process-recovery.test.ts \
  src/__tests__/issue-recovery-actions.test.ts \
  src/services/recovery/successful-run-handoff.test.ts \
  src/__tests__/heartbeat-task-drain-admission-release.test.ts \
  src/__tests__/attention-service.test.ts \
  src/__tests__/heartbeat-comment-wake-batching.test.ts
npx tsc --noEmit -p tsconfig.json
```

Live reproduction: the exact stuck state (issue `in_progress`, latest
run `interrupted` with `scheduledRetryAttempt` 2, the "Bounded retry
exhausted" lifecycle event, sweeper `skipped` every tick) was observed
on a managed instance running current master before this change was
written.

## Risks

- Behavior change is limited to runs whose transient budget is already
spent, which previously produced no action at all. Nothing new is
retried; the change only adds the escalation, so no retry loop can be
introduced.
- The board-owned action spawns no run. Resolving it (restore to the
owner) re-dispatches through the existing recovery-action routes, the
same flow as every other stranded escalation.
- Native-runtime and legacy-reconciliation predecessors are untouched:
they return before the exhaustion check.

## Model Used

Claude (Anthropic) — `claude-fable-5-1`, extended thinking, tool use
(Claude Code CLI). The incident diagnosis and the choice to escalate
rather than re-dispatch were steered by the maintainer.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-25 13:17:48 -07:00
Devin FoleyandPaperclip bd6caf51bb fix: preserve restore failure results and stop unsafe retries (#14035)
Preserve agent output and earlier execution errors when workspace restore fails. Report the restore phase and confirmed saved-plan links. Require verified repair before retrying unsafe archives, while preserving approval states and the retry budget.

Verified with full CI, 506 focused regression tests, and Greptile 5/5 with all review threads resolved.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-25 11:34:31 -07:00
Devin FoleyandPaperclip 5b09d66183 fix(heartbeat): keep drain-time wakes queued and retry interrupted disposition handoffs (#14028)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The server's heartbeat scheduler queues, admits, and recovers agent
runs.
> - A control plane arms an instance task drain before it restarts the
process, so no new run starts mid-restart.
> - Two recovery paths treated that transient hold as a permanent
verdict: a wake that arrived during a drain was written as `skipped` and
never replayed, and a corrective disposition run that the restart
interrupted counted as an exhausted attempt, so the issue went to
`blocked`.
> - Both outcomes leave an assigned issue with no run and no path until
a person notices.
> - This pull request keeps drain-time wakes in the durable queue and
gives an interrupted corrective run the same bounded transient retry any
interrupted run gets.
> - The benefit is that a graceful restart never strands or blocks an
issue by itself.

## Linked Issues or Issue Description

No existing issue found. I searched open PRs that touch the task drain
(#13527 adds opt-in termination of active runs; #13965 handles deferred
wakes after a stale cancellation). Neither replays drain-time wakes or
changes the corrective-handoff escalation.

**What happened?**

1. An operator accepted a plan confirmation while the instance was in a
task drain (a control plane armed it before a deploy restart).
`enqueueWakeup` wrote the assignee's continuation wake with `status:
"skipped"` and `heartbeatSkip.reason: "task_drain"`. The drain is
process-local and cleared on the restart seconds later, but nothing
replays skipped wakes. The issue stayed in `todo` with no run for 20
hours until a person retried it by hand. The stranded-issue sweeper did
not help: a `todo` issue whose latest run succeeded is treated as
deliberately handed back.
2. On another issue, the watchdog correctly raised
`successful_run_missing_state` and queued the single corrective handoff
run. A graceful server shutdown (the same deploy restart) interrupted
that run with `server_shutdown_interrupted`. On the next boot,
`reconcileStrandedAssignedIssues` saw a corrective run at
`handoffAttempt >= maxHandoffAttempts` and escalated the issue to
`blocked` with a board-owned `missing_disposition` recovery action. The
agent never got to finish one corrective attempt.

**Expected behavior**

A task drain holds admission, not the request. A wake that arrives
during a drain must run once the drain lifts or the process restarts. An
interrupted corrective run is not evidence that the agent could not
choose a disposition; it should be retried like any interrupted run, and
escalate only when that retry budget is spent or a finished attempt
still leaves no disposition.

**Steps to reproduce**

1. Assign an agent an issue and `POST /api/instance/task-drain` with a
Cloud control assertion (or call `startTaskDrain({})` in a test).
2. Trigger any wake for that agent (assignment, comment, or an accepted
`request_confirmation`).
3. Observe the `agent_wakeup_requests` row: `status = skipped`, `reason
= heartbeat.scheduling_suppressed`, `payload.heartbeatSkip.reason =
task_drain`. Stop the drain or restart: no run is ever created for that
wake.
4. For the second case: let a `finish_successful_run_handoff` run be
interrupted by SIGTERM during a graceful shutdown, then boot. The issue
moves to `blocked` on the startup sweep without any retry.

**Paperclip version or commit**

master at e2f1a66aa7 (2026-09-25).

**Deployment mode**

Managed cloud instance behind a control plane that arms task drains
before deploy restarts; the code paths are the same for any operator who
uses the drain endpoint.

## What Changed

- `enqueueWakeup` no longer writes a wake as `skipped` when the only
suppression is `task_drain`. The wake and its queued run land in the
durable queue; `startNextQueuedRunForAgent` and `executeRun` already
refuse to admit work while the drain is active, and the boot-time
`resumeQueuedRuns` pass picks it up after the restart.
`worktree_instance` and `database_restore_in_progress` suppression still
write `skipped`: those holds are not transient restarts.
- `reconcileStrandedAssignedIssues`: when the latest run is a corrective
successful-run handoff at its attempt cap **and** that run is
`interrupted`, the sweeper schedules the bounded transient retry (via
`enqueueStrandedIssueRecovery` → `scheduleRecoveryRetry`) instead of
escalating. The retry keeps the handoff context, so it is still the
corrective run. If the retry budget is spent, or a finished attempt
still leaves no disposition, escalation proceeds exactly as before.
Failed corrective runs (adapter or provider failures) are unchanged:
they still escalate immediately.
- New sweep counter `successfulRunHandoffRetried`, included in the
startup and periodic recovery log lines.
- Tests: `heartbeat-task-drain-admission-release.test.ts` gains a case
that a wake during a drain is queued, held while the drain is active
(quiescence still reports true), and runs to completion once the drain
lifts. `heartbeat-process-recovery.test.ts` gains two cases: an
interrupted corrective run is retried with its handoff context and the
issue stays `in_progress`; an interrupted corrective run with a spent
transient budget escalates to `blocked` with the usual recovery action
evidence.

## Verification

```
pnpm -r --filter './packages/**' build
cd server
npx vitest run src/__tests__/heartbeat-process-recovery.test.ts            # 296 passed
npx vitest run src/__tests__/heartbeat-task-drain-admission-release.test.ts \
  src/__tests__/heartbeat-worktree-suppression.test.ts \
  src/__tests__/heartbeat-scheduling-suppression.test.ts \
  src/__tests__/heartbeat-task-drain.test.ts \
  src/__tests__/instance-settings-routes.test.ts \
  src/services/recovery/successful-run-handoff.test.ts \
  src/__tests__/issue-recovery-actions.test.ts \
  src/__tests__/heartbeat-comment-wake-batching.test.ts \
  src/__tests__/attention-service.test.ts                                    # 226 passed
npx tsc --noEmit -p tsconfig.json                                           # clean
```

Live reproduction of the first defect: after this change was written, a
board wake issued against a draining managed instance (running current
master) was again recorded as `skipped` with `heartbeatSkip.reason =
task_drain`, which is the exact row the new test asserts no longer
appears.

## Risks

- A wake that lands during a drain now waits in the queue instead of
being dropped. `getTaskDrainStatus().pendingWakes` counts only in-flight
enqueue promises, so quiescence is unchanged and a control plane waiting
for the drain is not held longer. If a drain is stopped without a
restart, the queued run starts on the next scheduler tick.
- The interrupted-handoff retry reuses the existing transient retry
budget (two attempts) and the existing `hasActiveExecutionPath` skip, so
a sweep cannot double-schedule. Native-runtime corrective runs still
return `null` from the recovery enqueue and escalate as before.
- Self-hosted instances that never call the drain endpoint see no change
on the first path; the second path only changes behavior for corrective
runs interrupted by a graceful shutdown.

## Model Used

Claude (Anthropic) — `claude-fable-5-1`, extended thinking, tool use
(Claude Code CLI). Human-directed: diagnosis of the two production
incidents, the fix design, and review were steered by the maintainer.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-25 11:28:54 -07:00
Devin FoleyandPaperclip e2f1a66aa7 feat: support default-hidden experimental settings (#13980)
## Thinking Path

> - Paperclip is the open source control plane for AI-agent companies.
> - Operators can hide settings that their users must not change.
> - The server shares the effective restrictions with the UI and
settings API.
> - An explicit list must change whenever a new experimental flag is
added.
> - This pull request adds a wildcard with named exceptions to the
existing setting.
> - Core applies the policy to its current catalog, so new flags stay
hidden automatically.

## Linked Issues or Issue Description

Related: #11823 introduced settings visibility. #13907 added workspace
isolation visibility. I searched related PRs and issues and found no
duplicate wildcard implementation.

**What existing behavior does this improve?**

Operator control of experimental setting visibility through
`PAPERCLIP_HIDDEN_SETTINGS`.

**Subsystem affected**

Shared settings policy and its existing server health and mutation
consumers.

**Current behavior**

Operators must name every hidden experimental toggle. A new Core flag
can become visible until the operator updates that list.

**Proposed behavior**

`instance.experimental.*` hides current and future experimental toggles.
Entries such as `!instance.experimental.enableEnvironments` leave named
controls available. Explicit hidden keys and the hidden parent page take
precedence over exceptions.

**Reason and benefit**

Operators can maintain a short list of allowed controls instead of a
second copy of Core's full feature catalog.

**Breaking changes**

Existing explicit lists and unset configuration keep their behavior. The
new syntax is opt-in. Older images ignore it, so operators must retain
explicit restrictions until those images are upgraded. Visibility does
not change feature values.

## What Changed

- Expand the wildcard into concrete catalog keys in the shared parser.
- Limit exceptions to known experimental controls and preserve explicit
restrictions in either input order.
- Test a synthetic future catalog addition, duplicate and invalid
entries, API rejection, same-value echoes, and the effective health
payload.
- Document the syntax and the transition for deployments with mixed
image versions.

## Verification

- Targeted parser, future-catalog, health, and settings-route tests
pass: 95 tests across four files.
- `pnpm -r typecheck` passes, including Rust checks, with the installed
Cargo directory on PATH.
- `pnpm build` passes.
- All current-head CI gates pass, including the full test shards, Rust,
build, browser E2E, and canary dry run:
https://github.com/paperclipai/paperclip/actions/runs/36083292578. One
unchanged runtime-exposure cold-start test passed on its first retry.
- The full local `pnpm test:run` did not pass on macOS/Node 25: the
first server group reported 13,356 passed, 18 failed, and 99 skipped,
with six failed files (including two failed suite setups). Failures were
in unchanged runtime/company skill cache, chat/email connector fixtures,
embedded-Postgres setup, and workspace cleanup tests. A standalone
filesystem probe reproduced the read-only-directory rename permission
failure. Missing connector fixture paths, database startup failures, and
two integration assertions also occurred; the remaining local groups
were not reached after this group failed. The corresponding CI lanes all
pass. These local failures are not claimed as fixed by this PR.
- Browser suites were not run because this changes the shared policy,
not UI rendering or browser workflows. Health payload and route tests
cover the shared UI/API contract.

## Risks

A malformed exception remains hidden and is reported as unknown.
Exceptions cannot override an explicit hidden toggle or parent page.
Older images ignore wildcard syntax; keep their explicit list during a
mixed-version rollout. No schema or feature-value changes are included.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, tool use, and code
execution. The exact runtime variant and context-window size are not
exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-25 08:07:50 -07:00
DottaandPaperclip 1dbffb4c04 fix(daytona): clean up sandbox allocation after create failure (#13979)
## Thinking Path

> - Paperclip manages AI agents and their execution environments.
> - The Daytona provider creates sandboxes before it returns a lease.
> - Daytona can allocate a sandbox, then time out while waiting for it
to start.
> - The SDK throws without returning the allocated sandbox handle.
> - Paperclip previously lost the resource identity and could leave a
running sandbox behind.
> - This pull request keeps an exact creation identity and deletes a
matching failed allocation.

## Linked Issues or Issue Description

Related: #13882. Searches for Daytona create cleanup and orphan fixes
found no duplicate for this path.

**What happened?**

In [Grok qualification campaign
36080870743](https://github.com/paperclipai/paperclip/actions/runs/36080870743),
Daytona sandbox creation exceeded 300 seconds. No native runner or
provider session started. The SDK threw, but an ownership-filtered
provider lookup found the allocated sandbox still running. The
disposable resource was then deleted separately. The original attempt
and its failure remain retained.

**Expected behavior**

A failed create must retain enough identity to clean up an allocated
sandbox. Cleanup must never delete another attempt's resource. If the
provider cannot confirm cleanup during sandbox lease acquisition, the
host must retain a durable cleanup record and retry the exact owned
allocation after restart.

**Steps to reproduce**

1. Have Daytona allocate a sandbox for an environment acquisition.
2. Make the SDK throw while waiting for startup, before it returns the
sandbox handle.
3. Before this fix, acquisition fails without deleting the allocated
sandbox.
4. The regression tests reproduce this failure and verify exact-owner
deletion.

**Paperclip version or commit**

Observed at `1b0551bb7c8de3c54f4bee64dbe2c88328b3645e`. The same
creation path exists on master; this fix starts from `efce9356b`.

## What Changed

- Assign each cold create a unique provider name and creation-attempt
label.
- On create failure, look up that exact name and verify every ownership
label before deletion.
- Wait for provider deletion with bounded lookup and deletion calls.
- Preserve the original create error after successful cleanup. Report
both errors and the provider name when cleanup cannot be confirmed.
- Treat a missing lookup after an uncertain create as unconfirmed
cleanup: a delayed provider request might still create the sandbox.
- Transfer only validated ownership fields through the acquire-lease RPC
error; provider exceptions and arbitrary error data are not serialized.
- Persist a pending-cleanup lease before the host retries deletion.
Reuse the existing cleanup sweep and durable spool, retaining the
original environment scope after its row is deleted.
- Fence retries by company, environment, run, provider account, unique
creation name, and every ownership label. Missing provisional-name
lookups remain unresolved.
- Journal an ownership-verified provider ID before host-driven deletion.
Its absence then confirms cleanup after a lost deletion reply or final
database update; failed journal writes block deletion.
- Keep ownership labels separate from the mutable input Daytona modifies
during create.
- Report confirmed host-side deletion accurately while still rejecting
the failed acquisition.
- Add protocol, malformed-evidence, cross-scope denial,
controller-restart, and deleted-environment regressions alongside the
original bounded cleanup tests.

## Verification

- SDK suite: 92 passed. Daytona plugin suite: 277 passed; six opt-in
live tests skipped. Environment-runtime suite: 103 passed, then all four
focused observation/recovery variants passed after adding a
foreign-observation case. SDK and provider TypeScript checks and `git
diff --check` pass.
- Regressions failed before the respective fixes: missing durable
ownership, misleading successful-cleanup message, and SDK mutation of
the ownership labels.
- Controlled real-Daytona proof at
`d2ac1c95b32ca64daf039c0427137657f351c0e0`: create a real sandbox;
inject an SDK failure and a failed immediate lookup; persist the
validated ownership envelope; have a child journal the observed provider
ID before deleting; deliberately drop its successful deletion receipt;
reconcile absence from another fresh process and confirm it with a
separate provider read. Passed, with no manual cleanup and zero model
calls. The live proof uses an owner-only envelope file; actual host
database/spool recovery is tested separately in the environment-runtime
suite.
- The preceding controlled attempt failed because the real SDK mutated
the labels object. That failure is retained. The harness deleted its
exact owned sandbox, and a fresh lookup confirmed absence. The fix
copies the SDK input labels separately from the ownership snapshot.
- [Repository
CI](https://github.com/paperclipai/paperclip/actions/runs/36095925399)
passes on this exact head, with a clean 5/5 review. The unchanged Cursor
remote-command test passed its one bounded rerun after a 10-second
timeout; the same test had already passed on the combined tree. The
original failure is retained. Earlier `199f0852` CI and
immediate-deletion proof are retained, without being promoted to the new
source.
- Combined-source verification uses temporary PR #13990 against the Grok
feature branch. No Docker or Rust build runs on the developer laptop.

## Risks

Creation failures now add up to ten seconds for lookup and fifteen
seconds for deletion confirmation. The SDK still owns the initial
creation timeout. If a request materializes only after the failed
lookup, its durable pending record remains eligible for later cleanup; a
never-observed creation name is never reported deleted from a 404. A
previously observed provider ID can be reconciled as deleted. The host
and SDK must be deployed together. Durable handoff applies to
plugin-backed sandbox lease acquisition used by heartbeat and runner
login. Probe and custom-image interactive setup retain immediate cleanup
and explicit failure reporting, but do not use this lease-recovery path.
A worker crash before delivering the failure envelope still cannot be
recovered through this mechanism. A cleanup error never returns a lease.
Successful creation behavior is unchanged except for the provider name
and ownership label.

## Model Used

OpenAI GPT-6 through Codex, with repository tools and code execution.
The exact serving identifier and context-window size are not exposed in
this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-25 10:06:57 -05:00
DottaandPaperclip c96c751f39 fix(heartbeat): claim task ownership with queued runs (#13973)
## Thinking Path

> - Paperclip manages agents and their tasks.
> - The scheduler can start several runs for one agent.
> - Each task still needs one execution owner.
> - Task creation and recovery can queue assignments for the same task.
> - The scheduler previously marked both runs running before starting
either executor.
> - This change claims the run and task ownership in one transaction.
> - Tasks with different IDs retain concurrent execution.

## Linked Issues or Issue Description

Refs #13882. Related queue work: #11830 and #13965; those address
different admission and cancellation paths.

**What happened?**

The Grok subscription Product campaign exposed two assignment runs for
one task. Task creation and periodic recovery started runs within 200
ms. Both acquired Daytona sandboxes. One failed before provider work
with `paperclip_runner_attachment_staging_not_authorized` because the
other owned the task. Fixture cleanup then found an active session. The
original failed result is retained in [campaign
36071063537](https://github.com/paperclipai/paperclip/actions/runs/36071063537).

**Expected behavior**

Only the task execution owner may start provider setup. Competing queued
work must wait. Different tasks may use the agent's available slots.

**Steps to reproduce**

1. Set an agent's concurrent run limit to two.
2. Queue task-creation and recovery assignment runs for the same task.
3. Resume the queue while holding adapter execution open.
4. Before this fix, both runs become running. Only one has the task
execution lock.

**Paperclip version or commit**

Reproduced on master `8781f06a8` and in the Grok campaign at `4196a4cd`.

## What Changed

- Use the same company-scoped task ownership gate for assignments,
direct comments, and queued comments.
- Leave competing work queued while another live run owns the task.
Transfer a terminal pointer only after the tracked executor and durable
environment leases/finalization settle, including cleanup on another
controller.
- Commit the running state and task execution owner together.
- Preserve the dedicated review path and the native replacement checkout
guard.
- Add sixteen database regressions for assignment/comment orderings,
unrelated concurrent tasks, and cross-controller cleanup fences.

## Verification

- Live subscription Product E2E planning passed on its first attempt in
both Daytona (351,197 ms) and local execution (268,883 ms), with 6/6
matchers and cleanup passing in each, on combined source
`1b0551bb7c8de3c54f4bee64dbe2c88328b3645e`
([campaign](https://github.com/paperclipai/paperclip/actions/runs/36080870743)).
The local screenshots and persisted state show the same plan revised,
the first approval rejected, the revised approval accepted, and one
completion. The full campaign remains unqualified because of a separate
startup retry and EC2 Spot interruption.
- The new same-task regression failed before the fix: two running rows
instead of one.
- All seven focused concurrency cases passed. Three comment cases failed
before the shared gate was added. Five cross-controller cleanup cases
failed before the durable gate was added. A warm-retention regression
also failed before its successful release receipt was admitted;
missing/failed receipts and a different retention policy stay blocked.
- `node ../node_modules/vitest/vitest.mjs run
src/__tests__/heartbeat-stale-queue-invalidation.test.ts` from `server`:
48 passed. Native cleanup admission adds one passing test; task-drain
release adds two. Total focused checks: 51 passed.
- `git diff --check`: passed.
- Current-head
[CI](https://github.com/paperclipai/paperclip/actions/runs/36078495577)
passes repository typecheck, tests, Rust checks, build, and browser
suites. The unrelated repository-form browser shard passed one bounded
rerun without assertion or code changes; its original navigation failure
remains retained. No local Docker or Rust build was used.
- The failed test's exact two sandboxes were checked: one was already
absent; the other was deleted after its company, environment, and run
labels matched.

## Risks

The claim transaction adds a task row lock. It follows task-before-run
lock order. A competing run stays queued until the current owner
releases the task. The change does not alter attachment authorization,
provider permissions, schemas, or public APIs. Review is 5/5 with all
threads resolved on `23c0d13b6`. All CI checks pass on that exact head.
Live Grok requalification will use the combined runner and scheduler
source.

## Model Used

OpenAI GPT-6 through Codex, with repository tools and code execution.
The exact serving identifier and context-window size are not exposed in
this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

### Live regression verification

The Daytona plan/revise/accept lifecycle in [Grok campaign
36080870743](https://github.com/paperclipai/paperclip/actions/runs/36080870743)
passed on its first attempt at combined source
`1b0551bb7c8de3c54f4bee64dbe2c88328b3645e`, in 351,197 ms, with all six
outcome matchers and cleanup passing. The earlier competing-run failure
remains retained. This is one live planning result; the complete Grok
matrix and repeated qualification are separate gates, and this campaign
also has separately retained infrastructure failures.

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-25 10:06:05 -05:00
DottaandPaperclip 8781f06a87 feat(connections): enable MCP aggregators by default (#13964)
## Thinking Path

> - Paperclip helps people manage AI agents for work.
> - Connections let those agents use external services with explicit
access rules.
> - Zapier, Arcade, Composio Connect, and Executor already have setup
and runtime support.
> - Their experimental switch still blocks discovery and setup by
default.
> - This pull request removes those gates and the Settings toggle.
> - Users can connect these providers without enabling an experiment.

## Linked Issues or Issue Description

Refs #13755. Refs #13941.

**What existing behavior does this improve?**
Apps browsing, inline setup, and agent connection search for the four
MCP aggregators.

**Current behavior**
An instance must enable the MCP aggregators experiment before users or
agents can start setup.

**Proposed behavior**
All four providers are available by default on local and managed
instances. Old stored and managed values still parse but cannot disable
them.

## What Changed

- Remove the aggregator gates from Apps, inline setup, server setup, and
agent search.
- Remove the Settings toggle and its UI hook.
- Retain the old setting key only for upgrade compatibility. Normalize
it to true and ignore managed overrides, as Apps already does.
- Replace opt-in fixtures with default-on coverage. Test old false
values, all four setup flows, provider choice, and the removed toggle.
- Update current connector guidance and remove the opt-in from the
runner acceptance fixture.

## Verification

- 306 focused tests passed across eight files: shared remote MCP
contracts; server remote MCP lifecycle, aggregator fallback, settings
normalization, and managed overlay; UI Apps browsing, setup, and
experimental settings.
- Server and UI TypeScript checks passed.
- UI token gates and `git diff --check` passed.
- The full local suite was not run, per the maintainer's instruction.
All 54 CI checks passed; two checks were skipped. One unrelated
workspace-preview readiness timeout passed on one failed-shard retry.
- The setup fixtures use simulated MCP responses. This change does not
claim new live provider acceptance.

## Risks

- Existing instances now show all four providers, even if the old flag
was false. This is intentional.
- External provider choice, credentials, company isolation, agent
grants, and tool policies still apply. Showing a connector does not
authorize an external account.
- No data migration is required. The compatibility key keeps old managed
configuration documents valid.
- Historical Zapier live acceptance remains incomplete in the existing
evidence report. The maintainer explicitly requested the default-on
rollout for all four existing providers; the report records that scoped
exception.

## Model Used

OpenAI GPT-6 (`gpt-6-astra`) through Codex, with reasoning, repository
tools, and test execution. The context window size is not exposed in
this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-24 17:31:32 -05:00
DottaandPaperclip aa8fc86331 feat(connections): prefer native apps and ask users to choose external providers (#13941)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Connections give agents governed access to external services.
> - Native connections should remain the first choice for a supported
app.
> - Other apps may be available through an external MCP provider.
> - The user must know which external provider handles the connection
and choose it before setup.
> - This pull request adds ranked alternatives and server-authored
instructions to connection search.
> - Agents can follow the returned instructions while Paperclip
validates saved choices and access.

## Linked Issues or Issue Description

Related: #13879, which fixed inline MCP provider setup. This PR adds
discovery and provider selection on top of that work.

**Subsystem affected**

Cross-cutting: shared connection contracts, server search and intent
services, native runtime, CLI, inline setup UI, and evals.

**Problem or motivation**

An agent cannot offer a clear external-provider choice when Paperclip
has no native connection for an app. Adding provider-specific branches
to the core prompt would make those instructions harder to maintain.

**Proposed solution**

Prefer a native connection. Otherwise return verified alternatives in
Composio, Arcade, Executor, Zapier order. Include an external-service
disclosure, a question with None, and the next instruction in the search
result. Validate the saved human choice before creating a selected
fallback setup card. Reuse existing provider accounts and verify
underlying app access separately.

**Alternatives considered**

Do not silently choose a provider. Do not claim that broad execution
tools prove support for every app. Reuse existing questions and
connection intents rather than add another connection model.

**Roadmap alignment**

Extends the existing MCP Tool Gateway & Apps and Agent evals & feedback
capabilities. The MCP aggregators experiment remains the gate. No
duplicate provider-routing PR was found in the public search.

## What Changed

- Add a dated support index and authorized cached-tool evidence for
external routes.
- Return provider questions and next-step instructions from
`connections_search`.
- Preserve pending choices and declines across continuation. Validate
company, task, agent, human, app, and current route eligibility.
- Carry the selected app into new setup and account reuse, validate
explicit provider requests against persisted human messages, and
distinguish provider readiness from app authorization.
- Sync native, MCP, REST, and CLI contracts. Keep core agent
instructions provider-neutral.
- Add production-component Storybooks, focused database tests, and three
real-agent browser eval cases.
- Record the plan, observed failures, fixes, passing evidence, and
acceptance limits.

## Verification

- Latest head `586f0e6cd`: 54 checks passed, 2 skipped; Greptile 5/5 and
all review threads resolved.

- After rebasing on master `18dac1e1e`: 64 focused shared, validator,
route-contract, and database tests passed; server typecheck passed.
- Embedded-browser test drive on the rebased head: native Jira card,
HubSpot external-provider question, Arcade account reuse, one actual MCP
read against a local synthetic fixture, reload persistence, and None
preventing further calls. A real OpenAI-backed agent performed discovery
and continuation.
- UX observation: the agent initially combined mutually exclusive
request fields; the server rejected it and the agent recovered without
changing access. This extra retry remains visible in the transcript.
- After rebase: 23 focused eval grader/catalog tests, affected
TypeScript checks, token gates, production UI build, and Storybook build
passed. Full local tests are intentionally excluded at the maintainer's
request.
- Before rebase: four browser/real-agent attempts passed: native Jira,
None, and reuse of the second provider on two Codex profiles.
- Browser evals used an isolated deterministic MCP fixture through the
real Paperclip gateway. They do not prove production compatibility with
all four providers.
- Review `Apps / Connections / Provider choice` in Storybook. Choose
Arcade, continue through Access, and verify the app name,
external-service disclosure, and URL configuration.
- The detailed verification report is
`doc/connections/2026-09-23-aggregator-routing-verification.md`.

## Risks

- The public support index is finite and can age. Account capability and
app authorization still require verification after selection.
- Existing installed-tool permissions remain in effect. Provider choice
is not a new execution permission boundary. An early Mini attempt
skipped search; clearer provider-neutral instructions made the targeted
rerun pass. This is not a measured reliability rate.
- Explicit requests skip provider confirmation only when a clear
persisted human message or saved provider choice supports them. Other
phrasing falls back to confirmation; the agent query alone is not
consent. These routes do not add tool permissions.
- No database migration or legacy Composio broker is added.
Real-provider acceptance remains separate from fixture proof.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, code
execution, and browser tools. The exact deployment variant and
context-window size are not exposed in this session. Product evals
separately used the repository's primary Codex and Codex Mini profiles;
those agents supplied test behavior, not independent provider
compatibility proof.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-24 16:36:49 -05:00
DottaandPaperclip 18dac1e1ef feat(connections): add experimental memory providers and remote MCP access (#13942)
## Thinking Path

> - Paperclip manages AI agents and their work.
> - Connections give agents governed access to external tools.
> - Agents need durable memory across tasks and execution environments.
> - Mem0, Zep, Supermemory, Cognee, and Honcho provide hosted memory
tools.
> - This pull request adds their setup flows behind an experimental
toggle.
> - It also delivers assigned MCP tools through the native remote Codex
runner.
> - Operators can connect a provider once and use the same governed
tools locally or in Daytona.

## Linked Issues or Issue Description

**Problem or motivation**

The Apps catalog lacks a complete set of memory providers. Remote native
Codex agents also need access to assigned managed MCP tools without
receiving provider credentials.

**Proposed solution**

Add five memory connectors behind the disabled-by-default Experimental
memory connectors setting. Use the existing connection setup and
permissions UI. Default their tools to Allowed. Preserve company
boundaries, operator permission changes, provider scopes, and audit
attribution.

**Alternatives considered**

Direct provider credentials in each sandbox would duplicate setup and
bypass the managed gateway. The remote runner instead uses its existing
protocol channel to call the gateway on the server.

**Roadmap alignment**

This maintainer-requested experiment supports the Memory / Knowledge and
Connected Apps roadmap areas. It adds provider connections without
introducing a separate memory UI. Related connector authoring
documentation is tracked in #13692; no duplicate memory-provider
implementation was found.

## What Changed

- Add provider definitions, official branding, and the experimental
setting for all five providers.
- Use OAuth for Zep and Supermemory, API credentials for Mem0 and
Honcho, and a bundled Cloud API bridge for Cognee with no runtime
downloads or subprocesses.
- Default memory tools to Allowed and classify destructive actions
explicitly.
- Fix personal remote credential resolution and propagate provider tool
errors.
- Relay assigned managed MCP tools to remote native Codex through the
runner protocol. Recheck current authority for each call and rotate
stale tool contracts.
- Add 21 Storybook states and complete OAuth walkthrough fixtures.
- Document provider research, sanitized tool inventory, and live local
and Daytona proof.

## Verification

- Passed workspace typecheck: `pnpm -r typecheck`.
- Passed production build: `pnpm build`.
- Full local `pnpm test:run`: 13,273 passed, with failures from
process/readiness timeouts under parallel load. Reran all 16 affected
suites with one worker: 456 passed, leaving two macOS `/var` versus
`/private/var` path assertions. Both passed with `TMPDIR=/private/tmp`.
No test failures remain unverified. Latest-head remote CI passes all 54
checks (two optional Storybook jobs skipped). Greptile is 5/5 with no
unresolved findings.
- Latest Cognee gateway regression: 71 passed, including public
deployment without a runtime host and immediate recovery after a
provider error. Bundled bridge tests: 25 passed.
- Browser setup and real agent tasks exercised all five providers. Mem0,
Cognee, Zep, and Honcho have successful store/retrieve proof.
- Supermemory now has scoped read/write consent. Local storage and real
Daytona write, document read, and semantic recall passed; indexing
completion was verified before claiming success.
- All five providers were exercised through a real Daytona sandbox and
its native runner MCP relay. Provider credentials remained on the
server. The final bundled Cognee bridge also passed a fresh Daytona
store/recall run. All disposable sandboxes were removed and verified
absent after testing.
- Storybook is rebuilt and contains the experimental toggle, catalog,
setup, permissions, and error states. The Zep and Supermemory access
steps advance correctly.

## Risks

- Provider OAuth scopes and plan limits remain independent of Paperclip
tool permissions. An Allowed tool can still be rejected by the provider.
- Providers can queue memory indexing; save acceptance does not prove
that semantic recall is ready.
- Remote tool contracts must stay synchronized with current connection
authority. Regression tests cover revocation and stale contracts.
- This change has no database migration. Existing connections remain
usable when the experimental catalog toggle is disabled.

## Model Used

OpenAI Codex, GPT-6 Astra (`gpt-6-astra`), with reasoning, repository
edits, code execution, and browser automation. The exact context window
size is not exposed in this session. Live acceptance agents used
`gpt-5.6-sol`.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-24 15:54:38 -05:00
Devin FoleyandPaperclip 2914501c08 test(chat): retire fixture leases before recovery sweeps (#13952)
Retire verifying chat fixtures and their synthetic endpoint leases after service shutdown. Add a regression that recovers a stale receipt while preserving another company's lease.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-24 11:10:25 -07:00
DottaandPaperclip b0155a681a feat(slack): connect Paperclip conversations and scheduled messages (#13920)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Slack conversations use the same tasks and agents as the Paperclip
board.
> - A board reply must reach that Slack conversation and let the agent
continue the work.
> - An assigned agent also needs its Slack tools during normal tasks and
scheduled routines.
> - Both paths must keep the linked user's authority, delivery rules,
and conversation history.
> - This pull request adds those paths and reduces setup friction for
Slack bots.

## Linked Issues or Issue Description

**Subsystem affected**

Server orchestration, Slack connector tools, shared contracts, and chat
setup UI.

**Problem or motivation**

Replies entered in Paperclip did not provide a complete round trip to
the linked Slack thread. Slack and board wakeups could select different
model sessions for the same task. Agents also lacked their assigned
Slack tools outside Slack-origin work, which prevented a routine from
sending its responsible user a briefing. Inviting a bot could leave the
new channel disabled.

**Proposed solution**

Mirror human board messages with author attribution and route the agent
result to the same thread. Use the same session key across both entry
points. Supply Slack tools to the connection's assigned agent in normal
tasks and routines, using the current responsible user's verified link.
Enable newly invited channels while preserving explicit disabled
choices. Add a browser-agent setup prompt to the Slack wizard.

**Alternatives considered**

A separate Slack scheduler or task dispatcher would duplicate existing
Paperclip workflows. Reusing the connection owner's identity would grant
the wrong authority. Replaying old channel history could start
unintended work. This change uses ordinary task wakeups, routine
dispatch, and Slack's original invitation mention event instead.

**Roadmap alignment**

Extends the shipped Scheduled Routines and governed Apps capabilities.
It does not add a separate task lifecycle. Related work: #13828 and
#13809. Related test stabilization: #13877. The existing plugin
Slack-control proposals are separate from this built-in connector
change.

## What Changed

- Queue human Paperclip messages for the original Slack thread with
display-name attribution and stable delivery identities. Require the
author’s current linked Slack identity and recheck access before
delivering messages or agent replies.
- Apply pause, dependency, cancellation, and closed-workspace guards
before explicit Board sends request work and again when the durable
outbox dispatches it.
- Route agent results back to Slack and preserve model-session
continuity, including replies that reopen completed tasks.
- Resolve assigned Slack connections for normal agent tasks and
routines. Recheck the responsible user's link, membership, and
permissions at execution.
- Add `slack_open_dm` for the responsible user's bot DM and request the
`im:write` scope.
- Enable newly discovered invited channels. Keep explicit OFF choices
and normal admission and deduplication rules.
- Add a copyable Slack setup prompt for a computer-use agent, with
Storybook coverage. Share the prompt-button component with GitHub.
- Update Slack tool documentation and runtime instructions.
- Stabilize the mobile project browser test by waiting for the final
canonical route before editing, preserving all persistence assertions.

## Verification

- Live staging: invited the bot after the first mention. The channel
became enabled and the bot answered that original mention.
- Live staging: a normal Paperclip reply appeared in Slack with author
attribution. The agent completed the calculation and replied once in the
original thread and in Paperclip.
- Live staging: a codeword entered in Slack was recalled from Paperclip.
A following Slack calculation used the result from the Paperclip turn.
Run metadata confirmed the same model session for both entry points.
- Live staging: a scheduled routine used `slack_open_dm` and
`slack_post_message` to deliver one DM. The existing app was reinstalled
with `im:write`. The test routine was paused after verification.
- Before the master merge: 397 focused feature tests passed. The
continuity fix passed all 76 issue comment/update route tests and six
focused route/integration cases. Typecheck, build, and token gates
passed.
- Review fixes: 47 focused integration cases passed, covering link
revocation/replacement, private membership removal, guarded outbox
dispatch, concurrent workers, lost scheduler responses, a real
one-connection pool, exact reply provenance, and attachment retries. All
99 issue-comment route tests and the Slack catalog browser test passed.
- Full local typecheck, production build, token gates, and
module-boundary checks passed. The full local test command passed 25,915
tests before a 15-second timeout in
`issue-thread-interaction-routes.test.ts`; that entire suite passed on
isolated rerun (81 tests). Remaining serialized coverage is provided by
the current-head CI shards.
- An unchanged Cursor adapter test hit its 10-second limit in CI; all
five tests in that file passed on a local rerun in 3.11 seconds, and the
failed CI shard passed on its single retry.
- The preview-server readiness test passed a local rerun (28 tests). The
mobile-project readiness fix passed three repetitions of both browser
tests (6/6).
- Final commit `64ac0d9897f4353375996f1b1b38e5040bdeb0a0`: all CI gates
passed, including all eight browser shards, all server/chat suites,
typecheck, build, runner checks, and security checks. Greptile reviewed
this exact commit at 5/5; all review threads are resolved.

## Risks

- Human messages on a Slack-linked task now publish to its Slack thread.
The task banner states this behavior. Incoming Slack messages and
internal agent bookkeeping must not echo back.
- Normal tasks and routines can now use the assigned bot. Authority
remains bound to the current responsible user's link; it does not fall
back to the connection owner. Revocation, private-context limits, and
queued-write checks still apply.
- Existing Slack apps need `im:write` and a reinstall to open DMs. Other
existing capabilities remain available without that scope.
- New invited channels default to enabled. Explicit disabled choices
remain disabled. Channels created by bot tools still require a person to
enable responses.
- No database migration or new provider credentials are required.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, code execution, GitHub CLI,
and browser tools. The runtime does not expose a more specific model
build identifier or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-24 10:51:45 -05:00
Devin FoleyandPaperclip 8ee8f1fd6e ci: retire recurring public cloud image builds (#13827)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Core publishes standard images and source verification for
downstream services.
> - Managed services can now compose private images from the signed
standard image.
> - Core still builds a second public cloud image on every master push
and release.
> - That duplicate producer consumes build capacity and retains an
obsolete readiness contract.
> - This pull request retires recurring cloud publication while
preserving the standard producer and rollback artifacts.

## Linked Issues or Issue Description

Refs #13797 and #13789. Related: #12856 changes image dependency
packaging; it does not retire this producer.

**What existing behavior does this improve?**

Core's recurring Docker publication and Cloud readiness workflow.

**Current behavior**

Master pushes call the legacy cloud publisher from Cloud readiness.
Release tags and manual Docker runs call it too. Canary promotion also
requires the legacy image.

**Proposed behavior**

Publish standard Core images and retain `Cloud source verified v1`. Let
downstream services build their managed image. Keep explicit commit
previews and existing images available.

## What Changed

- Remove `docker-cloud.yml`, its master and release callers, and its
unused cache selector.
- Remove the legacy image/migrator wait and `Cloud deployable v1` job.
Keep the full source verification workflow and exact source-proof name.
- Make canary promotion inspect and promote the standard image only.
- Preserve signed standard-image publication, direct migrator
publication, and explicit `release.yml` previews. The preview path still
uses the Dockerfile `cloud` target.
- Update workflow, preview, build-stamp, and packaging tests. Exercise
the promotion shell with mocked registry commands, including
missing-image and missing-tag cases.
- Document frozen legacy aliases, consumer requirements, preview
compatibility, and rollback retention.

## Verification

- All 377 workflow tests pass: `node --test
.github/scripts/tests/*.test.mjs`.
- All 129 release-registry tests pass: `pnpm test:release-registry`.
- Focused source-proof, standard-image, preview, and workflow tests
pass: 256 tests.
- Focused image packaging/build-stamp tests pass: 16 tests.
- Actionlint passes on all three changed workflow files. `git diff
--check` passes.
- Full local `pnpm build` and `pnpm -r typecheck` pass.
- The policy follow-up updates an old assertion that required the
removed readiness job. All 37 source-proof/release-workflow tests pass
locally.
- Full local `pnpm test:run` did not complete successfully while the Mac
ran out of disk space. No full-suite pass is claimed. Removed 1.2 GiB of
generated Cargo output from this isolated worktree with `cargo clean`.
GitHub CI passed on the final head: 52 successful checks and 2 optional
skips.
- Fresh Greptile review for `4f5fe1951f0bd7f7739cf6655d395ff78f1ed944`:
**5/5**, successful current-head check, zero review threads.
- September 23 refresh: the unchanged PR head merges cleanly with
current master `db8f8fe5b73a2697684a30261b0d306a9c631aba`. In an
isolated temporary worktree, all 377 workflow tests and 29
release/preview tests pass on the combined tree. `git diff --cached
--check` passes.
- Refreshed Actionlint workflow validation passes with ShellCheck
disabled. Full Actionlint reports the same 10 existing ShellCheck
diagnostics as master, with no added diagnostics. No source changes or
new PR commits were needed.
- The full local build/typecheck and current-head Linux CI results above
remain the verification for the unchanged PR head. They were not rerun
for this metadata-only refresh. No image publication or tenant
deployment was initiated for this refresh.

## Risks

**Deployment prerequisite satisfied (September 23):** The combined
cleanup release is deployed to staging and production, and production
Support is verified. Active managed-fleet automation uses standard-image
composition. Explicit immutable previews remain supported by the
retained preview publisher. This PR is ready for maintainer review; keep
auto-merge disabled and wait for explicit merge authorization.

- A consumer still selecting `Cloud deployable v1` will stop advancing
at the last legacy-ready commit. Confirm active automatic consumers use
the standard-image composition contract before merge.
- Legacy cloud release-channel aliases stop advancing. Standard
self-hosted aliases continue.
- This PR deletes no registry images, cache tags, migrators,
credentials, or runner infrastructure. Existing immutable releases
remain usable for rollback.
- Explicit legacy previews remain for commit-specific operator
deployments. Retiring that compatibility path requires a separate
consumer migration.
- These changes affect CI publication, not database schema or
application behavior.

## Model Used

OpenAI Codex, GPT-6. The runtime does not expose a more specific model
identifier or context-window size. Used repository inspection,
reasoning, code editing, shell tools, and test execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-23 19:35:19 -07:00
DottaandPaperclip a6448cd060 fix(runner): tolerate Codex account notifications during active turns (#13902)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The Rust runner connects task runs to Codex app-server sessions.
> - Codex sends account notifications on the connection during
authentication updates.
> - These notifications have no task, thread, or turn identity.
> - The runner treated a valid account update as an invalid task event
and stopped the run.
> - This pull request classifies account updates as connection
information while preserving task identity checks.
> - Sandbox images can also contain an older runner with compatible
metadata. The server must compare its bytes before reuse.
> - Slack conversations use the deployed runner and can complete when
Codex refreshes account state.

## Linked Issues or Issue Description

Related: #13853. That change fixes the supported Codex version range.
This fixes a separate notification failure after the version check
succeeds.

**What happened?**

An active Codex task stopped with `thread_binding_mismatch` when the
provider emitted `account/updated` without a thread ID. Slack showed
that the agent stopped before completing its turn.

**Expected behavior**

Connection-level account updates must not stop a task or acquire task
authority. Account details must not appear in task output.

**Steps to reproduce**

1. Start a task with the native Codex runner.
2. Emit `account/updated` during the turn, with `authMode` and
`planType` but no thread ID.
3. Observe that the old runner rejects the notification and fails the
turn.

**Paperclip version or commit**

Reproduced after #13853. The fix is based on `b648d8cdd`.

**Deployment mode**

Cloud staging with a remote Codex runner and a Slack chat connection.

## What Changed

- Classify `account/updated` and `account/login/completed` as connection
information.
- Reject account notifications that contain execution identity fields.
- Reuse the existing bounded diagnostic path without publishing account
payloads.
- Test both notification types during two consecutive turns. Preserve
existing identity rejection tests.
- Compare preinstalled sandbox runner bytes with the controller
artifact. Stage the deployed binary after a mismatch, failed checksum,
or timeout. Keep exact retained artifact reuse.

## Verification

- `cargo test --release --locked -p paperclip-runner-core`: passed, 582
test executions; two existing ignored cases.
- `cargo test --release --locked -p paperclip-runner-core --test
codex_provider`: 89 passed; two existing ignored cases.
- `cargo fmt --all` and `git diff --check`: passed.
- Repository-wide `pnpm -r typecheck`: passed. Server typecheck also
passed after the artifact-selection change.
- Server executor suite: 456 tests passed, including preinstalled digest
match, mismatch, checksum failure, and timeout.
- Repository-wide Vitest and build are running.
- Deployed `5648d90d5dd762c5a8b697face269f6eea9cd59d` to one staging
stack and verified its serving SHA.
- Retried the failed conversation through the Slack UI. The agent
replied and the task completed.
- Sent a fresh Slack mention asking which bots belong to the channel.
Verified successful `slack_members` and `slack_user` calls, a successful
run, and the delivered Slack answer.
- Sent a follow-up in the same Slack thread without another mention. The
agent returned the requested names from context, with one final reply.

## Risks

- Only two known connection-level notification methods change behavior.
Unknown authoritative methods and malformed or conflicting identities
still fail closed.
- An older sandbox runner now requires one upload when its bytes differ
from the controller artifact. Remote OS and architecture checks still
apply.
- No schema, credential, permission, Codex version-range, or UI changes.
The Codex minimum remains 0.149.0.

## Model Used

OpenAI GPT-6 (`gpt-6-astra`) through Codex, with reasoning, code
editing, terminal tools, and browser verification. The context-window
size is not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-23 15:43:46 -07:00
Michael Nguyen 30182c704c feat(telemetry): propose connector telemetry events behind proposal markers (#13582)
Adds connection.created / connection.updated / connection.invoked as proposed-telemetry wrappers (unregistered; the TelemetryClient cannot queue or send them) plus the server-side emitter wiring and tests. No telemetry is transmitted by this change. Proposal: #13578. Coordinated release plan and review package: PAP-24 (Step A.1).
2026-09-23 22:32:40 +00:00
Devin FoleyandPaperclip 6681c71b40 fix(ci): stabilize chat startup and close the initial live-update gap (#13895)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Browser tests check chat state across navigation, reload, and agent
runs.
> - Runtime tests check that a service is ready before Paperclip
publishes its address.
> - CI for #13891 failed in these paths, then passed attempt 3 with the
same code.
> - The browser traces stopped during development asset startup. A
separate sidebar assertion used an unstable focus path through the rich
editor.
> - This PR gives browser tests fresh built assets and a direct keyboard
path to the star button. It adds service-worker reload coverage under
CPU throttling.
> - A later browser failure exposed a real reload race: comments can
change between the first query and the first live subscription. The UI
now refreshes active queries when that subscription opens.
> - Runtime fixtures now have separate registry state and better failure
evidence. The readiness deadline remains unchanged.
> - A later CI run exposed a wall-clock backoff assertion and three
authorization cases sharing one test lifecycle. The PR anchors the
assertion to transport time and separates the cases.

## Linked Issues or Issue Description

Refs #13891. Related runtime ownership and cleanup work: #11791, #11389,
#11278.

Evidence: [original CI run, attempt
3](https://github.com/paperclipai/paperclip/actions/runs/35910335089/attempts/3).
Attempt 3 passed both affected shards without code changes. This PR is
separate from the wake-payload change, which has since merged.

| Failure | Diagnosis and classification |
| --- | --- |
| Sidebar blank page and retry text missing after reload in CI | Both
traces show a blank document before React startup. Vite connects, but
the failing page makes no application API requests. The service worker
forwards the unbundled development module graph. The retry response is
already stored and its process-adapter run succeeded before reload. This
places the failure in browser bootstrap, not wake payload or reply
persistence. The exact reason the development module graph stopped is
not established by the retained trace. The harness now serves built
assets, and the new test covers startup and controlled reload at 4x CPU
throttling. |
| Star opacity remains zero on macOS | Reproduced on unchanged master
`f55759942b`. The old test clicked the rich editor, focused the star,
then used Tab and Shift+Tab. An instrumented baseline run captured the
sequence: Tab moved from the star to the next sidebar link, then the
editor bundle called `focus()` on its contenteditable before Shift+Tab.
That key reached the composer, where it is a work-mode shortcut. This
confirms a pending editor selection update stole focus; it was not the
browser skipping sidebar buttons. The assertion did not prove the star
retained focus. The test now moves the pointer away, focuses the
preceding sidebar link, presses Tab once, and asserts both actual focus
and opacity. This is test synchronization and keyboard traversal, not a
demonstrated CSS defect. |
| Runtime readiness exceeds 10 seconds | The old error only says `fetch
failed`. There is no child startup output in that failure, so it cannot
distinguish slow process startup from a refused or stalled probe. It
passed unchanged locally and took 3.416 seconds in attempt 3. Resource
contention is plausible but unproved. This PR does not claim a proven
historical runtime root cause: it isolates fixture registry/log files,
checks ports before spawn and after stop, checks live backends at
publication, and preserves transport errors, probe count, elapsed time,
and fixture startup timestamps for the next occurrence. |
| Later CI: recovery reply disappears after reload | Product
synchronization bug, distinct from the blank-page bootstrap failure.
Reproduced locally with a trace: the comments request started at
`21:56:07.443`, the server saved the reply at `.520`, and the first live
subscription started at `.541`. The reload fetched a successful run and
its complete log, but missed the comment event. The provider refreshed
after reconnects only. It now refreshes active queries on the first
connection too. A deterministic regression test fails before the fix. No
browser assertion was changed. |
| Later CI: Slack backoff and authorization tests | The 30-second
backoff assertion required more than 25 seconds to remain when it read
the saved action. CI spent 10.311 seconds in the test, exceeding that
five-second allowance. A local six-second read delay reproduces the
failure; the new transport-anchored lower and upper bounds pass the same
fault injection. The neighbouring authorization test ran three
independent fixtures in one test and timed out at 15 seconds. Each case
now has its own fixture cleanup and the normal per-test deadline, so
earlier cases do not remain active during later global worker sweeps. No
specific production slow call was established. |
| Self-hosted runner loses communication or shuts down | The original
lost-communication failure has no assertion. On final-head [attempt
1](https://github.com/paperclipai/paperclip/actions/runs/35925901613/attempts/1),
Build, Runner Vitest 1/2, and chat 2/3 ran on three separate fleet
instances. All received a runner shutdown signal at `22:04:55 UTC`,
within 35 milliseconds, then cancellation. Server shard 6/12 received
the same shutdown signal one minute later. All four were Spot
`m7i-flex.xlarge` instances in `us-east-1a`. No test assertion or build
error preceded those stops. This is infrastructure interruption; the
reason the fleet stopped the runners is not available in job logs. |

## What Changed

- Build the browser fixture UI into the static server's preferred
directory, `server/ui-dist`, and disable Vite middleware. Ignore these
generated assets. Start the source CLI directly from the repository
root, as required by the CLI invocation safety contract.
- Keep all existing chat assertions. Add first takeover and three
service-worker-controlled reloads under CPU throttling. Assert that the
page uses built module assets and that opening it creates no chat task.
- Use forward keyboard traversal from the Zeta link to its star. Check
focus before checking the reveal style.
- Give each runtime exposure test a temporary Paperclip home and restore
environment state after process cleanup.
- Check that reserved ports are free before spawn, serve the fixture
response before exposure, and become free after stop.
- Include the nested fetch error, probe count, and elapsed time in
readiness failures. Add a unit test for this diagnostic contract.
- Refresh active queries when the first live connection opens, covering
events missed during initial page loading. Keep reconnect toast
suppression unchanged.
- Measure Slack retry timing from the transport attempt and recovery
completion. This checks the full provider-requested delay without
spending a small wall-clock allowance on unrelated processing.
- Run each Slack authorization-revocation scenario as a separate test,
with cleanup between cases. All assertions remain.
- Document the browser fixture's build and serving mode. No configured
assertion deadline, readiness deadline, retry count, or skip was added.

## Verification

- Final-head [Linux CI, attempt
2](https://github.com/paperclipai/paperclip/actions/runs/35925901613/attempts/2):
**green**. All 52 check runs passed; two conditional checks were
skipped. The legacy Snyk status also passed. No pending or failed checks
remain.
- `pnpm -r typecheck` passed.
- `pnpm build` passed again after the final UI fix.
- Focused runtime suites: 38 passed, 3 existing platform skips.
- CLI invocation safety suite: 39 passed.
- Slack timing negative control: inserting a six-second delay before
reading the saved action fails the old assertion. All four revised
authorization/backoff cases pass with that same delay. The diagnostic
delay is not committed.
- Live update suites: 92 passed, including the new regression. UI
typecheck and token gates passed.
- Recovery browser negative control: the unchanged tests reproduced the
missing reply (1 failed, 9 passed). After the first-connection fix, all
six recovery paths passed twice (12 passed). Browser assertions and
deadlines are unchanged.
- Server typecheck passed after the Slack test adjustment.
- `GITHUB_WORKFLOW=PR pnpm test:run:general -- --group general-chat
--shard-index 0 --shard-count 3`: 342 passed. The other 682 tests belong
to the remaining shards; collection verified exact coverage.
- Final direct-source CLI launch: all five chat session tests passed.
- `PAPERCLIP_E2E_PORT=32993 PAPERCLIP_PLAYWRIGHT_CHANNEL=chrome pnpm
exec playwright test --config tests/e2e/playwright.config.ts
tests/e2e/agent-chat-sessions.spec.ts --repeat-each=3`: 15 passed,
including nine controlled reloads at 4x CPU throttling.
- `GITHUB_WORKFLOW=PR pnpm test:run:general -- --group
general-server-without-chat --shard-index 10 --shard-count 12`: 55 files
passed, 867 tests passed, 21 existing skips. An earlier run could not
initialize PostgreSQL because this Mac exhausted its System V
shared-memory slots. After reclaiming the orphaned segment from this
task's stopped browser server, the full shard passed.
- On final commit `1f1fafc08d`, [browser
4/8](https://github.com/paperclipai/paperclip/actions/runs/35925901613/job/107400837947)
passed all 16 tests; [browser
8/8](https://github.com/paperclipai/paperclip/actions/runs/35925901613/job/107400838155)
passed all 23 tests; [server
11/12](https://github.com/paperclipai/paperclip/actions/runs/35925901613/job/107400838328)
passed 887 tests with one existing skip. The readiness lifecycle case
took 1.842 seconds. Chat 1/3 also passed. All four jobs interrupted by
runner shutdowns passed unchanged on their single rerun.
- Greptile reviewed `1f1fafc08d`: **5/5**, with no unresolved comments.
- Full local `pnpm test:run` was attempted and stopped after confirming
failures outside this patch: a sibling `skills` directory shadows
bundled Slack/AgentMail skills; macOS rejects rename of read-only
skill-cache directories (`EACCES`, also reproduced in an isolated
filesystem probe); and the host exhausts PostgreSQL System V
shared-memory slots. Focused reruns confirmed these limits. The complete
Linux CI run is the repository-wide verification; the full local run is
not green.
- Baseline evidence: the original sidebar test failed on unchanged
master; a development-mode run at 4x CPU throttling passed six selected
cases, so CPU pressure alone did not reproduce the CI bootstrap stall.

## Risks

- Default browser tests now exercise the shipped static UI. They no
longer implicitly cover Vite middleware or HMR; use the development
server for those checks.
- The new browser startup test uses Chromium CDP, matching the only
configured browser project.
- Opening a live subscription now causes one active-query refresh to
close the initial event gap. This adds startup API reads but no
recurring poll.
- Runtime behavior and deadlines are unchanged except for error details.
The historical readiness stall remains unconfirmed; a green rerun alone
cannot establish its cause.
- Local verification runs on macOS. The runtime lifecycle checks also
passed on Linux CI.

## Model Used

- OpenAI GPT-6 through Codex. The session identifies the model family as
GPT-6; an exact served model ID and context window size are not exposed.
Used reasoning, repository inspection, shell tools, code editing, and
test execution. No subagents were used.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#123` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass — targeted suites passed;
full local host limits are listed above
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-23 15:22:21 -07:00
DottaandPaperclip 24429024e7 feat: add Fireflies connector and summary-ready routines (#13890)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Apps gives agents governed access to external tools through stored
credentials.
> - Routines start work when an external service sends an event.
> - Fireflies provides meeting transcripts and summaries through an
official hosted MCP server.
> - This PR adds that connection and accepts signed meeting events
through the shared app webhook flow.
> - Agents can review completed meetings with the same permissions and
audit records as other work.

## Linked Issues or Issue Description

**Problem or motivation**

Operators need agents to read Fireflies meetings and start follow-up
work when a summary is ready. The Apps catalog lacks Fireflies. The
shared app webhook flow needs to accept its signed deliveries.

**Proposed solution**

Use the official Fireflies MCP endpoint with OAuth or a vaulted bearer
API key. Extend the existing Another app or script flow with signed
webhook support. Verify the raw-body signature and pass the JSON payload
as external data. Select Meeting Summarized in Fireflies. Deduplicate
identical signed deliveries, including setup deliveries.

**Alternatives considered**

A separate REST connector would duplicate the governed MCP path.
Polling, legacy V1 payloads, and automatic provider-side webhook
registration are outside this change.

**Roadmap alignment**

This extends the existing MCP Tool Gateway & Apps and Scheduled Routines
surfaces. It adds a provider to those systems. It does not introduce a
second integration framework.

**Additional context**

A GitHub search found no existing Fireflies issues or PRs. Provider
references and verification limits are in
`doc/connections/FIREFLIES.md`.

## What Changed

- Add the official Fireflies catalog definition, generated registry,
provider evidence, and branded artwork.
- Reuse Access → Connect, dynamic discovery, Permissions, vault storage,
policy, and audit behavior.
- Classify Fireflies sharing, movement, and access revocation as writes.
- Preserve Off and Ask first restrictions during OAuth reauthorization
and API-key replacement. New actions retain normal defaults.
- Add `app_webhook` authentication to the shared Another app or script
flow. Accept bearer tokens or raw-body HMAC-SHA256. Preserve earlier
`fireflies_hmac` triggers and revision snapshots for compatibility.
Existing text columns need no migration.
- Verify `X-Hub-Signature` or `X-Hub-Signature-256` against the exact
request body. Preserve generic event payloads and deduplicate identical
signed requests.
- Keep the routine wizard generic. Show one webhook URL and secret in
Another app or script. Keep all new app webhook event names
provider-neutral. Keep provider setup instructions in the connector
documentation.
- Pass generic webhook JSON to the task in an explicit external-data
block, capped at 16,384 characters. Keep strict meeting validation for
existing legacy Fireflies triggers.

## Verification

- Feature implementation commit `0882dc8a1`: all 54 CI checks passed;
two conditional Storybook checks skipped. This includes full tests,
typecheck, build, browser E2E, canary dry run, and security checks.
Greptile rated this commit 5/5; all review threads are resolved.
- Full local `pnpm -r typecheck`, `pnpm build`, and token gates passed
on the final code. Targeted connector, gateway, webhook, revision, and
UI suites passed during implementation. After the provider-neutral
follow-up, all 84 app-webhook and routine-service tests passed; the
final payload-to-task assertion also passed in the 72-test routine suite
and a clean-config rerun.
- The long local `pnpm test:run` invocation started before the final
edits and was stopped after the final-commit CI suites passed. It
reported one generic webhook test failure while those files were
changing; that test and the entire routine suite passed on the final
source, including a clean-config reproduction. The interrupted local run
is not counted as a full-suite pass.
- In the embedded browser, completed official OAuth consent and
discovered 20 live actions. Real meeting listing, transcript retrieval,
and summary/action-item retrieval succeeded as the selected agent.
Turning a live read Off blocked its test; catalog refresh preserved the
restriction.
- Embedded-browser Another app or script setup, back/save/resume, narrow
layout, and a signed synthetic Fireflies delivery succeeded. The UI
reported authentication passed without creating a task. Fixtures cover
signature tampering, malformed requests, ordinary app event names,
duplicate/setup deliveries, rotation, revisions, pause/archive, and
company isolation.
- Existing MCP browser suite: 8 passed and 2 provider-dependent cases
skipped. Branding checks passed; connector artwork and webhook setup
were checked at desktop/mobile widths and in light/dark modes.
- An unauthenticated POST to a correctly formatted public webhook URL
reached the staging tenant verifier through the existing Cloud gateway.
- A real Fireflies webhook delivery remains unverified. A staging
callback is available for the operator walkthrough. Live API-key
authorization, credential expiry, and a new meeting's summary completion
were not tested against the provider. Fixtures cover these protocol and
lifecycle paths where applicable.

- Storybook follow-up `c54174faa`: 27 production-component stories cover
every UI change, with a source-to-story map in the connector
documentation. Static Storybook build, UI typecheck, token gates, and
Playwright checks for all stories and the mobile footer pass. All PR
checks passed for this Storybook follow-up; Greptile reviewed
`c54174faa` at 5/5.

## Risks

- Fireflies may change its hosted MCP tools or OAuth behavior. Tool
discovery stays dynamic. Experimental search/fetch tools are not
required.
- Public webhook setup requires HTTPS and a separate signing secret.
Fireflies normally emits events for meetings owned by the configuring
account.
- Reauthorization touches shared MCP permission code. Regression tests
cover existing restrictions, new actions, connection removal, and other
gateway callers.
- Webhook receipt grants no tool access. The routine agent still needs
an authorized Fireflies connection.

- New generic triggers rely on provider event subscriptions. Without a
sender-supplied idempotency key, changed request bytes count as a new
event. Existing legacy Fireflies triggers retain summary-only filtering
and per-meeting deduplication.

## Model Used

OpenAI Codex, model `gpt-6-astra`. Used reasoning, repository editing,
code execution, and embedded-browser testing. The runtime did not expose
a context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-23 17:11:16 -05:00
Devin FoleyandPaperclip 4b8ec588f3 Stop duplicating wake context in adapter environments (#13891)
## Thinking Path

> - Paperclip manages agent work and preserves task context.
> - Built-in adapters already include wake context in the agent prompt.
> - They also copy the full wake JSON into a process environment
variable.
> - A large environment entry can prevent the agent from starting with
`spawn E2BIG`.
> - This change removes the duplicate environment entry and uses the
existing prompt delivery.
> - The agent keeps its context without extra file transport or new
history limits.

## Linked Issues or Issue Description

Refs #13144, #13860, #13872, #13793.

Large wake payloads can exceed the operating system limit for one
environment entry. The launch-envelope fix in #13793 handles the outer
transport but leaves that child environment entry intact.

Credit to @nickyleach for the prompt-only approach in #13144. This PR
applies that part on current master. It does not include that PR's
30-item history limits or recovery-history endpoint. Those behavior
changes can be reviewed separately from the process launch fix.

## What Changed

- Stop exporting `PAPERCLIP_WAKE_PAYLOAD_JSON` in the shared ACP engine
and all ten built-in adapter writers.
- Ignore configured values of the retired variable so saved adapter
settings cannot restore the oversized entry. Also drop inherited copies
in Hermes, which builds its environment directly.
- Keep scalar runtime variables, existing prompt rendering, continuation
history, resume deltas, gateway bodies, and Hermes JSON template
variables.
- Document the prompt delivery contract and the migration for custom
instructions that read the retired variable.
- Test large local and sandbox child-process launches, fresh and resumed
ACP turns, SDK delivery, and configured-variable filtering.

## Verification

- `pnpm -r typecheck` passed.
- Focused adapter utility, ACP, Codex child-process, and Cursor Cloud
suites: 340 tests passed.
- Hermes execution and prompt tests: 18 tests passed using its package
Vitest configuration.
- The child-process tests deliver over 128 KB of context through stdin
and check the complete text. The ACP test retains 50 complete messages
and 50 completed actions, then checks the resumed delta.
- `pnpm build` passed.
- `pnpm test:run` was attempted, then stopped after it reproduced ten
macOS runtime-skill-cache permission failures (also reproduced on
unchanged master) and one HTTPS backfill test failure. The HTTPS test
passed when rerun unchanged on this branch and master. The complete
local suite was not completed; Linux CI provides the full-suite gate.
- CI is green on c3b8191f97: 53 checks passed and two skipped. Attempt 3
passed the remaining shards without source changes. Earlier attempts hit
a lost runner, preview readiness, and chat browser failures. The preview
test passed unchanged locally. Local chat tests passed 3/4; the
remaining sidebar focus-style assertion also fails on unchanged master.
No tests were weakened or skipped to obtain the passing rerun.
- The staged diff passed `gitleaks stdin --redact` and `git diff
--check`.

## Risks

- Custom instructions or scripts that read `PAPERCLIP_WAKE_PAYLOAD_JSON`
must use the wake payload in the prompt instead. The variable is absent
even for small wakes.
- This removes one cause of `E2BIG`. Legacy CLI paths for Gemini, Grok,
Kimi, Pi, and Hermes still pass prompts as arguments and retain their
existing argument-size limits. Other large environment variables also
remain subject to OS limits.
- No new context truncation, API route, database migration,
authorization change, or production rollout is part of this PR. Existing
prompt windows and resume rendering remain unchanged.

## Model Used

OpenAI GPT-6 through Codex, with code execution and repository tools.
The exact model variant and context window size are not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-23 13:50:10 -07:00
DottaandPaperclip f55759942b fix(runner): accept compatible Codex releases from 0.149.0 (#13853)
Keep the minimum fixed at 0.149.0 until a deliberate maintainer change. Accept stable versions below 0.157.0 separately from the reproducible install pin.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-23 11:27:43 -07:00
DottaandPaperclip 7944ed3d97 fix(runner): preserve hire runtime safety and first-activity timing (#13852)
## Thinking Path

> - Paperclip is the open source control plane for companies of AI
agents.
> - Native runner agents need governed tools, durable runtime state, and
useful execution evidence.
> - A first activity trace waited 53.467 seconds even though tool
activity took 6.274 seconds; provider input arrived before the server
API call executed.
> - Native agents also need a safe way to hire teammates without asking
the model to rebuild runtime configuration.
> - This pull request separates the observed ACP input-stream window
from the actual server `tool.execute` span and adds a server-owned
native hire contract.
> - The benefit is clearer latency evidence and safer native teammates
with existing approval, auth, and company boundaries preserved.

## Linked Issues or Issue Description

Related Daytona provenance work is in
[#13814](https://github.com/paperclipai/paperclip/pull/13814). No
duplicate public PR was found for this combined timing and native-hire
change.

**What existing behavior does this improve?**

Native runner agents can use governed tools and request hires. The
server did not expose a safe native hire operation that reused the
caller's validated runtime settings. First-activity traces also mixed
provider input timing with server tool execution timing.

**Current behavior**

A native hire must construct a separate runner configuration. Full
configuration copying could expose paths, instructions, secrets, or
sessions. Timing evidence could make a provider or MCP identity join
appear proven when the trace did not contain that join.

**Proposed behavior**

The native `hire_agent` operation accepts identity and persona inputs.
The server sends `adapterType: "paperclip_runner"` with
`inheritRuntimeFrom: "caller"`, then copies only validated provider,
model, permission, lifecycle, and bounded execution settings. It
inherits and validates the default environment, derives the managed AI
binding through existing normalization, preserves approval and
permissions, and creates fresh child instructions. Caller secrets,
paths, prompts, and sessions are excluded.

Provider events now include the optional boolean `inputUpdated`, with
Rust forwarding support. Timing evidence separately records the ACP
input-stream window and the actual server `tool.execute` activity. It
does not claim a provider or MCP join without matching evidence.

**Reason and benefit**

Native agents can hire teammates that start with the caller's approved
execution policy. Operators retain company boundaries, auth rules,
approval gates, and requalification. Reviewers can distinguish provider
streaming time from server API execution time when diagnosing
first-activity delays.

**Breaking changes**

None for existing hires or tool calls. `inheritRuntimeFrom` is optional
and only applies to same-company native agent callers. Conflicting
explicit runtime settings are rejected. The provider event field is
optional for existing producers.

## What Changed

- Added the native `hire_agent` protocol action, catalog entry, API
contract, and runner authority checks.
- Added `inheritRuntimeFrom: "caller"` validation and a closed native
runtime inheritance allowlist.
- Preserved managed AI binding normalization, default-environment
validation, approval snapshots, permissions, requalification, and fresh
child instructions.
- Added provider `inputUpdated` schema support and Rust forwarding.
- Added first-activity and server tool timing evidence with conservative
identity-join handling.
- Added route, authority, provider-event, sidecar, API, catalog, and
Rust-focused tests.
- Kept private Honeycomb links, raw traces, and local result paths out
of this description.

## Verification

Focused checks passed:

- 458 timing/session checks.
- 61 native hire inheritance checks.
- 20 hire authority checks.
- 1,741 API checks.
- 106 catalog checks.
- 54 provider sidecar checks.
- 12 Rust provider checks.

Live R2 and R3 each passed 45 checks across 6 runs (361,135 ms for R2).
R1 stopped at missing Docker image setup. The final trace is available
at
https://ui.honeycomb.io/paperclip/environments/test/datasets/paperclip/result/BiMypLNvmiB?tab=traces.
Latest-head CI passed all required build, typecheck, Rust, static,
Vitest, serialized-server, workspace, chat, and E2E jobs. The focused
local checks listed above passed; the broad local suite was not run
before the live evaluation, while CI provides the full repository
verification.

## Risks

- Timing fields describe separate observed windows. They do not prove a
provider or MCP owner without a valid trace join.
- The inheritance allowlist must stay synchronized with native runner
configuration fields.
- Approval snapshots include resolved safe inherited settings and should
be reviewed when native configuration fields change.
- The focused local suite is narrower than the full repository suite;
latest-head CI covers the broader repository checks.

> Roadmap review: `ROADMAP.md` places this work within Paperclip's
bring-your-own-agent direction. It extends existing native runner hiring
and observability behavior.

## Model Used

OpenAI GPT-6 (exact serving model ID is not exposed), with extended
reasoning and repository tool use; GPT-5.6 Luna assisted with focused
implementation and verification work.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-22 23:16:11 -05:00
Devin Foley 8c6cc7dccf feat(cloud): sync primary-company archive state with the Cloud control plane (#13837)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Paperclip Cloud runs managed instances and controls their lifecycle
from a control plane.
> - Inside the app, a person can archive a company. On a managed
instance, the Cloud-pinned primary company is the whole organization.
> - Today that archive stays local. The control plane does not learn
about it, so it keeps the instance running and shows the organization as
live.
> - This pull request notifies the control plane when the primary
company crosses the archived boundary, and adds two control endpoints so
the control plane can verify the state and undo the archive during a
restore.
> - The benefit is that an archived organization stops running and shows
as archived, and a restore brings the company back without manual steps.

## Linked Issues or Issue Description

No public issue exists. Description follows the enhancement template:

**What existing behavior does this improve?**

Company archive on Cloud-managed instances. Archiving the primary
company pauses agents and cancels runs, but the hosting control plane
never learns about it.

**Subsystem affected**

Server: company service, cloud-control middleware, instance routes.

**Current behavior**

A person archives the primary company. The instance keeps running. The
Cloud portfolio still shows the organization as live. Unarchive after a
Cloud restore requires manual steps inside the product.

**Proposed behavior**

The company service rings a Cloud lifecycle doorbell when the primary
company is archived or unarchived. The control plane verifies the state
through `GET /api/instance/lifecycle` before it acts. During a restore,
the control plane calls `POST /api/instance/lifecycle/unarchive-primary`
to bring the company back. Self-hosted instances are not affected.

**Reason and benefit**

An archived organization should not keep running, and its hosting
console should show it as archived. The verified read-back keeps the
doorbell a hint: a forged or duplicated ring cannot change state.

## What Changed

- New `services/cloud-lifecycle-sync.ts`: fire-and-forget doorbell `POST
{cloudOrigin}/v1/tenant/lifecycle-changed` with bounded retries. It runs
only on cloud-managed instances, and only for the derived primary
company id. It never blocks or fails the company mutation.
- `services/companies.ts`: ring the doorbell after a committed archive
or unarchive transition, in both the `update()` status-patch path and
`archive()`.
- New `GET /api/instance/lifecycle`: reports the primary company id, its
status, and how many other companies are not archived. Bound to a
`lifecycle:read` Cloud control assertion; board members can also read
it.
- New `POST /api/instance/lifecycle/unarchive-primary`: idempotent
unarchive of the primary company as a system actor. Bound to
`lifecycle:unarchive-primary`; requires instance admin otherwise.
- `middleware/cloud-control.ts`: a closed endpoint→method→action table
replaces the single hardcoded endpoint. Each assertion still authorizes
exactly one action on one endpoint.
- `services/cloud-instance.ts`: the primary-company id derivation moves
here as the single definition; `middleware/auth.ts` delegates to it. The
derivation itself is unchanged.

## Verification

- `pnpm exec tsc --noEmit` in `server/` is clean.
- `pnpm exec vitest run src/__tests__/cloud-lifecycle-sync.test.ts
src/__tests__/cloud-control-task-drain.test.ts
src/__tests__/instance-settings-routes.test.ts
src/__tests__/companies-service.test.ts
src/__tests__/cloud-tenant-company-provisioning.test.ts` — all pass.
- New tests cover: doorbell env-gating, retry bounds, no-throw contract;
the read-back status and sibling count; idempotent unarchive and the
admin gate; non-cloud 404s; control-assertion action binding for the new
endpoints, including cross-action and wrong-method rejection; and a
companies-service test that proves the doorbell rings exactly on
archived-boundary transitions.

## Risks

- Self-hosted instances see no behavior change: without a Cloud signal
the doorbell is a no-op and the endpoints answer 404.
- The doorbell is advisory by design. The control plane verifies through
the read-back before it acts, so a lost or duplicated ring cannot
corrupt state.
- The unarchive endpoint reuses the existing `companyService.update`
path, so agent reactivation and activity logging behave exactly like an
in-product unarchive.

## Model Used

- Claude Fable 5 (`claude-fable-5`), via Claude Code CLI, extended
thinking and tool use enabled.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-22 17:49:15 -07:00
Devin FoleyandPaperclip be6f49a425 feat(runner): refresh shared coding harness runtimes (#13838)
## Thinking Path

> - Paperclip runs agents through local adapters and the native runner.
> - Both paths must use the same installed provider CLI.
> - New models require current harness releases.
> - The runner still pins Codex 0.153.4, Claude SDK 0.3.263, and
OpenCode 1.18.29.
> - Changing the image alone would fail the runner's exact version and
executable checks.
> - This pull request updates those dependencies, integrity checks,
controller checks, and image pins together.
> - Shared installations can then run the current models without a
task-time download.

## Linked Issues or Issue Description

Refs #13829, which updates model choices and reasoning controls.
Searches found no open PR that updates these runtime pins.

**Current behavior**

The shared provider pack ships old CLIs. Claude Code 2.1.263 cannot run
Opus 5.5, which requires 2.1.280. Remote controllers reject provider
packs whose versions differ from their declared pins.

**Proposed behavior**

Use Codex 0.156.0, Claude Agent SDK 0.3.280 / Claude Code 2.1.280, and
OpenCode 1.18.32 throughout the runner. Keep the reviewed ACP bridge
patches and one shared CLI installation per provider.

**Reason and benefit**

Current harnesses support the new model IDs while preserving executable
verification and remote provider-pack compatibility checks.

## What Changed

- Update dependency overrides, the Codex ACP package patch, runtime
profiles, and remote controller pins.
- Verify the new Claude Linux x64 and macOS arm64/x64 executables and
Codex Linux x64 executable against integrity-verified npm archives.
- Refresh OpenCode version checks, fixtures, and the runner
configuration label.
- Refresh the eval image's Grok, Gemini, Kimi, Cursor, and GitHub CLI
pins and archive hashes. Hermes remains current at 0.19.0.
- Refresh the build-time lock digest from clean pnpm 9.15.4 resolution.
Leave lockfile commits to repository automation.
- Document model compatibility and the separation between CLI runtimes
and patched ACP bridges.

## Verification

- `pnpm -r typecheck` and `pnpm build` passed.
- Rust workspace release tests passed.
- Package/patch and OpenCode binary-materialization contract tests: 11
passed.
- Real Codex 0.156.0 startup-ownership and paginated session-resume
probes passed with isolated synthetic homes and no model turn.
- Codex app-server `thread/start` preserved `gpt-6-sol` and
`gpt-6-luna`; no `turn/start` was sent. An unauthenticated built-in
catalog does not include those account-served entries.
- Installed Claude integrity probes passed for `claude-opus-5-5` and
`claude-fable-5-1`.
- `pnpm --filter @paperclipai/paperclip-runner
test:opencode:qualification` passed with the actual OpenCode 1.18.32
executable under Node 24 and Node 25. The loopback provider exercise
covers health/version, session creation/read/delete, SSE, and a
completed async prompt.
- `pnpm check:token-gates` passed.
- The targeted runner suite passed 130 tests. Three macOS failures in
snapshot module lookup and OpenCode final-message selection also
reproduce on the unchanged base; Linux CI will provide the platform
check.
- [Final Linux
CI](https://github.com/paperclipai/paperclip/actions/runs/35798076399):
all gates passed. Four jobs needed one retry after their CI workers
received shutdown signals. The PR has 55 successful checks, two skipped
checks, Greptile 5/5, and no unresolved review threads.
- Changed runner configuration UI tests: 5 passed.
- Full macOS `pnpm test:run` reached 13,094 passing server tests, 84
skipped, and 18 failures before the wrapper stopped. Failures involved
skill-cache publication permissions, missing bundled connector skills in
the worktree, and a conversation-reset timing case. The 10 cache
permission failures reproduce on the unchanged base; both
conversation-reset cases passed on a targeted retry. The wrapper did not
reach its later workspace/serialized groups locally; Linux CI covers
those groups.
- The local Docker daemon did not respond, so no local Docker build was
run. No billable model requests were made.

## Risks

- Deploy the matching controller and provider pack together. Older
controllers enforce their previous exact pins.
- Current upstream CLIs can change behavior. Existing protocol tests and
isolated real Codex probes cover the integration boundaries;
authenticated model inference is not part of these checks.
- ACP bridge package versions and executable digests stay unchanged
because their executable bytes are unchanged. Only the underlying
CLI/SDK dependencies move.
- No schema migration. Revert the runtime and image pins together to
roll back.

## Model Used

OpenAI GPT-6 via Codex, with repository tools, code execution, and web
research. The exact serving model ID and context window were not exposed
by this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass for the changed surfaces
and real-executable probes; full macOS-suite limitations are listed
above
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-22 17:02:29 -07:00
Devin FoleyandPaperclip cdf04a33fa feat(adapters): refresh current coding models and reasoning controls (#13829)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Its adapters supply model catalogs and reasoning controls to agent
setup.
> - Several provider releases are missing from the fallback catalogs.
> - Some newer models also have effort levels that the UI does not
offer.
> - Operators need the exact supported IDs and controls when discovery
is unavailable.
> - This pull request updates the existing adapters from current
provider documentation.
> - Operators can select current coding models without entering custom
IDs.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

Model selection and reasoning controls across the existing coding-agent
adapters.

**Subsystem affected**

Claude, Codex, Grok, Gemini, Cursor, Kimi, and OpenCode adapters; model
discovery tests; agent creation and editing.

**Current behavior**

The catalogs omit Opus 5.5, GPT-6 Sol/Luna, Grok 4.7/4.6/4.5, current
Gemini Flash models, and several Cursor/Kimi choices. Bedrock has
obsolete IDs. The UI omits supported effort levels and saves Grok effort
under a key the runtime does not read.

**Proposed behavior**

Offer verified current model IDs and model-specific efforts. Remove
retired Gemini 2.0 choices. Keep configured defaults and saved model
IDs. Keep runtime discovery for account-specific choices.

**Reason and benefit**

Catch up with provider releases through September 22, 2026. Correct the
picker and runtime controls together.

**Breaking changes**

No database or API change. Gemini 2.0 options leave the picker after
their June 1 shutdown. Existing saved IDs remain unchanged. Corrected
Bedrock catalog IDs do not rewrite saved configuration.

**Additional context**

Supersedes the separate GPT-6 Sol PR #13830. Fable 5.1 was already
merged in #12730, and GPT-6 Astra in #12851. The Grok 4.6/4.5 proposal
#11324 was closed and parked by its author. This change retains the
default-sentinel fix from #12062. Related discovery proposals #13127 and
#13565 do not supply these catalog and effort updates. Searches found no
open PR for the additional model IDs.

See [the dated
audit](https://github.com/paperclipai/paperclip/blob/feat/claude-opus-5-5/doc/adapter-model-audit-2026-09-22.md)
for exact scope, primary sources, runtime observations, and
account-specific limits. This updates existing adapters and does not
duplicate planned core work.

## What Changed

- Add Opus 5.5 for direct Claude and Bedrock, with a Claude Code 2.1.280
gate. Correct and extend Bedrock model IDs.
- Add GPT-6 Sol/Luna and Fast mode. Offer Ultra for Astra/Sol and
GPT-5.6 Sol/Terra, and Max for both Luna generations.
- Add Grok 4.7/4.6/4.5, expose supported Extra High effort, and save
Grok edits under `reasoningEffort`.
- Add Gemini Flash 3.8/3.7/3.6/3.5, Flash Lite 3.5/3.1, and 3 Flash
Preview. Remove retired 2.0 choices.
- Add the current documented Cursor fallback models, including Fable
5.1, Composer 2.5, and Muse Spark 1.3.
- Refresh OpenCode fallback IDs used in remote environments from its
installed provider registry.
- Add Kimi K3 256K. Update the existing coding alias to K2.8 Preview and
enable its CLI effort settings.
- Use model-specific Claude/Grok efforts in creation and editing. Clear
unsupported effort when switching models.
- Add catalog, CLI/ACP forwarding, compatibility, and UI persistence
coverage. Record the audit and sources.

## Verification

- Latest head `6e63c9ef53b54ba869cd4fb431a8570bebe289f4`: 53 CI checks
passed, 2 skipped. This includes full workspace typecheck, build, and
all test shards. Greptile is 5/5 with zero unresolved threads. GitHub
reports no merge conflicts.
- 340 focused tests passed across adapter metadata, CLI/ACP arguments,
Claude version checks, Kimi effort, Grok execution, server model
discovery, and UI effort selection/persistence.
- `pnpm --filter @paperclipai/adapter-claude-local --filter
@paperclipai/adapter-codex-local --filter
@paperclipai/adapter-grok-local --filter
@paperclipai/adapter-gemini-local --filter
@paperclipai/adapter-kimi-local --filter
@paperclipai/adapter-cursor-local --filter
@paperclipai/adapter-opencode-local typecheck` — passed. The same
filters with `build` passed.
- `pnpm check:token-gates` and `git diff --check` — passed.
- Full workspace and UI typechecks were attempted locally. They stop on
existing missing `three` dependencies in `packages/shared/src/cliplab`.
- Full `pnpm test:run` and `pnpm build` were not run locally. Worktree
creation exhausted disk space, so a clean dependency install is not
feasible on this host. Focused checks reuse existing dependencies. CI
supplies full workspace verification.
- No provider inference was run. Account-specific runtime model lists
were inspected where available.
- Manual check: select the new models in agent setup and editing.
Confirm Luna has Max but no Ultra, Grok 4.7 has Extra High, and Fable
5.1 has Extra High/Max. Save Grok effort and confirm
`adapterConfig.reasoningEffort` contains the selection.

## Risks

- Catalog presence does not grant account access. Older CLIs and
restricted accounts can reject a model. Opus 5.5 has an explicit upgrade
check.
- Higher effort can increase cost and latency. Existing agent defaults
are unchanged.
- Cursor fallback IDs come from public model documentation; the local
account exposed no live catalog. Runtime discovery still adds
account-specific variants.
- Kimi effort remains supported only on its explicit CLI engine. This
does not add effort support to its default ACP engine.
- Saved obsolete Bedrock or retired Gemini IDs are not migrated
automatically.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, tool use, and code
execution. The exact deployment ID and context window are not exposed to
this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-22 16:56:00 -07:00
Devin FoleyandPaperclip 7badae6981 feat(plugins): add an optional organization switcher slot (#13832)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Plugins can add UI surfaces to the board.
> - Organization navigation is still fixed in the host sidebar.
> - A distribution needs a supported way to supply its own organization
menu.
> - This pull request adds one optional React slot with host-owned
navigation controls.
> - The built-in menu stays available when the optional contribution
cannot render.

## Linked Issues or Issue Description

**Subsystem affected**

Plugin SDK, server capability validation, and sidebar UI.

**Problem or motivation**

An installed plugin cannot replace the organization switcher without
editing the host menu. Existing sidebar and overlay slots do not provide
this replacement surface.

**Proposed solution**

Add an `organizationSwitcher` slot that requires `ui.sidebar.register`.
Pass display state, an icon renderer, and navigation/logout callbacks.
Keep the built-in menu for absent, ambiguous, missing, failed, or
unsupported contributions.

**Roadmap alignment**

This keeps distribution UI in plugins and adds a small host contract. It
does not add an account system or change company authorization. The
maintainer requested this extension.

Related prior menu changes: #12788, #10917, and #10850. No duplicate
replacement-slot PR was found.

## What Changed

- Add the slot to shared validation, SDK types, and server capability
checks.
- Wrap both sidebar menu variants with the optional replacement.
- Resolve selection against the current account query before mounting,
and reset replacement state on account or company changes.
- Add host-specific component props and a fallback to `PluginSlotMount`.
- Document the React-only contract and its trust boundary.
- Report runtime-supervisor fixture startup details when CI readiness
fails.

## Verification

- `pnpm -r typecheck` and `pnpm build` passed.
- `pnpm check:token-gates` passed.
- `pnpm exec vitest run --project @paperclipai/ui`: 6,580 tests passed.
- Targeted manifest, replacement, and built-in menu tests: 26 passed,
including the incoming-account selection regression.
- Installed a local test contribution into an isolated server. CLI
inspection reported `ready`. Browser checks covered the loaded
production UI, keyboard dismissal, current-organization selection, and
an expired remote session. Remote account responses were fixtures.
- `pnpm test:run` was attempted. Its first server phase passed 8,323
tests but failed in 36 files due to embedded PostgreSQL startup and
filesystem permission errors on this Mac. Later phases did not run. The
current Linux CI run is green: 54 checks passed and two were skipped,
including build, typecheck, and browser gates. See
https://github.com/paperclipai/paperclip/actions/runs/35796722770.
- The earlier runtime-supervisor readiness failure did not reproduce
locally. The complete affected shard passed locally: 58 files and 843
tests. The six supervisor tests also passed on Node 24.21.0 with CI
flags. Added fixture startup diagnostics for the selected Node
executable and listener port. The affected Linux shard then passed all
843 tests. The original root cause remains unconfirmed; no production
runtime behavior or timeout was changed.

## Risks

- Plugin UI remains trusted same-origin code. Display props do not
authorize account requests.
- A replacement can change navigation behavior. The host retains the
built-in menu when discovery or rendering fails and keeps logout/session
cleanup host-owned.
- No database migration. Existing menus and portfolio behavior remain
available.
- Local full-suite verification is limited by the environment failures
listed above.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository inspection, code
execution, and browser testing. The runtime does not expose a more
specific model ID or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-22 16:39:11 -07:00
DottaandPaperclip a10702a878 feat(slack): add governed tools for Slack-origin tasks (#13828)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Chat connectors let people start and continue agent tasks from other
services.
> - A Slack conversation needs access to its surrounding discussion and
Slack collaboration tools.
> - The agent must use the linked requester's access and keep private
material within its permitted audience.
> - This pull request adds Slack tools through the existing connector
contribution and approval framework.
> - People can ask an invited bot to read a discussion, create follow-up
tasks, and collaborate in Slack.

## Linked Issues or Issue Description

**Subsystem affected**

Chat connectors, connector runtime, tool gateway, and connection
Settings/Access.

**Problem or motivation**

Slack-origin tasks can receive messages but cannot inspect the rest of a
channel or act through the originating bot. People must paste context or
configure a separate integration.

**Proposed solution**

Supply typed Slack tools and a bundled skill only to the originating
task and assigned agent. Resolve the linked requester on the server.
Check bot and requester access before reads and writes. Use existing
durable actions and approvals. Retrieved messages remain source
material.

**Alternatives considered**

Slack's user-OAuth MCP server does not replace the customer-created chat
bot. An unrestricted Web API proxy would not provide suitable permission
or publication boundaries.

**Roadmap alignment**

This extends the existing MCP Tool Gateway & Apps work with a provider
contribution. It does not add a task dispatcher or a separate Slack task
lifecycle. Related: #11144 covers generic per-user MCP grant execution;
this change binds Slack bot operations to chat-origin tasks.

## What Changed

- Add 39 typed Slack tools, a method/scope matrix, a bundled skill, and
shared native/HTTP execution.
- Bind tools to company, endpoint, task, run, assigned agent, and
admitted linked requester. Check membership and revocation on each call
and before queued writes.
- Add paginated reads, bounded history search, source links, messages,
file uploads, reactions, pins, bookmarks, topics, canvases, lists, and
approved channel operations.
- Restrict private-source publication, including automatic replies and
uploaded deliverables. Keep other people's bot DMs inaccessible.
- Reuse action receipts, idempotency, approvals, and reconciliation.
Suppress an identical explicit-send/final-reply duplicate. Return
governed results through their verified originating conversation.
- Add endpoint-bound personal search OAuth storage and lifecycle. Keep
native real-time search disabled until a runtime meets Slack's
transient-result requirements. Current runtimes use bounded history
search.
- Show capabilities, scope upgrades, and personal search authorization
in Settings/Access and Storybook. Document provider and runtime limits.

## Verification

- Current head `0eb21cba4`: CI checks pass and Greptile is 5/5 with no
unresolved findings. One unchanged rapid-callback timing test passed on
a single CI retry.
- Approval presentation regressions cover board-comment precedence and
exact Slack publication; the expanded database assertion passed in CI.
The local PostgreSQL startup probe later became unavailable, so that
final assertion was verified in CI. Slack setup and failed-run retry
browser tests also passed locally.

- Full workspace typecheck and build passed. Server typecheck/build
passed again after the approval routing fix.
- Broad local suites passed in separate groups: server 12,958 tests, UI
6,555, shared 770, skills catalog 20, and other workspace packages
2,652. CLI and serialized server checks passed after environment/timeout
retries. These are composite results, not one uninterrupted green
full-suite invocation.
- PostgreSQL authority regression covers admitted identity,
cross-company/task/agent rejection, recovery, retained-session
revocation, OAuth refresh/disconnect races, approval execution, exact
publication lineage, retries, uncertain sends, and duplicate
suppression.
- Gateway/response regressions cover separate-origin approval batches
and durable continuation. Focused provider, access, search, native
runtime, route, and AgentMail regressions pass.
- Storybook capability, missing-scope, OAuth configuration,
authorization, and disconnect states were inspected in the browser.
- Live staging: read a channel decision and full thread, create exactly
two assigned backlog tasks, add a reaction, paginate discovery to
exhaustion, and return bounded search matches with source links and
coverage.
- Live staging: create/edit/read a canvas and list, inspect the canvas
in Slack, post/edit one message, and create a channel only after
approval. New channels remain disabled for responses.
- Live staging: read a response-disabled channel from the requester's
DM; writes to that channel were denied. The test setting was restored.
- Final live retest passed: explicit file upload and exact content
read-back; approved deletion of only the disposable bot message;
continuation confirmation returned to the original Slack thread without
repeating the action.
- Optional OAuth, private multi-user boundaries, native RTS, and CLI
provider execution are not fully live-qualified. The staging agent
initially supplied malformed tool arguments; valid arguments succeeded,
and the tool/skill descriptions now emphasize UUID write keys.

## Risks

- Existing Slack apps must add scopes and reinstall for new
capabilities. Provider plans and document permissions can still restrict
operations.
- Instances need an independent `PAPERCLIP_TOOL_ACTION_SIGNING_SECRET`
for governed tool actions. The staging instance was configured with
explicit operator approval; fleet provisioning is a separate gap.
- Native RTS is not exposed on current transcript-retaining runtimes.
Bounded history scans are deliberately reported as incomplete. Inline
file reads support text/canvas content up to 256 KiB; other types return
metadata.
- Private document edits fail closed when the full audience cannot be
verified. Uncertain effects other than posts/uploads require inspection
instead of blind retries.
- Shared approval-delivery code now separates outcomes by source run to
preserve origin boundaries. No database migration is required.
- A separate completion-validator gap remains when the agent cites a
prior run's registered artifact during finalization. It asked for
registration again even though Slack delivery was confirmed. This change
does not add a connector-specific task-completion policy.

## Model Used

OpenAI GPT-6 through Codex, with repository tools, code execution, and
browser testing. The exact deployed model identifier and context-window
size were not exposed in the session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-22 15:29:07 -05:00
DottaandPaperclip a959e47508 fix(apps): reduce Google Chat scopes and block unread filters (#13820)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Apps give agents controlled access to external services.
> - Google Chat uses OAuth profiles and reviewed MCP tools.
> - Those profiles request membership and read-state access that the
supported feature set does not need.
> - Removing read-state access also requires us to block unread search
filters, including on existing connections.
> - This pull request reduces both OAuth methods and enforces the
reduced search contract before dispatch.
> - Users retain conversation lookup, message history, ordinary search,
and approved message sending.

## Linked Issues or Issue Description

**What happened?**

Both Google Chat profiles request membership and read-state scopes. The
supported tool set does not include membership listing or read-state
updates. Message search still advertises an unread filter. Related work:
Refs #12619.

**Expected behavior**

Managed and customer-owned OAuth request only the scopes needed for
supported features. Unsupported unread filters fail clearly before any
provider call. Existing cached catalogs and broader grants must not
bypass that policy.

**Steps to reproduce**

Start Google Chat OAuth from either connection method and inspect the
requested scopes. Inspect the message-search tool schema, then submit a
search with `searchParameters.isUnread` set to true or false.

**Paperclip version or commit**

The scope change is based on master at `110d176fc`.

**Deployment mode**

Managed Cloud and self-hosted instances with Google Chat Apps enabled.

## What Changed

- Remove `chat.memberships.readonly` and `chat.users.readstate.readonly`
from shared profiles and all four Chat connection methods.
- Hide unsupported read-state fields and instructions in agent and board
Test tool schemas.
- Reject explicit unread filters, including false, null, snake-case
fields, and encoded filter objects, before provider dispatch.
- Recheck previously approved calls and support existing profile-bound
and URL-only Chat connections.
- Add scope, signed broker request, OAuth URL, allowlist, schema, and
dispatch regression tests.
- Document coordinated app/broker rollout, existing-grant reconnects,
and the remaining deployment checks.

## Verification

- All seven focused OAuth and Chat gateway test files pass: 498 tests on
the rebased branch.
- `pnpm -r typecheck` and `pnpm build` pass, using pinned pnpm 9.15.4.
- The complete sharded CI test matrix passes, including general server,
Chat, workspace, serialized server, Runner, and all eight browser e2e
shards. The duplicate unsharded local `pnpm test:run` was stopped after
CI passed; it did not complete locally.
- `git diff --check` passes.
- `node scripts/ingest-app-definitions.mjs` succeeds and leaves the
branch unchanged. Google Workspace JSON is the durable reviewed input
used by the generator.
- All current-head CI checks pass at
`7ee755371714dc036fff1c7da844776fee3f2ec1`, including typecheck, build,
and canary dry run. Greptile is 5/5 with zero unresolved threads. The
generator concern was withdrawn after review of the source and
regeneration evidence.
- No production deployment or live Google consent test was performed.
After coordinated deployment, verify reduced consent scopes, normal
search/history, message sending, and rejection of unread filters.

## Risks

- Coordinate deployment with the companion Cloud broker scope change.
Mixed versions can reject exact-scope requests.
- Existing tokens are not narrowed or revoked. Grants with old scopes
need new consent. Do not revoke a shared Google client to migrate one
profile.
- Explicit unread filters now return an error instead of being sent to
Google. Ordinary search and the approved send tool remain available.
- No database, UI, lockfile, or workflow changes.

## Model Used

OpenAI Codex, a GPT-5-based coding agent, with tool use and code
execution. The exact runtime model ID and context window were not
exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-22 15:12:30 -05:00
DottaandPaperclip 9a76390cfa test(runner): improve blank-page diagnostics and infrastructure coverage (#13824)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Runner E2E tests verify real tasks and retain evidence for failures.
> - Some exposure tests assumed that port 42000 was free.
> - A blank task page could also fail without enough browser startup
evidence.
> - This pull request tests occupied ports and task reloads, and records
private startup diagnostics.
> - These changes make test failures easier to reproduce and explain.

## Linked Issues or Issue Description

Refs #13815. Report publication already received its production fix in
#13750; this PR adds a regression check for that workflow.

**What happened?**

Three exposure tests assumed that the allocator would select port 42000.
The synthetic failure disappeared when another process occupied that
port. A prior Daytona run also retained an empty task page after
navigation, but its evidence did not record pending modules or
service-worker control.

**Expected behavior**

Exposure tests must exercise the intended failure on the actual assigned
port. Browser failure evidence must distinguish an empty root from
loaded content. A saved task must remain usable after navigation and
reload.

**Steps to reproduce**

Run the exposure regression with the base port pair marked unavailable.
Run the browser-support tests with an unresolved entry module. Open a
saved task under the service worker, navigate to the same URL, and
reload it.

**Paperclip version or commit**

Based on master at a68f3d8e3.

## What Changed

- Make synthetic exposure failures use the assigned port. Test both free
and occupied base pairs.
- Record document readiness, root children, service-worker control,
pending module paths, and recent module error or 304 statuses in private
runner evidence.
- Add browser tests for pending startup, failed module responses, and
diagnostic reset after navigation.
- Add a provider-free browser test for persisted task content and the
composer across three navigation/reload cycles.
- Guard trusted report-job lockfile resolution before frozen install and
AWS credential setup.
- Document the new evidence and test commands.

## Verification

- `pnpm test:e2e:runner:unit`: 443 tests passed.
- `pnpm test:e2e:runner:browser-support`: 10 tests passed.
- Exposure tests: 28 passed; 3 platform-specific skips.
- Harness typecheck, workspace typecheck, and full build passed.
- Task-reload browser test passed against a throwaway local instance. It
checks three navigation/reload cycles with a controlling service worker.
- All PR verification suites passed, including server, chat, workspace,
serialized server, Rust, runner, build, typecheck, and browser shards.
The existing sidebar navigation case showed an empty page on the first
CI attempt and passed on one retry.
- The monolithic local `pnpm test:run` was started, then stopped after
CI completed the equivalent suites. Additional local browser
reproduction attempts also hit embedded-Postgres startup failures; the
focused checks listed above completed successfully.
- Greptile: 5/5 on the current head, with no findings.

## Risks

This change affects tests and private test evidence. It does not change
production prompts or runtime behavior. Module paths omit queries; the
collector reads no response bodies or headers. The existing sanitizer
and publication allowlist still apply. A 304 response does not cause a
test failure.

The blank-page cause remains unconfirmed. The first CI attempt
reproduced it in an existing sidebar navigation case: the trace shows an
empty page after a service-worker-mediated 304 response for the large
editor module. The single retry passed. This PR adds coverage and
diagnostics; it does not claim to fix that intermittent symptom.

## Model Used

OpenAI Codex, GPT-6, with repository tools, code execution, and browser
tests. The exact deployed model variant and context-window size are not
exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-22 15:04:13 -05:00
Devin FoleyandPaperclip 92d4868e79 fix(server): isolate run errors and redact runtime capability headers (#13826)
## Thinking Path

> - Paperclip manages AI agents and their work.
> - Operators use Sentry to investigate failed runs and server errors.
> - Run reports attach a task ID, run ID, error code, and adapter
fingerprint.
> - The server skips Sentry's OpenTelemetry setup to preserve its
separate tracing and privacy settings.
> - Without an async context manager, a scope mutation can attach old
run data to later errors.
> - The HTTP logger also retains a runtime credential capability header.
> - This change isolates run metadata and redacts that header so
diagnostics identify failures without leaking credentials.

## Linked Issues or Issue Description

Refs #13446 and #13719.

**What happened?**

After a terminal run failure, an unrelated server exception can inherit
that run's tags, context, and fingerprint. Sentry then groups a database
error with an earlier adapter failure. The real SDK reproduces this with
the application's `skipOpenTelemetrySetup: true` setting. HTTP request
logs also retain the `x-paperclip-github-capability` header, which must
be treated as a credential.

**Expected behavior**

Run metadata belongs to the terminal run event. Later exceptions must
not inherit it. Every genuine error must still be captured. Runtime
capability headers must be redacted on success and failure logs.

**Steps to reproduce**

1. Initialize the optional Sentry SDK with the application's options and
an in-memory transport.
2. Capture a terminal run failure.
3. Capture an unrelated exception.
4. Inspect the second event. Before this fix, it contains the first
run's identity and fingerprint.
5. Send a request with a fixture runtime GitHub capability header.
Before this fix, HTTP logs retain the fixture value.

## What Changed

- Pass tags, context, and fingerprint directly to `captureException`
instead of mutating the ambient scope.
- Preserve the existing run fields, grouping keys, ordinary exception
capture, and privacy settings.
- Test two run identities interleaved with unrelated exceptions against
the real optional SDK.
- Update the capture contract tests and document event-local run
metadata.
- Redact the runtime GitHub capability header through the existing HTTP
logger policy. Test successful, denied, and failed requests.
- Add a dedicated GitHub-hosted CI check that installs the exact
optional SDK version declared in `server/package.json`. It fails if the
real-SDK regression would be skipped. The SDK stays outside the
workspace and production dependency graph.

## Verification

- The real-SDK regression failed before the fix because the unrelated
event contained `contexts.run_failure`.
- Five focused suites passed: 123 tests, including all optional SDK
tests. Suites: `run-failure-sentry-real-sdk.test.ts`,
`run-failure-sentry.test.ts`, `sentry.test.ts`,
`run-failure-report.test.ts`, and `http-log-redaction.test.ts`. A custom
in-memory transport prevented outbound Sentry delivery.
- All three new header-redaction cases failed before the policy fix and
passed afterward.
- The dedicated CI command passed locally with
`PAPERCLIP_REQUIRE_SENTRY_TEST_SDK=1` and the audited SDK available
through `NODE_PATH`.
- Server TypeScript check passed with a scratch configuration that
resolves this checkout's workspace packages. The existing dependency
links point to another checkout.
- `node scripts/check-module-boundaries.mjs` and `git diff --check`
passed.
- Gitleaks and a separate private-data scan passed before push.
- Full local workspace typecheck, test, and build were not run. The
machine has less than 2 GiB free and those commands include Rust builds.
Full PR CI must pass before merge.
- The dedicated real-SDK GitHub check passed with 1 test executed and no
skips: https://github.com/paperclipai/paperclip/actions/runs/35774449002
- Greptile reviewed c9db03bcab at 5/5. Its only thread is resolved. Full
PR CI passed on that same head:
https://github.com/paperclipai/paperclip/actions/runs/35774449020

## Risks

Small change to error attribution. Unrelated errors may now form their
correct Sentry groups instead of reopening a prior run group. No errors
are filtered or suppressed. No tracing is enabled and no new event
fields are added. No schema or runtime-execution changes. HTTP logs
retain their request and status diagnostics while masking the capability
value. The new SDK job has read-only permissions, no secrets, and an
in-memory Sentry transport.

## Model Used

OpenAI GPT-6 (Codex), with tool use and code execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub references)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-22 12:53:31 -07:00
DottaandPaperclip a68f3d8e35 fix(runner-e2e): align Daytona image and provider pack provenance (#13814)
## Thinking Path

> - Paperclip is an open source app people use to manage AI agents for
work.
> - The Daytona runner image provides the native runner and its provider
package.
> - The controller also sends a provider package to native Daytona cells
when the image package does not match.
> - A stale image and a package from another source revision caused a
1.8 GB upload before Claude could run.
> - This pull request refreshes the reviewed lock checksum and documents
how to reuse the exact package from an immutable image.
> - The benefit is a reproducible setup path and clear evidence when
image and package provenance do not match.

## Linked Issues or Issue Description

**What happened?**

A Daytona native Claude run used image source revision
`45c99a0d06cbd5b04982b06b79de321149930ac5` with a controller provider
package from revision `294853dc...`. Runtime verification rejected the
image package and staged a large package upload before model execution.
A hosted campaign also failed during image setup because the Dockerfile
expected lock checksum `d7d96cf0...` while the resolved lockfile
checksum was
`4b796c312833ebf2be4c38228babc0292c76774fb40d43bd18b54bd6b205753d`.

Related public work reviewed:
[#12795](https://github.com/paperclipai/paperclip/pull/12795),
[#12862](https://github.com/paperclipai/paperclip/pull/12862), and
[#12887](https://github.com/paperclipai/paperclip/pull/12887).

**Expected behavior**

The Daytona image and controller provider package must come from the
same verified build. A package extracted from the immutable image must
pass the existing manifest, source revision, lockfile, binary, bridge,
and artifact checks before a native Claude run starts.

**Steps to reproduce**

1. Set `PAPERCLIP_E2E_DAYTONA_IMAGE` to the old immutable image digest.
2. Set `PAPERCLIP_RUNNER_REMOTE_PROVIDER_PACK_PATH` to a package built
from a different source revision.
3. Run a native Claude Daytona cell.
4. Observe provider package verification failure followed by the large
staging upload.
5. Build the image with the stale Dockerfile lock checksum and observe
the checksum failure.

**Paperclip version or commit**

`3b8df3dcd6f99e99277faa45f3351989edb5c239`.

**Deployment mode**

Daytona native runner E2E.

**Install method**

Built from source.

**Agent adapter(s) involved**

Claude Code through the native ACPX runner.

**Database mode**

Not database-related.

**Access context**

Not applicable to the setup failure.

## What Changed

- Refreshed `PAPERCLIP_RUNNER_LOCK_SHA256` in
`docker/daytona-runner/Dockerfile` to the resolved lockfile checksum.
- Added a local guide for extracting the provider package from an
immutable verified image with Docker.
- Documented the required provenance checks and the expected
manifest-matched runtime log.
- Documented that cold package upload coverage must remain separate from
recovery coverage.
- Pinned the preview-service test guest to the test runner’s Node
executable and logged guest startup and bound ports for readiness
diagnostics.

## Verification

- 442 focused E2E tests passed.
- Typecheck passed.
- `pnpm build` passed.
- Five provider-pack reuse tests passed.
- Six Daytona image contract tests passed.
- Fixed Daytona campaign
[35740613581](https://github.com/paperclipai/paperclip/actions/runs/35740613581)
passed.
- The overall hiring campaign
[35739993219](https://github.com/paperclipai/paperclip/actions/runs/35739993219)
failed because of an unrelated Mini metadata failure; its Claude cell
passed.
- Full local `pnpm test:run` was attempted but did not complete. The
isolated Postgres install was repaired and its 15-test probe passed.
- After the fixture change, all seven preview reservation tests passed
locally and CI server shard 6/12 passed on `20234f75f`. This removes
login-shell Node resolution variance; the exact cause of the earlier
CI-only timeout is not established.
- CI run
[35766034635](https://github.com/paperclipai/paperclip/actions/runs/35766034635)
passed on `20234f75f`. All server, runner, browser, build, and typecheck
gates passed.
- The signoff browser case initially failed waiting for an approver run.
All five signoff tests passed locally without changes; the one allowed
CI retry passed all 19 shard tests. This is recorded as an intermittent
failure, not a demonstrated product fix.
- Greptile reviewed `20234f75f`: 5/5, no actionable findings.
- [Follow-up
report](https://pages.paperclip.ing/runner-daytona-hiring-20260922/)
includes timings, evidence links, and the remaining hiring configuration
failure.
- Review the immutable image source revision and extracted
`provider-pack.json` before another paid recovery run.

## Risks

- The Dockerfile checksum gate intentionally fails when the resolved
lockfile changes. A future dependency change must refresh the reviewed
checksum with the image change.
- The local extraction guide requires Docker and a pullable immutable
image.
- An image built from an older source revision can still fail runtime
manifest verification. The guide does not bypass that check.
- The change does not alter runner prompts, approval policy, or recovery
behavior.

## Model Used

OpenAI Codex using the primary GPT-6 backend; the exact backend
deployment ID is not exposed. Repository analysis and code execution
used tool access. Assistance also came from OpenAI gpt-5.6-luna.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-22 13:49:00 -05:00
Devin FoleyandPaperclip 5f1100e3b3 refactor(server): remove retired operator UI snippet injection (#13789)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Operators can extend its interface through trusted plugin UI
contributions.
> - The server also accepts executable HTML through two legacy
environment settings.
> - This older path bypasses the plugin installation and lifecycle
model.
> - This pull request removes snippet injection from static and
development pages.
> - Operators must migrate existing integrations before upgrading.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

Retire the operator HTML injection path from the server. Related public
changes: #13168, #13245 and #13496 introduced the legacy settings;
#13646 supplies the generic plugin host contract.

**Current behavior**

Managed instances append operator-supplied HTML or a decoded script body
to every page. The same integration can use supported trusted plugin UI
slots.

**Proposed behavior**

Ignore both retired settings. Serve normal branded static and
development HTML, and keep plugin contributions unchanged.

**Breaking changes**

Installations that rely on `PAPERCLIP_CLOUD_UI_SNIPPET` or
`PAPERCLIP_CLOUD_UI_SNIPPET_B64` must migrate before upgrading. The
maintainer-owned staging and production deployments have completed the
migration prerequisite. Other operators must migrate their integrations
before adopting this change.

## What Changed

- Remove the snippet injector and its static/dev rendering integration.
- Remove injector-specific tests and retain a regression that old
settings no longer change served HTML.
- Replace setup instructions with a retirement and plugin-migration
note.

## Verification

- Rebased onto current master; `pnpm exec vitest run
server/src/__tests__/static-index-html.test.ts
server/src/__tests__/vite-html-renderer.test.ts`: 5 tests passed.
- With repository-pinned Rust/Cargo installed, full `pnpm -r typecheck`
and `pnpm build` pass after rebase.
- Previous full local `pnpm test:run` encountered unrelated macOS
runtime-cache rename `EACCES` errors and a missing AgentMail skill path;
a focused reproduction confirmed 5 failures / 95 passes. That is not a
passing full-suite result. The unchanged focused suites, build and
typecheck were repeated after rebase; the full local suite was not
repeated. All 54 refreshed GitHub checks/contexts passed on
`bd62bff63f9d7980bfd10e54cb0102693d27ecfb`; fresh Greptile is 5/5 with
no unresolved threads.
- No UI component styling, database or API contract changed.
- Maintainer approved the remaining rollout and cleanup. Staging snippet
retirement and sleep/wake verification are complete. Production
migration and removal of the legacy settings are complete for serving
tenant instances. Remaining old warm inventory is excluded from new
signups until configuration reconciliation completes. The maintainer
authorized upgrading the remaining old deployments and clearing their
pins.

## Risks

- Removing the settings disables integrations that still depend on them;
operators outside the completed maintainer rollout must migrate before
upgrading. This is an intentional behavior change, documented at the
existing setup-doc path.
- Keep a previous image and its configuration for rollback. Existing
browser tabs need a refresh to unload already-injected code.
- Plugin UI remains trusted same-origin code. This does not add a
security sandbox or change ordinary branding.

## Model Used

- OpenAI GPT-6 (Codex; exact deployment variant and context-window size
are not exposed in this session). Reasoning, repository inspection,
local code execution and GitHub tools.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused rendering tests;
unrelated local full-suite failures are disclosed above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All refreshed checks pass on
`bd62bff63f9d7980bfd10e54cb0102693d27ecfb`
- [x] Fresh Greptile is 5/5 on the current head, with no unresolved
findings
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-22 11:18:46 -07:00
DottaandPaperclip 3d78e3a4ec fix(runner): keep warm sessions alive with managed GitHub access (#13815)
## Thinking Path

> - Paperclip manages AI agents and their work.
> - The native Runner keeps a live provider process between task turns.
> - Managed GitHub access used a token tied to one run.
> - A new run forced Paperclip to replace that process to replace its
token.
> - This PR gives the session a stable credential transport and binds
each operation to the active run.
> - The agent can keep its process while Paperclip checks current
identity and grants.

## Linked Issues or Issue Description

Follow-up to #13738. Related credential-rotation work: #11770 and #8208
use process replacement for other adapter credentials; this change
applies to managed GitHub access in the native Runner.

**What happened?**

A configured GitHub connection forced a warm native provider process to
close at each new run. The saved conversation survived, but the live
process did not.

**Expected behavior**

Keep the warm provider process. Resolve GitHub access for the current
run when each command starts. Deny access while idle or after the run
ends.

**Steps to reproduce**

1. Configure managed GitHub access for a native Runner agent with a warm
session.
2. Complete a turn, then send another message to the same task.
3. Observe the provider process close with the reason `warm native
session configuration changed`.

**Paperclip version or commit**

Reproduced on master `8326e33ad`. Rebased onto `e3d8fb087` before
submission.

**Deployment mode**

Local and remote native execution, including the sandbox callback
bridge.

## What Changed

- Move configured native GitHub transport and launcher ownership from
the run to the provider session.
- Bind the broker only after the executor acquires session ownership.
Clear that binding when the run exits.
- Keep the shared live-run, identity, grant, and trust-policy checks for
each credential request.
- Reject wrong scopes, idle requests, and credential responses that
arrive after their run binding changes.
- Retire transport and launcher files with the provider session. Keep
anonymous commands available if bridge startup fails.
- Add red/green executor tests, real subprocess and callback-bridge
tests, and database checks. Update the runtime documentation.

## Verification

- Before the fix, both new local and remote warm-session reuse tests
failed.
- After the fix, 435 targeted tests passed across the executor, broker,
launcher, token, and database suites.
- A real long-lived test process kept the same PID and original
environment across two runs, including through the production callback
bridge on local test processes.
- Server typecheck and TypeScript compilation passed.
- Full workspace typecheck and build passed. Server typecheck passed
again after the review fix.
- The fallback-logging regression failed before the fix; all 9 broker
tests pass afterward.
- The exact chat sidebar browser scenario passed locally. The initial CI
timeout showed failed Vite module downloads; all eight browser shards
pass on the latest commit.
- All 53 latest-head checks passed, including the full CI test matrix
and security checks (two unrelated conditional checks skipped).
- The duplicate full local test run was stopped after CI passed; it is
not claimed as a completed local pass. Targeted local tests, workspace
typecheck/build, and the browser scenario passed.
- Greptile reviewed the latest commit at 5/5 with no unresolved
findings.
- No fresh paid provider or Daytona campaign has run for this change.

## Risks

- The broker now lives as long as the provider session. Tests cover idle
denial, late cleanup, late responses, shutdown, and failed startup.
- Its in-memory authority does not survive a controller restart.
Existing checkpoint and process-recovery rules still apply.
- Raw GitHub credentials remain confined to individual command
processes. The session transport token cannot select a different task,
agent, company, or run.
- No database migration or public API change.

## Model Used

OpenAI Codex, GPT-6, with reasoning, terminal tools, and code execution.
The exact serving model ID and context-window size are not exposed in
this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-22 13:07:13 -05:00
DottaandPaperclip 74a9730acb fix: continue native agent chats after worker loss (#13813)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agent Chat uses native workers to run Claude and Codex
conversations.
> - A worker crash leaves a cleanup hold because its provider did not
acknowledge suspension.
> - A new user message must not reuse that unverified session or repeat
old tool calls.
> - The existing continuation path can preserve history and start a
fresh session, but local cleanup ownership remained held.
> - This pull request verifies the stopped local owners and releases
only their cleanup hold for a new user turn.

## Linked Issues or Issue Description

Refs #13775.

**What happened?**

After a native worker crashed, both providers retained cleanup
quarantine. A saved plan survived, but the conversation could not
produce another answer.

**Expected behavior**

Once the old worker and provider process groups have stopped, a new user
message can continue in a fresh session with the saved work and prior
action history.

**Steps to reproduce**

Run the opt-in `agent-chat-qualification` suite with case
`worker-crash-retry` on native Codex and native Claude. The fixture
saves a plan, kills the exact worker through a Linux pidfd, releases a
local read-only brief, and sends a new message.

## What Changed

- Verify the exact local worker stop receipt, provider identity
receipts, released leases, and retained state before retiring a native
cleanup hold.
- Recheck process liveness and state before admission. Keep the old run
and durable session files intact.
- Use the existing explicit conversation continuation path. Generic
Retry remains blocked for cleanup quarantine, including on the old
failed-run marker after a successful continuation.
- Extend the live oracle to require a successful fresh session, correct
predecessor context, unchanged plan, one original message, and one
answer containing a reference introduced after the crash.
- Add physical-proof and database-backed admission tests. Document the
precise qualification scope.

## Verification

- Live Product E2E: **2/2 passed**, **2/2 cleanup passed**, with real
native `gpt-5.6-sol` and `claude-sonnet-5`, Chromium, server, database,
and public APIs.
- Core recovery proof source:
`3592b04c2bc76e23795fcdf964720e38a409dc4d`. Suite definition version 8:
`9867367994d81a0c726956d91f2c7fddab6417a12f41b5cef3f7e62f3be417da`.
- [Core recovery
campaign](https://github.com/paperclipai/paperclip/actions/runs/35741746990)
· [Public evidence
report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35741746990-1/).
- Both cells verify the real crash boundary, blocked generic Retry,
unchanged saved plan, one original prompt, one fresh successor with
predecessor context, and one run-attributed answer containing the
post-crash reference. Billing coverage is partial because the crashed
runs did not report complete usage; missing cost is not zero cost.
- [Master
baseline](https://github.com/paperclipai/paperclip/actions/runs/35737364443):
both providers stopped at quarantine, with cleanup passing.
- [Follow-up baseline without the
fix](https://github.com/paperclipai/paperclip/actions/runs/35738638866):
Codex produced a complete red result. Claude reached the same error, but
its artifact upload was canceled.
- [Complete Claude baseline with the same version-8
definition](https://github.com/paperclipai/paperclip/actions/runs/35740555177):
red at cleanup quarantine, cleanup passed, source
`6479a90c5754044356b39a9278b9a3e92ce8e55e`.
- Earlier candidate attempts remain retained:
[first](https://github.com/paperclipai/paperclip/actions/runs/35738449214)
passed Codex and found a Claude fixture wait race;
[second](https://github.com/paperclipai/paperclip/actions/runs/35740408065)
exposed the normalized session-open receipt mismatch. Both corrections
are in the final source.
- Targeted server suites: 570 passed before the final two additional
receipt regression cases. The physical-proof suite, including those
cases, passed 46/46. Eval oracle and catalog: 40 passed. Server and
Product E2E typechecks passed.
- **All 54 PR checks passed on final head `db6f775df`**, including
repository typecheck, build, tests, browser shards, and canary dry run.
The canary job required one retry after its runner received a shutdown
signal. On the earlier core proof head, two timing-sensitive tests
passed in isolation and on a single CI retry.
- Final UI regression checks: 5 passed; UI typecheck and token gates
passed. Updated eval oracle/catalog: 40 passed; eval typecheck passed.
- Final version-9 two-provider campaign: **2/2 passed, 2/2 cleanup
passed**, including the browser assertion that the quarantined
historical run never regains Try again. [Final
campaign](https://github.com/paperclipai/paperclip/actions/runs/35747416013),
source `db6f775dfff405e1514ec02fedb0450d42c7dad2`, definition hash
`bf5abf1cc45cb6dad4e082fbf818b8fa0f4d8c282776a7e98762b22919eadab7`.
[Final public evidence
report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35747416013-1/).
Final screenshots and retained state inspected for both providers; cost
coverage remains partial.

## Risks

- Missing or conflicting stop evidence keeps the conversation blocked.
This change does not kill an unverified process.
- This qualifies new user input after local worker loss. It does not
enable automatic replay, exact-session recovery, remote crash recovery,
or native onboarding defaults.
- Old action outcomes remain part of the continuation. A process exit is
not proof that an action did not happen.
- No schema migration or production prompt change.

## Model Used

OpenAI Codex, GPT-6. The runtime does not expose a more specific model
identifier or context-window size. Used reasoning, repository
inspection, code editing, shell tools, and test execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-22 13:01:45 -05:00
Devin FoleyandPaperclip 110d176fc9 fix: preserve chat message bindings in review recovery (#13818)
## Thinking Path

> - Paperclip manages AI agents and their work.
> - External chat messages can start work on tasks that remain in review.
> - Paperclip queues one recovery run when a task loses its review path.
> - That recovery keeps the chat source but loses the admitted message IDs.
> - The authorization check then rejects the recovery before execution starts.
> - This change retains the message IDs so the existing check can verify current access.

## Linked Issues or Issue Description

Refs #13809.

**What happened?**

A successful external-chat run can leave an active task in review without a maintained review path. Its automatic recovery then fails with `reviewed_chat_execution_binding_not_authorized`. The recovery context retains `chat:slack` but drops `wakeCommentIds`.

**Expected behavior**

An eligible recovery should retain its admitted message references and pass a new authorization check. Missing or revoked access must still prevent execution.

**Steps to reproduce**

1. Finish a chat run whose task remains in review with no maintained review path.
2. Build the bounded review recovery from that run's context.
3. Dispatch the recovery through the reviewed-chat authorization check.

The new regression tests fail before this patch. Related PR #13809 handles answered conversations that become idle. This patch handles recovery when the conversation remains active.

## What Changed

- Retain the admitted message batch for external-chat review recovery. Derive the current comment reference from that batch.
- Share the existing supported-provider selector between recovery and run-bound chat authorization. AgentMail remains on its separate email inbox path.
- Keep the existing authorization check. Do not copy prior checkout, authorization, session, or prompt state.
- Test Slack and Discord recovery, missing batches, revoked access, and changed task or company bindings.
- Document the recovery authorization contract.

## Verification

- The new tests reproduced the missing-message failure before the fix.
- Six focused suites passed: 257 tests covering review recovery, reviewed-chat authorization, issue liveness, Slack lifecycle, comment-wake batching, and external-chat waits. PostgreSQL integration tests ran against a temporary local PostgreSQL database.
- Focused test files: `review-path-recovery.test.ts`, `heartbeat-reviewed-chat-binding.integration.test.ts`, `heartbeat-issue-liveness-escalation.test.ts`, `slack-conversation-lifecycle.test.ts`, `heartbeat-comment-wake-batching.test.ts`, and `external-chat-wait.integration.test.ts`.
- Server TypeScript check passed with scratch configuration that resolves this checkout's workspace packages. Existing dependency links point to another checkout; the default check reports stale shared-type errors.
- `node scripts/check-module-boundaries.mjs` and `git diff --check` passed.
- `git diff | gitleaks stdin --redact --no-banner` passed. The diff was also checked for private identifiers and user data.
- [Full PR CI](https://github.com/paperclipai/paperclip/actions/runs/35760757895) passed on the latest commit, including build, workspace typecheck, general and serialized tests, Rust checks, and all eight browser shards. The unchanged local-service readiness test and agent-chat page-load assertion passed when their failed shards were retried. The local-service suite also passed locally (6 tests). The first, superseded run lost a chat runner; its stuck browser job was cancelled to unblock the current run.
- Full local workspace typecheck, tests, and build were not run. The machine has about 2 GiB free, and those commands include Rust builds. The full PR CI checks passed before merge.

## Risks

Small context-construction change. Dispatch still checks current execution ownership and access for every admitted message. The recovery remains bounded to one attempt per consumed path. No schema changes. Existing failed runs are not retried by this patch.

## Model Used

OpenAI GPT-6 (Codex), with tool use and code execution.

## Checklist

- [x] I have included a thinking path that traces from project context to this change
- [x] I have specified the model used (with version and capability details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before requesting merge

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-22 10:46:13 -07:00