Commit Graph
2130 Commits
Author SHA1 Message Date
nickyleachandPaperclip ba5c01e08b merge: bring retry module up to current master
Keep the retry module wired through the heartbeat service after the service split.

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-09 15:05:09 +00:00
nickyleachandPaperclip 7da508f433 refactor(server): merge retry module with current heartbeat service
Move retry scheduling into the retry module after the heartbeat service split. Keep the source run issue scope on each retry.

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-09 15:04:37 +00:00
DottaandPaperclip 835a022936 perf(server): wake delivery queues from transaction-aware work signals (#15625)
## Thinking Path

> - Paperclip manages AI-agent work and its delivery to users and
agents.
> - Five delivery queues scan the database even when empty.
> - Each queue already has durable rows; empty polling wastes idle
hosting capacity.
> - Producers can register intent inside their existing database
transaction.
> - Central transaction tracking can wake consumers after the outer
commit without extra caller wrappers.
> - A lost response needs reconciliation against PostgreSQL before an
empty queue is safe to leave idle.
> - This change uses one scheduler for outstanding work and lets
settled, empty queues go quiet.

## Linked Issues or Issue Description

Supersedes the closed #15614. Related idle-safety work: #15522 and
#15599.

**What existing behavior does this improve?**

Delivery scheduling for feedback exports, chat completions, connection
continuations, question answers, and tool-action receipts.

**Current behavior**

Feedback scans every five seconds. Four delivery queues scan on the
heartbeat interval, normally 30 seconds. Startup and some event paths
also scan them.

**Proposed behavior**

Each producer awaits one intent registration before its queue write. The
existing `createDb` transaction boundary tracks nested savepoints and
wakes the consumer after outer settlement. Startup scans restore durable
work. Pending rows, failures, and unresolved transactions retain retry
deadlines. Empty, settled queues have no timer or recurring scan.

**Reason and benefit**

Normal delivery starts after commit. Central tracking removes
caller-specific post-commit plumbing. PostgreSQL transaction status
resolves lost responses without new tables or a permanent uncertainty
latch.

**Breaking changes**

No API or schema changes. The notification scope is one DB owner and its
dedicated child pools. Direct SQL, separate roots, and other processes
require an explicit wake/recovery integration. This path requires
PostgreSQL 14+ transaction-information functions; embedded PostgreSQL
uses 18.

## What Changed

- Add one named-deadline scheduler and app-owned coordinator for five
existing queues.
- Instrument `createDb().transaction()` centrally, including nested
savepoints. Dedicated child connections share their owner's signal
scope.
- Register work before the five queue insert paths. The first
registration obtains the outer XID; unrelated transactions issue no
extra SQL. Tool receipt insertion gains a small transaction around its
existing insert.
- Query `pg_xact_status` only for rejected, tracked transactions. Probe
before scanning the queue. Retain the idle hold while PostgreSQL still
reports an in-progress transaction or the probe fails. Release it after
settlement and reconciliation.
- Remove unconditional delivery recovery scans and caller-specific
completion post-commit actions. Retain the existing task-scoped
completion activity fast path and durable claims.
- Coalesce wakes without postponing an earlier deadline. Retry
outstanding rows and failures; suppress dispatch during idle drain while
allowing transaction and queue reconciliation; suppress both during warm
standby. Cancel deadlines and await active delivery sweeps on shutdown.
- Serialize feedback flushes. Vote routes return after saving. Uploads
have a 30-second deadline and shutdown cancellation; unfinished exports
remain recoverable.
- Document the writer contract and update the dated sleep inventory.

## Verification

- Final head: `a4609956841107ad60c4cb50fe2acb05636b1363`.
- Passed: all final-head hosted checks, with 54 successes and two
intentional skips. The full test matrix, typecheck, build, canary, and
aggregate verification passed in [run
37940948440](https://github.com/paperclipai/paperclip/actions/runs/37940948440).
- Passed: Greptile 5/5 on the same head, with a successful check run and
zero unresolved review threads. The branch is mergeable.
- Passed locally after the final rebase: 50 coordinator, scheduler, and
startup tests; server typecheck; server build.
- Passed locally before the final import-only rebase: repo-wide
typecheck and build; 237 PostgreSQL producer/delivery and database
signal tests; feedback, native-question, and idle-safety regression
tests.
- Real PostgreSQL tests hold transactions open after an injected client
response loss, then commit or abort. They verify that an empty scan
cannot clear an in-progress transaction. Another test terminates a real
backend and verifies recovery without replay.
- Coordinator tests cover an hour with no empty-queue timers,
pending/error retries, worker replacement, concurrent writes,
earliest-deadline preservation, rollback during idle drain, committed
work during drain, standby, and shutdown.
- Local `pnpm test:run` was started and stopped after embedded
PostgreSQL startup failures appeared. It did not finish and is not
reported as a pass. Later local database reruns had skipped suites;
those skips are not claimed as verification. The earlier PostgreSQL
tests did execute and pass, and the final hosted full matrix passed.

## Risks

- One XID query is added per transaction that registers delivery work.
Nested transactions share that ID. Registration must be awaited before
writing; new enqueue paths must follow this contract.
- Work signals are local to a DB owner and its dedicated child pools.
Independent roots, direct SQL, or other processes are not observed.
Startup scans recover already committed rows, but do not fence late
transactions from a previous process. Cross-process ownership and
database failover require separate work.
- A database outage or genuinely unresolved transaction keeps an idle
hold until status can be reconciled. No timer clears uncertainty by
assumption.
- These changes remove five empty polling loops. They do not implement
whole-instance sleep, database teardown, or an external waker. The
earliest deadline covers only the registered queues.
- Waiting tool reviews retain their existing retry cadence until their
receipts can be delivered.
- Feedback vote responses no longer wait for the remote upload. Sharing
consent and the saved vote response are unchanged.
- Real hosting-provider sleep and cost savings have not been measured.

## Model Used

OpenAI Codex, GPT-6. Used reasoning, repository inspection, code
editing, and command execution. The runtime did not expose the exact
model revision or context window.

## Checklist


- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-09 10:01:26 -05:00
DottaandPaperclip 0ca4b7c4c0 refactor(heartbeat): extract run cancellation and cleanup (#15685)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The heartbeat service admits, executes, and settles agent runs.
> - Earlier extractions separated workspace, preparation, state, retry,
recovery, queue, and lifecycle handling.
> - Cancellation and resource release still live inside the main
service.
> - These operations share execution ownership and saved-comment
continuation rules.
> - This pull request moves them into one run-control module and adds
boundary tests.
> - Reviewers can inspect cancellation and cleanup separately from
adapter execution.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

The structure and testability of heartbeat cancellation and cleanup.

**Subsystem affected**

server/ — orchestration services.

**Current behavior**

Stop, pause and budget cancellation, environment lease release, and
saved-comment delivery live in `heartbeat.ts`.

**Proposed behavior**

Keep these operations in `server/src/services/heartbeat/run-control.ts`.
Bind their database and existing callbacks with
`createHeartbeatRunControl`.

**Reason and benefit**

This removes 1,115 lines from `heartbeat.ts`, leaving 10,443 lines. The
new module has 1,360 lines. Related cancellation and cleanup rules stay
together in one file.

**Breaking changes**

None. Function bodies, service methods, public helpers, status
conditions, and side effects keep their behavior.

**Additional context**

Continues the merged lifecycle extraction in #15672. Related
cancellation work: #15212. Searches of open and closed issues and PRs
found no duplicate run-control extraction. This does not add a roadmap
feature.

## What Changed

- Move run cancellation, pause and budget cancellation, lease release,
issue-lock release, and saved-comment resumption into
`heartbeat/run-control.ts`.
- Supply lifecycle, queue, recovery, and cleanup callbacks through an
explicit dependency interface.
- Keep executor and cancellation maps shared by all service instances.
Preserve callback construction order with forwarding functions.
- Keep the existing helpers available from `heartbeat.ts`.
- Add 15 tests for inert construction, cancellation gates, shared Stop
barriers, failed termination, warm-resource retention, pending wake
batches, project scope, and stale budget enforcement.
- Document the module boundary in `doc/DEVELOPING.md`.

## Verification

- Before extraction: **298 tests passed across six existing cancellation
and cleanup suites**.
- After extraction: **313 tests passed across seven files**, including
the 15 new module tests and real PostgreSQL coverage.
  ```sh
pnpm exec vitest run server/src/services/heartbeat/run-control.test.ts
server/src/__tests__/heartbeat-native-runner-cancellation.test.ts
server/src/__tests__/heartbeat-run-terminalize-before-release.test.ts
server/src/services/run-cancellation.test.ts
server/src/services/adapter-execution-control.test.ts
server/src/services/explicit-native-continuation.test.ts
server/src/__tests__/native-cancellation-request.integration.test.ts
  ```
- Wider queue, recovery, reassignment, comment delivery, and native
cancellation verification: **139 tests passed across nine files**,
including the final module tests. The recovery suite first hit its
20-second database setup timeout while local builds were running. After
removing verified orphaned PostgreSQL headers, its isolated run passed
**380 of 381 cases**. One teardown assertion exceeded its one-second
wait for a successor to settle; that exact case passed in a filtered
rerun. Combined focused coverage is **818 distinct passing tests across
16 files**. The module suite appears in both runs.
  ```sh
pnpm exec vitest run server/src/services/heartbeat/run-control.test.ts
server/src/__tests__/heartbeat-comment-wake-batching.test.ts
server/src/__tests__/heartbeat-process-recovery.test.ts
server/src/__tests__/heartbeat-native-cleanup-admission.test.ts
server/src/__tests__/heartbeat-lock-release-on-reassignment.test.ts
server/src/services/heartbeat/queue.test.ts
server/src/services/heartbeat/recovery.test.ts
server/src/services/heartbeat/retries.test.ts
server/src/services/heartbeat-stop-metadata.test.ts
server/src/services/native-runtime/native-cancellation-request.test.ts
  ```
  ```sh
pnpm exec vitest run
server/src/__tests__/heartbeat-process-recovery.test.ts
pnpm exec vitest run
server/src/__tests__/heartbeat-process-recovery.test.ts -t "does not
adopt unrelated queued comments for a non-coalescing recipient after
Stop"
  ```
- Full `pnpm -r typecheck` and `pnpm build` passed on head
`b26efd5f9de390ef8a675f84980f49f7fb7799c2`.
- Full local `pnpm test:run` was started on this head and stopped after
the complete CI suite passed. It did not finish and is not counted as a
full local pass.
- All **54 final-head checks passed**; two optional Storybook checks
were skipped. This includes the full general and serialized test matrix,
browser tests, typechecks, build, Runner verification, and canary dry
run. The complete recovery suite passed in CI. [CI
run](https://github.com/paperclipai/paperclip/actions/runs/37945389261).
- Greptile completed on head `b26efd5f9de390ef8a675f84980f49f7fb7799c2`
with **5/5 and no actionable findings**. There are no review threads or
merge conflicts.
- Structural comparison confirms all 21 moved function bodies, the
cancellation options type, remaining service logic, and all 140 existing
public exports are preserved.

## Risks

Moving closures can break callback construction order or shared Stop
barriers. The factory starts no work and receives the same process-wide
maps and existing callbacks. Tests cover concurrent Stop callers and
failed termination. Native authority, status predicates, provider
receipts, cleanup gates, and budget enforcement keep their original
bodies. This PR has no schema or API changes.

## Model Used

OpenAI GPT-6 via Codex. The exact serving model ID and context window
were not exposed in this session. Used reasoning, repository inspection,
shell tools, and code execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-09 09:59:42 -05:00
DottaandPaperclip 49bbfe3092 fix(server): register native runner transport in HTTP startup (#15682)
Register native Runner transport in the shared HTTP server constructor. Add startup regression tests and custom-launcher guidance. Verified with live OpenCode/OpenRouter tasks and complete CI coverage.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-09 09:53:42 -05:00
DottaandPaperclip 0b80ea17ef refactor(heartbeat): extract run lifecycle handling (#15672)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The heartbeat service admits, executes, and settles agent runs.
> - Earlier extractions separated workspaces, preparation, state,
retries, recovery, and queue admission.
> - Run status, progress, liveness, and completion handling still live
in the main service.
> - These operations share status predicates, event writes, and
completion policies.
> - This pull request moves them into one lifecycle module and adds
boundary tests.
> - Reviewers can inspect lifecycle policy separately from adapter
execution.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

The structure and testability of heartbeat run lifecycle handling.

**Subsystem affected**

server/ — orchestration services.

**Current behavior**

Run status writes, progress, events, liveness, completion handoffs,
issue-comment finalization, and runtime settlement live inside
`heartbeat.ts`.

**Proposed behavior**

Keep these operations in
`server/src/services/heartbeat/run-lifecycle.ts`. Bind the database and
reporting, recovery, and wakeup callbacks with
`createHeartbeatLifecycle`.

**Reason and benefit**

This removes 2,682 lines from `heartbeat.ts`, leaving 11,558 lines. The
lifecycle module has 2,982 lines and keeps related policy together in
one file.

**Breaking changes**

None. Existing function bodies, service methods, public exports, status
conditions, and side effects keep their behavior.

**Additional context**

Continues #15626 and the earlier merged heartbeat extractions. A search
of open and closed issues and PRs found no duplicate lifecycle
extraction. This does not add a roadmap feature.

## What Changed

- Move status transitions, run events and progress, liveness
classification, completion handoffs, plan-resume failure reporting,
issue-comment finalization, and runtime/cost settlement into
`heartbeat/run-lifecycle.ts`.
- Supply reporting, recovery, and admission callbacks through an
explicit dependency interface. Preserve construction order with
forwarding callbacks.
- Keep adapter execution, cancellation, and terminal telemetry reporting
in the main service.
- Add 16 tests for inert construction, continuation gates, late
progress, status races, native ownership, usage metadata, revoked wakes,
event sequencing, and redaction.
- Update the existing publication-order source assertion to read the
extracted event writer. Keep its persist-before-publish check.
- Document the lifecycle module boundary in `doc/DEVELOPING.md`.

## Verification

- Before extraction: 87 tests passed across seven existing suites.
- After extraction: 103 tests passed across eight files, including the
16 new module tests and real PostgreSQL coverage.
  ```sh
pnpm exec vitest run server/src/services/heartbeat/run-lifecycle.test.ts
server/src/__tests__/heartbeat-run-status-payload.test.ts
server/src/__tests__/heartbeat-run-event-sequencing.test.ts
server/src/__tests__/heartbeat-run-terminalize-before-release.test.ts
server/src/__tests__/heartbeat-run-lease-release-terminalization.test.ts
server/src/__tests__/heartbeat-cost-accounting.test.ts
server/src/__tests__/heartbeat-issue-liveness-escalation.test.ts
server/src/__tests__/run-liveness.test.ts
  ```
- Additional recovery, cancellation, plan-resume, summary, and handoff
coverage: **595 tests passed across 14 files**. The module suite appears
in both runs. With the publication-order suite below, combined focused
coverage is **694 distinct tests across 22 files**, with no skips.
  ```sh
pnpm exec vitest run server/src/services/heartbeat/run-lifecycle.test.ts
server/src/services/heartbeat/recovery.test.ts
server/src/services/heartbeat/retries.test.ts
server/src/services/heartbeat/queue.test.ts
server/src/services/recovery/successful-run-handoff.test.ts
server/src/services/recovery/review-path-recovery.test.ts
server/src/services/heartbeat-run-runtime-status.test.ts
server/src/__tests__/heartbeat-process-recovery.test.ts
server/src/__tests__/heartbeat-native-runner-cancellation.test.ts
server/src/__tests__/heartbeat-native-cleanup-admission.test.ts
server/src/__tests__/heartbeat-accepted-plan-workspace-refresh.test.ts
server/src/__tests__/heartbeat-runtime-state.test.ts
server/src/__tests__/heartbeat-run-summary.test.ts
server/src/__tests__/heartbeat-context-summary.test.ts
  ```
- The publication-order and lifecycle suites passed together: **28 tests
across two files**. The first CI run exposed an assertion that still
read the old event-writer source path. That assertion now reads the
lifecycle module.
  ```sh
pnpm exec vitest run
server/src/services/chat-publication-reconciliation.test.ts
server/src/services/heartbeat/run-lifecycle.test.ts
  ```
- After resolving the import conflict with the Pi Runner integration and
rebasing onto `57e977be7`: **44 tests passed across four files**,
including the updated source assertion, lifecycle module, run-status
payloads, and accepted-plan workspace refresh.
  ```sh
pnpm exec vitest run
server/src/services/chat-publication-reconciliation.test.ts
server/src/services/heartbeat/run-lifecycle.test.ts
server/src/__tests__/heartbeat-run-status-payload.test.ts
server/src/__tests__/heartbeat-accepted-plan-workspace-refresh.test.ts
  ```
- Full `pnpm -r typecheck` and `pnpm build` passed on final head
`bca87df089c541450932f6b678ca776942d267b2`. Full local `pnpm test:run`
was restarted on this head and stopped after the entire CI suite passed.
It did not finish and is not counted as a full local pass.
- The earlier full local run had one transient public MCP Cloud
bootstrap failure. All **82 public MCP tests passed in isolation**. The
first local run was stopped before the source-assertion fix and rebase.
- All **54 latest-head checks passed**; two optional Storybook checks
were skipped. One general-test shard passed all 1,314 tests but failed
during runner temporary-directory cleanup with `EACCES`. One rerun
passed. [CI
run](https://github.com/paperclipai/paperclip/actions/runs/37940997283).
- Greptile completed on final head
`bca87df089c541450932f6b678ca776942d267b2` with **5/5 and no actionable
findings**. There are no review threads or merge conflicts.
- Structural comparison confirms 63 module function bodies and ten
declarations match the originals. Remaining service function bodies and
all 140 public exports are preserved.

## Risks

Moving closures can break callback construction order or status
settlement. The factory receives explicit callbacks and starts no work
during construction. Compare-and-set predicates, native ownership holds,
revoked-wake guards, event sequencing, redaction, and completion
policies keep their original bodies. This PR has no schema or API
changes.

## Model Used

OpenAI GPT-6 via Codex. The exact serving model ID and context window
were not exposed in this session. Used reasoning, repository inspection,
shell tools, and code execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-09 09:26:45 -05:00
DottaandPaperclip 57e977be72 feat: integrate Pi 1.0 into the experimental Runner (#14921)
## Thinking Path

> - Paperclip manages AI agents and their work.
> - The experimental Runner owns provider processes and durable
sessions.
> - Pi needs working task execution and human controls.
> - The five-PR stack must preserve changes already on master.
> - Each layer now carries the complete integrated source for a safe
sequential fallback.
> - This PR belongs to native GitHub stack #15602, ending at #14956.

## Linked Issues or Issue Description

Refs #14436, #14631, #14743 and #14956.

Ship Pi 1.0 through the experimental Paperclip Runner. The five PRs are
#14921, #14922, #14923, #14924 and #14956. The user authorized the
complete merge after checks pass. Existing `pi_local` execution is
unchanged. Accounting and wider provider/platform qualification remain
deferred.

## What Changed

- Recover missing final replies after workspace finalization changes
owners, using accepted-turn evidence without rerunning work or granting
external-chat publication.
- Preserve the admitted Pi instruction root across warm runs, while
retaining changed-root rejection.
- Give Pi a bounded 15-second default shutdown grace so stop, drain
acknowledgement and durable suspension can complete. Explicit deadlines
and other providers retain their existing behavior.
- Integrate the Pi 1.0 runtime and master contracts.
- Use Pi profile 22. Preserve explicit caller-selected models and exact
native thinking levels. Keep Pi's wrapper, helper, extension and
question/control behavior unchanged from the qualified profile-19
runtime.
- Preserve master's Dot lifecycle and consent fields, configured task
environment, status guards and current Codex/Claude dependency versions.
Cursor stays qualified. Copilot stays pending; profile 17 binds the
changed shared protocol validation sources.
- Exclude general AWS IAM credentials from Pi static/custom provider
bindings and selected task projections; preserve the provider-scoped
Bedrock bearer key. Profile 21 is retained as historical provenance.
Rust and cloud install probes use the current declaration.
- Patch bundled brace-expansion 5.0.9 to the exact official 5.0.12
payload. Pin the patch and complete runtime closures. Include the patch
in normal installed setup tooling. Keep the upstream Pi shrinkwrap as
provenance and permit only this exact security correction.
- Include current attestation files in the Docker build context. Keep
the repository lockfile unchanged from master. CI and private image
builds resolve manifest changes before their frozen installation.

## Verification

- Full local `pnpm -r typecheck` passes, including Runner Rust, server
and UI. Focused integration checks pass: 194 Runner
admission/environment tests, 63 profile/credential tests with one
expected skip, 152 Dot/UI configuration tests, and Pi transcript/notice
tests.
- Full local `pnpm build` passes on the final source.
- Fresh final-source checks pass: all 698 Rust workspace tests (32
binaries), 156 credential/profile/controller tests with one expected
skip, Runner TypeScript typecheck, and 20 package/setup/sandbox tests.
- The profile-21 Pi materializer passes on the native host with the
official pinned Node 24.21.0 and its npm. It verifies all 150 locked
packages, the patched dependency and the exact closure. Setup/package
bundle tests and UI token gates pass.
- The old hashes were reproduced for all three supported targets before
calculating the patched graph. New closure hashes are darwin-arm64
`282022db10150c6632b3444df421342e7d534bdf5d5fb1097a2e79d0625a2bcf`,
darwin-x64
`64e251e19009f755c0b04f73ce2138246faab71a961b0f13d75ebfcc34bef12e`, and
linux-x64
`713b1fdff42fb56a1518bdc084f181d70bee8ebadc3e4b1d76321ed9108c8410`.
Independent native platform execution is separate from graph identity
reproduction.
- Historical cloud qualification remains unchanged: all seven core cases
pass on shipping source `10dc43c9ec65d88c2f782d62afb296d09494f215`,
harness `1a4408a48cfb5a1f094a311141c257c92cd7a893`, image
`sha256:5b3a775b383591bda1b0c1889e509acc70ce7f37c53f09733c81d59037f02280`,
and accepted Sonnet 4.6/low fixture. All 215 canonical files and all
seven cleanup checks pass independent verification. These are profile-19
results and are not relabeled as fresh profile-22 runs.
- Current Pi digest:
`sha256:e92078bee3c23bec4100aa589013a44613d054cd686826534025d8019e9f39a9`.
[The readiness
plan](https://github.com/paperclipai/paperclip/blob/codex/pi-production-readiness/doc/plans/2026-10-02-pi-production-readiness.md)
preserves campaign and failed-attempt provenance.
- Merge only after every PR's current-head CI and fresh review pass.
Linux CI covers the full suites, build and browser tests. The local
embedded Postgres API-authority suite cannot start on this macOS/Node 26
host, so Linux CI must confirm that suite.

### Fresh profile-22 core qualification — 2026-10-08

All seven accepted core cases pass canonically on Pi profile 22, with
`openrouter/anthropic/claude-sonnet-4.6` and native-confirmed low
thinking. This model is a fixture; production accepts the caller's
explicit Pi provider/model.

Runtime/install source: `3241a992f2a7703e59e97ed0fd3e5d6405de4401`.
Frozen accepted harness: `1a4408a48cfb5a1f094a311141c257c92cd7a893`.
Immutable cloud image:
`ghcr.io/paperclipai/paperclip-daytona-runner@sha256:506f22db7edd78f37c0c40bec1cc084af1850455026dbf467194bfbb8fcef141`.
Pi digest:
`sha256:e92078bee3c23bec4100aa589013a44613d054cd686826534025d8019e9f39a9`.

[Hosted Linux image and clean-install
verification](https://github.com/paperclipai/paperclip/actions/runs/37868328023)
passes, including all 20 source-bound archives, normal CLI/Pi setup,
companion import and the production pack reader. This exact installation
source includes the latest master integration and the corrected Pi warm
instruction-root fence. Full local typecheck/build and current-head
hosted CI verify the final stack. All 13 focused real-root regressions
pass. The full local executor suite passed 662 tests; 15 database tests
could not start the Mac embedded PostgreSQL service. Hosted Linux CI
passes the full required verification and E2E checks. These fresh
results keep their own source identity; profile-19 results remain
historical.

| Core path | Canonical campaign | Retained archive SHA-256 |
| --- | --- | --- |
| File edit, validation, download and Done |
`pi-core22-replyfix-0-1791511228` | 23 files;
`a473e8603a3dd4737863291f8d3d1e392391f0b16d433c3e0e0e9d8baf7a97b0` |
| Pending question and controller restart |
`pi-core22-replyfix-1-1791511376` | 33 files;
`6b829c4eb74e1f32a89c692a4ae7130dbfc1c6d3cf13915effe2103d9e242c8e` |
| Three-turn session/process/workspace continuity |
`pi-core22-replyfix-2-1791511587` | 23 files;
`7a87021f8f9a3fdd3c58bb4467f8d82c635e3ea4795d6e75f144d9aa14818df8` |
| Four typed questions and browser reconnects |
`pi-core22-replyfix-3-1791511881` | 42 files;
`9e31755252be1f4f9cb0626c984c142d4d1ae5f5bee3a7af08444db8d12c280a` |
| Plan approval and completion | `pi-core22-replyfix-4-1791512031` | 22
files;
`a0383ce1aab38e7b5a25ce0e9dd3bebea5c037ebd96ae6b29dae19015da2ae2c` |
| Same-turn steering and permission denial |
`pi-core22-replyfix-5-1791512261` | 39 files;
`c929b8c7070f0b66aedc17e65ca46e6beab1e363926ac9f7e2a75fb250f05949` |
| Stop during pending permission | `pi-core22-replyfix-6-1791512390` |
33 files;
`7f0a58ae0f4d5bfc76149435f4e322537089c5bd16e7ffe9b5ad71f10a621a07` |

All 215 canonical files (28714587 bytes) are independently
hash-verified. All seven cleanup grades pass, with no owned runtime
process or temporary root after each case. Automatic retries are zero.
The owned cloud host stopped normally after retention. The prior
profile-22 warm attempt remains failed and separately retained: archive
SHA-256
`1e54eba5ec72b50cee1534b23d1d1d4f21a090006b8a64501ba70db972abfde5`. Its
original canonical classification is preserved. Diagnosis reproduced a
product bug comparing an agent-files root against an unset
checkpoint-only field. The fix stores the admitted physical root
separately from the adopted per-run collection capability. The real-root
regression fails before the fix and passes afterward, including
rejection of a changed physical root. Fixture, grader, model and all
seven accepted case IDs are unchanged; this fresh campaign tests
final-reply publication after file registration first. The intermediate
restart attempt also remains failed and retained: archive SHA-256
`5dcaefdf1d17cf4cd54fd4cf810f45e736667392339b8ce7caf08bb4e225277f`. Its
original canonical classification is preserved. Pi resumed, wrote the
verified answer and completed its task; exact runner suspension was
proven, but idle stop consumed about 5.2s and left under 3s for the
drain acknowledgement. The Pi-only default shutdown grace is now 15s,
preserving a full 5s drain round trip and a finite suspension reserve.
Explicit caller deadlines, other provider defaults, literal drain
receipts and exact suspension identity checks remain unchanged. The
timing regression fails before this correction and passes afterward; all
18 focused settlement tests and Runner typecheck pass. The final-source
file attempt is also preserved as failed (`candidate_failure`), archive
SHA-256
`db6767b6773ea618997927ac77bdb005a5ac81492c7b9c0ffbc900449f829bc9`.
Native edit, validation, exact downloadable artifact and Done/succeeded
all passed, and the exact final reply was durably recorded. A workspace
recovery owner completed before the live heartbeat reached presentation,
leaving that reply absent from task chat. Recovery now materializes only
a completed final reply from the accepted turn of an ordinary internal
Done task, preserving issue/run/contract binding, suppression,
external-chat authorization and same-run deduplication. The database
regression covers the generated file-preparation receipt, suppression,
unapproved external continuation and replay. Server typecheck and all 49
response-selection tests pass; hosted Linux verifies the database
regression because embedded PostgreSQL cannot start on this Mac.

The delayed-final-answer database regression passes on [the final
root-source Linux server
shard](https://github.com/paperclipai/paperclip/actions/runs/37868262553/job/113628594152),
alongside 1,108 passing tests. The first root Runner shard had one
unchanged durable-resume test exceed its 5-second timeout; the identical
top-source shard and the isolated exact test passed. One rerun of that
failed job and its required aggregate passed without source or test
changes. The original failed job log and the single-rerun receipt remain
retained.

### October 9 merge verification

Current merge head: `5a8fe63512a7166aaef5cf50065a25008aa8b44b`. All
current-head checks pass, including `ci / verify` and `ci / e2e`;
exact-head Greptile review is 5/5 with no unresolved threads. Current
master conflicts are resolved. The user authorized the maintainer
override of the code-owner review gate after these checks. The seven
retained live core cases remain bound to source
`3241a992f2a7703e59e97ed0fd3e5d6405de4401` and its recorded cloud image.

## Risks

- The security correction changes the dependency closure and profile
identity. Old sessions must reopen on the new profile. Exact identities
and credential bindings fail closed.
- The runner remains experimental and requires explicit selection.
Legacy Pi Local is unchanged. Caller model IDs pass through; the E2E
model is a fixture.
- Accounting and the broad platform/provider matrix remain deferred.
This merge does not publish a release or deploy a service.

## Model Used

OpenAI GPT-6 through Codex assisted with reasoning, repository
inspection, editing and tool use. The exact serving ID and context
window are not exposed in this session. Final live qualification uses Pi
1.0.0 with `openrouter/anthropic/claude-sonnet-4.6` and native-confirmed
low thinking.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-09 08:56:41 -05:00
DottaandPaperclip 6f9d0a56ba fix: resolve installed Codex and preserve npm host dependencies (#15555)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Installed agents need their own runtime dependencies.
> - Native Codex startup and browser login could depend on a global CLI.
> - npm bundles do not inherit workspace dependency overrides or
patches.
> - Codex native packages must come from npm for the consumer host.
> - This PR fixes executable resolution and the required npm dependency
layout.
> - Usable installed CLI versions can run without a numeric version
gate.

## Linked Issues or Issue Description

Refs: #15422. This is a prerequisite for the later Codex-default change.

**What happened?**

Packaged native Codex startup and browser login could require a global
Codex CLI. The server's vendored runner lacked its own declared Codex
bridge. A bundled wrapper could retain producer-host binaries or select
the legacy adapter's separate platform version.

**Expected behavior**

Prefer the installed Codex executable. Accept usable older or newer
versions. Use the selected host's PATH when the dependency is absent.
Keep the patched JavaScript graph. Let npm install official native
packages for the consumer platform. Grant sandbox reads only to the
selected vendor resources.

**Steps to reproduce**

1. Prepare a clean npm installation without a global Codex command.
2. Start native Codex or its browser login.
3. Inspect the selected executable, installed host package, and sandbox
resource paths.

**Paperclip version or commit**

Frozen master base: `1881894973a2b25838d8abed9bd8aeebc3af4441`.

**Deployment mode**

Source and packaged self-hosted installation.

## What Changed

- Share installed Codex command resolution across native startup, direct
evals, and browser login. Accept usable version differences. Preserve
explicit commands, recorded sessions, remote host boundaries, and the
Linux ARM64 legacy login path. Return safe errors for missing
executables and failed login terminals.
- Declare the server's Codex bridge. Retain the patched JavaScript
dependency graph in npm packages. Strip Codex native payloads. Declare
official optional host packages on the published server manifest so
native and legacy versions remain separate. Preserve consumer esbuild
platform dependencies.
- Adjust resource lookup for npm's separate platform packages. Retain
package identity, path containment, resolver ownership, and narrow
sandbox reads. Keep exact version and digest checks for explicit ACPX
artifact qualification.
- Extend existing packaging, installed-consumer, login, selection,
recovery, and integrity tests. Document the retained behavior.

The diff is now 23 files, 1,709 additions, and 53 deletions. The prior
diff had 35 files and about 3,700 changed lines. Removed work is
preserved on `codex/runner-packaging-full-snapshot` at
`6f3060beaa842ffcee21af058370d3cab5f571e9`.

Removed from this PR: release workflows and assembly, provider-pack
changes, Docker materialization, Git installer changes, extra login
HOME/working-directory isolation, and unrelated CI fixture repairs. This
PR does not change agent defaults, stored runner choices, UI, schema, or
provider qualification.

## Verification

- Final candidate: `dde37d7ed0121c10b60b8801eda80f8dc17909ad`. [All 47
ordinary CI jobs
passed](https://github.com/paperclipai/paperclip/actions/runs/37842288835)
on attempt 1, including typecheck, tests, build, E2E, runner checks, and
the installed-consumer canary. [Fresh Greptile
review](https://github.com/paperclipai/paperclip/pull/15555#issuecomment-6059581546)
is 5/5 on this head; no unresolved threads remain. Human approval is
still outstanding.
- Reused focused controls passed with Node 24: 71 initial narrowed
checks, then 56 affected Codex/selection/eval checks after the relative
PATH correction. Runner TypeScript no-emit, syntax, and diff checks
passed. The existing fixture reproduces the original relative PATH
failure and verifies working-directory selection, empty entries,
ordering, and absolute launch. One CLI entrypoint test was initially
blocked by sandbox IPC and passed with its existing local socket
allowed. Login HOME, config-directory assignments, and working directory
match master.
- [The actual clean Linux npm
consumer](https://github.com/paperclipai/paperclip/actions/runs/37842288835/job/113535523044)
passed on the final candidate's CI integration. All 17 Paperclip
tarballs omit native Codex payloads. Official npm host packages retain
their own integrity and `inBundle=false`. Native Codex selects 0.160.0;
the legacy closure retains 0.156.1. Package admission, command leases,
narrow native sandbox resources, consumer hooks, preserved lock, and
offline lifecycle controls passed. Provider calls were zero. No new
workflow is added.
- Actual official Codex 0.156.1 passed on the final committed source on
macOS ARM64: installed dependency preference, absolute and relative PATH
fallback, narrow resource lookup, and app-server initialize/initialized.
Executed module hashes match the candidate. One scripts-disabled
install, three version probes, and one handshake completed in 26
seconds; owned files and process group were removed. No login,
account/model request, or provider task ran.
- CI checked out `dfda8e708e87306d22ed735d78bdfcbe770e99b7` on base
`65b558180533039a891ee0cd1ccab9988aa79adc`. Its 13 upstream paths do not
overlap this PR's 23 paths or alter packaging/Codex inputs. The consumer
report records producer `7b8e94c08657b5ddc265946459762340f125726c`,
after the existing canary staged a generated-lock-only commit (one file,
three insertions). These identities are kept separate; raw
generated-lock bytes were not retained.
- No local Docker or Rust build, paid provider turn, merge, or
deployment. Later PRs must prove live onboarding and production cloud
packaging before changing defaults.

## Risks

- npm must install optional host dependencies. Missing dependencies
still return errors. Paperclip tarballs do not pre-bundle Codex
executables.
- The repository requires CI-owned lockfile updates. PR CI resolves the
changed manifest and stages its own producer lockfile. A raw source
Docker build with `--frozen-lockfile` must wait for the existing master
lockfile bot to merge its refresh, or use a disposable resolved
checkout. No Docker build or deployment is qualified by this PR.
- Ordinary Codex startup accepts version differences. Actual protocol or
login failures remain errors. Explicit ACPX artifact checks retain their
release pins.
- The published server delegates platform installation outside its
bundled JavaScript graph. Focused negative controls reject unsafe
package metadata and paths. The hosted consumer test passed on this
candidate’s CI integration.
- Existing agents retain their stored runner choices. There is no data
migration or automatic upgrade. Release pipeline and platform
qualification work remain separate prerequisites for later defaults.

## Model Used

OpenAI Codex, GPT-6, with reasoning, tool use, code execution, and
parallel agents. The exact serving snapshot and context window are not
exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-09 08:19:51 -05:00
DottaandPaperclip 2b688a8fe3 feat(connections): add verified MCP providers and setup fixes (#15621)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Apps give agents governed access to external tools.
> - Several provider MCP servers lacked a supported catalog entry or
failed during setup.
> - Real browser tests identified specific registration, session, and
form defects.
> - This pull request adds seven catalog entries and fixes the shared
paths used by eleven verified providers.
> - Users can connect these providers through the existing Apps flow and
control each action.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

The Apps catalog and remote MCP setup, discovery, and action tester.

**Subsystem affected**

Shared app definitions, the server tool services, and the Apps UI.

**Current behavior**

Seven providers lack a catalog entry. Airtable can select an advertised
client-metadata flow that fails. Calendly rejects the registration name.
Tavily needs an initialized session before a call. Firecrawl exposes
invalid defaults for optional nested form objects. HTTPS setup can
display a callback that differs from the server callback.

**Proposed behavior**

Add Calendly, Exa, Firecrawl, GSC Wizard, Parallel Search, Tavily, and
Windsor.ai. Preserve Airtable, Linear, Make, and PostHog. Use reviewed
provider options through the shared MCP and OAuth paths. Display the
callback supplied by the server. Omit untouched optional object inputs.

**Reason and benefit**

Eleven providers passed bounded reads through the real Paperclip browser
action tester. This PR includes that verified set and its required
shared fixes. Unfinished providers remain outside this change. AgentMail
keeps its existing integration.

**Breaking changes**

No database migration. Existing OAuth credentials and action policies
keep their ownership and access rules. Airtable's reviewed method
re-registers a retained client-metadata binding through DCR. Tavily's
reviewed method initializes sessions before dispatch. The existing
customer, managed, and Vercel OAuth gateway paths now refresh on
upstream 401 and return `oauth_refreshed_retry_required` (409),
requiring an explicit caller retry. They do not replay the rejected call
automatically; the next invocation initializes a fresh credential-scoped
session when required.

Related catalog work: #15545. This branch preserves current master
catalog entries and does not add another Google Workspace integration.
Open and closed provider PRs and public MCP issues were searched. No
duplicate for this verified set was found. The change extends the
shipped Connected Apps roadmap item.

## What Changed

- Add seven catalog entries with official provider branding and reviewed
permissions.
- Update Airtable, Linear, and Make metadata and show Make in Apps.
- Add authoritative provider source overrides that use the existing
generation and review checks.
- Add the reviewed Airtable DCR option and compatible Calendly
registration names.
- Initialize Tavily sessions before discovery and calls, including
retained connections. Keep credential-scoped session caching and single
call dispatch.
- Use the server's callback URL in the OAuth setup UI.
- Omit untouched optional object defaults; keep supplied empty strings,
false, zero, and object values intact in the action tester, with strict
required-child validation.
- Require an explicit retry after OAuth refresh instead of replaying a
tools/call with missing session headers.
- Record all eleven successful reads, write policies, auth-method
limits, and qualification caveats. Omit credentials and private account
data.
- Find chat connector cards through catalog search in browser tests, so
pagination does not hide Slack or other later entries.

## Verification

- Real browser qualification: eleven bounded reads passed. Catalog
refresh and reload proof are recorded in
`doc/connections/verified-mcp-qualification-2026-10-08.md`.
- Provider writes and actual agent-adapter sessions were not run.
Alternative documented auth methods and expiry refresh remain
unqualified.
- `pnpm -r typecheck`: passed.
- `pnpm build`: passed.
- Focused shared/catalog tests: 41 passed. Source override tests: 3
passed. Catalog regeneration fixture suite: 9 passed. Latest form
regression suites: 31 passed. Gateway suite: 38 passed. Other focused UI
suites passed before these last fixes.
- UI token gates and both token sync checks: passed.
- The catalog regeneration fixture includes the authoritative provider
overrides; its full suite passes. The focused OAuth socket case passed
on isolated rerun.
- Complete local UI suite: 7,866 passed. CLI suite: 511 passed, 6
skipped. Complete shared suite: 892 passed.
- The unchanged AgentMail/ClickUp discovery fallback suite passes all 45
cases. An OpenAI login test passed on isolated rerun.
- Local broad database coverage is limited by macOS PostgreSQL
shared-memory exhaustion (`shmget: No space left on device`, not disk
space). The broad run was interrupted after diagnosis; no host settings
or other running services were changed.
- CI found that the Slack browser test assumed its catalog card was on
the first page. The test now uses catalog search; the Slack case passes
locally. Adjacent GitHub and iMessage cases could not start locally
because embedded PostgreSQL initialization failed before browser
assertions.
- Final commit
[`febe4520e`](https://github.com/paperclipai/paperclip/commit/febe4520e13d4b3a4121eafc3d22540c2b11379d):
all 54 checks passed, with none pending or failed. [CI
run](https://github.com/paperclipai/paperclip/actions/runs/37847182065)
passed typecheck, build, all general and serialized test suites, all
eight browser shards, runner checks, canary dry run, and the aggregate
verification gate.
- Greptile reviewed that exact final commit at 2026-10-08 21:33 UTC and
returned 5/5 with no outstanding findings or unresolved threads.
- PR UI preview against the existing local test server: Calendly
catalog/setup observed; Firecrawl read passed after expanding More
options with nested optional inputs untouched. This retest proves
frontend behavior, not the revised OAuth backend against a live
provider.

<details>
<summary>Browser evidence</summary>

![Calendly
catalog](https://raw.githubusercontent.com/paperclipai/paperclip/codex/verified-mcp-connectors/doc/connections/verified-mcp-qualification-2026-10-08/calendly-catalog.jpg)

![Calendly
setup](https://raw.githubusercontent.com/paperclipai/paperclip/codex/verified-mcp-connectors/doc/connections/verified-mcp-qualification-2026-10-08/calendly-setup.jpg)

![Firecrawl
read](https://raw.githubusercontent.com/paperclipai/paperclip/codex/verified-mcp-connectors/doc/connections/verified-mcp-qualification-2026-10-08/firecrawl-read.jpg)

</details>

## Risks

- Provider registration and consent behavior can change. Airtable's DCR
option is explicit and keeps the existing issuer, resource, redirect,
and PKCE checks.
- Tavily uses initialized sessions without automatic call retries. After
a successful OAuth refresh following 401, callers now receive a
retry-required result. A later explicit retry can still require
reauthorization if the provider rejects the refreshed token. Provider
writes are not live-qualified.
- Optional object cleanup affects the shared action tester. Regression
tests cover absent objects, required children, defaults, and populated
values.
- GSC Wizard's underlying Google data scopes and Google flow completion
were not independently verified. Its account reported paid/trial
metadata of unknown origin. No purchase was performed.
- No database migration, new AgentMail integration, or change to
existing action grants is included.

## Model Used

OpenAI GPT-6 through Codex assisted with implementation, research, tool
use, and code execution. The root backend model ID and context window
are not exposed in this session. The cheaper subagents used OpenAI
`gpt-6-luna`. No model context size is inferred.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass — targeted checks and
complete UI/CLI/shared suites passed; the broad local database run was
blocked by the macOS PostgreSQL startup limitation documented above. The
full database/workspace suite passed in CI.
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-09 08:10:46 -05:00
DottaandPaperclip f6d2a6b79e refactor(heartbeat): extract queue admission and wakeup dispatch (#15626)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The heartbeat service admits and dispatches agent runs.
> - Earlier extractions separated workspaces, preparation, state,
retries, and restart recovery.
> - Queue admission and wakeup dispatch still occupy more than 4,200
lines in the main service.
> - Their transaction, budget, ownership, and concurrency checks must
stay together.
> - This pull request moves those operations into one queue module and
adds boundary tests.
> - Reviewers can inspect queue policy separately from adapter
execution.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

The structure and testability of heartbeat queue admission and wakeup
dispatch.

**Subsystem affected**

server/ — orchestration services.

**Current behavior**

Wake coalescing, queued claims, dispatch, daily limits, and native wake
intents live inside `heartbeat.ts`.

**Proposed behavior**

Keep their behavior in `server/src/services/heartbeat/queue.ts`. Bind
lifecycle callbacks and shared ownership with `createHeartbeatQueue`.

**Reason and benefit**

This removes 4,243 lines from `heartbeat.ts`, leaving 14,201 lines. The
queue module has 4,576 lines and keeps related policy together in one
file.

**Breaking changes**

None. The service methods, public exports, database transactions, locks,
and execution gates keep their existing behavior.

**Additional context**

Continues #15622 and the earlier merged heartbeat extractions. A search
of open and closed PRs found no duplicate queue extraction. Related
#15600 adds a new experimental scheduler; this change preserves the
current scheduler policy. This does not add a roadmap feature.

## What Changed

- Move wake coalescing and batching, queue claims, daily caps,
concurrency and priority checks, timer admission, queue resumption, and
native status wake intents into `heartbeat/queue.ts`.
- Pass lifecycle callbacks and the existing process-wide execution and
wakeup sets into the factory. Preserve construction order with three
forwarding callbacks.
- Keep adapter execution, cancellation, and status writes in the
service.
- Add nine tests for inert construction, scheduling suppression, shared
wake tracking, policy wiring, claim release, newer ownership, and budget
gates.
- Document the queue module boundary in `doc/DEVELOPING.md`.

## Verification

- Before extraction: 133 tests passed across nine existing suites.
- After extraction: **150 tests passed across 12 files**, including all
nine new module tests and real PostgreSQL coverage.
  ```sh
pnpm exec vitest run server/src/services/heartbeat/queue.test.ts
server/src/__tests__/heartbeat-queued-run-claim-isolation.test.ts
server/src/__tests__/heartbeat-start-lock.test.ts
server/src/__tests__/heartbeat-stale-queue-invalidation.test.ts
server/src/__tests__/heartbeat-scheduling-suppression.test.ts
server/src/__tests__/heartbeat-comment-wake-batching.test.ts
server/src/__tests__/heartbeat-running-followup.test.ts
server/src/__tests__/heartbeat-dependency-scheduling.test.ts
server/src/__tests__/heartbeat-auto-checkout.test.ts
server/src/__tests__/heartbeat-archived-company-guard.test.ts
server/src/__tests__/heartbeat-task-drain-admission-release.test.ts
server/src/__tests__/heartbeat-worktree-suppression.test.ts
  ```
- Additional control and recovery coverage: **390 tests passed across
five files**. Combined focused coverage is **540 passing tests across 17
files**, with no skips.
  ```sh
pnpm exec vitest run
server/src/__tests__/heartbeat-process-recovery.test.ts
server/src/__tests__/heartbeat-native-cleanup-admission.test.ts
server/src/__tests__/heartbeat-native-status-context.test.ts
server/src/__tests__/heartbeat-task-drain.test.ts
server/src/services/chat-control-admission-retry.test.ts
  ```
- Full `pnpm -r typecheck` and `pnpm build` passed.
- Full local `pnpm test:run` was started and then stopped after the full
CI suite passed. It did not finish and is not counted as a full local
pass.
- All **54 current-head checks passed**; two optional Storybook checks
were skipped. This includes the full general and serialized test matrix,
typechecks, build checks, Runner verification, and canary dry run. [CI
run](https://github.com/paperclipai/paperclip/actions/runs/37847175452).
- Greptile completed on head `cc0001961a052a716d81c136929f89312e0b5dee`
with **5/5 and no actionable findings**. There are no review threads or
merge conflicts.
- Structural comparison confirms 31 moved function bodies and eight
declarations match the originals. Remaining service function bodies and
all 140 public exports are preserved.
- An initial local run skipped database tests because the host had
exhausted its PostgreSQL shared-memory slots. After verified stale
segments were cleared, all 150 focused tests ran and passed with no
skips.

## Risks

Moving closures can break shared ownership or callback construction
order. The factory receives the original sets and explicit callbacks. It
starts no work during construction. Transaction boundaries,
compare-and-set predicates, budget gates, cleanup holds, and lock
ordering stay intact. This PR has no schema or API changes.

## Model Used

OpenAI GPT-6 via Codex. The exact serving model ID and context window
were not exposed in this session. Used reasoning, repository inspection,
shell tools, and code execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-09 07:59:19 -05:00
Devin FoleyandPaperclip 1c4ce44e13 fix: retain bounded workspace transfer failure evidence (#15637)
## Thinking Path

> - Paperclip manages AI agents and the work they produce.
> - Sandbox work must return to the host when an agent run ends.
> - A failed transfer must keep its source available for recovery.
> - Some transfer errors lose their useful fields before restore
diagnostics reach Sentry.
> - This change records fixed transfer stages, failure kinds, and
numeric RPC codes.
> - The next failure can identify the operation that failed without
exposing workspace contents.

## Linked Issues or Issue Description

**What happened?**

A workspace transfer can fail after model execution succeeds. Several
command, archive validation, and download errors become plain errors.
The existing RPC envelope then reports `unknown`, with no transfer stage
or command exit code. A host RPC timeout also loses its code from
restore diagnostics.

**Expected behavior**

Keep bounded evidence through the provider, worker RPC, restore result,
and Sentry context. Keep the same exception, error message, retry rules,
and source retention policy. A transfer timeout must remain separate
from the model execution timeout flag.

**Steps to reproduce**

Return a nonzero exit code from the sandbox archive command, return a
per-file download error, exceed a tar listing limit, or time out the
sync-out RPC. Observe the missing fields in the saved restore
diagnostic. The added tests use local fixtures for these cases.

Related public work: #15479 preserves the source after restore failure.
#15481 carries bounded sync-out diagnostics through RPC. This change
supplies missing producer evidence and extends that same envelope. I
checked the roadmap and searched open PRs for duplicate transfer
diagnostic work.

## What Changed

- Add optional transfer stage and failure kind fields to the existing
diagnostic envelope. Capture command exits and listing deadlines at
their producers.
- Record only known codes from typed RPC errors at the sync-out
boundary. Revalidate every field before persistence and Sentry
projection.
- Keep annotations private to each outbound request and restore
settlement. Preserve frozen error identity and prevent evidence from
leaking across concurrent or later calls.
- Document the fields. Cover producer failures, quota limits, old
workers, concurrent error reuse, and redaction through the real Sentry
SDK.

## Verification

- Focused SDK, provider, host, persistence, and real Sentry tests: 449
passed, no skips. Independent review also ran the focused contracts and
all provider tests. The optional Sentry SDK is required for this check,
so the contract tests cannot skip.
- A compiled Daytona transfer and compiled SDK worker pass a local RPC
round trip. This uses the development TypeScript loader for workspace
dependency exports. The transfer and capture modules are compiled
JavaScript.
- `pnpm -r typecheck` and `pnpm build` passed. The excluded Daytona
package also passed `tsc --noEmit`.
- The initial local `pnpm test:run` stopped in the general-server group:
60 suites could not initialize embedded PostgreSQL because the isolated
install skipped its Darwin library-link setup. The existing package
postinstall repairs this; `initdb --version` now passes. All 60 affected
suites then passed: 1,342 tests, no skips. No local full aggregate pass
is claimed.
- Independent source review covers privacy, concurrency, packaging, and
recovery behavior.

## Risks

- These diagnostics help identify future failures. They do not establish
or repair the cause of a past transfer failure.
- New fields are optional. Older workers remain compatible. Unknown
values are omitted.
- The change adds a small request-local diagnostic scope. It keeps the
original thrown errors, archive confinement, cleanup order, timeout
values, retry count, and source retention policy.
- No schema or deployment changes are required. No live provider actions
are part of validation.

## Model Used

OpenAI Codex, based on GPT-6, with repository inspection, code
execution, and independent agent review. The exact deployed model ID and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-08 19:27:39 -07:00
Devin FoleyandPaperclip 9b624a110e fix: checkpoint database backups before idle sleep (#15629)
## Thinking Path

> - Paperclip coordinates agent work and preserves company state.
> - Hosted instances can stop compute when admission and all work
sources are quiet.
> - The periodic database backup timer currently blocks every automatic
sleep.
> - Disabling backups would remove recovery points.
> - This change offers an explicit final-backup checkpoint before sleep.
> - A verified archive and restart marker preserve recovery while
compute is stopped.

## Linked Issues or Issue Description

**What happened?**

An otherwise idle instance with automatic backups enabled always reports
background work. Operators cannot reclaim idle compute without disabling
those backups. The work inventory also queries the database before
checking known local blockers.

**Expected behavior**

An operator can enable checkpoint mode. The server may authorize sleep
only after all other work is quiet and a fresh backup is verified under
the same owned hold. Backups resume on restart with the existing
retention policy.

**Steps to reproduce**

Acquire a bounded idle drain on an instance with no application work and
automatic backups enabled. Read its owned idle safety report. It reports
present even when all other work is clear. With this change and
`PAPERCLIP_DB_BACKUP_IDLE_CHECKPOINT_ENABLED=1`, a successful final
backup can permit sleep; backup failures still refuse it.

**Paperclip version or commit**

Reproduced from master at 4265cb3a2b.

**Deployment mode**

Hosted single-process instances with persistent backup storage and
serialized sleep/wake operations.

Searched open backup and sleep PRs and issues. Related: #15522 and
#15599 establish owned idle holds and plugin drains. #15358 proposes
general backup health/retention validation; this change verifies only
opted-in sleep checkpoints and their restart catch-up.

## What Changed

- Add an off-by-default checkpoint flag. Keep automatic backups enabled.
- Reject known local blockers before database work.
- Finish a new backup during each owned safety check, including final
revalidation.
- Verify the complete gzip archive and sync its file and directory
before pruning older recovery points.
- Use unique checkpoint filenames so repeated checks cannot replace an
earlier verified archive.
- Keep backup work counted through dump, validation and persistence.
Changed owners, expired holds, concurrent work and errors refuse sleep.
- Persist a restart marker before the dump. Attempt a fresh backup on
restart before idle eligibility, then resume the normal timer. Only a
verified fresh backup clears the marker.
- Restrict managed checkpoint reads to verified Cloud control actors.
Tenant instance-admin elevation does not grant backup authority.
- Keep archive verification outside engine fallback so failures retain
their original cause.
- Document that recurring duplicate archives pause while asleep. The
final checkpoint and historical archives remain on persistent storage;
retention does not prune while asleep.

## Verification

- Focused idle, spool, backup and route checks: 137 tests passed,
including restoration into a separate PostgreSQL database.
- Backup library and checkpoint tests: 15 passed, including rejection
before historical backups can be pruned.
- `pnpm -r typecheck` and `pnpm build` passed. Follow-up server
typecheck and database build cover the verification callback.
- Review follow-up: 51 backup/restore/startup tests, 122 route/safety
tests, 32 signed-control tests, and server typecheck passed. Startup
cases exercise both flag states and retry after failure. The final
test-cleanup adjustment passes all 35 startup tests.
- The full local run hit four ancestor-directory skill-fixture failures
and one tool-route failure. All five passed from an isolated checkout at
949210 (3 targeted chat/tool checks and the complete 39-test email
suite). The redundant local run was stopped after final-head hosted CI
passed. No full local pass is claimed.
- Added-line secret/private-reference review and `git diff --check`
passed.
- [Final-head hosted
CI](https://github.com/paperclipai/paperclip/actions/runs/37853380365)
passed, including typecheck, build, all general and serialized test
shards, all eight E2E shards and the canary dry run. Apex is 5/5 on
5b1de388ef, with no new findings and all review threads resolved. The
branch is conflict-free. Live staging activation awaits merge and a
canonical image; no production configuration or deployment changed.

## Risks

- The flag defaults off. This is for a single process whose host
controls every writer and preserves its backup volume across sleep.
Independent database writers require a different backup owner.
- Safety reads perform a backup only after a valid owned drain and all
other checks pass. Slow backups can outlive the caller timeout; they
remain counted until their work settles and cannot authorize sleep after
hold expiry.
- Initial and final reads each take a fresh backup. Hosts must allow
sufficient request time while retaining shutdown headroom.
- Retention settings are unchanged. Sleeping instances stop producing
duplicate hourly archives and retain their last checkpoint. Filesystem
data, encryption keys and off-provider disaster recovery remain separate
responsibilities.
- No schema migration. Disable the checkpoint flag to restore the
blanket backup blocker; automatic backups and an existing restart marker
remain active.

## Model Used

OpenAI Codex, GPT-6. Used reasoning, repository inspection, code editing
and command execution. The runtime did not expose an exact model
revision or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have described the issue in-PR following the bug report template
- [x] I have not referenced internal/instance-local Paperclip issues or
links
- [x] My branch name describes the change and contains no internal
identifier
- [x] Focused and isolated local checks pass; final-head hosted CI is
green (local full-run limitation documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation
- [x] I have considered and documented risks
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all review comments before requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-08 15:58:55 -07:00
Devin FoleyandPaperclip 4265cb3a2b fix(ci): cache native server integration builds in release verification (#15619)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Release verification must run the same server integration coverage
as PR verification.
> - Three server suites use real Rust Runner binaries.
> - PR checks already run them with the shared Rust dependency cache,
but release checks cold-build them inside ordinary server test shards.
> - A cold Dot Runner build can consume its entire setup deadline before
any test runs.
> - This pull request gives release verification a required cached
native integration lane and preserves every test.

## Linked Issues or Issue Description

**What happened?**

On master commit `1881894973a2b25838d8abed9bd8aeebc3af4441`, [Cloud
readiness run
37837810707](https://github.com/paperclipai/paperclip/actions/runs/37837810707)
failed in server shard 7. Dot Runner's `cargo build --release` exceeded
its 300-second child-process deadline. The other 1,640 tests in that
shard passed. The failed setup then tried to remove an undefined
temporary path and emitted a second error. Cargo output was captured as
an opaque buffer, which obscured build progress.

**What did you expect to happen?**

Build fixture binaries before tests in a lane with the existing Rust
dependency cache. Run all native integration tests and make their result
required for release readiness. A setup failure should keep its original
diagnostic.

**Steps to reproduce**

Run release verification on a clean runner. The existing
`general-server-without-chat` group retains the three Cargo-backed
suites outside the PR workflow. The Dot suite can time out during a cold
release build. Run the new workflow and partition tests against the
previous source to reproduce five routing and coverage failures
deterministically.

**Version / commit**

Observed at `1881894973a2b25838d8abed9bd8aeebc3af4441`; this change is
based on `ed6abbf158b`.

**Deployment mode**

GitHub Actions Release and Cloud readiness verification. No application
runtime or deployment action changes.

Related work: #15581 also edits Dot integration tests for onboarding
behavior. It does not repair release test routing. Searches found no
open PR for this failure.

## What Changed

- Add an explicit server test group that excludes the dedicated chat and
native suites. Preserve the existing PR and local test groups.
- Run all three native server suites in one required matrix lane with
the existing trusted Rust dependency cache.
- Build debug and release fixture binaries in a visible step with a
10-minute limit before tests start. Keep Cargo freshness checks and the
existing test deadlines.
- Stream Cargo diagnostics and clean up safely when Dot setup stops
before creating its temporary directory.
- Test complete, non-overlapping partitions for Release, Cloud
readiness, other callers, and existing PR/local groups. Check cache
restrictions, build order, and the source verification dependency.

## Verification

- Passed 44 workflow and partition tests: `node --test
scripts/__tests__/run-vitest-stable-shard.test.mjs
scripts/__tests__/release-verify-workflow.test.mjs`.
- The same tests fail in five relevant cases against the previous
workflow and selector. They pass after this correction.
- An independent review repeated all 44 tests successfully.
- Passed `actionlint .github/workflows/release-verify.yml`, `git diff
--check`, and a secret scan.
- Passed all 381 workflow policy tests: `node --test
'.github/scripts/tests/*.test.mjs'`.
- Passed full `pnpm build` in a clean worktree.
- Built debug and release fixture binaries, then passed all 32 tests in
the three real native integration suites: `pnpm test:run:general --
--group general-server-native-runner` (49.5 seconds, no skips).
- Passed full `pnpm -r typecheck`.
- Injected a synthetic Cargo setup failure. The suite reports that
failure without the secondary undefined-path cleanup error.
- Exact-head CI completed: 53 successful checks, including all server
shards, both native Runner lanes, build, typecheck, browser tests, and
the canary dry run. Two optional Storybook checks were skipped.
- Started the duplicate local `pnpm test:run` aggregate and stopped it
after exact-head CI passed. No completed local aggregate result is
claimed. All test processes owned by that run exited.
- Greptile scored the final head
`ffe5ab5a28af2dfea9db1c1952c652f070bf52f8` at 5/5 with no review
threads.
- `Superagent Supply Chain Scan` is neutral, not passed: it only
supports exact dependency pin replacements and cannot verify structural
edits to `.github/workflows/release-verify.yml`. It reported no
annotations. Independent source review, Greptile, actionlint, and the
workflow policy tests cover this structural change.
- GitHub still requires a code-owner review for the workflow files.
There are no merge conflicts.

## Risks

- Adds one CI matrix job, which increases concurrent runner demand. The
existing cache writer and trust restrictions stay in place.
- Missing dependencies still compile from the lockfile. A build that
exceeds its step limit fails visibly; failed tests still block source
verification.
- Workflow execution uses the new lane after merge. PR tests validate
the workflow contract and run the existing native test lane.
- The supply chain scanner cannot analyze this workflow structure. Its
neutral result is an explicit coverage limit; it is not counted as a
successful scan.
- No runtime, schema, deployment, publication, credential, or
test-timeout changes.

## Model Used

OpenAI Codex, based on GPT-6, with repository inspection, code
execution, and independent agent review. The exact deployed model ID and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-08 14:33:10 -07:00
DottaandPaperclip 2b4a1d7073 feat(dot): add guided invitations and persistent agent connections (#15581)
Add guided external-agent invitations with live Dot connection checks and ongoing access until revoked. Keep expired credentials invalid and enforce permissions for private avatar uploads. Preserve Runner recovery and remote launch boundaries.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 16:26:41 -05:00
DottaandPaperclip 04de3132c2 refactor(heartbeat): extract restart recovery and lease cleanup (#15622)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The heartbeat service coordinates agent runs and their recovery.
> - The earlier refactors separated workspaces, run preparation, run
state, and retries.
> - Restart recovery and lease cleanup still occupy more than 2,000
lines in the main service.
> - These operations share process ownership and database claims that
must stay intact.
> - This pull request moves that group into one recovery module and adds
boundary tests.
> - Reviewers can now inspect recovery separately from queue admission
and run execution.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

The structure and testability of heartbeat restart recovery and lease
cleanup.

**Subsystem affected**

server/ — orchestration services.

**Current behavior**

Hot-restart adoption, native restart recovery, shutdown draining, orphan
reaping, and lease cleanup live inside `heartbeat.ts`.

**Proposed behavior**

Keep the same recovery behavior in
`server/src/services/heartbeat/recovery.ts`. Bind the database and
lifecycle callbacks with `createHeartbeatRecovery`.

**Reason and benefit**

This removes 2,181 lines from the main service. The extracted module is
about 2,400 lines. It keeps related recovery operations together in one
file.

**Breaking changes**

None. The public service methods, database claims, cleanup fences, and
ownership checks stay the same.

**Additional context**

Continues the merged heartbeat extractions in #15568, #15573, #15578,
#15591, #15601, and #15617. A search of open and closed PRs found no
duplicate recovery extraction. This does not add a roadmap feature.

## What Changed

- Move hot-restart snapshots and adoption, native restart recovery,
shutdown draining, orphan reaping, and lease sweeps into
`heartbeat/recovery.ts`.
- Pass lifecycle callbacks and the existing shared execution sets into
the factory. Preserve module-wide cleanup single-flight state.
- Set the service's existing shutdown flag through a callback at the
same point in shutdown preparation.
- Add 14 focused tests for construction, shutdown ordering, empty
selective drains, ownership races, shared state, retry timers, and
background execution tracking.
- Wait for the existing execution-drain barrier in the Stop recovery
test instead of polling for one second.
- Document the module boundary in `doc/DEVELOPING.md`.

## Verification

- New module and additional cleanup coverage: **57 tests passed** across
five files.
  ```sh
pnpm exec vitest run server/src/services/heartbeat/recovery.test.ts
server/src/__tests__/heartbeat-native-cleanup-admission.test.ts
server/src/__tests__/heartbeat-run-terminalize-before-release.test.ts
server/src/__tests__/heartbeat-run-lease-release-terminalization.test.ts
server/src/services/hot-restart.test.ts
  ```
- Final recovery coverage on the rebased branch: **514 tests passed
across 11 files**, including all 14 new module tests.
  ```sh
pnpm exec vitest run server/src/services/heartbeat/recovery.test.ts
server/src/__tests__/heartbeat-process-recovery.test.ts
server/src/__tests__/heartbeat-pending-cleanup-sweep.test.ts
server/src/__tests__/heartbeat-orphaned-active-lease-sweep.test.ts
server/src/__tests__/native-session-resumption.test.ts
server/src/__tests__/heartbeat-task-drain-admission-release.test.ts
server/src/shutdown.test.ts
server/src/__tests__/heartbeat-native-cleanup-admission.test.ts
server/src/__tests__/heartbeat-run-terminalize-before-release.test.ts
server/src/__tests__/heartbeat-run-lease-release-terminalization.test.ts
server/src/services/hot-restart.test.ts
  ```
- Full `pnpm -r typecheck` and `pnpm build` passed, including after
rebasing onto master.
- Full local `pnpm test:run` was started again after the rebase. It was
stopped after the full CI suite passed. It did not finish and is not
counted as a full local pass. An earlier local AI-connection run hit an
HTTP-test timeout; all corresponding CI coverage passed.
- Structural comparison confirms 28 moved functions and 15 moved
declarations. The only function-body adaptation is the shutdown flag
callback. Remaining service functions and all 140 public exports match
master.
- All **54 current-head checks passed**; two unrelated Storybook checks
were skipped. This includes the full general and serialized test matrix,
typechecks, build checks, Runner verification, and canary dry run. [CI
run](https://github.com/paperclipai/paperclip/actions/runs/37843023706).
- Greptile completed on head `01fe39818d483e05f940b9fe842162a08e9b9e8e`
with **5/5 and no actionable findings**. There are no review threads or
merge conflicts.

## Risks

The main risk is breaking shared execution ownership or changing
recovery order while moving closures. The factory receives the original
execution sets and lifecycle callbacks. It does no work during
construction. Database transactions, compare-and-set predicates, cleanup
deadlines, and native ownership gates stay intact. This PR has no schema
or API changes.

## Model Used

OpenAI GPT-6 via Codex. The exact serving model ID and context window
were not exposed in this session. Used reasoning, repository inspection,
shell tools, and code execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-08 16:15:01 -05:00
DottaandPaperclip 428b617fc9 test(evals): qualify native question resume paths (#15616)
## Thinking Path

> - Paperclip lets agents pause tasks for human answers and continue the
work.
> - Native questions can yield a run or pause inside the provider's
running turn.
> - Those paths have different run identities and different forms.
> - The documentation-placement test requires a semantic three-turn
journey.
> - A valid provider question therefore needs separate behavioral
qualification.
> - This pull request adds an opt-in two-answer journey with exact
source and resume checks.
> - Qualification exposed a Vite startup crash in idle-handler
traversal, so the branch also distinguishes Connect mount paths from
Express Route objects.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

Product E2E evaluation of native task questions and answer delivery.

**Current behavior**

The semantic documentation case rejects the provider adapter's optional
Other field before answering. Its three-run expectation also cannot
prove a provider question that resumes within the same run.

**Proposed behavior**

Keep that semantic test intact. Add a separate opt-in case that
classifies each recorded question by authoritative identities. Accept
the optional Other field only for a verified provider question. Require
both user answers, the correct continuation, and one saved final
document.

Related: #15564 added strict lifecycle checks. #15522 introduced
idle-request tracking; this PR corrects its Connect/Vite compatibility
after reproducing the startup crash. #14591 changes ACP question drafts
and provider support; this PR has no overlapping files with that change.

## What Changed

- Add the explicit-only `question-resume` suite for native Codex and
ACPX Claude. Reuse the existing neutral choice-then-text task request.
- Bind each question to its company, task, agent and run. Require an
applied semantic tool receipt or the exact provider request identity.
- Require a provider answer to continue the same run. Require a semantic
answer to start one new run with the matching interaction and source-run
wake fields.
- Preserve question forms and both real board answers through the final
checkpoint. Reject missing inputs, duplicate answers, extra runs and
substituted identities.
- Keep the existing lifecycle, output, task ownership and budget checks.
Limit the new case to one attempt and one to three runs.
- Add mutation calibrations and document the separate qualification
claims. Keep production guidance, scheduling, provider code and the
historical semantic case unchanged.
- Fix startup with Vite middleware: treat Connect string mount paths
separately from Express Route objects when traversing idle-request
handlers. A real-Vite regression verifies startup and retained async
work through an idle hold.

## Verification

- 81 focused continuation, question and calibration tests passed before
review. After the reference-preservation correction, all 50
question/resume tests pass, including two new negative cases that failed
against the previous grader.
- Final eval support suite passes: 1,921 Vitest tests and 128 Node
tests, with one intentional skip.
- Eval and workspace typechecks pass. Workspace build passes.
- Catalog discovery finds only the two declared opt-in cells. Existing
continuation remains 23 cells; default paid scope is unchanged.
- Original [campaign
37831726391](https://github.com/paperclipai/paperclip/actions/runs/37831726391),
source `1ac4b2873466896ff2892f49be82904927b6c57d`, failed during server
startup before any test or provider run. Its original
`transient_infrastructure` result and `cleanup: not_started` are
preserved. Local reproduction identifies the Connect string route
traversal crash; the question grader was never reached.
- Setup-corrected [campaign
37833872898](https://github.com/paperclipai/paperclip/actions/runs/37833872898)
measures source `9781a772389d41421df55759004676e1b905e512` with trusted
workflow `8e59efc50b162126f33816841e06509634eba65c`. The workflow bytes
are unchanged from the first dispatch. Same single selected native
Claude cell, one attempt, at most three runs, ten-minute cell deadline,
and 1,000-cent company/agent hard stops. Original result **PASS**, 24/24
checks and cleanup pass. Exactly three successful native Claude runs:
two applied Paperclip `request_human_input` questions and one finishing
run. Both real board answers persist, exact interaction/source-run
response wakes match, one revision-1 task document contains Afternoon
and the supplied reference, and final task state is done with no active
lock, retry, recovery, monitor, pending interaction or child task. No
model retry.
- Current review head `3c08cb7d26e7f9ac16469eeb07d1f31d828ec02e` differs
from the measured source only in the question/resume grader and its
tests (11 insertions, 3 deletions). The saved document must contain the
complete submitted reference, including its prefix and punctuation.
Separate provider-free replay of the retained successful campaign passes
all 24 continuation/lifecycle checks under this stricter grader. The
original result and its measured source remain unchanged; no new
provider run was made.
- Observed paths were **semantic → semantic**. The provider built-in
question/optional Other path was not exercised by this new live attempt.
Its form acceptance has retained-original calibration; same-run
answer/resume has deterministic positive/negative calibration. This
neutral-task pass alone does not qualify the provider path; the
subsequent explicit bridge trial below is separate evidence. Neither
trial regrades the old Claude failure or establishes a causal
behavior/performance comparison.
- Subsequent explicit provider-path [campaign
37841107684](https://github.com/paperclipai/paperclip/actions/runs/37841107684)
measures the final PR source `3c08cb7d26e7f9ac16469eeb07d1f31d828ec02e`,
using trusted workflow `2a5f65c9ae501b49b2a38ce7edd04209b806c9d0`. Exact
cell: `continuation.runner-acpx-claude.local.provider-question-bridge`,
Claude `claude-sonnet-5`, suite fingerprint
`fc77203bf85fd99b06ebc53cd52982186c0ae0dc369fcc38a9ec3c737e9a31e5`.
**Original PASS, 15/15 checks, cleanup passed, one actual provider run
and zero eval rerolls.** The built-in question produced one choice plus
its optional Other companion. The browser selected the requested
reference, left Other blank, and submitted. Exact runtime
request/response IDs and the selected option match the saved
interaction; the same paused native run created one revision-1 document
and finished Done with no pending interaction, lock, retry, recovery,
monitor or child task. Independent evidence and runtime-receipt audits
pass; marked initial/final screenshots were visually checked. [Original
public
report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37841107684-1/index.html).
- This separate provider-directed trial qualifies one choice-answer
bridge and traversal of an empty optional Other field. It does not
establish natural tool selection, a typed Other answer, two consecutive
built-in questions, free-text-only input, restart recovery or general
reliability. No production, prompt or grader change was made for this
dispatch. The old failure and the separate neutral semantic-path result
retain their original grades.
- Preserve within-run friction: one `write_task_document` input-schema
rejection and two HTTP 409 “Document does not exist yet” responses
preceded a successful `write_document` call. These are three
unsuccessful tool calls inside the same run, not three extra provider
runs; no flawless document-tool behavior is claimed. Billing reports
token usage but remains unpriced/incomplete, so actual charges are
unknown and local/hosted runtime is unmetered. Both company and agent
hard stops were 1,000 cents, with one attempt, one selected cell,
concurrency one and a ten-minute cell deadline.
- Billing records all three model runs with token usage, but no priced
dollar receipt (`unpriced`, `complete: false`). Actual charges are
unknown; local/hosted runtime is unmetered. Numeric zero reported cost
does not mean free.
- Startup repair: real-Vite reproduction passes; 35 targeted
idle-tracking tests pass, 34 database-dependent checks skip because
embedded PostgreSQL is unavailable locally. Workspace typecheck and
build pass again after the repair.
- Full local repository database tests were not repeated because
embedded PostgreSQL was unavailable in the preceding workspace
verification. Full Linux CI now passes on the final review head.

- Final-head [CI run
37837804284](https://github.com/paperclipai/paperclip/actions/runs/37837804284)
passes. Complete current-head audit: 53 successful check-runs, two
intentional Storybook skips, and separate Snyk success. The ready
transition’s contributor and security checks also pass (security scan:
no findings). [Fresh
review](https://github.com/paperclipai/paperclip/pull/15616#issuecomment-6068092601)
is 5/5 on `3c08cb7d26e7f9ac16469eeb07d1f31d828ec02e`; zero unresolved
threads and no merge conflicts.

## Risks

This is a new behavioral definition. The neutral two-question trial
records semantic-tool use; the separate provider-directed trial records
one built-in question. They do not establish consistent tool selection,
documentation placement, default hiring, remote environments, crash
recovery or general reliability. The separate explicit bridge trial
covers one provider choice, an empty optional Other companion and
same-run completion. Typed Other, two consecutive built-in questions and
free-text-only provider input remain unqualified. The observed
document-tool schema rejection and two 409 responses are retained as a
separate follow-up, without attributing their cause. The historical case
and its original grades remain intact. The reference question must still
be text-only; the oracle does not turn an arbitrary extra question into
an optional companion. Missing or unfamiliar identity evidence fails
closed. The only production change distinguishes Connect string mount
paths from Express route objects during idle-handler traversal. Unknown
route objects still fail closed. No API, schema, migration or workflow
change is included.

## Model Used

OpenAI Codex, GPT-6-based, with tool use and code execution. The exact
serving model ID and context-window size are not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-08 16:00:01 -05:00
Devin FoleyandPaperclip 65b5581805 Keep proven AI account selection blockers out of Sentry (#15618)
Preserve trusted evidence from explicit AI account selection rejections so known setup blockers stay out of Sentry. Keep task failure, blocked state, recovery actions, HTTP behavior, and unexpected error reporting unchanged.

Validated with full typecheck/build, focused PostgreSQL lifecycle and reporting tests, independent reporting tests, and all PR CI checks.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 13:48:48 -07:00
DottaandPaperclip 2a5f65c9ae refactor(server): extract heartbeat retry scheduling (#15617)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Heartbeat orchestration starts runs and recovers from supported
failures.
> - The heartbeat service still contains more than 22,000 lines after
earlier extractions.
> - Retry scheduling has a clear boundary, but it shares lifecycle
callbacks with the service.
> - This pull request moves retry scheduling into one module with
explicit dependencies.
> - It preserves the scheduling rules, database writes, and lock order.
> - The benefit is a smaller service and focused tests for the retry
boundary.

## Linked Issues or Issue Description

**Current behavior**

`server/src/services/heartbeat.ts` owns retry schedules, contention
deferrals, due promotion, and retry-now requests. These operations sit
among execution and recovery code.

**Proposed behavior**

Move these operations into `server/src/services/heartbeat/retries.ts`.
Keep the public heartbeat API and existing helper exports. The service
supplies its lifecycle callbacks and worktree cutoff.

**Reason and benefit**

Extract one complete area in a separate PR. This removes 1,641 lines
from the main service without changing retry policy. Future retry
changes can use a smaller module and direct tests.

**Breaking changes**

None. Existing imports, error-class identity, retry metadata, and API
outcomes remain available.

**Additional context**

Refs #15601, the preceding run-state extraction. Related open retry
changes: #14930, #13975, #15157, #12587, and #13276. This PR moves the
existing implementation. Those policy changes remain separate.

## What Changed

- Add `createHeartbeatRetries(db, dependencies)` for bounded retries,
workspace and connection deferrals, shared-workspace holder checks, due
promotion, and retry-now requests.
- Keep status transitions, event publication, issue-lock release,
plan-resume reporting, and worktree settings in the heartbeat service.
- Preserve all 140 existing public exports. Re-export the same helpers,
constants, and workspace-busy error class.
- Add 14 tests for construction, callback order, cancellation races,
concurrent scheduling, committed promotion, and retry-now scope.
- Document the module boundary in `doc/DEVELOPING.md`.

## Verification

- Existing focused suite before the move: 84 tests passed.
- Expanded focused suite after the move: 243 tests passed across 11
files. Command:
`pnpm exec vitest run server/src/services/heartbeat/retries.test.ts
server/src/__tests__/heartbeat-retry-scheduling.test.ts
server/src/__tests__/heartbeat-workspace-busy.test.ts
server/src/__tests__/heartbeat-model-fallback-warning.test.ts
server/src/__tests__/heartbeat-running-followup.test.ts
server/src/__tests__/heartbeat-issue-rewake-throttle.test.ts
server/src/__tests__/issue-scheduled-retry-routes.test.ts
server/src/modules/run-dispatch`
- After rebasing onto the latest master, 61 retry and plugin-idle tests
passed across four files.
- `pnpm -r typecheck` passed again on the rebased branch.
- `pnpm build` passed again on the rebased branch.
- `pnpm test:run` was started locally. The long serial run was stopped
after full CI passed; it is not claimed as a complete local pass.
- Full CI passed on commit `69e76f3eec75b0703570f1fce37ebfe9c63fa3ac`:
54 successful checks, two intentional Storybook skips, zero failures.
This includes general tests, serialized server tests, browser tests,
runner checks, typecheck, build, and canary verification.
- The setup-timeout review fix passed all 14 module tests and the server
typecheck.
- Greptile completed a fresh review of commit
`69e76f3eec75b0703570f1fce37ebfe9c63fa3ac`: 5/5, no open findings.
- A structural comparison confirmed 22 unchanged moved functions and one
unchanged three-line string reader, 16 unchanged declarations, one
unchanged API body, and all 140 preserved exports. The remaining service
body is unchanged after accounting for the new binding and removed
declarations.

## Risks

- A missing lifecycle callback or an eager call during construction
could change retry behavior. The dependency contract and direct tests
cover these boundaries.
- Concurrent scheduling must keep the issue-first lock order and reuse
one successor. The original transaction bodies are unchanged. PostgreSQL
tests cover concurrent calls from separate factory instances.
- Promotion must publish after commit and preserve worktree cutoffs.
Direct tests cover both paths.
- No schema, route, adapter, or retry-policy changes are included.

## Model Used

OpenAI GPT-6 via Codex. The exact serving model ID and context window
were not exposed in this session. Used reasoning, repository inspection,
shell tools, and code execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-08 15:38:48 -05:00
Devin FoleyandPaperclip 1881894973 fix: prevent inbox archive deadlocks during completion (#15615)
Take a company-scoped parent lock before writing archive state and reuse the caller transaction. Preserve newer archive timestamps and attribution when a request waits, so completed tasks remain archived.

Verified with real PostgreSQL race, rollback, company isolation, concurrency, timestamp and visibility tests, independent review, full CI and Greptile 5/5.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 13:13:16 -07:00
Devin FoleyandPaperclip 3d1d5294d9 Drain idle plugin workers before automatic sleep (#15599)
## Thinking Path

> - Paperclip runs autonomous work through agents, schedules, and
plugins.
> - Automatic idle sleep must preserve accepted work and unfinished
cleanup.
> - Every enabled plugin currently blocks sleep, including unused
providers.
> - An enabled flag or package version cannot prove that a worker is
idle.
> - This change adds a live worker drain under the existing owned hold.
> - Idle workers can allow sleep while unknown work stays protected.

## Linked Issues or Issue Description

Refs #15522. Related: #15391 adds plugin readiness for agent admission;
this change concerns instance sleep and does not replace that contract.

**Subsystem affected**

Plugin worker lifecycle and automatic idle sleep.

**Problem or motivation**

A workspace with no pending work cannot sleep when any plugin is
enabled. Removing that check alone would lose accepted RPCs, background
tasks, or cleanup after a caller timeout.

**Proposed solution**

Require a live `onIdleDrain` handshake from each worker. Close admission
in both processes for the exact owner and expiry. Count accepted work
until completion. Continue checking durable work separately.

## What Changed

- Add bounded worker holds, exact-owner release, automatic expiry, and
an abort signal for plugin-owned background work.
- Count host and worker requests through their real completion receipts.
Keep timed-out work counted. Check active notifications and terminal
routes.
- Accept enabled plugins only when their current workers provide
matching runtime receipts. Missing workers, old SDKs, crashes, invalid
replies, and unknown cleanup still prevent sleep.
- Let unused Daytona workers opt in. Once a worker contacts the
provider, it remains a blocker for that process lifetime. This
restriction avoids treating its existing timeout and terminal-close
behavior as a cleanup receipt.
- Document the plugin author contract. No schema or user-facing API is
added.

## Verification

- All hosted CI checks passed on 6ffb682894, including all server,
runner, workspace, build, typecheck, and end-to-end checks. One
unrelated OpenCode transport test hit a five-second timeout and passed
on rerun.
- `pnpm -r typecheck` passed.
- `pnpm build` passed.
- Focused Vitest: 293 tests passed across worker admission, real
child-process RPC, durable idle checks, and Daytona.
- Adjacent worker manager and duplex tests: 119 passed.
- Real child-process tests cover late worker completion, a host write
after caller timeout, stale owners, closed admission, and release.
- SDK tests cover hook failures, busy workers, expiry without another
request, and late acknowledgements after expiry.
- Initial long local run: 17,202 passed, 10 failed. Five failures
involved parent-directory skill fixtures; four of those were already
reproduced and passed in an isolated worktree. One HTTP test had a
socket hang-up; all 66 tests in that file passed on rerun. The other
four used new tests with runtime modules loaded before the follow-up
edits; all pass in the fresh focused run above. A clean build and
full-suite rerun at the final commit are running in an isolated
worktree. No full-suite pass is claimed.
- Follow-up: 44 focused tests passed, including repeated owned holds
without an intervening normal request. Worker status changes now
invalidate an in-progress scan.
- Apex follow-up: 13 idle tests, 119 adjacent worker tests, 71
instance-settings route tests, and server typecheck passed. A real
queued notification completes its write while preparation is in flight;
synchronous and async session callback failures are logged.
- Diff whitespace and added-line secret/private-reference scans passed.

## Risks

- Plugins are trusted code. A hook must account for work outside SDK
RPCs and keep it quiescent until its signal aborts. There is no
package-name or manifest-only exemption.
- This first change targets unused workers. Used Daytona workers remain
awake until remote cleanup receipts are complete. Restarting does not
bypass durable lease and recovery checks.
- Unknown completion remains a blocker, including after a worker crash.
This can retain cost but cannot authorize sleep from a timeout alone.
- Enabled jobs, schedules, credentials, integrations, and other durable
work remain blockers. Durable wake ownership is separate follow-up work.
- Rollback restores the blanket plugin blocker. No data migration is
required.

## Model Used

OpenAI Codex, GPT-6. Used repository inspection, reasoning, code
editing, and command execution. The runtime did not report an exact
model revision or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either linked existing public issues or described the issue
in this PR
- [x] I have not referenced internal or instance-local issues or links
- [x] My branch name describes the change and contains no internal
identifiers
- [x] Focused local tests pass; full CI is green (long local rerun
status is recorded above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation
- [x] I have considered and documented risks
- [x] All Paperclip CI gates are green
- [x] Greptile Apex is 5/5 on 6ffb682894 with no open findings
- [x] I will address all review comments before requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-08 12:57:32 -07:00
DottaandPaperclip 4fe45bf8fe refactor(heartbeat): extract run retrieval and session state (#15601)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Heartbeat runs retain task sessions, usage totals, and bounded run
data.
> - After the preparation extraction, the main service still has over
24,000 lines.
> - Run reads and session operations share one database and have no
dispatch dependencies.
> - This pull request moves that group into one module and preserves its
callers.
> - The benefit is a smaller run engine and one place to maintain saved
run and session state.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

It improves the structure and test coverage of heartbeat run retrieval
and session state.

**Current behavior**

`server/src/services/heartbeat.ts` has 24,033 lines on the merged base.
It mixes run projections, session queries, resume rules, compaction, and
usage helpers with execution and recovery.

**Proposed behavior**

Move that group into `server/src/services/heartbeat/run-state.ts`. The
main file loses 1,770 lines and ends at 22,263 lines. Keep current
behavior, public exports, and service methods. This continues #15568,
#15573, #15578, and #15591.

**Reason and benefit**

One domain extraction makes useful progress toward a few manageable
files. Run reads and saved session state can be reviewed together
without searching the run engine. A search found no duplicate open
extraction. Related session changes include #15142, #14661, and #14670;
this PR does not apply their proposed behavior changes.

**Breaking changes**

None. Existing public helpers retain their identity. Service method
names and results stay the same.

## What Changed

- Move bounded run projections, encoding checks, and run/event reads
into `heartbeat/run-state.ts`.
- Move task session persistence, explicit resumes, reset rules,
compaction, and usage/billing helpers into the same module.
- Bind database operations through `createHeartbeatRunState(db)`.
Construction does no database work. The encoding cache stays local to
each service instance.
- Keep execution order, status transitions, session-goal recovery, cost
accounting writes, dispatch, and cancellation in `heartbeat.ts`.
- Preserve all 140 exports. A syntax-tree comparison confirms that 43
moved function bodies, 22 declarations, and nine service method bodies
are unchanged. The copied private string helper is unchanged.
Surrounding orchestration is unchanged apart from the factory binding.
- Add eight tests for legacy export identity, database binding, encoding
cache isolation, scoped resumes, cumulative usage, compaction, and stale
or cancelled conversation session writes.
- Document the module boundary in `doc/DEVELOPING.md`.

## Verification

- Before and after extraction, eight existing suites pass: 235 tests.
They cover workspace/session helpers, runtime state, cost accounting,
ledger attribution, task reset, timer reset, run lists, and run privacy.
- The new module suite passes: eight tests, including five real
PostgreSQL cases. Total focused coverage: 243 tests across nine suites.
- Commands: `pnpm exec vitest run
server/src/services/heartbeat/run-state.test.ts` and `pnpm exec vitest
run server/src/__tests__/heartbeat-runtime-state.test.ts
server/src/__tests__/heartbeat-cost-accounting.test.ts
server/src/__tests__/heartbeat-ledger-billing-code.test.ts
server/src/__tests__/heartbeat-task-session-reset.test.ts
server/src/__tests__/heartbeat-workspace-session.test.ts
server/src/__tests__/heartbeat-timer-wake-session-reset-pf4.test.ts
server/src/__tests__/heartbeat-list.test.ts
server/src/__tests__/heartbeat-run-privacy-routes.test.ts`.
- `pnpm --filter @paperclipai/server typecheck` passes.
- `pnpm -r typecheck` and `pnpm build` pass.
- The full local `pnpm test:run` did not complete and was stopped with
SIGINT after it reported an OAuth scope health-check failure and a
workspace-runtime timeout. All six OAuth variants and the workspace case
pass in isolation. A single rerun of the 413-case tool-access suite
passed the OAuth case but hit a different Notion callback timeout; that
Notion case also passes in isolation. The full local suite is not
claimed as passing.
- Greptile is 5/5 with no findings or inline comments on head
`2ca0908cfc112dd6ae26e2565cb3611b5bec52ba`. Its exact-head check
completed successfully.
- CI remains blocked on hosted-runner assignment after more than ten
minutes. [`ci / Select trusted
runner`](https://github.com/paperclipai/paperclip/actions/runs/37828454507/job/113487287178)
is queued for `ubuntu-latest` with no runner assigned. All seven
completed review/security checks pass; two optional Storybook checks are
skipped. The full CI matrix has not started. It must complete before
merge.

## Risks

- Wiring errors could bind reads to the wrong database or share the
encoding cache. Direct factory tests cover database and cache isolation.
Real PostgreSQL suites cover query and transaction behavior.
- Session writes must retain their issue row locks and generation
fences. The original function bodies are unchanged. New database tests
reject stale and cancelled conversation writes and clears.
- Existing SQL_ASCII safeguards and bounded projections remain in place.
No schema, API contract, migration, lockfile, or workflow changes are
included.

## Model Used

OpenAI GPT-6 via Codex. The exact serving model ID and context window
were not exposed in this session. Used reasoning, repository inspection,
shell tools, and code execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-08 14:53:48 -05:00
Devin Foley 53105d5830 feat(apps): add Gauge connection (#15596)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents reach external services through the Apps catalog. Each
catalog entry is a reviewed `AppDefinition` that connects a provider's
hosted MCP server to Paperclip's shared vault, grants, policies,
gateway, and audit trail.
> - Gauge (`withgauge.com`) measures how AI answer engines mention and
cite a brand. Its hosted MCP server also gives SEO, analytics and ads
reports, and runs a content pipeline that can publish to a connected
CMS.
> - Gauge is not in the catalog. Operators must use the generic "Connect
your own MCP server" flow. That flow has no branding, no organization
guidance, and no warning about live CMS publishing.
> - The connector playbook supports this provider with existing
definition fields. The server supports dynamic client registration with
a reviewed `mcp` scope, and it accepts an organization API key as a
bearer header.
> - This pull request adds the Gauge definition, its artwork, the
research and permission-review ledger rows, documentation, and
deterministic tests.
> - The benefit is a branded, governed Gauge connection with browser
sign-in, an API-key option, and a clear publish warning.

## Linked Issues or Issue Description

**Problem or motivation**

Marketing and growth teams use Gauge to track their brand in AI answers
and to run content workflows. They want their Paperclip agents to read
visibility, keyword and traffic data, and to prepare content. Gauge is
not in the Apps catalog. Operators must paste the MCP URL into the
generic remote-MCP flow, which gives no branding, no method guidance,
and no warning that content tools can publish live.

**Proposed solution**

Add a catalog-only Gauge connection built from the connector playbook.
Browser sign-in uses Gauge's dynamic client registration and requests
only the reviewed `mcp` scope. The user selects one Gauge organization
on the consent screen. An optional method sends a customer-created
organization API key as an `Authorization: Bearer` header. Both methods
warn the operator to set publish actions to Ask first. Every discovered
tool stays governed by the normal per-action policies.

**Alternatives considered**

A plugin was not needed because no custom UI, tables, workers, or
webhooks are involved. The identity scopes that Gauge also advertises
(`openid`, `profile`, `email`, `organizations`) are not requested,
because Gauge selects the organization on its consent screen. A
Gauge-specific `classifyRisk` rule was not added, because the tool names
cannot be seen without an account. The generic rule already classifies
`publish`, `create` and `update` tools as writes.

**Roadmap alignment**

This extends the existing self-serve remote-MCP connection catalog and
does not overlap planned core work.

## What Changed

- Added the `gauge` provider to `scripts/ingest-app-definitions.mjs`
(category `analytics`, API-key placement, guidance, warnings,
description) and regenerated
`packages/shared/src/app-definitions/gauge.json` and the generated
registry.
- Added the Gauge row to the self-serve MCP research ledger with
`dcr_or_api_key` auth and risk tier S3.
- Added permission reviews for `gauge/mcp-oauth` (explicit scope `mcp`,
from Gauge's live authorization-server metadata) and `gauge/mcp-api-key`
(provider key), with evidence links.
- Added Gauge's mark (`ui/public/brands/apps/gauge.png`, the avatar of
Gauge's official GitHub organization) and the brand manifest entry.
- Added gallery copy for the Gauge card.
- Added `doc/connections/GAUGE.md` (endpoints, scopes, administrator
setup, capabilities and policy, manifest, brand provenance, validation
hook) and linked it from the connections README and the permission
audit.
- Tests: Gauge definition shape, store visibility and artwork, URL
recognition, reviewed scope, bearer-header placement, the connect form's
sign-in default and API-key gating, and the pinned catalog counts.

## Verification

- `pnpm exec vitest run packages/shared/src/app-definitions.test.ts
packages/shared/src/app-definitions-url.test.ts
ui/src/pages/apps/AppsConnect.test.tsx ui/src/pages/apps/Browse.test.tsx
ui/src/lib/app-brand-assets.test.ts
ui/src/pages/apps/AppLogo.brand-assets.test.tsx
server/src/__tests__/tool-access-service.test.ts` — 738 passed.
- `node scripts/check-app-brand-assets.mjs` and `node --test
scripts/app-brand-validation.test.mjs` — passed.
- `pnpm --filter @paperclipai/shared typecheck`, `pnpm --filter
@paperclipai/server typecheck`, `pnpm --filter @paperclipai/ui
typecheck` — clean.
- `node scripts/ingest-app-definitions.mjs --definitions-only` produces
no Gauge drift.
- Manual: on a local instance, open Apps → Browse and confirm the Gauge
card and icon. Open `/apps/connect?source=gauge` and confirm that
sign-in is the default and that the API-key method is under Advanced.
- Live metadata probe on 2026-10-08: an unauthenticated `initialize` on
`https://app.withgauge.com/mcp` returns 401 with `resource_metadata`.
Both `.well-known` documents return the recorded endpoints and scopes.
Dynamic client registration succeeds. The authorize endpoint accepts
`scope=mcp` and rejects an unknown scope with HTTP 400.

## Risks

- Low risk to existing providers: the change is additive catalog data
plus tests. The generated registry only gains one import.
- Gauge content tools can publish to a connected CMS, including live,
and a bulk keyword update replaces each prompt's keyword list. Both
methods warn the operator, and the API-key helper text tells operators
to set publish actions to Ask first.
- Gauge API keys are not scoped and reach the whole organization.
Paperclip cannot narrow an issued key.
- The permission-review ledger records live proof for both methods as
not run. The lifecycle checklist in `doc/connections/GAUGE.md` still
needs a documented pass with a Gauge organization.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- Claude Opus 5.5 (`claude-opus-5-5`, 1M context) in Claude Code, with
tool use (shell, file editing, web research).

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-10-08 12:37:06 -07:00
DottaandPaperclip 0d57b9b98e refactor(heartbeat): extract run preparation (#15591)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The heartbeat service prepares context and configuration before
agent runs.
> - The workspace extraction left the main file at about 27,000 lines.
> - Preparation helpers form another large boundary with explicit inputs
and one database dependency.
> - This pull request moves those helpers into one module and preserves
their callers.
> - The benefit is a 24,000-line orchestration file and one place to
maintain run preparation.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

It improves the structure and test coverage of heartbeat run
preparation.

**Current behavior**

`server/src/services/heartbeat.ts` mixes context and configuration
preparation with execution, queueing, and recovery in 27,048 lines.

**Proposed behavior**

Move preparation into
`server/src/services/heartbeat/run-preparation.ts`. The main file loses
3,048 lines and ends at 24,000 lines. Keep the existing imports and
behavior. This continues #15568, #15573, and #15578.

**Reason and benefit**

A larger domain extraction makes useful progress toward a few manageable
files. Context, identity, environment, and tool preparation can be
reviewed together without searching the run engine.

**Breaking changes**

None. Existing public helpers and the configuration-incomplete error
class retain their identity.

## What Changed

- Move wake payloads, comments, attachments, skill mentions, adapter
environment configuration, and MCP/tool setup into
`heartbeat/run-preparation.ts`.
- Bind issue context, pinned routine snapshots, organization rows, and
responsible-user resolution through `createHeartbeatRunPreparation(db)`.
- Keep factory construction free of database work. Keep queueing,
dispatch, retries, cancellation, and execution order in `heartbeat.ts`.
- Preserve all 140 existing exports. A syntax-tree comparison confirms
that the 47 extracted function bodies and 25 moved declarations are
unchanged. The private string normalizer is also unchanged.
- Add six tests for legacy export identity, independent database
binding, company scope, pinned routine configuration, wake author
selection, company-owner fallback, and credential preflight.
- Document the preparation boundary in `doc/DEVELOPING.md`.

## Verification

- Passed: 183 focused tests across ten suites, including the six new
tests. Command: `pnpm exec vitest run
server/src/services/heartbeat/run-preparation.test.ts
server/src/__tests__/heartbeat-context-summary.test.ts
server/src/__tests__/heartbeat-agent-session-message.test.ts
server/src/__tests__/heartbeat-runtime-mcp-servers.test.ts
server/src/__tests__/heartbeat-runtime-skills.test.ts
server/src/__tests__/heartbeat-project-env.test.ts
server/src/__tests__/heartbeat-responsible-user-invariant.test.ts
server/src/__tests__/heartbeat-comment-wake-batching.test.ts
server/src/__tests__/run-secret-redaction.test.ts
server/src/__tests__/low-trust-red-team-routes.test.ts`.
- One runtime-skills case reported a Postgres deadlock in its after-test
`TRUNCATE` cleanup. A focused rerun passed both runtime-skills cases.
The other nine suites passed in the combined run.
- Passed: `pnpm -r typecheck`.
- Passed: `pnpm build`.
- Passed: all 54 GitHub checks on
`8d0cb9a251b10be5735fa13035e8a31e50b132e7`. The two optional Storybook
jobs were skipped. The complete CI test matrix is green.
- The full local `pnpm test:run` was started and then stopped after the
complete CI test matrix passed. It had no failures reported before
stopping and did not complete locally.
- Greptile: 5/5 on the same commit, with no inline comments or
actionable findings.

## Risks

The extraction crosses context, identity, credential, and tool-access
module boundaries. Existing company filters, secret redaction, low-trust
rules, and error classes stay intact. The new factory captures only the
service database and does no database work during construction. Existing
integration tests cover responsible-user authority, credential
boundaries, MCP tokens, and quarantine. New tests cover the factory
wiring and legacy identity. A syntax-tree comparison confirms that
orchestration changed only to bind the extracted loaders.

## Model Used

OpenAI GPT-6 via Codex. The exact serving model ID and context window
were not exposed in this session. Used reasoning, repository inspection,
shell tools, and code execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-08 13:47:29 -05:00
dependabot[bot] 9f4acc8f14 build(deps): bump @modelcontextprotocol/sdk from 1.30.0 to 1.31.0 (#15361)
Bumps
[@modelcontextprotocol/sdk](https://github.com/modelcontextprotocol/typescript-sdk)
from 1.30.0 to 1.31.0.
<details>
<summary>Release notes</summary>
<p><em>Sourced from <a
href="https://github.com/modelcontextprotocol/typescript-sdk/releases">@​modelcontextprotocol/sdk's
releases</a>.</em></p>
<blockquote>
<h2>1.31.0</h2>
<h2>Upgrade notes</h2>
<ul>
<li>Stored OAuth tokens and client information now include an
<code>issuer</code> field. Storage that rejects unknown fields needs to
allow it.</li>
<li>Pass <code>expectedIssuer</code> when constructing
<code>ClientCredentialsProvider</code>,
<code>PrivateKeyJwtProvider</code> or
<code>StaticPrivateKeyJwtProvider</code>. Constructing them without it
is deprecated.</li>
</ul>
<h2>What's Changed</h2>
<ul>
<li>[v1.x] Bind stored OAuth credentials to the authorization server
that issued them by <a
href="https://github.com/maxisbey"><code>@​maxisbey</code></a> in <a
href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/pull/2888">modelcontextprotocol/typescript-sdk#2888</a></li>
<li>chore: bump version to 1.31.0 by <a
href="https://github.com/claude"><code>@​claude</code></a>[bot] in <a
href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/pull/2890">modelcontextprotocol/typescript-sdk#2890</a></li>
</ul>
<p><strong>Full Changelog</strong>: <a
href="https://github.com/modelcontextprotocol/typescript-sdk/compare/1.30.1...1.31.0">https://github.com/modelcontextprotocol/typescript-sdk/compare/1.30.1...1.31.0</a></p>
<h2>1.30.1</h2>
<h2>What's Changed</h2>
<ul>
<li>[v1.x] fix(server): read HTTP request bodies with a size limit and
bound JSON-RPC batch length by <a
href="https://github.com/maxisbey"><code>@​maxisbey</code></a> in <a
href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/pull/2717">modelcontextprotocol/typescript-sdk#2717</a></li>
<li>fix(auth): preserve resource URI without trailing slash (<a
href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/issues/1968">#1968</a>)
by <a
href="https://github.com/MukundaKatta"><code>@​MukundaKatta</code></a>
in <a
href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/pull/1972">modelcontextprotocol/typescript-sdk#1972</a></li>
<li>chore: bump version to 1.30.1 by <a
href="https://github.com/claude"><code>@​claude</code></a>[bot] in <a
href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/pull/2848">modelcontextprotocol/typescript-sdk#2848</a></li>
</ul>
<h2>New Contributors</h2>
<ul>
<li><a
href="https://github.com/MukundaKatta"><code>@​MukundaKatta</code></a>
made their first contribution in <a
href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/pull/1972">modelcontextprotocol/typescript-sdk#1972</a></li>
</ul>
<p><strong>Full Changelog</strong>: <a
href="https://github.com/modelcontextprotocol/typescript-sdk/compare/1.30.0...1.30.1">https://github.com/modelcontextprotocol/typescript-sdk/compare/1.30.0...1.30.1</a></p>
</blockquote>
</details>
<details>
<summary>Commits</summary>
<ul>
<li><a
href="https://github.com/modelcontextprotocol/typescript-sdk/commit/4b0051f400219f8d8855f9a5433c6df35f15a639"><code>4b0051f</code></a>
chore: bump version to 1.31.0 (<a
href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/issues/2890">#2890</a>)</li>
<li><a
href="https://github.com/modelcontextprotocol/typescript-sdk/commit/51ad4f03190dc5c84ab8b1a25f9b78b277be0dc7"><code>51ad4f0</code></a>
[v1.x] Bind stored OAuth credentials to the authorization server that
issued ...</li>
<li><a
href="https://github.com/modelcontextprotocol/typescript-sdk/commit/289ac2c3af7e1536160e80414296b175171a1a87"><code>289ac2c</code></a>
chore: bump version to 1.30.1 (<a
href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/issues/2848">#2848</a>)</li>
<li><a
href="https://github.com/modelcontextprotocol/typescript-sdk/commit/12b425678a76cd54b0452a2ccf1e5dc7740f73ef"><code>12b4256</code></a>
fix(auth): preserve resource URI without trailing slash (<a
href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/issues/1968">#1968</a>)
(<a
href="https://redirect.github.com/modelcontextprotocol/typescript-sdk/issues/1972">#1972</a>)</li>
<li><a
href="https://github.com/modelcontextprotocol/typescript-sdk/commit/a9f6eb709b85459d01e8f2a9e881fef2621756c1"><code>a9f6eb7</code></a>
[v1.x] fix(server): read HTTP request bodies with a size limit and bound
JSON...</li>
<li>See full diff in <a
href="https://github.com/modelcontextprotocol/typescript-sdk/compare/1.30.0...1.31.0">compare
view</a></li>
</ul>
</details>
<br />

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-10-08 11:05:20 -07:00
Devin FoleyandPaperclip 317ed367d4 Handle incomplete task policies and preserve review limits (#15593)
Apply the reviewed change for Handle incomplete task policies and preserve review limits.

Validation: local typecheck/build and focused regression tests, passing exact-head CI, Greptile 5/5, and independent source review. The PR records full-suite evidence and any local environment limitations.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 10:20:32 -07:00
Devin FoleyandPaperclip 3d0e74e743 Keep proven external connection failures out of Sentry (#15590)
Apply the reviewed change for Keep proven external connection failures out of Sentry.

Validation: local typecheck/build and focused regression tests, passing exact-head CI, Greptile 5/5, and independent source review. The PR records full-suite evidence and any local environment limitations.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 10:19:40 -07:00
DottaandPaperclip 75918997a8 refactor(heartbeat): extract workspace preparation and resolution (#15578)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The heartbeat service prepares workspaces and dispatches agent runs.
> - Its main file still has more than 30,000 lines after two small
extractions.
> - Workspace preparation forms a larger boundary with one database
dependency.
> - This pull request moves that code into one workspace module with
focused tests.
> - The benefit is a smaller orchestration file and one place to
maintain workspace policy.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

It improves the structure and test coverage of heartbeat workspace
preparation.

**Current behavior**

`server/src/services/heartbeat.ts` mixes workspace preparation, run
execution, and scheduling in 30,634 lines.

**Proposed behavior**

Move workspace preparation into
`server/src/services/heartbeat/workspaces.ts`. Keep the public imports
and run behavior stable. The main file loses about 3,540 lines in one
extraction. Related extractions: #15568 and #15573.

**Reason and benefit**

A domain-level extraction makes useful progress toward a few manageable
modules. Future workspace changes can be reviewed without searching the
whole run engine.

**Breaking changes**

None. The existing public exports and workspace validation error class
retain their identity.

## What Changed

- Move managed checkout preparation, workspace validation and reuse,
referenced project resolution, and session/workspace config freshness
into `heartbeat/workspaces.ts`.
- Bind run workspace resolution to the database through
`createHeartbeatWorkspaceResolver(db)`.
- Keep the checkout single-flight map at module scope so all callers
share pending materialization.
- Keep all 140 existing `heartbeat.ts` exports. The 84 moved function
bodies and 53 other moved declarations are unchanged in a syntax-tree
comparison.
- Add seven tests for legacy export identity, database binding, issue
selection, session fallback, concurrent checkouts, same-name
repositories, and retry after failure.
- Document the module boundary in `doc/DEVELOPING.md`.

## Verification

- Passed: seven focused workspace suites, 264 tests. Command: `pnpm exec
vitest run server/src/services/heartbeat/workspaces.test.ts
server/src/__tests__/heartbeat-workspace-session.test.ts
server/src/__tests__/heartbeat-referenced-projects.test.ts
server/src/__tests__/heartbeat-remote-referenced-projects.test.ts
server/src/__tests__/heartbeat-workspace-branch-containment.test.ts
server/src/__tests__/heartbeat-accepted-plan-workspace-refresh.test.ts
server/src/__tests__/heartbeat-project-env.test.ts`.
- Passed: `pnpm -r typecheck`.
- Passed: `pnpm build`.
- Passed: all 54 GitHub CI checks on
`14dec775c7fa83716432f73740bd697106fdef68`. The two optional Storybook
jobs were skipped. The full CI test matrix is green.
- The full local `pnpm test:run` hit a 15-second timeout in the first
`public-mcp.test.ts` case. A focused rerun passed all 80 public MCP
tests. The long local run was stopped after all CI test shards passed;
it did not complete locally.
- Greptile: 5/5 on the same commit, with no inline comments or
actionable findings.

## Risks

Module initialization and shared checkout state are the main extraction
risks. Legacy exports retain the same function and class objects. The
checkout map stays outside the database factory. Existing integration
suites cover company boundaries, workspace containment, reuse, and
referenced project authorization. New local Git tests cover concurrent
materialization and failure cleanup. The orchestration body is unchanged
except for the resolver factory binding.

## Model Used

OpenAI GPT-6 via Codex. The exact serving model ID and context window
were not exposed in this session. Used reasoning, repository inspection,
shell tools, and code execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-08 10:56:33 -05:00
Devin FoleyandPaperclip 3367b75ccc fix: fence accepted work and cleanup before idle sleep (#15522)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - A host can stop an idle instance to reduce unused compute.
> - Zero active runs do not prove that requests, saved work, or cleanup
are complete.
> - A client can disconnect while its request still writes, and failed
cleanup can remain only in memory or on disk.
> - This pull request holds new admission and checks accepted work,
cleanup, and persisted work under an owned drain.
> - The host gets an empty report only while those checks remain valid.

## Linked Issues or Issue Description

Refs #13413. That companion PR uses the same owned-hold vocabulary for
runtime services. This PR covers HTTP requests, scheduler work,
accounting, cleanup, and the instance work inventory. It does not
include the preview gateway changes.

**What existing behavior does this improve?**

The instance task-drain API reports process counters. It does not prove
that a host can safely stop the instance for idle sleep.

**Current behavior**

A quiet run set can coexist with an unfinished request, cleanup after a
completed run, future work, or a failed accounting write.

**Proposed behavior**

Provide a bounded owned idle hold. Block new ingress, track accepted
handler promises, inspect durable and local work, and return `none` only
when the same hold stays quiet through the checks. Keep normal
deployment drains compatible.

**Reason and benefit**

Hosts can identify eligible idle instances without treating a disconnect
or failed cleanup write as completed work.

## What Changed

- Add `purpose: "idle"`, a bounded TTL, unique owners, and owner-checked
release to task drain. Existing holds cannot be replaced by another API
request.
- Gate HTTP ingress before parsers, auth, webhooks, and MCP. Gate new
WebSocket upgrades. Track async handlers in nested Express routers and
error middleware until they settle, even after the response or client
disconnect.
- Count accepted live-event WebSocket authentication through settlement,
even after disconnects. Count detached built-in agent, managed-home and
runtime-service startup reconciliation after readiness.
- Keep health and control mutations tracked. Count control-request
authentication separately from the read-only report, including
concurrent user/company/membership writes.
- Pause new scheduler admissions during idle holds. Count work already
in flight, including database backups, and reject scans whose work
generation changes. Periodic backups block sleep without a host wake
schedule.
- Inspect accounting and orphan-cleanup spool directories without
skipping temporary or malformed entries. Retain orphan tokens through
queue splices, flush failures, and buffer overflow. Keep failed usage
capture counted until its database failure fence is written.
- Check persisted work across companies in a bounded read-only
transaction. Include deferred agent-file cleanup and saved watchdogs
whose watched issues are complete. Enabled plugins and unsupported
retained work remain blockers.
- Require the exact idle owner and completed startup tracking before
returning an empty report. Keep reports free of tenant details.
- Document the hosting protocol, retry behavior, conservative blockers,
and the remaining external provider-stop race.

## Verification

- Current head: `8d9a599b732c2047b8c671d4799065d7c90a3567`, rebased on
master `941a3fa991aeb97eb1ac390c65b7973b5f6de1ad`. Both heartbeat helper
extractions are preserved. GitHub confirms no merge conflicts.
- Focused heartbeat renderer/run-log, drain, control-auth, admission,
route and PostgreSQL inventory checks: 234 tests passed in 10 suites
after the rebase.
- The POST task-drain contract includes the expected `409` conflict
response. The 84 OpenAPI and instance-settings route tests and server
typecheck passed after that final documentation fix.
- Final accepted-upgrade/startup regression run: 104 tests passed in
five suites, including success and failure after readiness or
disconnect. Final server typecheck also passed.
- `pnpm -r typecheck`: passed.
- `pnpm build`: passed.
- `pnpm install --frozen-lockfile` and the PR diff secret scan passed.
No dependency or lockfile changes are added by this PR.
- The prior verification covered HTTP disconnects and early responses,
async error handlers, concurrent authentication writes, signed
bootstrap, saved watchdogs, backup promises, orphan cleanup, spools and
failed accounting fences. Those tests passed.
- The last full local `pnpm test:run` stopped in the general-server
phase with 16,658 passing tests and five skill/connector fixture
failures caused by an ancestor workspace skill directory. That full
local run preceded this rebase and has not been repeated for the import
conflict. Full current-head CI passed: 53 successful checks and two
conditional skips, with no failures.
- All eight review findings are fixed and their threads resolved,
including accepted upgrade authentication, detached startup writes and
the POST conflict contract. Current-head Apex review is 5/5 with no new
findings; all eight review threads remain resolved.
- No live provider stop or production change was performed.

## Risks

- The hosting controller must use the owned protocol and hold external
admission through its final validation and provider stop. A legacy drain
cannot authorize idle sleep. During an idle hold, new requests receive
503 with `Retry-After: 1`; the host must handle queueing or retry before
enabling this path.
- This is a single-process protocol. An unexpected restart after the
last validation can race an external provider stop. The host must
serialize deploy/wake/sleep operations and bind the validation to the
instance it stops. Multiple replicas need shared fencing.
- Some retained state conservatively prevents sleep, including every
enabled plugin. This PR does not promise that every inactive instance
becomes eligible.
- The HTTP adapter uses Express 5 router layers. Real Express tests
cover nested routes, errors and disconnects. New routes must register
before tracking is installed. Detached work must have durable state or
explicit work tracking.
- Unrecoverable in-memory cleanup debt keeps the instance awake until
reconciliation. This change does not make such debt survive an unplanned
process crash. No schema migration or provider configuration change is
included.

## Model Used

OpenAI GPT-6 via Codex, with reasoning, repository inspection, tool use
and code execution. The exact deployment variant/model ID and context
window are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run focused tests locally and they pass; the full local
fixture limitation and clean full CI result are documented above
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-08 08:35:45 -07:00
DottaandPaperclip 71af2fbc3b fix(slack): teach agents how people connect their accounts (#15576)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Slack connections let people talk to an agent with their own
Paperclip permissions.
> - Each bot has a saved command that starts account linking.
> - The Access page explains this flow, but Slack agent turn context
omitted it.
> - An agent could guess the command or confuse channel membership with
Paperclip access.
> - This pull request gives agents the saved command and account
confirmation instructions.
> - People can ask the agent how to join without changing access
controls.

## Linked Issues or Issue Description

**What happened?**

Slack agent context explained messages, tools, and questions. It did not
explain how a teammate can connect an account. The saved slash command
can differ from the current agent name.

**Expected behavior**

Give the saved bot command, such as `/research_ops connect`. The
teammate runs it themselves. They sign in through a private confirmation
link. A new company member must receive admin approval before account
confirmation. The connection manager can copy instructions from Access →
Invite people.

**Steps to reproduce**

1. Configure a Slack bot with a custom slash command.
2. Rename its assigned agent.
3. Inspect a fresh or resumed Slack task prompt. Before this change, it
contains no account invitation instructions or saved connect command.

**Paperclip version or commit**

Base: `d6df12cef69fcaf2d2fe66a393168931d5b8b4e7`.

**Deployment mode**

Applies to local and hosted Slack chat connections. Deterministic tests
used a local isolated database. The new model probe has not run against
a live Slack bot.

Related work: Refs #15413 and #13638. This fixes agent guidance for
their existing account-linking flow. It adds no new membership system or
invitation endpoint.

## What Changed

- Read only the saved public slash command from the company-scoped
conversation endpoint. Supply it only to the assigned agent for Slack
turns with an active or verifying connection.
- Validate the command with the shared Slack configuration schema. Refer
to Access → Invite people when it is missing or invalid. Never guess
from the current agent name.
- Explain personal account confirmation, link expiry, and company
membership approval on fresh and resumed turns. Distinguish channel
invitations from Paperclip access.
- Add deterministic guidance regressions, a manual invitation model
probe, and setup documentation.

## Verification

- Passed: 63 tests in `heartbeat-context-summary.test.ts` and
`heartbeat-chat-task-link.test.ts`.
- Passed: four real-heartbeat regressions in
`heartbeat-slack-invitation.test.ts`. They check the saved command after
an agent rename, a persisted resumed session, full and compact prompts,
missing-command fallback, inactive endpoints, and a different assigned
agent. They also reject caller-supplied command fields and exclude other
setup metadata.
- Passed: the isolated `chat-channels.integration.test.ts` case
`discovers a Slack connect identity without starting work or granting
access`. It covers the private link, duplicate connect requests, and
nonmember access requests without a membership grant.
- Passed: `pnpm --filter @paperclipai/server typecheck` and `pnpm
--filter @paperclipai/server build`.
- Passed: `pnpm test:slack-connector --list` and `git diff --check`.
- Passed: repository `pnpm -r typecheck`, `pnpm build`, and `pnpm
check:token-gates`. Typecheck required local IPC access for the
migration check.
- Passed: final-commit CI, with 54 successful checks and two optional
Storybook checks skipped. CI includes all server, chat, workspace,
serialized, Runner, and browser test shards, typecheck, build, and
canary dry run.
- The full local `pnpm test:run` was started. It was stopped after full
CI passed; it had no final local summary. The focused invitation checks
passed locally. Do not count the interrupted local run as a full-suite
pass.
- Greptile gave the exact final commit
`f37e52bde621ef7f5b2bb345074d53345f8e4ee9` a 5/5 score. Both review
findings were fixed and their threads resolved.
- The new `invite-person` model probe is manual. These deterministic
results do not establish a live model or Slack acceptance pass.

## Risks

- Model guidance cannot prove that a person joined. The existing
account-linking and membership checks remain authoritative.
- Legacy rows without a saved command use the Access page fallback.
Invalid command text is excluded from the prompt.
- The query selects only the public command. It does not expose
registration secrets, tokens, or personal confirmation links.
- No schema changes, new provider requests, permission grants, or
telemetry changes.

## Model Used

- OpenAI GPT-6 through Codex. The exact backend variant and context
window are not exposed in this session. Used reasoning, repository
inspection, code editing, and test execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-08 10:28:14 -05:00
941a3fa991 Pin Copilot native dependencies with the maintained lockfile refresh (#15572)
Pin the three optional GitHub Copilot 1.0.88 native packages for Runner and server using the maintained lockfile workflow. Synchronize the package contract and bound initial render readiness in the deliberately throttled browser fixture. Current-head CI and focused checks pass.

Co-Authored-By: Dotta <cryppadotta@users.noreply.github.com>
Co-Authored-By: lockfile-bot <lockfile-bot@users.noreply.github.com>
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 09:37:53 -05:00
DottaandPaperclip 87a7312cfd refactor(server): extract heartbeat run-log formatting (#15573)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Heartbeat orchestration records agent output and run events.
> - The heartbeat service still contains more than 30,000 lines after
the first extraction.
> - Log formatting has fixed limits and does not write run state.
> - These helpers form a small next step toward a more manageable
heartbeat service.
> - This PR moves them into the heartbeat folder without changing their
function bodies.
> - Direct boundary tests and existing caller tests check the output.

## Linked Issues or Issue Description

Refs: #15568.

**What existing behavior does this improve?**

This improves the structure and test coverage of heartbeat run-log
formatting.

**Subsystem affected**

server/ — orchestration services.

**Current behavior**

`heartbeat.ts` contains excerpt handling, payload size limits, and log
chunk formatting beside run orchestration.

**Proposed behavior**

Move these helpers and their constants into
`server/src/services/heartbeat/run-log.ts`.
Keep the same function bodies, output, limits, and public exports.

**Reason and benefit**

Run-log formatting has its own small module and focused tests.
This keeps the second extraction small enough to review on its own.

**Breaking changes**

None. The existing public import path and persisted output stay the
same.

Related run-log PRs: Refs: #6373, Refs: #8841. Those PRs change
redaction behavior. This PR only moves existing formatting code.

## What Changed

- Move 109 lines of helper functions and six constants into
`heartbeat/run-log.ts`.
- Move `appendExcerpt` and retain the existing public exports for
`boundHeartbeatRunEventPayloadForStorage` and `compactRunLogChunk` in
`heartbeat.ts`.
- Add 17 direct regression cases for size and depth limits, cycles,
shared references, immutable input, image omission, redaction order,
excerpt tails, and UTF-8 boundaries.
- Document both heartbeat extractions in `doc/DEVELOPING.md`.

## Verification

- Before extraction, four focused files passed with 55 tests.
- After extraction and the two new excerpt tests, the same four files
passed with 57 tests.
- Run `pnpm exec vitest run
server/src/services/heartbeat/run-log.test.ts
server/src/__tests__/heartbeat-run-log.test.ts
server/src/__tests__/heartbeat-list.test.ts
server/src/__tests__/redaction.test.ts`.
- The original and extracted function blocks match byte for byte after
adding the export keyword to `appendExcerpt`.
- Existing tests still import the public helpers from `heartbeat.ts`.
- `node scripts/check-module-boundaries.mjs` and `git diff --check`
passed.
- `pnpm -r typecheck` passed.
- `pnpm build` passed.
- The full suite passed in CI. The duplicate local `pnpm test:run` was
stopped after CI passed; it did not complete locally.
- All CI gates passed on commit
`4e65c666e3b5b46162b3a528b37fba8902a1d200`: 54 passing checks, two
skipped Storybook checks.
- Fresh Greptile review of the same commit: 5/5, no actionable or inline
findings. The PR has no merge conflicts.

## Risks

- Low risk. Moving code can cause an import or build error.
- The same redaction and adapter utility modules remain in use.
- Database writes, current-user redaction, live event delivery, and run
state remain in `heartbeat.ts`.
- The event schema and payload output do not change.
- No database, API, Telemetry, Observability, or UI contract changes are
required.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

OpenAI GPT-6 via Codex. The exact serving model ID and context window
were not exposed in this session. Used reasoning, repository inspection,
shell tools, and code execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-08 09:34:37 -05:00
DottaandPaperclip 51653becc2 refactor(server): extract heartbeat task markdown rendering (#15568)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Heartbeat orchestration supplies each run with task context.
> - The task Markdown renderer is inside a service with more than 31,000
lines.
> - The renderer formats task data and does not write run state.
> - It is a small first step toward a more manageable heartbeat service.
> - This PR moves the renderer without changing its function body or
public export.
> - Direct tests and existing caller tests check the output before and
after the move.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

This improves the structure and test coverage of heartbeat task context
rendering.

**Subsystem affected**

server/ — orchestration services.

**Current behavior**

`buildPaperclipTaskMarkdown` occupies 352 lines inside `heartbeat.ts`.
The service also owns run execution, scheduling, recovery, and
cancellation.

**Proposed behavior**

Move the renderer to `heartbeat/task-markdown.ts`.
Keep its function body and the existing `heartbeat.ts` export unchanged.
Leave orchestration in place for this first PR.

**Reason and benefit**

Prompt formatting has its own small source file and direct regression
tests.
Reviewers can check this one extraction before a later refactor.

**Breaking changes**

None. The input type, rendered output, and existing import path stay the
same.

Related renderer changes: Refs: #4732, Refs: #14030.
These PRs change prompt behavior. This PR only moves the current
renderer.

## What Changed

- Move the 352-line renderer to
`server/src/services/heartbeat/task-markdown.ts`.
- Import and re-export it from `heartbeat.ts`. Remove imports used only
by the renderer.
- Add seven direct regression cases. Cover empty context, ordered
comment-only wakes, input preservation, nested code fences, ancestor
limits, attachment-only wakes, and rejected plans.
- Keep the renderer and its direct tests in
`server/src/services/heartbeat/`. Document this folder as the home for
relevant later extractions in `doc/DEVELOPING.md`.

## Verification

- The focused suite passed before and after extraction, and after moving
to the heartbeat folder: four files, 85 tests.
- Run `pnpm exec vitest run
server/src/services/heartbeat/task-markdown.test.ts
server/src/__tests__/heartbeat-context-summary.test.ts
server/src/__tests__/heartbeat-chat-task-link.test.ts
server/src/__tests__/codex-local-execute.test.ts`.
- The extracted function matches the original function byte for byte.
The folder move only adjusts imports.
- `node scripts/check-module-boundaries.mjs` passed.
- `git diff --check` passed.
- The initial extraction passed local `pnpm -r typecheck` and `pnpm
build`.
- The folder update passed `pnpm --filter @paperclipai/server exec tsc
--noEmit` and `pnpm --filter @paperclipai/server build`.
- The initial extraction passed the full CI suite. Its duplicate local
`pnpm test:run` was stopped after CI passed; it did not complete
locally.
- The folder update passed all CI gates on commit
`e21e589b379d7ca2fae16c1dcd91cf2f604fd03f`: 54 passing checks, two
skipped Storybook checks. Three test shards passed after one retry
following simultaneous runner shutdowns; those failures had no failed
test assertions.
- Fresh Greptile review of commit
`e21e589b379d7ca2fae16c1dcd91cf2f604fd03f`: 5/5, no actionable or inline
findings.

## Risks

- Low risk. A moved module can change import resolution or expose an
import cycle.
- Existing caller tests still load the compatibility export from
`heartbeat.ts`.
- The same guidance constants and public task URL resolver remain in
use.
- No run-state writes, transactions, locks, shared process state, or
cleanup paths moved.
- No database, API, telemetry, or UI contracts changed.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

OpenAI GPT-6 via Codex. The exact serving model ID and context window
were not exposed in this session. Used reasoning, repository inspection,
shell tools, and code execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-08 09:06:04 -05:00
DottaandPaperclip f47614046d fix(slack): upload agent avatars directly during Cloud setup (#15566)
## Thinking Path

> - Paperclip helps people manage AI agents for work.
> - Slack setup creates a dedicated bot for an agent.
> - The setup should upload that agent's avatar.
> - The current upload asks Slack to fetch an image from the board
origin.
> - Cloud requires a tenant session at that origin, so Slack receives
HTTP 401.
> - This change renders the PNG on the server and uploads the file
directly.

## Linked Issues or Issue Description

Refs #15413.

**What happened?**

Automatic Slack setup left the app and bot with default icons on Cloud
staging. An unauthenticated request to the exact avatar URL returned
HTTP 401 with `tenant_session_required`. Local setup did not expose this
Cloud ingress requirement.

**Expected behavior**

New Slack bots receive the assigned agent's 512-pixel avatar with the
Paperclip dark background.

**Steps to reproduce**

1. Create a Slack app through automatic setup on a Cloud tenant.
2. Complete installation.
3. Inspect the bot avatar in Slack. The previous URL-based upload cannot
fetch the image without a tenant session.

## What Changed

- Render the assigned agent's preset PNG with the existing bounded
worker pool.
- Send PNG bytes as multipart `file` data to `apps.icon.set` instead of
passing a board URL.
- Keep the temporary token in the Authorization header. Let fetch set
the multipart boundary.
- Recheck management permission and credential-lease ownership after
rendering. Close the worker pool during chat service shutdown.
- Cover actual PNG dimensions, uploaded bytes, failure recovery, and
secret-safe responses. Update deployment documentation.

## Verification

- Passed: 84 focused tests across automatic Slack registration and
on-demand agent avatars.
- Passed: full repository build.
- Passed: full repository typecheck. All 54 current-head GitHub checks
passed, including the complete test matrix, all browser shards, build,
typecheck, canary, and security checks. Greptile completed on
`79e79f6fc` with 5/5 and no actionable findings or open review threads.
- The Cloud fetch failure was reproduced without browser credentials. No
Cloud access rule was changed.
- A real Slack upload with this new path still requires deployment and a
fresh automatic setup. Existing apps retain the manual avatar-upload
fallback.

## Risks

- The renderer can time out or Slack can reject the upload. Both
failures preserve the saved app and leave installation usable.
- The renderer adds a bounded, lazy worker pool to Slack registration.
Shutdown closes it.
- No migration, bot permissions, credential retention, or Cloud
authentication behavior changes.

## Model Used

OpenAI Codex, GPT-6 family, with reasoning, repository tools, code
execution, and browser inspection. The runtime does not expose a more
specific authoring model ID or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-08 08:12:18 -05:00
DottaandPaperclip 2f0c485dec fix(skills): ship the completion helper with the installed skill (#15554)
## Thinking Path

> - Paperclip manages work for AI agents.
> - Legacy agents use the Paperclip skill to save task status and
comments.
> - The skill names a script relative to the task workspace.
> - That script exists only in the Paperclip source repository.
> - Agents in other workspaces can hit a missing command or search for
it.
> - This PR ships the helper inside the skill and uses the installed
skill path.
> - The repository command remains available through a forwarding
wrapper.

## Linked Issues or Issue Description

Fixes #9527. Refs #15548 for the preceding runtime checkout guidance.
Related: #6052 addresses LF line endings for the repository helper; this
change addresses helper delivery and path resolution.

## What Changed

- Bundle the existing issue update helper with the Paperclip skill.
Preserve its HTTP checks, echoed-status check and two-attempt limit.
- Resolve the command from the installed skill directory. Use a verified
PATCH when that path is unavailable, without searching the filesystem.
- Keep the repository command as a wrapper that works from any
directory.
- Test shell execution and exact status/comment payloads through both
provider skill-home layouts, including paths with spaces.
- Add helper sources and existing verification tests to stock-harness
admission. Record an absent historical helper explicitly. Add the
missing declaration for the admission fingerprint export.

## Verification

- Complete directly affected source suites: 30 tests pass. They cover
skill delivery, preserved multiline comments and links, authentication
headers, empty responses, mismatched status, transient retries and
definitive rejections.
- Product E2E typecheck passes. Support suites: 1,835 Vitest tests pass,
one is skipped; 128 Node tests pass.
- Full local build and workspace typecheck pass.
- Full local repository tests are not claimed as passed. Embedded
PostgreSQL was unavailable in this worktree during the preceding task;
Linux CI will run the repository gates.
- The authorized matched Codex/Claude comparison is pending. It uses the
existing assigned-skill case and original oracle, one initial attempt
per profile and variant.
- CI and a fresh Greptile review are pending. Keep this PR in draft
until readiness gates complete.

## Risks

- Correct path resolution depends on the harness supplying the installed
skill path. The instructions use verified PATCH when that path is
unavailable.
- The helper still requires Bash, curl and jq. Its existing retry and
response-verification behavior is unchanged.
- Tests use the shared skill-directory symlink mechanism and an HTTP
fixture. Real provider completion behavior still requires the bounded
live comparison.
- This fix does not redesign native completion, legacy recovery or
ambiguous transport handling.

## Model Used

OpenAI Codex, GPT-6 family. The exact model build and context window are
not exposed in this session. Used code editing, shell tools and test
execution.

## Checklist


- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-08 07:21:30 -05:00
DottaandPaperclip 71cd0a2621 fix(skills): honor the current run harness checkout (#15548)
## Thinking Path

> - Paperclip manages work for AI agents.
> - The runtime claims eligible assigned tasks before it starts an
agent.
> - The wake tells the agent when the runtime already holds that claim.
> - The legacy skill still requires another checkout in every case.
> - This PR makes the skill honor the current task and run claim.
> - Manual checkout and server ownership checks remain in place for
other cases.

## Linked Issues or Issue Description

**Where is the issue?**

`skills/paperclip/SKILL.md`, in the scoped wake procedure and Step 5.

**What's wrong?**

The wake can say that the harness already checked out the issue. The
skill still tells the agent that it must call checkout. These
instructions conflict.

**Suggested fix**

Skip the second checkout only when the runtime wake explicitly confirms
the claim for this issue and run. Retain manual checkout when that
statement is absent or the agent selects another task. Refs #14948 for
the existing shared prompt reduction.

## What Changed

- Honor the explicit runtime claim in the scoped wake procedure and Step
5.
- Keep context reads, status writes, deliverable handling and conflict
rules.
- Add checks for normal and resumed wake text and excluded automatic
claims.
- Retain successful checkout HTTP activity for legacy stock-task evals.
Bind each receipt to the exact company, task, agent and run. Keep this
observation separate from the original task grades.

## Verification

- Checkout observation calibration: nine tests pass.
- Focused skill, wake and database ownership tests: in progress.
- Full repository build, typecheck and tests: in progress.
- Planned live comparison: the existing assigned-skill document case on
legacy Codex and Claude. One attempt per variant and profile. No
automatic retries. The baseline and candidate share the observation code
and task oracle.
- Live results are pending. This draft does not claim behavioral
qualification.

## Risks

- Agents may misread prompt guidance. The API still enforces ownership;
the text grants no new authority.
- The exception is specific to the current issue and run. It does not
remove ordinary legacy completion writes or authorize another task.
- Activity measures successful checkout HTTP calls. Failed attempts
require separate run-log inspection. Missing or mismatched observations
cannot count as zero calls.
- One trial per profile cannot establish general reliability, speed or
cost trends.

## Model Used

OpenAI Codex, GPT-6 family. The exact model build and context window are
not exposed in this session. Used code editing, shell tools and test
execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-08 06:38:09 -05:00
DottaandPaperclip d66acb7ac1 feat: automate Slack bot app setup and installation (#15413)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Chat connectors give each agent a customer-owned bot and task-backed
conversations.
> - Manual Slack setup requires app creation and copying durable
credentials.
> - Operators need a shorter setup that an assisting agent can use
safely.
> - This pull request creates the app through Slack's Manifest API and
installs it through OAuth.
> - Durable registration state supports recovery without creating
another app.
> - A four-screen wizard, automatic avatar upload, and OAuth account
linking reduce setup work.
> - Connector settings and per-turn tool guidance support daily use
after installation.

## Linked Issues or Issue Description

**Subsystem affected**

Native Slack bot setup, company secret storage, chat connector
management, and agent tool guidance.

**Problem or motivation**

New Slack bots require manual app creation and copying a signing secret
and bot token. Interrupted setup can create duplicate apps. The setup
and management screens contain unnecessary controls. Agents also need
guidance for native questions, files, thread replies, and governed Slack
actions.

**Proposed solution**

Use a temporary app-configuration access token to create a
customer-owned app. Save durable secrets in the vault. Bind OAuth to the
initiating actor, company, endpoint, registration revision, scopes, and
configured origins. Preserve manual and existing-app recovery. Link the
installing user's account, send a welcome DM, and advance from saved
server evidence. Keep request-URL recovery instructions available if
automatic connection detection waits.

**Alternatives considered**

The Slack CLI adds installation requirements. Socket Mode changes
transport. A shared Paperclip-owned app changes app ownership. These
alternatives are outside this change.

**Roadmap alignment**

This extends existing chat connectors and secrets capabilities. Related
public work: #14037 and #13954 cover Slack MCP prerequisites and user
OAuth. No duplicate bot-registration PR was found.

## What Changed

- Share one reviewed manifest builder between automatic registration and
manual setup.
- Add replay-safe migration 0318 and company-bound registration state
with vault references and uncertain-creation recovery.
- Add registration, installation, callback, and resume APIs with
short-lived, single-use OAuth state.
- Save installation credentials before downstream checks and preserve
bot identity constraints.
- Reduce automatic setup to four screens. Keep advanced app details,
manual recovery, and existing-app setup.
- Upload the agent avatar with the Paperclip dark background. Link the
OAuth installer's account and send setup DMs.
- Show agent and connector-owner avatars. Simplify settings, access, and
conversation screens.
- Discover joined Slack channels and enable them by default. Start a
task from a bare mention and admit same-thread follow-ups.
- Refresh Slack tool guidance each turn. Add native-form, file,
approval, and delivery regressions plus manual model probe definitions
and sanitized acceptance records.
- Update deployment/database docs, OpenAPI, redaction, removal cleanup,
production Storybook stories, and provider browser tests.
- Merge current master and move the registration migration after its
latest migration without rewriting published commits.

The completed Slack success view intentionally has a single centered
**Done** action and no **Save & exit**, as explicitly requested by the
product owner. `DESIGN.md` records this exception; unfinished setup
steps retain the aligned wizard footer.

## Verification

- Passed after the master merge: repository typecheck, full build,
Storybook build, design-token gates, module-boundary gates, and
migration generation.
- Passed: all 352 focused Slack deterministic tests and all 14 affected
provider browser tests. Browser tests use controlled provider fixtures
and a separate throwaway instance.
- Passed on current head `c5d01e0e2`: the complete GitHub test matrix
(general server, chat, all workspaces, serialized server, and Runner),
all eight browser shards, typecheck/release registry, build, canary dry
run, security checks, and policy gates. There are 52 passing checks and
no pending or failing checks.
- Greptile completed on the exact current head with 5/5 and no
actionable findings or open review threads.
- Local repair verification passed 93 focused tests, including same-app
reinstall after revocation and rejection of consent started before
revocation, the AgentMail browser journey, and repository typecheck.
Local build and Storybook build also passed. The redundant local
full-suite rerun was stopped after the complete current-head CI matrix
passed.
- Real Slack setup and agent replies were exercised in the authorized
isolated test drive during the setup iteration.
- The ten additional model probes were attempted with legacy
`codex_local`, `gpt-5.6-sol`: five passed, two failed, and three were
partly verified. Native runtime is not qualified. See
`server/src/services/connectors/slack/evals/2026-10-08-acceptance.md`
for evidence and limits.
- Passing model probes cover native forms, downloaded file bytes, bare
mentions with thread replies, explicit posts/reactions, and saved
approval denial.
- The controlled uncertain-write probe found wrong delivery-check IDs.
The canvas fallback attempt used an invented tool name. Search
pagination/native search, a private-source denied-tool receipt, and
distinct board/webhook origins remain unqualified.

Reviewer path: enable Chat connectors, start Slack chat setup, select an
agent, enter an app-configuration access token, and approve Slack
installation. Send a message to the bot and confirm that setup advances
to success. Inspect settings and allowed channels. See
`doc/connections/SLACK-AUTOMATIC-SETUP.md` for deployment and recovery.

## Risks

- Slack app creation has no provider idempotency guarantee. A timeout
after dispatch stays uncertain until the operator checks Slack.
- OAuth needs a stable public HTTPS board origin. Webhook ingress may
use a separate configured HTTPS origin. Workspace policy can delay
installation.
- Migration 0318 can replay safely on instances that applied the earlier
development migration.
- OAuth installation now links the installer to the initiating Paperclip
user. Identity checks and company access rules still apply.
- Joined channels now enable bot responses by default. Linked-user
authorization and per-action approval rules still apply.
- Model behavior has the documented delivery-check and canvas fallback
failures. A passing CI run does not establish that every model probe
passed.
- Removing the connection does not delete the customer's Slack app. No
new first-party telemetry is added.

## Model Used

OpenAI Codex, GPT-6 family, with reasoning, repository tools, code
execution, and browser verification. The runtime does not expose a more
specific authoring model ID or context-window size. The live bot probes
used OpenAI `gpt-5.6-sol` through `codex_local` in legacy mode.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-08 06:33:31 -05:00
DottaandPaperclip d0f69670db fix(runner): recover saved execution prompts after upgrades (#15518)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Native runs save an immutable execution context for restart
recovery.
> - The context includes the prompt text, revision, and content hashes.
> - The parser required that saved prompt to match the current release.
> - A server upgrade could reject a valid saved run before provider
recovery.
> - This pull request validates and preserves the saved prompt snapshot.
> - Routine prompt changes no longer need a catalog of past strings.

## Linked Issues or Issue Description

**What happened?**

A hot restart selected a dead native runner for same-run recovery.
Reading its saved v5 execution input failed with
`input.runtimeContext.prompt must match the fixed Paperclip prompt
revision`. The new controller accepted only v6.

**Expected behavior**

Recovery uses the saved prompt and validates its content hashes. It
preserves the same run and provider session without starting a duplicate
turn.

**Steps to reproduce**

1. Start a native run and save its execution input and provider
checkpoint.
2. Change the fixed execution prompt in the server release.
3. Stop the runner and recover the saved run with the new controller.
4. Observe that the old parser rejects the saved prompt before provider
recovery.

**Paperclip version or commit**

The v5-to-v6 prompt change was introduced in #15446. The defect also
reproduces on current master before this fix.

**Deployment mode**

Source-built server with the native runner.

Related work: #15446 added task-monitor guidance. The held prompt-size
experiment in #15489 changes prompt wording but does not add recovery
compatibility.

## What Changed

- Read the prompt text and revision from the saved execution snapshot.
- Treat the revision as non-empty metadata and preserve the exact saved
bytes.
- Validate the prompt SHA-256 and the aggregate context digest.
- Keep fresh-run builders on the current prompt constants.
- Test arbitrary saved prompts, malformed fields, altered text, stale
hashes, and aggregate drift.
- Test recovery parsing for Codex input versions v3-v5 and OpenCode,
ACPX Pi, and Dot v6 inputs.
- Extend the real-process restart suite with both the incident's v5 wire
fixture and a prompt unknown to this release.
- Run the restart recovery suite in the existing Rust-equipped PR lane,
where its runner and fake-provider binaries are built. Verify complete,
non-overlapping test coverage for PR, release, and local callers.
- Document recovery from saved snapshots without a historical prompt
catalog.

## Verification

- Red: the new contract regressions fail against the catalog-based
parser with the original prompt-validation error.
- Green: 53 focused contract and materialization tests pass.
- Red: the real-process unknown-prompt regression fails with master's
original parser at the saved-input recovery read after process loss.
- Green: all 15 real-process restart tests pass locally on the final
branch. The saved-prompt cases keep the run and provider session,
replace the PID, and record one `turn/start`.
- Local repository `pnpm -r typecheck` and `pnpm build` passed after
rebase on `89f09dad723766e5351953f0731b9aa5daada28d`. The 53 focused
tests also passed on that head.
- Red: the new test-roster checks fail against the old CI placement.
- Green: all 26 test-scheduling checks pass after moving the restart
suite.
- CI ran all 15 restart recovery tests with no skips on final head
`6719fc2bb7a31a0f72ea04c7e525a63dcc6f9105`. [Runner test
job](https://github.com/paperclipai/paperclip/actions/runs/37764782334/job/113271464966).
- Greptile scored 5/5 on that exact head with no actionable findings.
- The complete CI matrix passed on final head
`6719fc2bb7a31a0f72ea04c7e525a63dcc6f9105`: general/workspace tests,
serialized server suites, both runner Vitest lanes, Rust and static
checks, all browser shards, typecheck, build, and the canary dry run.
[CI
run](https://github.com/paperclipai/paperclip/actions/runs/37764782334).
- There are no unresolved review threads or merge conflicts.
- Reproduce focused tests with `pnpm --filter
@paperclipai/paperclip-runner exec vitest run
src/contracts/runtime-context.test.ts
src/contracts/native-execution.test.ts
src/drivers/runtime-context-materializer.test.ts`.
- Reproduce restart tests with `pnpm --filter
@paperclipai/paperclip-runner build:rust` followed by `pnpm exec vitest
run
server/src/services/native-runtime/native-runner-restart-recovery.integration.test.ts`.
They use temporary PostgreSQL, real runner processes, and a fake Codex
provider. They do not use paid inference.

## Risks

- The parser now accepts internally consistent saved prompt text that is
absent from the current source. Inputs must come from trusted server
persistence. Content hashes verify consistency; they do not authenticate
authorship.
- Existing execution-schema, ownership, checkpoint, provider,
permission, and session-compatibility checks still apply.
- This change validates the saved base prompt. It does not make all
additional code-generated instruction strings versioned.
- The process-level recovery proof uses Codex. Other provider coverage
verifies the shared input parser and retained provider configuration.

## Model Used

OpenAI Codex, GPT-6 family, with repository inspection, code editing,
and test tools. The exact serving model ID and context-window size are
not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-08 05:57:20 -05:00
DottaandPaperclip bf9dbd18a8 fix(runner): recover incompatible Codex models before launch (#15519)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The Runner prepares and saves each task's provider configuration
before launch.
> - A sandbox image can have a supported Codex CLI that is too old for
the selected model.
> - The current version check stops that task even when the image can
run a similar model.
> - This pull request selects a compatible model before it saves a fresh
execution.
> - The task continues with a visible warning, and recovery uses the
saved effective model.

## Linked Issues or Issue Description

- Refs #15053. This fixes the same CLI and model mismatch for fresh
remote Runner tasks. It does not change subscription onboarding.
- Refs #14721. This adds a bounded startup choice for Codex CLI
compatibility. It does not add user-defined fallback chains or quota
failover.
- Related PR: #15399 added the model-specific CLI checks that this
change preserves at launch.

## What Changed

- Check the preinstalled remote Codex CLI before saving a fresh
execution that uses a model with a verified CLI minimum.
- Select the closest compatible older model in the same class, then the
stable Runner default. Consider each candidate once.
- Save the effective model before checkpoint selection. Keep the
requested agent and task settings unchanged.
- Add a task warning and a local `runner.model_fallback` run-log event
with both models and the CLI version.
- Share executable discovery with the launch verifier. Preserve explicit
artifact and install settings, saved executions, and existing artifact
checks.
- Add regression tests and document the selection and warning behavior.

## Verification

- Red: the three new provider-configuration regression cases failed
before the fix. They kept the incompatible requested model.
- Red: both rejected-probe cases failed before the review fix. They now
defer to launch verification.
- Green: targeted provider configuration, remote preflight, task
warning, and launch-verifier tests pass (680 tests).
- `pnpm -r typecheck` passes.
- [Full
CI](https://github.com/paperclipai/paperclip/actions/runs/37713759744)
passes on `eb273fc32661526d90a864b42c73a7da61dd62da`: 47 successful
jobs, including all general and serialized Vitest groups, browser
shards, Runner verification, typecheck, build, and the canary dry run.
- Local extended verification: three serialized shards pass (115
suites). HTTP route tests hit intermittent 15-second timeouts. The
authorization suite passes on rerun (130 tests). The document suite
passes on both master and the PR head in the same isolated setup (6
tests). The local general run was stopped after the equivalent full CI
matrix passed.
- `pnpm build` passes.
- Server typecheck and build pass again after the probe-error review
fix.
- `git diff --check` and `pnpm check:module-boundaries` pass.
- Example: request `gpt-6.1-sol` on Codex 0.158.0 to use `gpt-6-sol`; on
0.156.0, use `gpt-5.6-sol`. The warning names the requested model,
effective model, and CLI version.
- This PR has not been deployed to staging.

## Risks

- A fallback can have different capabilities. The task warning makes the
substitution visible. Later runs can use the requested model after the
image CLI is updated.
- Preparation adds two remote commands for models with a verified CLI
minimum. Failed or invalid version probes retain the existing launch
checks.
- This only handles known CLI and model mismatches before a fresh
provider launch. Authentication, capacity, and artifact failures keep
their existing behavior.
- No database migration is required.

## Model Used

OpenAI GPT-6 through Codex, with tool use and code execution. This
session does not expose the exact deployment model ID, context window
size, or reasoning setting.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-08 04:57:26 -05:00
Devin FoleyandPaperclip 5717523b9e fix: retain the original repository for reused task workspaces (#15528)
Apply the reviewed change for fix: retain the original repository for reused task workspaces.

Validation: required local typecheck and tests, passing CI, Greptile 5/5, and independent review of the exact source head.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-07 19:52:30 -07:00
Devin FoleyandPaperclip 1d19f9b562 Distinguish failed Git inspection from missing worktree registration (#15525)
Apply the reviewed change for Distinguish failed Git inspection from missing worktree registration.

Validation: required local typecheck and tests, passing CI, Greptile 5/5, and independent review of the exact source head.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-07 19:05:04 -07:00
Devin FoleyandPaperclip 6b171c9616 Keep proven local workspace configuration conflicts out of Sentry (#15521)
Apply the reviewed change for Keep proven local workspace configuration conflicts out of Sentry.

Validation: required local typecheck and tests, passing CI, Greptile 5/5, and independent review of the exact source head.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-07 18:54:35 -07:00
DottaandPaperclip dd777f4b73 feat(dot): complete onboarding and expand governed Runner capabilities (#15414)
## Thinking Path

> - Paperclip manages AI agents, their work, and their permissions.
> - Paperclip Runner supplies the same admitted tool authority to each
provider.
> - The Dot provider in #15402 needs reliable onboarding and useful
agent capabilities.
> - An idle Dot could not start work, assign a human, read skills, or
produce a workspace artifact.
> - Pairing also relied on a second configuration save before ordinary
admission could work.
> - This pull request adds governed idle admission and shared Runner
tools, and completes pairing atomically.
> - Operators can test Dot with fresh data while keeping normal company,
approval, budget, and run ownership checks.

## Linked Issues or Issue Description

**Subsystem affected**

Paperclip Runner, dedicated Dot MCP access, OAuth onboarding,
experimental settings, and empty worktree startup. This PR builds on
merged provider PR #15402. It reuses the merged MCP gateway from #14846
and assistant connection work from #14933 and #15380.

**Problem or motivation**

An idle Dot could see assignments but could not act on a conversation
request until someone created a task first. Its Runner catalog could not
assign tasks to humans or read pinned skills and workspace files.
First-time OAuth discovery and pairing also needed browser fixes, and a
completed pairing did not persist its binding reference on the agent.

**Proposed solution**

Keep Dot within the existing Runner. Admit a visible agent-authored
intake task for idle requests. Add people, human assignment, cross-task,
skill, and optional sandbox workspace tools through shared authority.
Relay assigned app calls through the configured MCP gateway. Add lease
renewal and follow-up references. Save pairing and its configuration
revision atomically. Show prerequisites and provide a complete copy
prompt.

**Roadmap alignment**

This extends the experimental provider in #15402. It uses the existing
governed gateway, task model, skills, artifact path, and native Runner.
It adds no separate execution subsystem.

## What Changed

- Add a standalone OpenAI Dot agent choice with an independent
experimental opt-in. It works with the general Runner option off.
Require Assistant connections (MCP), authenticated sign-in, and public
HTTPS for pairing.
- Use Dot’s own option for company package import. Allow an unpaired Dot
configuration to save after external billing acknowledgement; task
admission still requires pairing. Prepare the shared dev binary for
Dot-only opt-in.
- Label saved Dot agents as OpenAI Dot. Use prerequisite-check copy in
setup and runtime configuration.
- Route Dot creation directly to pairing after explicit external billing
acknowledgement. Hide local CLI, model, and harness setup for Dot.
Preserve the shared Runner implementation and canonical API type. Other
Runner providers still require their general opt-in.
- Add an empty worktree option with fresh signing keys and no production
data copy.
- Fix public-client OAuth negotiation, discovery compatibility, and
optional separate browser authorization origin.
- Add one-use pairing consent preview, clear copied setup instructions,
and atomic binding persistence and cleanup.
- Add idle request admission with stable request IDs and normal
scheduling, permissions, budgets, and task ownership.
- Add identity and people discovery, human task assignment and
reassignment, and authorized cross-task comments and documents.
- Read assigned skill files from pinned manifests. Relay assigned app
calls through the merged gateway without exposing credentials.
- Add an off-by-default workspace bridge. Constrain paths and writes.
Run commands in a deny-by-default OS sandbox with no network or injected
credentials. Reserve mutations before effects and never blindly repeat
uncertain work.
- Add rolling lease renewal, task pagination, bounded operation limits,
and deduplicated follow-up references without comment bodies in
webhooks.
- Add an off-by-default attachment reading setting. Restrict reads to
files on the current assigned task. Verify size and hash, cache bounded
verified copies per run, paginate text or binary bytes, and recheck live
authority before returning.
- Keep file grants operator-owned. Reject agent self-grants across
configuration routes. Preserve attachment consent in create/import
forms. Close generic API file bypasses while retaining current-run
response snapshots and permitted uploads.
- Keep provider limits explicit. Do not inherit a Dot binding,
attachment permission, or workspace permission when hiring another
agent.
- Include the required Markdown format in cross-task document writes and
validate the API title limit. Verify real creation and revision
persistence.
- Restrict command execution to Linux bubblewrap with descendant
containment. macOS retains workspace file tools and artifact publishing,
while refusing command calls. Explain the platform limit in setup.
- Exclude Paperclip instance state from workspace files, uploads,
artifact publication, and sandbox commands. Protect nested directories
and case variants. Fail closed when the directory protection scan
exceeds 4,096 directories.
- Add static UI compression for slow public tunnels. Document setup, the
complete tool inventory, and qualification limits.

## Verification

- Merge preparation on `00ca2c75b` integrates merged base #15402 and
master `fc6304dfe`. The ancestry commit preserves the reviewed follow-up
source tree. The subsequent security fix excludes instance state from
file tools, uploads, artifact publication, and sandbox commands. It
preserves private task authorization and task monitors. It regenerates
the combined tool catalog, seeded catalog digest, and protocol manifest.
Workspace typecheck, full build, and UI token gates pass. All sixteen
real Dot broker cases and 93 company import cases pass after the review
fixes and creator-attribution test correction. The exported
human-assignment catalog and Unicode page boundaries are also fixed. All
ten catalog tests and nine workspace/skill bridge tests pass. All 35
workspace bridge and authority tests passed after the instance-state
fix, including real macOS commands in that intermediate version. The
subsequent document and descendant-containment fixes pass 44 focused
tests across bridge, authority, and setup UI, with six Linux command
cases skipped on macOS. Real cross-task documents pass the route
validator and persist two revisions. macOS refuses command execution and
does not advertise the tool. Full workspace typecheck and build, changed
server/UI typechecks, and token gates pass again on the final commit.
Greptile rates final head 00ca2c75b 5/5 and its completed check
concludes success; all review threads are resolved. All 58 current-head
checks are complete: 54 passed and 4 intentionally skipped. This
includes full tests, typecheck, build, Runner, browser E2E, release
verification, and Canary Dry Run. The older full local root attempt
finished with 16,602 passes, 93 skips, and ten failures across two
suites: it began before source edits and retained earlier imported
implementations while loading later tests. Both suites pass a clean
final-head rerun (25 passed, six Linux-only command cases skipped on
macOS). This mixed-source full attempt is not a final-head full-suite
pass; use the fresh CI evidence below.
- Full workspace typecheck and build pass for the standalone Dot change
on head `392d54f80`. UI token gates are clean. The full workspace
typecheck passed before the final importer UI edit; the changed UI
typecheck, build, and token gates pass again on the final commit. The
server build passes for its unchanged final source.
- Standalone selection, creation, settings, harness visibility,
inventory, and runtime admission checks pass. OAuth onboarding and the
real Rust/PostgreSQL broker suite pass 27 cases with the general Runner
option disabled. Dot and MCP revocation still block pairing; other
Runner providers remain disabled. Create, hire, conversion, and
inherited-hire route regressions pass 68 cases. The runtime selection
suite passes 20 cases. The setup UI suite passes 50 cases. Company
import and dev binary checks pass 97 cases. Import UI checks pass 29
cases, including Dot-only selection, preservation of imported Dot
agents, generic Runner fallback, canonical configuration serialization,
and required billing acknowledgement. A read of the synthetic instance
reports the native binary required with Dot enabled, the general rollout
disabled, and no persisted native work.
- The latest attachment, operator-consent, generic API file-guard, and
bridge checks pass 80 tests. Cache and native lifecycle checks pass 550
tests; create/import builders pass 45 tests; attachment setup UI checks
pass 16 tests. The real Rust/PostgreSQL Dot broker suite passes 15
cases.
- Earlier authority, human assignment, pinned skill, confinement,
cancellation, inheritance, onboarding, replay, OpenAPI, and gateway
regressions pass their focused rechecks. Workspace writes reject
concurrent stale hashes.
- In the live empty test-drive, turn the general Runner option off and
leave Dot and MCP on. Add agent shows a separate OpenAI Dot choice.
Create rejects missing billing acknowledgement, saves the canonical Dot
configuration, and opens pairing. The existing paired Dot remains ready
and passes its prerequisite checks. No new plugin pairing was needed for
this UI change. The final browser import preview keeps a Dot source
agent as OpenAI Dot, falls back a generic Runner agent while its switch
is off, and offers Dot independently.
- Real Dot completed OAuth, signed MCP Events readiness, and an
event-only assigned document task in an empty synthetic instance.
Pairing saved its binding without a second configuration save.
- Real Dot verified human assignment, cross-task comments and documents,
pinned skill reading, sandbox commands, lease renewal, follow-up input,
and a downloaded artifact whose bytes and hash matched its receipt. An
idle conversation request created an intake, created a task for its
human owner, continued after a definite missing-file read error, and
finalized Done with exit code 0.
- After operator approval, real Dot read a synthetic assigned-task
attachment and wrote its file-only random proof into an agent-authored
document. Its first test required accepting review because the existing
document tool removed the final newline.
- On head `2626c8f9b`, real Dot read a 13,849-byte synthetic attachment
in two pages, used `expectedSha256` on page two, saved exactly its final
random marker, verified readback, and finalized Done with exit code 0. A
direct inbox check was required for this attachment qualification; it
does not claim event-only delivery.
- Previous head `392d54f80` passes all 55 checks (53 passed, two
intentionally skipped), including the full test, typecheck, build,
native Runner, browser, and release verification gates. Greptile rates
this head 5/5; all review threads are resolved.
- Head `2626c8f9b` passed all 56 checks: 54 passed and two intentionally
skipped. This includes full typecheck, build, native Runner, server
tests, serialized suites, browser E2E, and Canary Dry Run. The unchanged
Cursor managed-runtime test exceeded its five-second deadline on the
first attempt; its focused local suite passed all eight cases, and the
single CI rerun plus dependent verify gate passed.
- A prior full local serial test attempt reported 19 failures (16,141
passed, 88 skipped) and stopped before later wrapper groups. It began
before the final source edits. Every failed suite has a passing fresh
recheck; the macOS snapshot stress case passes alone in 227 seconds. The
previous head `000fb9261` passed all 56 CI checks. PRs leave
`pnpm-lock.yaml` unchanged; CI and the refresh bot own dependency
resolution.

## Risks

- Dot remains off by default. Its own opt-in does not enable other
Runner providers. It still requires the MCP option; disabling Dot blocks
new work while keeping existing recovery and saved bindings.
- Dot requires a stable public HTTPS origin. Its base provider in #15402
is merged. Temporary tunnels are useful for testing but are not
permanent deployments.
- Idle intake creates a visible agent-authored task. It does not
fabricate a human message or bypass ordinary admission.
- Workspace commands require Linux bubblewrap with a private PID
namespace; deployment qualification is still needed. macOS file tools
and artifact publishing remain available, but commands are disabled
because sandbox-exec does not contain detached descendants. There is no
unrestricted fallback. Command protection fails closed above 4,096
workspace directories. The historical macOS command walkthrough does not
qualify the current Linux implementation.
- Provider model choice, token usage, cost, native thread control, and
global external stopping remain unavailable. Known Paperclip budget
gates still apply.
- The workspace bridge and attachment reader are separate opt-in
settings. Reading assigned task files sends their contents to OpenAI.
Revocation prevents future reads but cannot withdraw bytes already sent.
Files are read on request; automatic inbound attachment staging remains
disabled.
- An existing event registration can retain earlier instructions that
prohibit extra tasks. Dot asked for permission before a second intake
after the final event-only bootstrap; that extra no-nudge continuation
is not live-qualified under that earlier registration. The later
attachment qualification submitted normally through the composer. The
revised onboarding prompt describes the new idle entry point, and its
protocol polling path is tested.
- OpenAI event delivery can be delayed; a webhook acknowledgement is not
proof that Dot has begun work.
- A running Dot can retain an old plugin catalog after tool refresh.
Refresh and actual tool exposure must be checked before using new
top-level actions.
- Hosted and remote controller modes are not qualified.

## Model Used

OpenAI Codex, based on GPT-6. The exact deployment ID and context window
size are not exposed in this session. Capabilities used: reasoning,
repository editing, code execution, test inspection, and browser
control.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either linked existing issues or described the issue in-PR
following the relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused suites and all
fresh failure rechecks pass; the earlier full attempt is disclosed
above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-07 19:44:17 -05:00
DottaandPaperclip fc6304dfe5 feat(runner): add experimental OpenAI Dot provider over MCP Events (#15402)
## Thinking Path

> - Paperclip manages AI agents, tasks, permissions, and execution
budgets.
> - Paperclip Runner gives each provider the same admitted task and tool
authority.
> - OpenAI Dot runs outside the local process tree and needs
asynchronous work delivery.
> - The merged MCP gateway supplies OAuth consent and signed event
delivery.
> - A personal assistant grant cannot safely stand in for an assigned
agent.
> - This pull request adds a separate Dot agent connection and a durable
Rust Runner bridge.
> - The operator can assign work to Dot and inspect its accepted work,
tool receipts, and result.

## Linked Issues or Issue Description

**Agent or provider**

OpenAI Dot, as an experimental provider of the existing Paperclip Runner
adapter.

**Why this adapter is useful**

An operator can assign normal Paperclip tasks to an existing Dot. Dot
can read its mailbox, request work on an assigned task, use admitted
task tools, and submit a result. Paperclip keeps company scope,
checkout, approvals, known budget limits, and activity attribution.

**How the agent is invoked**

A dedicated `/mcp/runner` OAuth resource pairs one Dot grant with one
agent. A signed MCP mailbox event wakes Dot. Dot explicitly accepts the
assignment. The Rust Runner owns the durable turn and operation
receipts. The first release supports self-hosted instances with a local
Runner controller.

**Additional context**

This extends the merged public MCP gateway from #14846 and the assistant
invitation and device-consent work from #14933. This also integrates the
merged assistant tool and configuration expansion in #15380. Dot retains
its dedicated agent resource and cannot receive personal configuration
permission. The public assistant connection remains a personal
connection.

## What Changed

- Add a durable Rust Dot provider and its TypeScript Runner driver.
- Add closed PRP v3 external-provider operations and native execution
input v6.
- Add company-scoped pairing, mailbox, assignment, and operation
records.
- Reuse merged browser/device consent, client metadata verification,
webhook admissions, refresh, secret rotation, and warm-standby gates.
- Keep Dot scopes, issuer, grants, event workers, and tool access
separate from personal assistant access.
- Add Dot configuration, pairing, readiness, and consent UI. Keep agent
grants out of the personal Connections entry.
- Regenerate the Dot-only migration after master. Preserve published
gateway migrations. Make the new migration safe to reapply.
- Document setup, recovery, accounting limits, evidence, and remaining
account qualification.
- Reverify reconnect callbacks and wake outstanding work with a fresh
mailbox reference; preserve the existing assignment and operation
receipts.
- Clean up Dot bindings and waiting runs on OAuth revoke and
refresh-token replay. Old grants cannot revoke replacement bindings.
- Restore the pairing reference when an unsaved agent form is reopened;
document board-only pairing routes in OpenAPI.
- Accept a clean Rust exit after the acknowledged shutdown receipt.
Unexpected exits still require recovery.
- Clear the cached binding after a successful revoke so a failed
connection refresh cannot restore it.
- Add production-component Storybook states and screenshots for pairing
and connection review. All preview account data is synthetic.
- Persist normalized completion, serialize Dot turns and durable work
admission, and poll subscription readiness.
- Serialize mailbox writes and cursor reads; retain paused fence
acknowledgement without task authority.
- Authorize admitted review runs without changing the worker assignee.
Include the fenced assignment ID in production stop notices.

## Verification

- This PR integrates master `4a8178e9c`. Dot migration
`0317_messy_famine.sql` follows the published history and is safe to
reapply. The merge preserves the reserved migration connection,
batch-commit handling, private task checks, task monitors, and native
accounting.
- Local workspace typecheck, full build, and UI token gates pass. The
server typecheck passes after the review fixes. Database and native
executor regressions pass.
- All twelve real Rust/PostgreSQL Dot integration tests and twelve Dot
driver tests pass. The tests cover native document writing and
finalization, durable replay, queue admission, mailbox ordering,
admitted reviews, stale authority, production stop references, and
paused acknowledgements.
- Current head `d0e7e0626` passes all 57 checks: 53 pass and four are
intentionally skipped. This includes full typecheck, build, tests, Rust
Runner verification, browser E2E, release verification, and Canary Dry
Run. Greptile rates this exact head 5/5. All review threads are
resolved.
- The full local root test run is slower than the sharded CI run and has
not completed. The full CI test gates pass on the current commit.
Focused local regressions pass.
- Real-account pairing and event delivery on this base commit remain
unqualified. Live account and setup proof are recorded in the follow-up
#15414.

The following screenshots use synthetic preview data. They show the
production pairing component and do not qualify a real account or the
full agent setup journey.

![Synthetic pairing
preview](https://github.com/paperclipai/paperclip/blob/codex/dot-events-prototype/doc/screenshots/openai-dot-runner/pairing.jpg?raw=true)

![Synthetic connected
preview](https://github.com/paperclipai/paperclip/blob/codex/dot-events-prototype/doc/screenshots/openai-dot-runner/connected.jpg?raw=true)

## Risks

- This base adapter uses `PAPERCLIP_ENABLE_OPENAI_DOT=1` plus Public MCP
and Paperclip Runner. The separate experimental-settings follow-up in
#15414 replaces this environment flag with saved operator settings.
- Dot does not expose provider token usage or cost. The operator must
acknowledge external billing. Known Paperclip budget gates still apply.
- Cancellation fences Paperclip authority. It does not confirm that Dot
stopped all external activity.
- Assigned skill files and third-party MCP bindings are unsupported and
reject admission. There is no mounted workspace, model selector, or
provider thread identifier.
- Hosted agent-broker and remote controller deployments are not
qualified.
- The new migration follows the merged master history. Existing
prototype databases still need the normal master migration history
before this Dot-only migration.

## Model Used

OpenAI Codex, based on GPT-6. The exact deployment ID and context window
size are not exposed in this session. Capabilities used: reasoning,
repository editing, code execution, and test inspection.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #123` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub references)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-07 19:11:50 -05:00
DottaandPaperclip a7a244ab33 feat: add company decision models with permission and cost controls (#15473)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Optional product features need small, typed model decisions.
> - Each company needs to choose which shared API connection pays for
those decisions.
> - Calls must retain user and task permissions, budget limits, and cost
attribution.
> - This pull request adds a managed decision service, setup UI, and
request history.
> - Features can check availability cheaply and keep their existing
behavior when decisions are unavailable.

## Linked Issues or Issue Description

**Subsystem affected**

Server services, shared contracts, database accounting, Company
Settings, and Costs.

**Problem or motivation**

Paperclip has no common decision-model service. Adding provider calls
within each feature would duplicate credential access, permission
checks, and billing rules.

**Proposed solution**

Let a connection manager configure one company decision model. Support
OpenAI Decisions and Jev through OpenRouter. Provide a fixed setup test
and metadata-only history. Default company-sponsored background
decisions to on during setup, and preserve a saved off setting.

**Alternatives considered**

Per-feature credentials would duplicate existing connection management.
Personal overrides and provider fallback chains add permission and
billing complexity; they remain deferred.

**Roadmap alignment**

Reviewed ROADMAP.md and searched open PRs. This extends existing
connection access and budget accounting. Product features that call the
service remain outside this change. No matching decision-model service
PR was found.

## What Changed

- Add company settings, an internal `decisionModelService`, local
availability checks, and trusted human, agent/run, and system contexts.
- Pin Vercel AI SDK provider dependencies and adapt boolean, choice, and
ordered-score decisions for both providers. Bound requests and time;
disable paid retries.
- Add durable invocation metadata and agentless decision ledger charges.
Preserve fractional cents, pricing evidence, dispatch identity, and
unresolved billing holds.
- Share the company accounting lock and apply company, agent, and
project budgets. Settle charges once, retain unknown holds, and recover
interrupted calls without resubmission.
- Reuse connection setup and management UI. Add a Decisions view under
Costs, production-component Storybook coverage, database migration, and
service documentation.

## Verification

- Passed 166 current-code tests covering the decision service/provider,
setup component, Costs, OpenAPI, and every failure from the earlier
broad run. Coverage includes native SDK wire formats, refusals, billed
malformed responses, permission and secret-rotation races, identity
changes, concurrent budget admission, unresolved holds, agent/task
deletion, and stale setup feedback.
- Passed 170 existing connection, cost, budget, heartbeat-accounting,
and profile regression tests.
- Passed repository typecheck, production build, and design token gates
after integrating master. Verified the generated migration on a fresh
test database and upgraded the populated preview database from the
branch's earlier migration without losing settings or usage.
- Ran the required full `pnpm test:run`: its general phase completed
with 16,354 passed and 10 failures across five files while this branch
was still being updated. Every reported failure passes in the
current-code rerun; the serialized phase did not run after that failure.
The full GitHub CI suite passed on `b1b856a88`: general and serialized
tests, browser shards, runner checks, typecheck, production build,
packaging/canary, and policy gates. [CI
evidence](https://github.com/paperclipai/paperclip/actions/runs/37673106360).
- Passed the full-shell Storybook setup-to-history interaction test
again after integrating master. Greptile rates the final revision 5/5
with zero unresolved threads.
- Walked through the running app: empty setup, add each provider, save,
reload, run all three sample questions, inspect fractional charges in
history, switch provider while sponsorship is off, disable, and
reconnect. Checked mobile settings. These tests used the actual UI,
vault, server, SDKs, and database with simulated upstream responses.
- Live paid setup tests remain unverified: this environment has no
authorized OpenAI/OpenRouter credentials available. No mocked test is
presented as live provider evidence.

Reviewer journey: Company Settings → General → Decision model.
Add/select a shared API connection, save, run the billed sample, open
View usage, then disable decisions and verify Run test is disabled after
reload.

## Risks

- The SDK decision interface is experimental. Pinned versions and
wire-format tests limit upgrade drift.
- The migration allows agentless service charges and reservations.
Existing agent cost-reporting APIs still require an agent, and decision
receipts stay separate from run reconciliation.
- Timeouts can have unknown provider charges. Holds remain until an
audited accounting correction resolves them.
- OpenAI prices use a versioned Decisions rate snapshot; OpenRouter
costs use provider receipts. Unknown pricing is retained as unknown.
- A configured company authorizes background spending by default. Setup
explains this, and managers can turn it off.
- Live provider account/model availability still needs the two
credentialed acceptance checks.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository inspection, code
execution, and browser tools. The exact serving revision and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-07 15:53:13 -05:00
Devin FoleyandPaperclip f1c44b7b56 Preserve remote work when required workspace restore fails (#15479)
Persist exact source-retention obligations before run finalization and protect them across cancellation, restart and task changes. Require board-authorized repair evidence without replaying old work or changing current task ownership, state or locks. Hide an unavailable retry action and document operator recovery and retention costs.

Validated with 674 scoped regressions, 224 combined integration tests, full typecheck/build, independent safety reviews and green CI with Greptile 5/5. No historical file recovery is claimed.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-07 13:29:44 -07:00
50b1f95e79 fix(tool-gateway): keep MCP connection healthy on oversized/malformed responses (#15462)
<!-- Write all pull request text in Simplified Technical English
(ASD-STE100): short sentences, one instruction per sentence, simple
approved vocabulary, and the active voice. -->

## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Agents call remote MCP tools through the Paperclip tool gateway
> - The gateway keeps a health status for each remote MCP connection,
and it hides the tools of a connection that has `healthStatus = "error"`
> - One oversized, malformed, or invalid-JSON `tools/call` reply set the
whole connection to `error`
> - Health goes back to `ok` only after a successful call, but the
hidden tools prevent that call, so the connection stayed locked until a
person reconnected it
> - This pull request makes these reply errors fail only the one call,
and it starts a new MCP session for the caller on the next call
> - The benefit is that one large or bad reply no longer disconnects a
working connection for every agent

## Linked Issues or Issue Description

No public issue exists. Related open PRs fix other health downgrades in
the same function. They do not overlap with this change:

- Refs #11910 (timeouts and JSON-RPC errors)
- Refs #15325 (transport errors)

**What happened?**

An agent called a remote MCP tool that returned more than
`MAX_REMOTE_MCP_RESPONSE_BYTES` (1 MB). The gateway returned 502
`mcp_remote_response_too_large`. It also set the connection to
`healthStatus = "error"`. Then `connectedMcpConnectionFilter` hid all
tools of the connection (except `per_user` connections), and
`connections_search` showed the connection as `needs_user_action`. The
`malformed_response` and `invalid_json` errors from
`readMcpHttpResponse` did the same thing.

`readMcpHttpResponse` can also cancel the reply stream before the end.
The remote server can then close its MCP session. The gateway kept the
cached `mcp-session-id` for up to 30 minutes and cleared it only after a
404. The next calls then failed with "Remote MCP session expired".

**Expected behavior**

A reply that is too large or not correct fails only that call. The
connection stays healthy, and its tools stay visible. The next call uses
a new MCP session.

**Steps to reproduce**

1. Add a remote MCP connection with `mcpSessionRequired: true`.
2. Call a tool that returns a reply larger than 1 MB.
3. Look at the connection: `healthStatus` is `error`.
4. Start a new gateway session: the tools of the connection are not in
the list.

**Paperclip version or commit**

`master` at `99a9de9940bf5974352d9dbfbb2f21e62e89689f`

**Deployment mode**

All modes. The fault is in the server tool gateway.

## What Changed

- `server/src/services/tool-gateway.ts`: `too_large`,
`malformed_response`, and `invalid_json` from `readMcpHttpResponse` now
fail only the call. The caller gets the same 502 reason code as before.
The gateway does not call `markRemoteConnectionHealth(…, "error")` for
these errors. The `invalid_json` branch after `JSON.parse(body)` also
does not change health now, so all `invalid_json` paths are the same.
- The `mcp_remote_response_too_large` message now tells the caller to
request a smaller result, for example a narrower query or a smaller page
size. The error details now include `maxBytes`.
- `server/src/services/mcp-http.ts`: new `forgetMcpHttpSession()`. It
removes only the cached session for one scope and one credential set,
and only while that entry still holds the session ID that failed. It
does not touch other agents, sessions that are initializing, or a newer
session that replaced the failed one.
- The gateway calls `forgetMcpHttpSession()` after these reply errors,
so the next call from that caller initializes a new session. The
existing 404 "session expired" path now uses the same function. Before,
both paths cleared all sessions and all pending initializations on the
connection. That made a concurrent initialization by another agent fail
with "MCP connection changed while initializing", which the gateway
reported as a fetch failure and marked as a connection `error`.
- `forgetMcpHttpSessions(connectionId)` is not changed. Disconnect and
revocation still clear the full connection.

Why `malformed_response` and `invalid_json` are also per-call errors:

- Each error is about one reply body. The server was reachable, accepted
the credentials, and sent HTTP 2xx. A reply can be too large or bad
because of the tool and its arguments. That is not a fault of the
connection.
- The other malformed-reply checks in the same function (payload is not
an object, or has no `result`) already throw
`remote_mcp_malformed_response` and do not change health. Only the
reader-level errors changed health.
- Health recovers only after a successful call. An `error` status for
one bad reply therefore locks the connection until a person reconnects
it.
- Other failures (HTTP errors, fetch failures, timeouts, JSON-RPC
errors) still change health. This PR does not change them. #11910 and
#15325 address some of them.

## Verification

- `server/src/__tests__/tool-gateway.test.ts`: the recovery test now
runs for three replies: oversized, invalid JSON, and a reply without the
requested message ID. Each run uses a fake HTTP MCP server with
`mcpSessionRequired: true` that gives `session-N` for each `initialize`.
Each run checks that:
- the call fails with the correct 502 reason code (and, for the
oversized reply, the smaller-page hint and `maxBytes: 1000000`)
  - `healthStatus` stays `ok`
  - a new gateway session still lists the tool and can call it
  - the second `tools/call` uses `session-2`, not `session-1`
- New test: agent A gets an oversized reply while agent B initializes a
session on the same connection. Agent B's call completes, health stays
`ok`, and the two calls use `session-1` and `session-2`. With the old
connection-wide reset, this test fails: agent B gets 502
`mcp_remote_fetch_failed`.
- `server/src/__tests__/remote-mcp-protocol.test.ts`: new unit test.
`forgetMcpHttpSession()` keeps the session of a different identity, and
a late failure from an old session does not remove the newer session.
- Each regression check fails without its fix. Without the health
change, the test fails on `healthStatus: 'error'`. Without the session
reset, it fails with `['session-1', 'session-1']`.

```
cd server
npx vitest run src/__tests__/tool-gateway.test.ts src/__tests__/tool-gateway-service.test.ts src/__tests__/remote-mcp-protocol.test.ts
 Test Files  3 passed (3)
      Tests  132 passed (132)
```

- I did not run the full server `tsc --noEmit` locally because the
sandbox does not have sufficient memory. The CI typecheck covers it.

## Risks

- Low risk. The change affects only the error path of remote MCP
`tools/call`.
- A remote server that always sends bad replies now keeps `healthStatus
= "ok"`. Each call still fails with a clear 502 reason code and an audit
record, so the failure stays visible. Only the connection-wide hiding of
tools stops.
- After one of these errors, the next call from the same caller sends
one more `initialize` request. This adds one round trip.
- A 404 "session expired" now clears only the session of the caller that
got the 404. Before, it cleared the cached sessions of all agents on the
connection. If the remote server restarts, each agent now gets its own
"session expired" error one time and then initializes again. A 404 is
about one session, and the old connection-wide clear also cancelled
other agents' initializations.
- #11910 and #15325 change the same catch block. The PR that merges last
can have a small merge conflict.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- Claude Opus 5.5 (Anthropic), model ID `claude-opus-5-5`, 1M-token
context window.
- Run as an agent in Claude Code with tool use (shell, file edit, and
local test runs).

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes (no
user-facing documentation changes are necessary)
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

This PR replaces #15418. It keeps the same commits on a branch name
without an internal ticket ID, and adds a fix for the Greptile review on
#15418.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: devinfoley <139239+devinfoley@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 12:48:01 -07:00
DottaandPaperclip fd8c6b920a fix(native): resume connection tasks after approval decisions (#15471)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Native tasks can pause while a human decides whether to allow a
connection action.
> - The next turn needs both the saved decision and complete accounting
for the previous turn.
> - Cancellation could discard final usage, and complete direct Claude
API receipts could remain unpriced.
> - A stale blocked or review report could also request approval again
after the original card was declined.
> - This pull request retains shutdown accounting and rejects approval
waits bound to an already resolved action.
> - The benefit is reliable continuation with the existing budget and
approval controls.

## Linked Issues or Issue Description

Refs #15420. Related: #15312 addresses requester ownership during
dispatch. This change addresses receipt capture and final-response
validation.

## What Changed

- Retain usage events after native cancellation. Continue to reject late
provider messages and work.
- Drain same-turn accounting and terminal events for at most fifteen
seconds after a durable governed wait. Keep incomplete accounting
blocked.
- Estimate complete, unpriced, direct Anthropic API receipts for the
exact `claude-sonnet-5` model. Record the rate version and assumptions.
Use the one-hour cache-write rate when the receipt lacks cache TTL.
- Bind stale approval reports to exact interaction, action-request, or
invocation IDs in the same company, task, agent, and run. Cover blocked,
review, and response-wake reports. Keep independent reviews valid.
- Fence checkpoint and result writes after a controller detaches for
restart, including operations waiting for a database lock. Reject stale
successful returns before certifying accounting.
- Allow bounded subscription teardown only after retaining an actual
provider terminal.
- Journal the exact governed-wait trigger and disposition before
provider interruption. Recover that wait independently of a later saved
answer, replay retained accounting, and reject mismatched or unproven
terminal evidence.
- Give settling governed turns a bounded window before shutdown detaches
their controller.
- Keep fuzzy external app matches alongside installed capability matches
instead of forcing an unrelated provider question for a generic query.
- Return up to twenty exact active catalog tool names after an invalid
request, after eligibility checks; still reject the request without
granting access or creating an approval.
- Require retained provider terminal proof before settling a governed
wait, including when complete usage arrives before stream
closure/error/timeout. Retain harmless numbered cancellation events so
restart replay stays contiguous.
- Isolate accounting-test OpenCode config from the host plugin
directory.
- Add regression coverage and document the accounting, restart and
connection-search behavior.

## Verification

Current PR source: `0cf08efd75f8fb23f7989beda6dbda92587088bf`. The live
matrix below measured frozen `09a776bcb1f77422156f8be11e13f8c29f43e7f7`;
later review fixes are verified separately and do not relabel those
runs.

- Repository typecheck and build pass.
- Current runner runtime and cancellation suites: 191 tests pass. Six
new regressions cover stream end/error/timeout without provider-stop
proof and contiguous cancellation acknowledgement/request replay; all
six failed before the fix. The existing bounded cleanup case now
explicitly supplies terminal proof. All 31 adapter accounting tests pass
with isolated fixture config.
- New checkpoint-rebinding and approval-criterion suites: 65 tests pass;
runner HTTP integration: 29 tests pass. Unchanged executor/control-plane
suites: 640 tests pass; database-backed connection suites: 68 tests
pass.
- Current retained evidence verification covers 222 file hashes across
all fifteen original result artifacts. The three-case [published
report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37651896525-1/index.html)
and all eight screenshot hashes verify. The [three-case recovery
report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37662587356-1/index.html)
and all six screenshots also verify; the nine-result campaign did not
publish.
- Evaluation support: 1,805 Vitest tests pass, one skipped; 128 Node
tests pass. Eval typecheck and catalog discovery pass.
- Full local repository unit run was interrupted before the follow-up
edits after three tool-access failures and one runner HTTP failure.
Those failures pass in isolation; the 09a full run was interrupted after
one rapid Slack callback-ordering failure and seven skill-service
failures. All eight pass both isolated and with full-runner environment
settings, and all 74 skill-service tests pass together; the subsequent
full run reported two 15-second OpenCode accounting timeouts and was
stopped with exit 130 to apply review fixes. The timeouts reproduce
while copying this host’s 61 MB OpenCode config. All 31 tests pass after
isolating config inside each fixture without increasing timeouts or
changing assertions. A complete local full-suite pass is not claimed.
Repository-wide CI also passes on the final review-fix head in [run
37669185952](https://github.com/paperclipai/paperclip/actions/runs/37669185952).
Earlier database-skipped diagnostics and the older ENFILE run are
retained and are not full-suite passing evidence.
- Three-case live campaign
[37651896525](https://github.com/paperclipai/paperclip/actions/runs/37651896525)
passes all three original grades on frozen source 09a: 55/55 checks,
eight succeeded run records, complete accounting receipts, matching
checkpoint identities and no pending approvals. Campaign
[37653533353](https://github.com/paperclipai/paperclip/actions/runs/37653533353)
adds nine original passes (173/173 checks, eighteen succeeded records)
on the identical source. Its other three jobs failed before runner
assignment or any step while GitHub could not load the paid environment;
those original infrastructure failures are retained. Campaign
[37662587356](https://github.com/paperclipai/paperclip/actions/runs/37662587356)
completes only those unstarted cells: all three original grades pass
(57/57 checks, six succeeded run records). All fifteen exact cases now
pass on source 09a: eight FAIL → PASS, seven PASS → PASS, zero new
overall failures and zero pending pairs. Total current evidence: 285/285
checks and thirty-two succeeded run records, complete accounting
receipts, matching checkpoint identities, no pending approvals or retry
records. Earlier campaigns retain forty-seven additional run records and
two known same-run recovery attempts; actual provider-call counts and
invoices remain unknown. The nine-result campaign skipped publication
and its public URL returns 403; original artifacts remain retained.
Previous ba9 campaign
[37645656840](https://github.com/paperclipai/paperclip/actions/runs/37645656840)
completed 2 PASS / 1 FAIL: Claude restart/approval and Codex decline
pass, while OpenCode resumes but times out searching for exact tool
names and never creates the access card. That failure and incomplete
cancelled-run accounting remain preserved; the new catalog error
guidance targets this observed dead end. Campaign
[37642957312](https://github.com/paperclipai/paperclip/actions/runs/37642957312)
remains 0 PASS / 3 FAIL and exposed the now-corrected cross-run marker
leak and unknown-criterion approval gap. The original baseline remains 7
PASS / 8 FAIL, first repair 2 PASS / 3 FAIL, and second repair 0 PASS /
3 FAIL. All fifteen selected cases are qualified by their original
grades in this bounded trial. These live grades belong to 09a. Its
twelve governed-wait checkpoints retain matching same-turn terminal
fingerprints, but passing artifacts omit detailed event journals; the
later six adversarial regressions qualify the new terminal-proof and
replay guards separately. Final-head repository CI passes. [Fresh
Greptile
review](https://github.com/paperclipai/paperclip/pull/15471#issuecomment-6044369780)
is 5/5, confirms both findings are fixed, and reports no new actionable
issues. All review threads are resolved and the PR has no merge
conflicts.

## Risks

- Governed cancellation drains accounting for up to fifteen seconds.
Restart detachment gives a settling batch up to twenty seconds to
finish. An incomplete receipt or unproven provider terminal still
prevents successful qualification.
- Claude prices are estimates, not invoices. The estimate assumes
standard global API pricing and uses a conservative cache-write rate.
Unsupported models, billers, and billing modes remain unpriced.
- Approval identity matching must remain scoped to the current run and
the requested approval. It does not authorize execution of a declined
call.
- Original baseline and final candidate have different merged master
context. Exact-case outcomes are before/after observations, not isolated
causal attribution to this repair.
- No schema migration, fixture, oracle or grader change. Search-result
guidance now treats fuzzy external matches as suggestions. Existing app
authorization and provider-consent checks remain required.

## Model Used

- OpenAI GPT-6 through Codex, with code editing, terminal tools, and
test execution. The exact deployment identifier and context window are
not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (191 current runtime tests
and 31 accounting tests; interrupted full-suite history and CI coverage
are disclosed above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-07 14:43:07 -05:00
Devin FoleyandPaperclip 5405b1f46f Synchronize chat delivery tests with completed drains (#15480)
Expose the already tracked and caught deferred-work promise to test schedulers while preserving default setImmediate behavior. Assert Slack delivery ordering while the primary lease is held, then join the drains; update the existing GitHub test to assert the resulting quiescent queue.

Verified all 1,063 chat tests, full typecheck/build, negative synchronization control, independent review and exact-head CI. An unchanged browser shard retry passed and is documented. Greptile 5/5 with no unresolved comments.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-07 12:38:12 -07:00
Devin FoleyandPaperclip 465140596f Preserve bounded workspace sync diagnostics across RPC (#15481)
Preserve bounded error codes and HTTP/exit statuses across the environmentSyncOut worker RPC boundary, and revalidate that method-scoped envelope before attaching host restore diagnostics. Keep the original error and all recovery policy unchanged; do not transmit provider payloads or credentials.

Verified real RPC roundtrip and privacy regressions, 95 focused tests, 68 independent tests, full typecheck/build and all exact-head CI. Greptile 5/5 with no unresolved comments.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-07 12:22:53 -07:00
Devin FoleyandPaperclip 128c95837b Retain bounded orphan process-loss diagnostics (#15475)
Record bounded observer uptime, run and output ages, process-check observations and retry eligibility before orphan cleanup changes the evidence. Preserve existing recovery and reporting behavior and omit process IDs, raw paths and credentials.

Verified focused helper, actual reaper and real SDK regressions, full typecheck/build, exact-head CI and independent review. An unchanged Cursor timeout passed its isolated retry; unrelated local Slack timing failure is documented separately. Greptile 5/5 with no unresolved comments.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-07 12:05:21 -07:00