Commit Graph
4905 Commits
Author SHA1 Message Date
DottaandPaperclip 2922f65ba1 fix(hermes): pin v7 permission policy in the independent server launch
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 02:52:40 -05:00
DottaandPaperclip b0f54fc1a2 fix(runner): preserve Dot v6 while versioning Hermes input as v7
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 02:35:34 -05:00
DottaandPaperclip 9d33a4d8a4 fix(hermes): publish and enforce native answer limits
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 02:35:34 -05:00
DottaandPaperclip dc76966146 fix(runner): bound encoded Hermes attachment messages
Count serialized message, attachments, and metadata before turn admission. Keep TypeScript and Rust on a 7 MiB encoded content allowance within the existing encrypted frame bound. Validate before active-turn state changes, and prove accepted images and escaped documents traverse the secure PRP command frame.

Validation: 28 attachment/permission tests, two encrypted-frame regressions, the Rust encoded-admission regression, and Runner TypeScript compilation passed.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 02:35:34 -05:00
DottaandPaperclip ef7cf8c398 fix(hermes): preserve planning authority and native transcript results
Qualification exposed denied plan workflows, missing structured tool output, and cancellation replay in subsequent turns. Admit assigned control-plane operations through their existing semantic authority, project denied native calls by identity, and keep managed history restoration separate from transcript replay.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 02:35:34 -05:00
DottaandPaperclip 97dc4efd53 docs(runner): pass resolved lock digest in image commands
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 02:35:34 -05:00
DottaandPaperclip 2f15770a88 fix(hermes): require the resolved image lock digest
Require the caller-verified target lock checksum instead of an incompatible checked-in default. Document lock-bot ownership and the companion browser qualification PR.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 02:35:34 -05:00
DottaandPaperclip 05d6a8f6e2 fix(acpx): separate restored history from active turn events
Keep history reconstruction without publishing restored text and tools as live output. Version the changed Cursor contract as profile 16 and retain replay compatibility.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 02:35:34 -05:00
DottaandPaperclip 0c75f68414 fix(hermes): keep native candidate independently restorable
Reject replaced learned-state parents, include authenticated run attachment and the advertised native transport fixtures in this candidate. Keep paid browser campaign definitions in the qualification PR.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 02:35:34 -05:00
Dotta fdb94c95ef fix(hermes): release credentials and reproduce relocatable runtime assets 2026-10-08 02:35:34 -05:00
DottaandPaperclip 9d1512a169 fix(runner): include Hermes in Daytona image identity
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 02:35:34 -05:00
DottaandPaperclip ba9a809ea2 feat(runner): add native Hermes ACP qualification candidate
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 02:35:34 -05:00
DottaandPaperclip 424d9fced8 test(routines): retain the current master authority inventory
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 02:35:21 -05:00
DottaandPaperclip c7afea8dab fix(routines): reject lock contention before receipt mutations
Keep owned-routine lock acquisition nonblocking after run authority locks. A scheduled firing can finish, and the same mutation may retry without a saved partial receipt. The real concurrency regression is added; local PostgreSQL bootstrap failed before assertions, so its execution still requires CI. Server typecheck passes.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 02:28:27 -05:00
Dotta 79e5d38139 fix(routines): remap annotations and verify atomic schedule edits 2026-10-08 02:27:42 -05:00
DottaandPaperclip 49fdc5899e feat(runner): bind routine management to the live service
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 02:27:42 -05:00
Devin FoleyandPaperclip 5717523b9e fix: retain the original repository for reused task workspaces (#15528)
Apply the reviewed change for fix: retain the original repository for reused task workspaces.

Validation: required local typecheck and tests, passing CI, Greptile 5/5, and independent review of the exact source head.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
canary/v2026.1008.0-canary.4
2026-10-07 19:52:30 -07:00
Devin FoleyandPaperclip e4b39da6f6 Isolate ACPX admission deadline tests from lease ports (#15527)
Apply the reviewed change for Isolate ACPX admission deadline tests from lease ports.

Validation: required local typecheck and tests, passing CI, Greptile 5/5, and independent review of the exact source head.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-07 19:20:08 -07:00
Devin FoleyandPaperclip 1d19f9b562 Distinguish failed Git inspection from missing worktree registration (#15525)
Apply the reviewed change for Distinguish failed Git inspection from missing worktree registration.

Validation: required local typecheck and tests, passing CI, Greptile 5/5, and independent review of the exact source head.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-07 19:05:04 -07:00
Devin FoleyandPaperclip ed0c6958f8 Explain refused local Hermes gateway connections (#15520)
Apply the reviewed change for Explain refused local Hermes gateway connections.

Validation: required local typecheck and tests, passing CI, Greptile 5/5, and independent review of the exact source head.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-07 18:55:55 -07:00
Devin FoleyandPaperclip 6b171c9616 Keep proven local workspace configuration conflicts out of Sentry (#15521)
Apply the reviewed change for Keep proven local workspace configuration conflicts out of Sentry.

Validation: required local typecheck and tests, passing CI, Greptile 5/5, and independent review of the exact source head.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-07 18:54:35 -07:00
github-actions[bot]andlockfile-bot f4e6362ccc chore(lockfile): refresh pnpm-lock.yaml (#15512)
Auto-generated lockfile refresh after dependencies changed on master.
This PR only updates pnpm-lock.yaml.

Co-authored-by: lockfile-bot <lockfile-bot@users.noreply.github.com>
canary/v2026.1008.0-canary.2
2026-10-07 20:17:05 -05:00
DottaandPaperclip dd777f4b73 feat(dot): complete onboarding and expand governed Runner capabilities (#15414)
## Thinking Path

> - Paperclip manages AI agents, their work, and their permissions.
> - Paperclip Runner supplies the same admitted tool authority to each
provider.
> - The Dot provider in #15402 needs reliable onboarding and useful
agent capabilities.
> - An idle Dot could not start work, assign a human, read skills, or
produce a workspace artifact.
> - Pairing also relied on a second configuration save before ordinary
admission could work.
> - This pull request adds governed idle admission and shared Runner
tools, and completes pairing atomically.
> - Operators can test Dot with fresh data while keeping normal company,
approval, budget, and run ownership checks.

## Linked Issues or Issue Description

**Subsystem affected**

Paperclip Runner, dedicated Dot MCP access, OAuth onboarding,
experimental settings, and empty worktree startup. This PR builds on
merged provider PR #15402. It reuses the merged MCP gateway from #14846
and assistant connection work from #14933 and #15380.

**Problem or motivation**

An idle Dot could see assignments but could not act on a conversation
request until someone created a task first. Its Runner catalog could not
assign tasks to humans or read pinned skills and workspace files.
First-time OAuth discovery and pairing also needed browser fixes, and a
completed pairing did not persist its binding reference on the agent.

**Proposed solution**

Keep Dot within the existing Runner. Admit a visible agent-authored
intake task for idle requests. Add people, human assignment, cross-task,
skill, and optional sandbox workspace tools through shared authority.
Relay assigned app calls through the configured MCP gateway. Add lease
renewal and follow-up references. Save pairing and its configuration
revision atomically. Show prerequisites and provide a complete copy
prompt.

**Roadmap alignment**

This extends the experimental provider in #15402. It uses the existing
governed gateway, task model, skills, artifact path, and native Runner.
It adds no separate execution subsystem.

## What Changed

- Add a standalone OpenAI Dot agent choice with an independent
experimental opt-in. It works with the general Runner option off.
Require Assistant connections (MCP), authenticated sign-in, and public
HTTPS for pairing.
- Use Dot’s own option for company package import. Allow an unpaired Dot
configuration to save after external billing acknowledgement; task
admission still requires pairing. Prepare the shared dev binary for
Dot-only opt-in.
- Label saved Dot agents as OpenAI Dot. Use prerequisite-check copy in
setup and runtime configuration.
- Route Dot creation directly to pairing after explicit external billing
acknowledgement. Hide local CLI, model, and harness setup for Dot.
Preserve the shared Runner implementation and canonical API type. Other
Runner providers still require their general opt-in.
- Add an empty worktree option with fresh signing keys and no production
data copy.
- Fix public-client OAuth negotiation, discovery compatibility, and
optional separate browser authorization origin.
- Add one-use pairing consent preview, clear copied setup instructions,
and atomic binding persistence and cleanup.
- Add idle request admission with stable request IDs and normal
scheduling, permissions, budgets, and task ownership.
- Add identity and people discovery, human task assignment and
reassignment, and authorized cross-task comments and documents.
- Read assigned skill files from pinned manifests. Relay assigned app
calls through the merged gateway without exposing credentials.
- Add an off-by-default workspace bridge. Constrain paths and writes.
Run commands in a deny-by-default OS sandbox with no network or injected
credentials. Reserve mutations before effects and never blindly repeat
uncertain work.
- Add rolling lease renewal, task pagination, bounded operation limits,
and deduplicated follow-up references without comment bodies in
webhooks.
- Add an off-by-default attachment reading setting. Restrict reads to
files on the current assigned task. Verify size and hash, cache bounded
verified copies per run, paginate text or binary bytes, and recheck live
authority before returning.
- Keep file grants operator-owned. Reject agent self-grants across
configuration routes. Preserve attachment consent in create/import
forms. Close generic API file bypasses while retaining current-run
response snapshots and permitted uploads.
- Keep provider limits explicit. Do not inherit a Dot binding,
attachment permission, or workspace permission when hiring another
agent.
- Include the required Markdown format in cross-task document writes and
validate the API title limit. Verify real creation and revision
persistence.
- Restrict command execution to Linux bubblewrap with descendant
containment. macOS retains workspace file tools and artifact publishing,
while refusing command calls. Explain the platform limit in setup.
- Exclude Paperclip instance state from workspace files, uploads,
artifact publication, and sandbox commands. Protect nested directories
and case variants. Fail closed when the directory protection scan
exceeds 4,096 directories.
- Add static UI compression for slow public tunnels. Document setup, the
complete tool inventory, and qualification limits.

## Verification

- Merge preparation on `00ca2c75b` integrates merged base #15402 and
master `fc6304dfe`. The ancestry commit preserves the reviewed follow-up
source tree. The subsequent security fix excludes instance state from
file tools, uploads, artifact publication, and sandbox commands. It
preserves private task authorization and task monitors. It regenerates
the combined tool catalog, seeded catalog digest, and protocol manifest.
Workspace typecheck, full build, and UI token gates pass. All sixteen
real Dot broker cases and 93 company import cases pass after the review
fixes and creator-attribution test correction. The exported
human-assignment catalog and Unicode page boundaries are also fixed. All
ten catalog tests and nine workspace/skill bridge tests pass. All 35
workspace bridge and authority tests passed after the instance-state
fix, including real macOS commands in that intermediate version. The
subsequent document and descendant-containment fixes pass 44 focused
tests across bridge, authority, and setup UI, with six Linux command
cases skipped on macOS. Real cross-task documents pass the route
validator and persist two revisions. macOS refuses command execution and
does not advertise the tool. Full workspace typecheck and build, changed
server/UI typechecks, and token gates pass again on the final commit.
Greptile rates final head 00ca2c75b 5/5 and its completed check
concludes success; all review threads are resolved. All 58 current-head
checks are complete: 54 passed and 4 intentionally skipped. This
includes full tests, typecheck, build, Runner, browser E2E, release
verification, and Canary Dry Run. The older full local root attempt
finished with 16,602 passes, 93 skips, and ten failures across two
suites: it began before source edits and retained earlier imported
implementations while loading later tests. Both suites pass a clean
final-head rerun (25 passed, six Linux-only command cases skipped on
macOS). This mixed-source full attempt is not a final-head full-suite
pass; use the fresh CI evidence below.
- Full workspace typecheck and build pass for the standalone Dot change
on head `392d54f80`. UI token gates are clean. The full workspace
typecheck passed before the final importer UI edit; the changed UI
typecheck, build, and token gates pass again on the final commit. The
server build passes for its unchanged final source.
- Standalone selection, creation, settings, harness visibility,
inventory, and runtime admission checks pass. OAuth onboarding and the
real Rust/PostgreSQL broker suite pass 27 cases with the general Runner
option disabled. Dot and MCP revocation still block pairing; other
Runner providers remain disabled. Create, hire, conversion, and
inherited-hire route regressions pass 68 cases. The runtime selection
suite passes 20 cases. The setup UI suite passes 50 cases. Company
import and dev binary checks pass 97 cases. Import UI checks pass 29
cases, including Dot-only selection, preservation of imported Dot
agents, generic Runner fallback, canonical configuration serialization,
and required billing acknowledgement. A read of the synthetic instance
reports the native binary required with Dot enabled, the general rollout
disabled, and no persisted native work.
- The latest attachment, operator-consent, generic API file-guard, and
bridge checks pass 80 tests. Cache and native lifecycle checks pass 550
tests; create/import builders pass 45 tests; attachment setup UI checks
pass 16 tests. The real Rust/PostgreSQL Dot broker suite passes 15
cases.
- Earlier authority, human assignment, pinned skill, confinement,
cancellation, inheritance, onboarding, replay, OpenAPI, and gateway
regressions pass their focused rechecks. Workspace writes reject
concurrent stale hashes.
- In the live empty test-drive, turn the general Runner option off and
leave Dot and MCP on. Add agent shows a separate OpenAI Dot choice.
Create rejects missing billing acknowledgement, saves the canonical Dot
configuration, and opens pairing. The existing paired Dot remains ready
and passes its prerequisite checks. No new plugin pairing was needed for
this UI change. The final browser import preview keeps a Dot source
agent as OpenAI Dot, falls back a generic Runner agent while its switch
is off, and offers Dot independently.
- Real Dot completed OAuth, signed MCP Events readiness, and an
event-only assigned document task in an empty synthetic instance.
Pairing saved its binding without a second configuration save.
- Real Dot verified human assignment, cross-task comments and documents,
pinned skill reading, sandbox commands, lease renewal, follow-up input,
and a downloaded artifact whose bytes and hash matched its receipt. An
idle conversation request created an intake, created a task for its
human owner, continued after a definite missing-file read error, and
finalized Done with exit code 0.
- After operator approval, real Dot read a synthetic assigned-task
attachment and wrote its file-only random proof into an agent-authored
document. Its first test required accepting review because the existing
document tool removed the final newline.
- On head `2626c8f9b`, real Dot read a 13,849-byte synthetic attachment
in two pages, used `expectedSha256` on page two, saved exactly its final
random marker, verified readback, and finalized Done with exit code 0. A
direct inbox check was required for this attachment qualification; it
does not claim event-only delivery.
- Previous head `392d54f80` passes all 55 checks (53 passed, two
intentionally skipped), including the full test, typecheck, build,
native Runner, browser, and release verification gates. Greptile rates
this head 5/5; all review threads are resolved.
- Head `2626c8f9b` passed all 56 checks: 54 passed and two intentionally
skipped. This includes full typecheck, build, native Runner, server
tests, serialized suites, browser E2E, and Canary Dry Run. The unchanged
Cursor managed-runtime test exceeded its five-second deadline on the
first attempt; its focused local suite passed all eight cases, and the
single CI rerun plus dependent verify gate passed.
- A prior full local serial test attempt reported 19 failures (16,141
passed, 88 skipped) and stopped before later wrapper groups. It began
before the final source edits. Every failed suite has a passing fresh
recheck; the macOS snapshot stress case passes alone in 227 seconds. The
previous head `000fb9261` passed all 56 CI checks. PRs leave
`pnpm-lock.yaml` unchanged; CI and the refresh bot own dependency
resolution.

## Risks

- Dot remains off by default. Its own opt-in does not enable other
Runner providers. It still requires the MCP option; disabling Dot blocks
new work while keeping existing recovery and saved bindings.
- Dot requires a stable public HTTPS origin. Its base provider in #15402
is merged. Temporary tunnels are useful for testing but are not
permanent deployments.
- Idle intake creates a visible agent-authored task. It does not
fabricate a human message or bypass ordinary admission.
- Workspace commands require Linux bubblewrap with a private PID
namespace; deployment qualification is still needed. macOS file tools
and artifact publishing remain available, but commands are disabled
because sandbox-exec does not contain detached descendants. There is no
unrestricted fallback. Command protection fails closed above 4,096
workspace directories. The historical macOS command walkthrough does not
qualify the current Linux implementation.
- Provider model choice, token usage, cost, native thread control, and
global external stopping remain unavailable. Known Paperclip budget
gates still apply.
- The workspace bridge and attachment reader are separate opt-in
settings. Reading assigned task files sends their contents to OpenAI.
Revocation prevents future reads but cannot withdraw bytes already sent.
Files are read on request; automatic inbound attachment staging remains
disabled.
- An existing event registration can retain earlier instructions that
prohibit extra tasks. Dot asked for permission before a second intake
after the final event-only bootstrap; that extra no-nudge continuation
is not live-qualified under that earlier registration. The later
attachment qualification submitted normally through the composer. The
revised onboarding prompt describes the new idle entry point, and its
protocol polling path is tested.
- OpenAI event delivery can be delayed; a webhook acknowledgement is not
proof that Dot has begun work.
- A running Dot can retain an old plugin catalog after tool refresh.
Refresh and actual tool exposure must be checked before using new
top-level actions.
- Hosted and remote controller modes are not qualified.

## Model Used

OpenAI Codex, based on GPT-6. The exact deployment ID and context window
size are not exposed in this session. Capabilities used: reasoning,
repository editing, code execution, test inspection, and browser
control.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either linked existing issues or described the issue in-PR
following the relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused suites and all
fresh failure rechecks pass; the earlier full attempt is disclosed
above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.1008.0-canary.1
2026-10-07 19:44:17 -05:00
DottaandPaperclip fc6304dfe5 feat(runner): add experimental OpenAI Dot provider over MCP Events (#15402)
## Thinking Path

> - Paperclip manages AI agents, tasks, permissions, and execution
budgets.
> - Paperclip Runner gives each provider the same admitted task and tool
authority.
> - OpenAI Dot runs outside the local process tree and needs
asynchronous work delivery.
> - The merged MCP gateway supplies OAuth consent and signed event
delivery.
> - A personal assistant grant cannot safely stand in for an assigned
agent.
> - This pull request adds a separate Dot agent connection and a durable
Rust Runner bridge.
> - The operator can assign work to Dot and inspect its accepted work,
tool receipts, and result.

## Linked Issues or Issue Description

**Agent or provider**

OpenAI Dot, as an experimental provider of the existing Paperclip Runner
adapter.

**Why this adapter is useful**

An operator can assign normal Paperclip tasks to an existing Dot. Dot
can read its mailbox, request work on an assigned task, use admitted
task tools, and submit a result. Paperclip keeps company scope,
checkout, approvals, known budget limits, and activity attribution.

**How the agent is invoked**

A dedicated `/mcp/runner` OAuth resource pairs one Dot grant with one
agent. A signed MCP mailbox event wakes Dot. Dot explicitly accepts the
assignment. The Rust Runner owns the durable turn and operation
receipts. The first release supports self-hosted instances with a local
Runner controller.

**Additional context**

This extends the merged public MCP gateway from #14846 and the assistant
invitation and device-consent work from #14933. This also integrates the
merged assistant tool and configuration expansion in #15380. Dot retains
its dedicated agent resource and cannot receive personal configuration
permission. The public assistant connection remains a personal
connection.

## What Changed

- Add a durable Rust Dot provider and its TypeScript Runner driver.
- Add closed PRP v3 external-provider operations and native execution
input v6.
- Add company-scoped pairing, mailbox, assignment, and operation
records.
- Reuse merged browser/device consent, client metadata verification,
webhook admissions, refresh, secret rotation, and warm-standby gates.
- Keep Dot scopes, issuer, grants, event workers, and tool access
separate from personal assistant access.
- Add Dot configuration, pairing, readiness, and consent UI. Keep agent
grants out of the personal Connections entry.
- Regenerate the Dot-only migration after master. Preserve published
gateway migrations. Make the new migration safe to reapply.
- Document setup, recovery, accounting limits, evidence, and remaining
account qualification.
- Reverify reconnect callbacks and wake outstanding work with a fresh
mailbox reference; preserve the existing assignment and operation
receipts.
- Clean up Dot bindings and waiting runs on OAuth revoke and
refresh-token replay. Old grants cannot revoke replacement bindings.
- Restore the pairing reference when an unsaved agent form is reopened;
document board-only pairing routes in OpenAPI.
- Accept a clean Rust exit after the acknowledged shutdown receipt.
Unexpected exits still require recovery.
- Clear the cached binding after a successful revoke so a failed
connection refresh cannot restore it.
- Add production-component Storybook states and screenshots for pairing
and connection review. All preview account data is synthetic.
- Persist normalized completion, serialize Dot turns and durable work
admission, and poll subscription readiness.
- Serialize mailbox writes and cursor reads; retain paused fence
acknowledgement without task authority.
- Authorize admitted review runs without changing the worker assignee.
Include the fenced assignment ID in production stop notices.

## Verification

- This PR integrates master `4a8178e9c`. Dot migration
`0317_messy_famine.sql` follows the published history and is safe to
reapply. The merge preserves the reserved migration connection,
batch-commit handling, private task checks, task monitors, and native
accounting.
- Local workspace typecheck, full build, and UI token gates pass. The
server typecheck passes after the review fixes. Database and native
executor regressions pass.
- All twelve real Rust/PostgreSQL Dot integration tests and twelve Dot
driver tests pass. The tests cover native document writing and
finalization, durable replay, queue admission, mailbox ordering,
admitted reviews, stale authority, production stop references, and
paused acknowledgements.
- Current head `d0e7e0626` passes all 57 checks: 53 pass and four are
intentionally skipped. This includes full typecheck, build, tests, Rust
Runner verification, browser E2E, release verification, and Canary Dry
Run. Greptile rates this exact head 5/5. All review threads are
resolved.
- The full local root test run is slower than the sharded CI run and has
not completed. The full CI test gates pass on the current commit.
Focused local regressions pass.
- Real-account pairing and event delivery on this base commit remain
unqualified. Live account and setup proof are recorded in the follow-up
#15414.

The following screenshots use synthetic preview data. They show the
production pairing component and do not qualify a real account or the
full agent setup journey.

![Synthetic pairing
preview](https://github.com/paperclipai/paperclip/blob/codex/dot-events-prototype/doc/screenshots/openai-dot-runner/pairing.jpg?raw=true)

![Synthetic connected
preview](https://github.com/paperclipai/paperclip/blob/codex/dot-events-prototype/doc/screenshots/openai-dot-runner/connected.jpg?raw=true)

## Risks

- This base adapter uses `PAPERCLIP_ENABLE_OPENAI_DOT=1` plus Public MCP
and Paperclip Runner. The separate experimental-settings follow-up in
#15414 replaces this environment flag with saved operator settings.
- Dot does not expose provider token usage or cost. The operator must
acknowledge external billing. Known Paperclip budget gates still apply.
- Cancellation fences Paperclip authority. It does not confirm that Dot
stopped all external activity.
- Assigned skill files and third-party MCP bindings are unsupported and
reject admission. There is no mounted workspace, model selector, or
provider thread identifier.
- Hosted agent-broker and remote controller deployments are not
qualified.
- The new migration follows the merged master history. Existing
prototype databases still need the normal master migration history
before this Dot-only migration.

## Model Used

OpenAI Codex, based on GPT-6. The exact deployment ID and context window
size are not exposed in this session. Capabilities used: reasoning,
repository editing, code execution, and test inspection.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #123` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub references)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.1008.0-canary.0
2026-10-07 19:11:50 -05:00
DottaandPaperclip 4a8178e9cf fix(ui): stabilize composer model selection and refine effort slider (#15444)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Task composers let users select an agent, a model, and an effort
level.
> - The model label changes width when users move the effort slider.
> - This movement makes the control harder to use.
> - This pull request keeps the label fixed while the picker is open.
> - A larger slider gives clearer feedback during selection.

## Linked Issues or Issue Description

**What existing behavior does this improve?**
The shared model and effort picker in task composers.

**Current behavior**
The trigger shows each effort change while the picker is open. Different
label lengths move the control. The slider is small.

**Proposed behavior**
The trigger shows “Select model” while any picker view is open. It shows
the selected model and effort after the picker closes. The slider has a
thicker track and a larger thumb with hover and press feedback.

**Reason and benefit**
Users can change effort without moving the trigger. The larger thumb
gives clearer pointer feedback.

## What Changed

- Keep the trigger text fixed while the picker is open.
- Add slider size and feedback tokens.
- Remove the gray input outline and thumb border. Show keyboard focus on
the thumb.
- Keep a thumb focus outline when Windows high contrast suppresses
shadows.
- Keep native keyboard and pointer controls. Respect reduced motion.
- Test open and closed labels on desktop and mobile.

## Verification

- Picker tests: 27 pass. Composer settings and new-task tests: 86 pass.
- UI typecheck and UI build passed before the final CSS-only
high-contrast fix.
- Token gates pass.
- Full typecheck and build stop at installed Runner API type mismatches
outside these files.
- Revised Chromium checks pass for desktop pointer drag, no input
outline during drag, stable open label, keyboard input, and mobile
rendering. The earlier rendered-image check confirmed reduced-motion
behavior.
- The earlier full local suite did not finish. The approved visual
revision passed 113 focused tests. The final high-contrast fix passed 27
picker tests and token gates.
- Chromium high-contrast and reduced-motion verification shows a visible
thumb focus marker and working keyboard input.
- Open the model picker. Change effort. Confirm “Select model” stays
visible. Close the picker. Confirm the selected model and effort appear.

- All CI gates pass on final head
`7a64104e14251a354b06b7916d09d44692557c4f`, including full typecheck,
tests, build, and browser E2E. Greptile: 5/5 with no findings.

## Risks

- Low risk: shared UI only. Native range controls stay in place.
- Browser thumb rendering can differ between engines.

## Model Used

- OpenAI GPT-6 Codex. Code editing, tool use, and test execution. The
runtime does not expose an exact model variant or context window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.1007.0-canary.16
2026-10-07 21:52:19 +00:00
github-actions[bot]andlockfile-bot 378e6d95e1 chore(lockfile): refresh pnpm-lock.yaml (#15488)
Auto-generated lockfile refresh after dependencies changed on master.
This PR only updates pnpm-lock.yaml.

Co-authored-by: lockfile-bot <lockfile-bot@users.noreply.github.com>
2026-10-07 15:55:57 -05:00
DottaandPaperclip a7a244ab33 feat: add company decision models with permission and cost controls (#15473)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Optional product features need small, typed model decisions.
> - Each company needs to choose which shared API connection pays for
those decisions.
> - Calls must retain user and task permissions, budget limits, and cost
attribution.
> - This pull request adds a managed decision service, setup UI, and
request history.
> - Features can check availability cheaply and keep their existing
behavior when decisions are unavailable.

## Linked Issues or Issue Description

**Subsystem affected**

Server services, shared contracts, database accounting, Company
Settings, and Costs.

**Problem or motivation**

Paperclip has no common decision-model service. Adding provider calls
within each feature would duplicate credential access, permission
checks, and billing rules.

**Proposed solution**

Let a connection manager configure one company decision model. Support
OpenAI Decisions and Jev through OpenRouter. Provide a fixed setup test
and metadata-only history. Default company-sponsored background
decisions to on during setup, and preserve a saved off setting.

**Alternatives considered**

Per-feature credentials would duplicate existing connection management.
Personal overrides and provider fallback chains add permission and
billing complexity; they remain deferred.

**Roadmap alignment**

Reviewed ROADMAP.md and searched open PRs. This extends existing
connection access and budget accounting. Product features that call the
service remain outside this change. No matching decision-model service
PR was found.

## What Changed

- Add company settings, an internal `decisionModelService`, local
availability checks, and trusted human, agent/run, and system contexts.
- Pin Vercel AI SDK provider dependencies and adapt boolean, choice, and
ordered-score decisions for both providers. Bound requests and time;
disable paid retries.
- Add durable invocation metadata and agentless decision ledger charges.
Preserve fractional cents, pricing evidence, dispatch identity, and
unresolved billing holds.
- Share the company accounting lock and apply company, agent, and
project budgets. Settle charges once, retain unknown holds, and recover
interrupted calls without resubmission.
- Reuse connection setup and management UI. Add a Decisions view under
Costs, production-component Storybook coverage, database migration, and
service documentation.

## Verification

- Passed 166 current-code tests covering the decision service/provider,
setup component, Costs, OpenAPI, and every failure from the earlier
broad run. Coverage includes native SDK wire formats, refusals, billed
malformed responses, permission and secret-rotation races, identity
changes, concurrent budget admission, unresolved holds, agent/task
deletion, and stale setup feedback.
- Passed 170 existing connection, cost, budget, heartbeat-accounting,
and profile regression tests.
- Passed repository typecheck, production build, and design token gates
after integrating master. Verified the generated migration on a fresh
test database and upgraded the populated preview database from the
branch's earlier migration without losing settings or usage.
- Ran the required full `pnpm test:run`: its general phase completed
with 16,354 passed and 10 failures across five files while this branch
was still being updated. Every reported failure passes in the
current-code rerun; the serialized phase did not run after that failure.
The full GitHub CI suite passed on `b1b856a88`: general and serialized
tests, browser shards, runner checks, typecheck, production build,
packaging/canary, and policy gates. [CI
evidence](https://github.com/paperclipai/paperclip/actions/runs/37673106360).
- Passed the full-shell Storybook setup-to-history interaction test
again after integrating master. Greptile rates the final revision 5/5
with zero unresolved threads.
- Walked through the running app: empty setup, add each provider, save,
reload, run all three sample questions, inspect fractional charges in
history, switch provider while sponsorship is off, disable, and
reconnect. Checked mobile settings. These tests used the actual UI,
vault, server, SDKs, and database with simulated upstream responses.
- Live paid setup tests remain unverified: this environment has no
authorized OpenAI/OpenRouter credentials available. No mocked test is
presented as live provider evidence.

Reviewer journey: Company Settings → General → Decision model.
Add/select a shared API connection, save, run the billed sample, open
View usage, then disable decisions and verify Run test is disabled after
reload.

## Risks

- The SDK decision interface is experimental. Pinned versions and
wire-format tests limit upgrade drift.
- The migration allows agentless service charges and reservations.
Existing agent cost-reporting APIs still require an agent, and decision
receipts stay separate from run reconciliation.
- Timeouts can have unknown provider charges. Holds remain until an
audited accounting correction resolves them.
- OpenAI prices use a versioned Decisions rate snapshot; OpenRouter
costs use provider receipts. Unknown pricing is retained as unknown.
- A configured company authorizes background spending by default. Setup
explains this, and managers can turn it off.
- Live provider account/model availability still needs the two
credentialed acceptance checks.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository inspection, code
execution, and browser tools. The exact serving revision and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.1007.0-canary.15
2026-10-07 15:53:13 -05:00
Devin FoleyandPaperclip f1c44b7b56 Preserve remote work when required workspace restore fails (#15479)
Persist exact source-retention obligations before run finalization and protect them across cancellation, restart and task changes. Require board-authorized repair evidence without replaying old work or changing current task ownership, state or locks. Hide an unavailable retry action and document operator recovery and retention costs.

Validated with 674 scoped regressions, 224 combined integration tests, full typecheck/build, independent safety reviews and green CI with Greptile 5/5. No historical file recovery is claimed.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-07 13:29:44 -07:00
Devin FoleyandPaperclip cd40ebb95b Preserve ACP bridge terminal evidence before disconnect (#15478)
Drain terminal frames through backpressure and record bounded terminal evidence without waiting for log persistence. Guard later events and input failures against writes after socket end.

Validated with 217 focused tests, full typecheck/build, independent review and green CI with Greptile 5/5.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
canary/v2026.1007.0-canary.14
2026-10-07 13:01:22 -07:00
50b1f95e79 fix(tool-gateway): keep MCP connection healthy on oversized/malformed responses (#15462)
<!-- Write all pull request text in Simplified Technical English
(ASD-STE100): short sentences, one instruction per sentence, simple
approved vocabulary, and the active voice. -->

## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Agents call remote MCP tools through the Paperclip tool gateway
> - The gateway keeps a health status for each remote MCP connection,
and it hides the tools of a connection that has `healthStatus = "error"`
> - One oversized, malformed, or invalid-JSON `tools/call` reply set the
whole connection to `error`
> - Health goes back to `ok` only after a successful call, but the
hidden tools prevent that call, so the connection stayed locked until a
person reconnected it
> - This pull request makes these reply errors fail only the one call,
and it starts a new MCP session for the caller on the next call
> - The benefit is that one large or bad reply no longer disconnects a
working connection for every agent

## Linked Issues or Issue Description

No public issue exists. Related open PRs fix other health downgrades in
the same function. They do not overlap with this change:

- Refs #11910 (timeouts and JSON-RPC errors)
- Refs #15325 (transport errors)

**What happened?**

An agent called a remote MCP tool that returned more than
`MAX_REMOTE_MCP_RESPONSE_BYTES` (1 MB). The gateway returned 502
`mcp_remote_response_too_large`. It also set the connection to
`healthStatus = "error"`. Then `connectedMcpConnectionFilter` hid all
tools of the connection (except `per_user` connections), and
`connections_search` showed the connection as `needs_user_action`. The
`malformed_response` and `invalid_json` errors from
`readMcpHttpResponse` did the same thing.

`readMcpHttpResponse` can also cancel the reply stream before the end.
The remote server can then close its MCP session. The gateway kept the
cached `mcp-session-id` for up to 30 minutes and cleared it only after a
404. The next calls then failed with "Remote MCP session expired".

**Expected behavior**

A reply that is too large or not correct fails only that call. The
connection stays healthy, and its tools stay visible. The next call uses
a new MCP session.

**Steps to reproduce**

1. Add a remote MCP connection with `mcpSessionRequired: true`.
2. Call a tool that returns a reply larger than 1 MB.
3. Look at the connection: `healthStatus` is `error`.
4. Start a new gateway session: the tools of the connection are not in
the list.

**Paperclip version or commit**

`master` at `99a9de9940bf5974352d9dbfbb2f21e62e89689f`

**Deployment mode**

All modes. The fault is in the server tool gateway.

## What Changed

- `server/src/services/tool-gateway.ts`: `too_large`,
`malformed_response`, and `invalid_json` from `readMcpHttpResponse` now
fail only the call. The caller gets the same 502 reason code as before.
The gateway does not call `markRemoteConnectionHealth(…, "error")` for
these errors. The `invalid_json` branch after `JSON.parse(body)` also
does not change health now, so all `invalid_json` paths are the same.
- The `mcp_remote_response_too_large` message now tells the caller to
request a smaller result, for example a narrower query or a smaller page
size. The error details now include `maxBytes`.
- `server/src/services/mcp-http.ts`: new `forgetMcpHttpSession()`. It
removes only the cached session for one scope and one credential set,
and only while that entry still holds the session ID that failed. It
does not touch other agents, sessions that are initializing, or a newer
session that replaced the failed one.
- The gateway calls `forgetMcpHttpSession()` after these reply errors,
so the next call from that caller initializes a new session. The
existing 404 "session expired" path now uses the same function. Before,
both paths cleared all sessions and all pending initializations on the
connection. That made a concurrent initialization by another agent fail
with "MCP connection changed while initializing", which the gateway
reported as a fetch failure and marked as a connection `error`.
- `forgetMcpHttpSessions(connectionId)` is not changed. Disconnect and
revocation still clear the full connection.

Why `malformed_response` and `invalid_json` are also per-call errors:

- Each error is about one reply body. The server was reachable, accepted
the credentials, and sent HTTP 2xx. A reply can be too large or bad
because of the tool and its arguments. That is not a fault of the
connection.
- The other malformed-reply checks in the same function (payload is not
an object, or has no `result`) already throw
`remote_mcp_malformed_response` and do not change health. Only the
reader-level errors changed health.
- Health recovers only after a successful call. An `error` status for
one bad reply therefore locks the connection until a person reconnects
it.
- Other failures (HTTP errors, fetch failures, timeouts, JSON-RPC
errors) still change health. This PR does not change them. #11910 and
#15325 address some of them.

## Verification

- `server/src/__tests__/tool-gateway.test.ts`: the recovery test now
runs for three replies: oversized, invalid JSON, and a reply without the
requested message ID. Each run uses a fake HTTP MCP server with
`mcpSessionRequired: true` that gives `session-N` for each `initialize`.
Each run checks that:
- the call fails with the correct 502 reason code (and, for the
oversized reply, the smaller-page hint and `maxBytes: 1000000`)
  - `healthStatus` stays `ok`
  - a new gateway session still lists the tool and can call it
  - the second `tools/call` uses `session-2`, not `session-1`
- New test: agent A gets an oversized reply while agent B initializes a
session on the same connection. Agent B's call completes, health stays
`ok`, and the two calls use `session-1` and `session-2`. With the old
connection-wide reset, this test fails: agent B gets 502
`mcp_remote_fetch_failed`.
- `server/src/__tests__/remote-mcp-protocol.test.ts`: new unit test.
`forgetMcpHttpSession()` keeps the session of a different identity, and
a late failure from an old session does not remove the newer session.
- Each regression check fails without its fix. Without the health
change, the test fails on `healthStatus: 'error'`. Without the session
reset, it fails with `['session-1', 'session-1']`.

```
cd server
npx vitest run src/__tests__/tool-gateway.test.ts src/__tests__/tool-gateway-service.test.ts src/__tests__/remote-mcp-protocol.test.ts
 Test Files  3 passed (3)
      Tests  132 passed (132)
```

- I did not run the full server `tsc --noEmit` locally because the
sandbox does not have sufficient memory. The CI typecheck covers it.

## Risks

- Low risk. The change affects only the error path of remote MCP
`tools/call`.
- A remote server that always sends bad replies now keeps `healthStatus
= "ok"`. Each call still fails with a clear 502 reason code and an audit
record, so the failure stays visible. Only the connection-wide hiding of
tools stops.
- After one of these errors, the next call from the same caller sends
one more `initialize` request. This adds one round trip.
- A 404 "session expired" now clears only the session of the caller that
got the 404. Before, it cleared the cached sessions of all agents on the
connection. If the remote server restarts, each agent now gets its own
"session expired" error one time and then initializes again. A 404 is
about one session, and the old connection-wide clear also cancelled
other agents' initializations.
- #11910 and #15325 change the same catch block. The PR that merges last
can have a small merge conflict.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- Claude Opus 5.5 (Anthropic), model ID `claude-opus-5-5`, 1M-token
context window.
- Run as an agent in Claude Code with tool use (shell, file edit, and
local test runs).

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes (no
user-facing documentation changes are necessary)
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

This PR replaces #15418. It keeps the same commits on a branch name
without an internal ticket ID, and adds a fix for the Greptile review on
#15418.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: devinfoley <139239+devinfoley@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 12:48:01 -07:00
DottaandPaperclip fd8c6b920a fix(native): resume connection tasks after approval decisions (#15471)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Native tasks can pause while a human decides whether to allow a
connection action.
> - The next turn needs both the saved decision and complete accounting
for the previous turn.
> - Cancellation could discard final usage, and complete direct Claude
API receipts could remain unpriced.
> - A stale blocked or review report could also request approval again
after the original card was declined.
> - This pull request retains shutdown accounting and rejects approval
waits bound to an already resolved action.
> - The benefit is reliable continuation with the existing budget and
approval controls.

## Linked Issues or Issue Description

Refs #15420. Related: #15312 addresses requester ownership during
dispatch. This change addresses receipt capture and final-response
validation.

## What Changed

- Retain usage events after native cancellation. Continue to reject late
provider messages and work.
- Drain same-turn accounting and terminal events for at most fifteen
seconds after a durable governed wait. Keep incomplete accounting
blocked.
- Estimate complete, unpriced, direct Anthropic API receipts for the
exact `claude-sonnet-5` model. Record the rate version and assumptions.
Use the one-hour cache-write rate when the receipt lacks cache TTL.
- Bind stale approval reports to exact interaction, action-request, or
invocation IDs in the same company, task, agent, and run. Cover blocked,
review, and response-wake reports. Keep independent reviews valid.
- Fence checkpoint and result writes after a controller detaches for
restart, including operations waiting for a database lock. Reject stale
successful returns before certifying accounting.
- Allow bounded subscription teardown only after retaining an actual
provider terminal.
- Journal the exact governed-wait trigger and disposition before
provider interruption. Recover that wait independently of a later saved
answer, replay retained accounting, and reject mismatched or unproven
terminal evidence.
- Give settling governed turns a bounded window before shutdown detaches
their controller.
- Keep fuzzy external app matches alongside installed capability matches
instead of forcing an unrelated provider question for a generic query.
- Return up to twenty exact active catalog tool names after an invalid
request, after eligibility checks; still reject the request without
granting access or creating an approval.
- Require retained provider terminal proof before settling a governed
wait, including when complete usage arrives before stream
closure/error/timeout. Retain harmless numbered cancellation events so
restart replay stays contiguous.
- Isolate accounting-test OpenCode config from the host plugin
directory.
- Add regression coverage and document the accounting, restart and
connection-search behavior.

## Verification

Current PR source: `0cf08efd75f8fb23f7989beda6dbda92587088bf`. The live
matrix below measured frozen `09a776bcb1f77422156f8be11e13f8c29f43e7f7`;
later review fixes are verified separately and do not relabel those
runs.

- Repository typecheck and build pass.
- Current runner runtime and cancellation suites: 191 tests pass. Six
new regressions cover stream end/error/timeout without provider-stop
proof and contiguous cancellation acknowledgement/request replay; all
six failed before the fix. The existing bounded cleanup case now
explicitly supplies terminal proof. All 31 adapter accounting tests pass
with isolated fixture config.
- New checkpoint-rebinding and approval-criterion suites: 65 tests pass;
runner HTTP integration: 29 tests pass. Unchanged executor/control-plane
suites: 640 tests pass; database-backed connection suites: 68 tests
pass.
- Current retained evidence verification covers 222 file hashes across
all fifteen original result artifacts. The three-case [published
report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37651896525-1/index.html)
and all eight screenshot hashes verify. The [three-case recovery
report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37662587356-1/index.html)
and all six screenshots also verify; the nine-result campaign did not
publish.
- Evaluation support: 1,805 Vitest tests pass, one skipped; 128 Node
tests pass. Eval typecheck and catalog discovery pass.
- Full local repository unit run was interrupted before the follow-up
edits after three tool-access failures and one runner HTTP failure.
Those failures pass in isolation; the 09a full run was interrupted after
one rapid Slack callback-ordering failure and seven skill-service
failures. All eight pass both isolated and with full-runner environment
settings, and all 74 skill-service tests pass together; the subsequent
full run reported two 15-second OpenCode accounting timeouts and was
stopped with exit 130 to apply review fixes. The timeouts reproduce
while copying this host’s 61 MB OpenCode config. All 31 tests pass after
isolating config inside each fixture without increasing timeouts or
changing assertions. A complete local full-suite pass is not claimed.
Repository-wide CI also passes on the final review-fix head in [run
37669185952](https://github.com/paperclipai/paperclip/actions/runs/37669185952).
Earlier database-skipped diagnostics and the older ENFILE run are
retained and are not full-suite passing evidence.
- Three-case live campaign
[37651896525](https://github.com/paperclipai/paperclip/actions/runs/37651896525)
passes all three original grades on frozen source 09a: 55/55 checks,
eight succeeded run records, complete accounting receipts, matching
checkpoint identities and no pending approvals. Campaign
[37653533353](https://github.com/paperclipai/paperclip/actions/runs/37653533353)
adds nine original passes (173/173 checks, eighteen succeeded records)
on the identical source. Its other three jobs failed before runner
assignment or any step while GitHub could not load the paid environment;
those original infrastructure failures are retained. Campaign
[37662587356](https://github.com/paperclipai/paperclip/actions/runs/37662587356)
completes only those unstarted cells: all three original grades pass
(57/57 checks, six succeeded run records). All fifteen exact cases now
pass on source 09a: eight FAIL → PASS, seven PASS → PASS, zero new
overall failures and zero pending pairs. Total current evidence: 285/285
checks and thirty-two succeeded run records, complete accounting
receipts, matching checkpoint identities, no pending approvals or retry
records. Earlier campaigns retain forty-seven additional run records and
two known same-run recovery attempts; actual provider-call counts and
invoices remain unknown. The nine-result campaign skipped publication
and its public URL returns 403; original artifacts remain retained.
Previous ba9 campaign
[37645656840](https://github.com/paperclipai/paperclip/actions/runs/37645656840)
completed 2 PASS / 1 FAIL: Claude restart/approval and Codex decline
pass, while OpenCode resumes but times out searching for exact tool
names and never creates the access card. That failure and incomplete
cancelled-run accounting remain preserved; the new catalog error
guidance targets this observed dead end. Campaign
[37642957312](https://github.com/paperclipai/paperclip/actions/runs/37642957312)
remains 0 PASS / 3 FAIL and exposed the now-corrected cross-run marker
leak and unknown-criterion approval gap. The original baseline remains 7
PASS / 8 FAIL, first repair 2 PASS / 3 FAIL, and second repair 0 PASS /
3 FAIL. All fifteen selected cases are qualified by their original
grades in this bounded trial. These live grades belong to 09a. Its
twelve governed-wait checkpoints retain matching same-turn terminal
fingerprints, but passing artifacts omit detailed event journals; the
later six adversarial regressions qualify the new terminal-proof and
replay guards separately. Final-head repository CI passes. [Fresh
Greptile
review](https://github.com/paperclipai/paperclip/pull/15471#issuecomment-6044369780)
is 5/5, confirms both findings are fixed, and reports no new actionable
issues. All review threads are resolved and the PR has no merge
conflicts.

## Risks

- Governed cancellation drains accounting for up to fifteen seconds.
Restart detachment gives a settling batch up to twenty seconds to
finish. An incomplete receipt or unproven provider terminal still
prevents successful qualification.
- Claude prices are estimates, not invoices. The estimate assumes
standard global API pricing and uses a conservative cache-write rate.
Unsupported models, billers, and billing modes remain unpriced.
- Approval identity matching must remain scoped to the current run and
the requested approval. It does not authorize execution of a declined
call.
- Original baseline and final candidate have different merged master
context. Exact-case outcomes are before/after observations, not isolated
causal attribution to this repair.
- No schema migration, fixture, oracle or grader change. Search-result
guidance now treats fuzzy external matches as suggestions. Existing app
authorization and provider-consent checks remain required.

## Model Used

- OpenAI GPT-6 through Codex, with code editing, terminal tools, and
test execution. The exact deployment identifier and context window are
not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (191 current runtime tests
and 31 accounting tests; interrupted full-suite history and CI coverage
are disclosed above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-07 14:43:07 -05:00
Devin FoleyandPaperclip 5405b1f46f Synchronize chat delivery tests with completed drains (#15480)
Expose the already tracked and caught deferred-work promise to test schedulers while preserving default setImmediate behavior. Assert Slack delivery ordering while the primary lease is held, then join the drains; update the existing GitHub test to assert the resulting quiescent queue.

Verified all 1,063 chat tests, full typecheck/build, negative synchronization control, independent review and exact-head CI. An unchanged browser shard retry passed and is documented. Greptile 5/5 with no unresolved comments.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-07 12:38:12 -07:00
Devin FoleyandPaperclip 0514e8abd5 Keep Cursor managed-runtime test installation offline (#15484)
Intercept the Cursor installer in its local shell fixture and guard against accidental curl downloads. Preserve real managed-home archive operations and the default timeout; change no production code.

Verified 46 related tests, safe negative mutation, independent review and all exact-head CI. Greptile 5/5 with no unresolved comments.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-07 12:38:06 -07:00
Devin FoleyandPaperclip 572cec0344 fix(ui): present missing costs as an informational notice (#15459)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The Costs dashboard shows spending and gaps in cost data.
> - Historical usage can have token counts without a recorded price.
> - The red notice implies that the user must fix those records.
> - This pull request shows the missing costs as a neutral note below
the totals.
> - Users can see the limit of the totals without a false call to
action.

## Linked Issues or Issue Description

Refs #14997.

**What happened?**

The Costs dashboard shows a red notice when a selected period contains
usage without prices. It also shows “0 runs await accounting” when no
runs are pending. Historical pricing gaps can therefore look like
current errors that need cleanup.

**Expected behavior**

Show a neutral explanation below the totals. State that totals include
only known costs. Show pending runs separately and only when the count
is positive.

**Steps to reproduce**

1. Open Costs for a period with unpriced usage and no pending runs.
2. Observe the red notice above the totals.
3. Apply this change. Confirm that the notice is muted and below the
totals, with no zero-count pending message.

**Paperclip version or commit**

Reproduced on master after #14997.

**Deployment mode**

Board UI in both the embedded Activity page and the standalone Costs
page.

## What Changed

- Move the cost-data notice below the summary tiles and use the muted
text token.
- Explain unavailable prices with “Totals include known costs only.”
- Separate pending runs from missing prices. Omit zero counts and use
singular or plural copy as needed.
- Update the existing UI tests for both page variants, including
missing-only, pending-only, and fully accounted states after refresh.

## Verification

- `pnpm --dir ui exec vitest run src/pages/Costs.test.tsx`: 21 tests
passed on the final commit. Both page variants cover mixed,
missing-only, pending-only, and fully accounted states.
- `pnpm check:token-gates`: passed on the final commit.
- `pnpm build`: passed locally.
- `pnpm -r typecheck`: passed locally.
- Full Linux CI passed on `0c23544c0f17e6641cc0c8de6ed498a731748bde`,
including browser tests, server tests, shared-package tests, build, and
typecheck. [CI
run](https://github.com/paperclipai/paperclip/actions/runs/37661883464).
- Greptile Apex: 5/5 on the final commit, with no unresolved review
threads.
- Local full-suite limit: `pnpm test:run` first failed because the fresh
checkout lacked embedded PostgreSQL library links. A runtime-skill
fixture also selected an unrelated parent directory, and a load test
timed out. After setup repair, the database suite (70 tests), skill
suite (3 tests), email suite (39 tests), and load suite (4 tests) passed
individually. The subsequent full local rerun was stopped after the
complete Linux CI run passed. This is not a claim that a full local test
invocation passed.

## Risks

Low risk. This changes presentation only. The note remains visible and
retains its accessible status role. Cost calculations, stored records,
and budget enforcement do not change. The copy does not assume that all
missing prices are historical. No schema migration or operator
documentation change is needed.

## Model Used

OpenAI Codex, based on GPT-6, with reasoning, repository inspection,
code editing, and test execution. The exact served model identifier and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-07 12:34:12 -07:00
Devin FoleyandPaperclip 465140596f Preserve bounded workspace sync diagnostics across RPC (#15481)
Preserve bounded error codes and HTTP/exit statuses across the environmentSyncOut worker RPC boundary, and revalidate that method-scoped envelope before attaching host restore diagnostics. Keep the original error and all recovery policy unchanged; do not transmit provider payloads or credentials.

Verified real RPC roundtrip and privacy regressions, 95 focused tests, 68 independent tests, full typecheck/build and all exact-head CI. Greptile 5/5 with no unresolved comments.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-07 12:22:53 -07:00
Devin FoleyandPaperclip 128c95837b Retain bounded orphan process-loss diagnostics (#15475)
Record bounded observer uptime, run and output ages, process-check observations and retry eligibility before orphan cleanup changes the evidence. Preserve existing recovery and reporting behavior and omit process IDs, raw paths and credentials.

Verified focused helper, actual reaper and real SDK regressions, full typecheck/build, exact-head CI and independent review. An unchanged Cursor timeout passed its isolated retry; unrelated local Slack timing failure is documented separately. Greptile 5/5 with no unresolved comments.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-07 12:05:21 -07:00
Devin FoleyandPaperclip 43b0aff54b Keep pre-dispatch secret configuration blockers out of Sentry (#15472)
Keep confirmed pre-dispatch missing-secret setup blockers out of Sentry while preserving failed runs and owner recovery actions. Runtime, provider and ambiguous failures remain reportable.

Verified full exact-head CI, focused local regressions, workspace typecheck and independent review; Greptile 5/5 with no unresolved comments.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-07 12:02:29 -07:00
DottaandPaperclip ae6f95ed7a feat: add personal primary agents (#15470)
Add a personal primary agent per company and user. Initialize it from the first human-created agent, expose profile-only switching with confirmation, and use it after recent choices for task and Chat defaults.

Persist authenticated preferences, preserve lifecycle and membership rules, keep selections out of shared audit events, and document the API contract. Include the reviewed Storybook surfaces and regression coverage for concurrent choices, onboarding, cross-device updates, and browser journeys.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-07 13:53:46 -05:00
Devin FoleyandPaperclip 06484b3c41 fix: preserve conversation retries through execution cleanup (#15463)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Recovery schedules bounded retries after a provider disconnects.
> - A stopped run can still hold its environment lease while cleanup
runs.
> - Retrying before that lease is released cancels the new run before it
starts and spends another retry.
> - Restoring the task can also send the worker repair instructions from
an already resolved recovery action.
> - This pull request preserves the waiting retry and removes settled
recovery instructions from later wakes.
> - The task can continue after cleanup without an operator repairing
the same incident again.

## Linked Issues or Issue Description

**What happened?**

A legacy conversation run disconnected while it was doing ordinary work.
Its environment cleanup took longer than the retry delay. Two retries
were cancelled before dispatch with `execution_reconciliation_required`.
Those cancellations exhausted the failure budget. After an operator
restored the task, the wake still told the original worker to repair the
runtime and hand the task back to itself.

**Expected behavior**

Cleanup waits preserve the pending attempt. The same retry can continue
after ownership is released, subject to all current gates. Once a
recovery action is resolved or cancelled, subsequent task wakes omit its
repair instructions.

**Steps to reproduce**

1. Fail a legacy conversation run while its environment lease remains in
`pending_cleanup`.
2. Schedule a bounded retry and run promotion before cleanup releases
that lease.
3. Repeat the scheduler sweep. Before this fix, retries promote and then
cancel without starting.
4. Resolve a stranded-task recovery action and build the restored task
wake with that action ID. Before this fix, the wake still includes the
settled repair instructions.

**Paperclip version or commit**

Reproduced with database regressions against `ceabc3bc880` on master.

Related: #15019 restores a skipped assignment handoff after lease
release. #15235 filters stale handoff evidence in the recovery sweep.
This change preserves an existing scheduled conversation retry and
corrects restored wake content. It does not create a new handoff wake.

## What Changed

- Keep an unstarted legacy conversation retry on the same durable row
while prior execution ownership remains active. Recheck after 30 seconds
without increasing retry accounting.
- Return a queued retry to scheduled state if it encounters that hold at
the claim gate. Retain its issue claim and publish the status change.
- Record one local lifecycle diagnostic per blocking run. Remove that
wait marker on promotion.
- Include recovery action metadata only while the referenced action is
active or escalated.
- Return an explicit `waiting` response and the saved schedule when
Retry now meets cleanup. Show the wait inline without a false success or
disabled button.
- Add database, rendered-prompt, route, and UI regressions. Document the
execution and run-log contracts.

## Verification

- Five cleanup and restored-wake regressions fail against the original
production code. The Retry now route and UI regressions also fail before
their correction.
- Related retry, dispatch, stale-queue, and recovery suites: 321 tests
pass across seven files. All ten focused cleanup/restored-wake cases
pass after rebase. The final dispatch adjustment passes all 46 adapter
tests.
- Retry now routes and affected UI suites: all 46 tests pass. The tests
cover repeated clicks, the saved schedule, unchanged accounting,
promotion after release, and no false success or error state.
- `pnpm -r typecheck`, `pnpm build`, `pnpm build-storybook`, and `pnpm
check:token-gates` pass.
- Full local `pnpm test:run` was started and then stopped after the
final commit passed all GitHub CI test shards. No complete local
full-suite result is claimed; CI supplies the complete test result for
the final commit.
- Final head `6f1058a332c039e33c4f002b296d20a5554e760e`: all 55 GitHub
checks are green or intentionally skipped. Apex review is 5/5 after two
reviews, with no unresolved threads. The PR has no merge conflicts.

## Risks

The wait applies only to unstarted legacy conversation retries. Native
runs and non-conversation execution keep their existing recovery rules.
Cleanup must actually release ownership before execution can resume. The
wait does not fix a cleanup service that never finishes. Promotion and
dispatch still enforce cancellation, reassignment, pause, budget, and
reconciliation gates. No schema change is required. The Retry now
response adds a `waiting` outcome; the shared contract and all three UI
controls handle it.

## Model Used

OpenAI Codex based on GPT-6, with repository analysis, tool use, and
local code execution. The runtime does not expose the precise serving
model ID or context window.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused suites; complete
final-head test coverage in CI)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-07 11:15:49 -07:00
db0bee19fa docs(readme): add Superagent badge and move Star History rank badge to its own line (#15460)
- Add the Superagent security badge to the README badge row, with light and dark variants via <picture>.
- Move the Star History global rank badge into its own centered paragraph below the badge row, with spacing.

Co-Authored-By: Amy (Opus) via Paperclip <claude@users.noreply.github.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 10:40:09 -07:00
9fb955e9a2 fix(codex): gate models on the Codex CLI floor before the ChatGPT backend rejects them (#15399)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The Codex adapter and the native runner run Codex in local hosts and
in managed sandboxes, and the model catalog lists the models an operator
can select
> - The Codex backend accepts a new model with ChatGPT sign-in only from
a recent enough Codex CLI. An older CLI fails every turn with `The
'<model>' model is not supported when using Codex with a ChatGPT
account.`
> - After #14942 added `gpt-6.1-sol`, an operator selected it for an
agent in a managed Daytona sandbox. The sandbox image still shipped
Codex 0.156.0. The connection test and runs failed with that sentence,
which reads like an account problem
> - Paperclip only checked the Codex compatibility window (`>=0.149.0
<0.161.0`), so 0.156.0 passed and the failure surfaced from the backend
without a cause
> - This pull request records the verified Codex CLI floor for each
gated model, compares the installed `codex --version` with that floor in
the environment Test and in the remote runner, and names the backend
rejection when it still happens
> - The benefit is a precise, actionable message ("gpt-6.1-sol requires
Codex CLI 0.159.0 or newer; detected 0.156.0; promote a sandbox image
with Codex 0.160.0") instead of an opaque 400, and no doomed hello probe

## Linked Issues or Issue Description

Refs #14942

**Bug: selecting `gpt-6.1-sol` in a managed sandbox fails with a ChatGPT
account error.**

**Steps to reproduce**
1. Promote a sandbox image that ships Codex CLI 0.156.0.
2. Create a Codex agent with a ChatGPT sign-in connection, select
`gpt-6.1-sol`, and run the environment Test or a task in that sandbox.

**Expected behavior**
The Test or the run tells the operator that the Codex CLI in the sandbox
is older than the model needs.

**Actual behavior**
The Test reports `codex_hello_probe_failed` with
`{"type":"error","status":400,"error":{"type":"invalid_request_error","message":"The
'gpt-6.1-sol' model is not supported when using Codex with a ChatGPT
account."}}`. A run fails with "The selected model is not supported by
the current ChatGPT connection."

**Evidence**
- Codex 0.156.1 and older are rejected for `gpt-6.1-sol` with ChatGPT
sign-in: https://github.com/openai/codex/issues/49396
- Codex 0.159.0 and 0.159.2 are accepted:
https://github.com/openai/codex/issues/49464 and
https://github.com/decolua/9router/issues/4471 (same error, fixed by
raising the client version identity from 0.155.0 to 0.159.0)
- Codex 0.157.0 release notes: "Add GPT-6 Sol and Luna to the model
catalog": https://github.com/openai/codex/releases/tag/rust-v0.157.0
- GPT-6.1 Sol is included for Plus, Pro, Business, Enterprise, and Edu
with ChatGPT sign-in: https://learn.chatgpt.com/docs/models

## What Changed

- `packages/adapters/codex-local/src/index.ts`: add
`minimumCodexCliVersionForModel` (`gpt-6.1-sol` → 0.159.0, `gpt-6-sol`
and `gpt-6-luna` → 0.157.0, with the sources above),
`parseCodexCliVersionOutput`, `codexCliVersionAtLeast`, and
`CODEX_CHATGPT_MODEL_REJECTION_RE`. Models without a verified floor
return `null`, so nothing probes them. The agent configuration doc
describes the floors.
- `packages/adapters/codex-local/src/server/cli-version.ts` (new): run
`codex --version` where the run would execute it (local, SSH, or
sandbox), compare it with the model floor, and build the
`codex_cli_version_compatible` / `codex_cli_version_incompatible`
checks. Map the backend rejection sentence to
`codex_hello_probe_model_rejected` with the detected CLI version and a
hint that separates a stale CLI from a plan that does not include the
model.
- `packages/adapters/codex-local/src/server/test.ts` (CLI lane Test):
run the version check before the hello probe when the model has a floor.
Skip the hello probe when the CLI is too old
(`codex_hello_probe_skipped_cli_version`). When the probe still fails
with the backend sentence, report `codex_hello_probe_model_rejected`
instead of the generic `codex_hello_probe_failed`. The new codes do not
match the auth-failure patterns, so a managed connection is not
invalidated.
- `packages/adapters/codex-local/src/server/acp.ts` (ACP lane Test): for
remote targets, run the same version check against the shared `codex`
the ACP server spawns.
- `server/src/services/native-runtime/native-session-executor.ts`: after
the compatibility-window check, compare the remote Codex with the
configured model's floor and fail with
`runner_remote_provider_artifact_incompatible: <model> requires Codex
<floor> or newer with ChatGPT sign-in, received <version> from the
sandbox image; promote a sandbox image with Codex 0.160.0 or configure
PAPERCLIP_RUNNER_REMOTE_CODEX_NPM_SPEC=...`. When a preinstalled Codex
fails this check and an npm spec is configured, the existing fallback
installs the pinned release.
- Tests: adapter metadata, CLI-lane Test (too old, compatible, no floor,
backend rejection), ACP-lane Test (too old, compatible, no floor), and
the remote runner floor (nine version/model cases).
- `doc/adapter-model-audit-2026-10-02.md`: record the floors and the
reason.

No catalog entry, default model, saved agent configuration, pin,
lockfile, or image definition changes.

## Verification

Local (Node 25.9, pnpm 9.15.4 via corepack):

```sh
pnpm --filter @paperclipai/adapter-codex-local typecheck
pnpm --filter @paperclipai/adapter-codex-local exec vitest run            # 81 passed
pnpm --filter @paperclipai/server exec vitest run \
  src/services/native-runtime/native-session-executor.test.ts \
  src/services/native-runtime/codex-runtime-compatibility.test.ts -t Codex   # 84 passed
```

Also run locally: the full `native-session-executor.test.ts` file (533
passed) and the full Codex adapter suite (81 passed).

Follow-up commit `9fd4d3b25` (Greptile P2: the ACP-lane version probe
ignored the agent's configured env): the probe now receives the
adapter's string-valued `env` entries, the same ones
`buildCodexAcpConfig` hands to remote ACP runs, so a `PATH` override
selects the same `codex` for the Test as for the run. New test covers a
`PATH` + `CODEX_HOME` override and a dropped non-string entry. Re-run on
that head: `pnpm --filter @paperclipai/adapter-codex-local typecheck`
clean; full Codex adapter suite 503 passed (31 files).

Not run here: `pnpm -r typecheck` for `server` (the direct `tsc
--noEmit` was killed by the sandbox memory cap; the adapter package
typecheck passes and CI covers the server), `pnpm build`, browser
suites, and a live sandbox probe. No UI files changed, so token gates do
not apply.

Manual check for a reviewer: set an agent's Codex model to
`gpt-6.1-sol`, point its environment at a sandbox whose `codex
--version` prints `codex-cli 0.156.0`, and run the environment Test. The
result contains `codex_cli_version_incompatible` with `Detected Codex
CLI 0.156.0.` and no hello probe check. With Codex 0.160.0 the Test
contains `codex_cli_version_compatible` and runs the hello probe.

## Risks

- A floor is enforced for all authentication modes, because the runner
does not know the credential method at verification time. An OpenAI API
key on Codex 0.157.0 or 0.158.0 with `gpt-6.1-sol` is now rejected
before launch. Paperclip pins Codex 0.160.0 everywhere, so this only
affects installs that lag the pin, and the message names the fix.
- The 0.159.0 floor for `gpt-6.1-sol` is the oldest stable release
verified to work; 0.157.0 and 0.158.0 were not verified either way. If
OpenAI accepts an older client, the floor can be lowered in one table.
- The `codex --version` probe adds one short process run to the Test
only for models with a floor.
- Rollback: revert this pull request. No data or configuration migrates.

## Model Used

Claude Fable 5.1 (`claude-fable-5-1`), Anthropic, with extended
thinking, tool use, and web research, operating as a Paperclip agent
through Claude Code.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Bender (Fable) <noreply@paperclip.ing>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
canary/v2026.1007.0-canary.11
2026-10-07 10:37:55 -07:00
ceabc3bc88 fix(e2e): wait on the server's own deferral signal in the signoff wakeup helper (#15345)
## Thinking Path

> - Paperclip coordinates work for AI agents through server-managed
tasks and heartbeats
> - Signoff policy tests exercise stage transitions that wake the next
participant
> - A stage transition can defer a wake while the previous run still
holds the issue execution lock
> - The test helper used a fixed wait and could fail before the server
released that lock
> - This pull request waits on the server deferral signal and verifies
the returned run identity
> - The benefit is a stable test with clear failure details and no
repeated wake request

## Linked Issues or Issue Description

**What happened?**

The signoff policy end-to-end spec failed intermittently with `No
issue-bound heartbeat run became available for agent <id>`. The server
had deferred the wake while another run held the issue execution lock.

**Expected behavior**

The test helper waits for the named blocking run to finish, then reads
and verifies the new run for the requested agent and issue.

**Steps to reproduce**

1. Run `npx playwright test --config tests/e2e/playwright.config.ts
tests/e2e/signoff-policy.spec.ts`.
2. Exercise a signoff stage transition while the previous stage run
still holds the issue execution lock.
3. Confirm that the helper waits on the server deferral signal and
returns only a matching run.

This change relates to [PR
#11299](https://github.com/paperclipai/paperclip/pull/11299), which also
touches the signoff policy end-to-end spec.

## What Changed

- Call the wakeup endpoint that reports the server deferral reason.
- Wait for the named blocking run instead of using a fixed delay.
- Verify that each returned run belongs to the requested agent and
issue.
- Bound read and wait loops and report the last server signal and lock
fields on failure.
- Keep one wakeup request per helper call.

## Verification

- `npx playwright test --config tests/e2e/playwright.config.ts
tests/e2e/signoff-policy.spec.ts` passed locally.
- All five signoff policy cases passed locally.
- The spec keeps all 60 assertions.
- The diff contains only `tests/e2e/signoff-policy.spec.ts`.
- The full CI end-to-end shards must pass before merge.

## Risks

- The helper now depends on the wakeup endpoint's deferral fields.
- A server change to those fields can fail the test with a named
diagnostic.
- The change affects end-to-end test behavior only.

## Model Used

OpenAI GPT-5 Codex, exact runtime model `gpt-5-codex`, with tool use and
code review assistance. The runtime does not expose a separate
context-window value.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: nickyleach <331803+nickyleach@users.noreply.github.com>
Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.1007.0-canary.10
2026-10-07 08:57:56 -07:00
DottaandPaperclip f669194298 fix(runner): propagate configured environment to future turns (#15451)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Native runners start provider processes and enforce a separate tool
environment policy.
> - The server resolves task environment bindings at each run boundary.
> - Fixed launch allowlists dropped custom variables after resolution.
> - Warm process reuse and a blanket fingerprint exclusion also hid
configuration changes.
> - This pull request carries a bounded, server-selected list of task
variable names through each process boundary.
> - The benefit is that later turns and tool commands receive configured
values while ambient host secrets stay excluded.

## Linked Issues or Issue Description

**What happened?**

A native Codex agent could not use configured task credentials. Adding
the variables in Settings did not fix the next turn. The provider
sanitizer, runner launch, Rust provider launch, and shell policy each
used fixed allowlists. Effective config fingerprints also excluded all
variables with the `PAPERCLIP_` prefix.

**Expected behavior**

Explicitly bound task variables must reach provider processes and tool
commands. Added, changed, and removed values must take effect at the
next run boundary. A running turn keeps its original configuration.

**Steps to reproduce**

1. Start a native Codex task without a custom environment binding.
2. Add a fake `PAPERCLIP_PAGE_BUCKET` value and a fake Pages credential
binding in Settings.
3. Continue the task and inspect the tool environment.
4. Before this fix, those values are absent even from a fresh runner
launch.

**Paperclip version or commit**

Reproduced at `b31558064`. The fix is rebased on current master.

**Deployment mode**

Self-hosted server with a native process runner.

Related changes: [the legacy Codex MCP environment
fix](https://github.com/paperclipai/paperclip/pull/13321) and [ambient
server-secret
exclusion](https://github.com/paperclipai/paperclip/pull/12870). These
affect different launch paths. This change preserves their credential
boundaries.

## What Changed

- Capture scoped task bindings after resolution, before managed provider
credential injection. Mint and validate the names-only projection at
native dispatch before host inheritance. Legacy adapters retain their
previous environment limits.
- Strip user-supplied projection markers from agent, environment,
project, and routine config.
- Carry selected values through the Codex, ACPX, OpenCode, runnerd, and
Rust subprocess launch boundaries.
- Add selected names to native and ACPX Codex shell include lists. Keep
selected values out of command arguments, including selected bootstrap
values.
- Reject malformed projections, reserved authority and loader names,
missing values, null bytes, and oversized input.
- Replace a retained native process when projected values change. Keep
unchanged processes reusable.
- Fingerprint custom namespaced variables while excluding known
generated runtime variables.
- Document next-run behavior and add regression coverage.

## Verification

- Red: the permanent reproduction failed at four launch/tool boundaries
and the namespaced fingerprint check. Two control checks passed.
- Green: the initial regression plus existing Codex environment and
shell tests passed (39 tests).
- Server config resolution, fingerprints, and native-session suites
passed (653 tests), including addition, rotation, removal, and unchanged
warm-session reuse.
- Rust regression tests passed. A real shell child received added and
rotated values, then lost them after removal. Unselected host variables
stayed absent.
- `pnpm -r typecheck` and `pnpm build` passed after rebase. The full
`pnpm test:run` was attempted but could not complete: fresh embedded
PostgreSQL databases fail during bootstrap on this macOS host. An
isolated suite and a disposable native `initdb` probe reproduced the
failure before test execution. `shmget` reports `No space left on
device` because the host has exhausted shared-memory IDs. This is not
disk exhaustion. The run was stopped after confirming the external setup
failure. CI results will be recorded separately.
- Runner boundary suites: 361 tests passed. Two process-launch errors
during concurrent binary staging passed on isolated rerun.
- Rust Codex provider and process supervisor integration suites: 97
passed, 2 intentionally ignored.
- ACPX shell and selected-bootstrap argv regressions: 3 failed before
the fix, then all 55 relevant tests passed.
- Review compatibility regression: 129-variable and large-value legacy
configurations failed before the correction and passed after moving
native-only validation to dispatch.
- Final review head `3f11e8d25`: repository typechecks and production
build passed. CI completed its implementation checks; one general-server
shard hit SQL `40P01` in `heartbeat-runtime-skills.test.ts` during its
`beforeEach` table truncate (1,199 tests passed in that shard). The
shard passed on its single rerun. All CI checks for this head are green.
Greptile reviewed this head at 5/5 with no new actionable findings; the
compatibility thread is resolved.
- No live provider credentials or model calls are required by these
tests.

## Risks

- Configured task credentials now reach the tools they were configured
for. The controller selects names only after existing scope and
secret-binding authorization.
- TypeScript and Rust validate the same bounded projection. Their
reserved-name rules must stay aligned.
- A changed projection replaces an idle provider process. Unchanged
values preserve reuse. Active turns retain their original environment.
- No database migration or API schema change.

> This fixes existing runner configuration behavior. It does not add a
new roadmap capability.

## Model Used

- OpenAI GPT-6 (Codex), with reasoning, local code execution, and
repository tools. The session does not expose a more specific API model
identifier or context window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-07 10:16:49 -05:00
DottaandPaperclip 0830737030 feat(ui): add CSV table previews to task file tabs (#15445)
## Thinking Path

> - Paperclip helps people manage AI agents and inspect their work.
> - Task file tabs show the files that agents produce.
> - Markdown has a readable view, but CSV files show only source text.
> - People need to scan CSV reports without a separate app.
> - This pull request adds a table view to attachment and workspace file
tabs.
> - View buttons let people inspect the original source beside download.

## Linked Issues or Issue Description

**Subsystem affected**

ui/ — task file tabs.

**Problem or motivation**

CSV exports are difficult to scan as source text. Related work:
https://github.com/paperclipai/paperclip/pull/14297 added text
attachment tabs.

**Proposed solution**

Render CSV as a table by default. Add row numbers, sticky headers, and
record counts. Keep rendered and raw icons beside download in task tabs.

**Roadmap alignment**

This is a small improvement to existing task file previews. It does not
add a new core subsystem.

## What Changed

- Add a shared CSV parser and table preview.
- Handle quoted commas, escaped quotes, UTF-8 BOM, CRLF, and multiline
cells.
- Show at most 500 data rows and 100 columns. Show a notice for larger
files. Keep the complete download.
- Add CSV attachment and workspace previews. Put workspace tab view
controls in the header.
- Add parser and interaction tests, Storybook examples, and product
documentation.

## Verification

- 28 focused tests pass across the parser, attachment tabs, and file
viewer.
- UI typecheck, UI build, and token gates pass.
- Playwright checked desktop and mobile rendered → raw → rendered flows
using the real attachment preview component in Storybook. Screenshots
are recorded on the driving task.
- Repository-wide typecheck and build were attempted. Both require
cargo, which is absent in this environment.
- The earlier repository-wide local test run stopped after a server test
race failure. The prior commit passed all GitHub CI jobs.
- Review fixes add source-line navigation coverage and CSV boundary
tests. All 22 file-viewer tests pass. UI typecheck and token gates pass.
- All 54 active GitHub checks pass on the final commit. Two Storybook
checks are skipped. Greptile gives 5/5 with no actionable findings.
- The final version preserves HTML previews from master and uses the
shared toolbar control. All 35 focused tests, UI typecheck, UI build,
and token gates pass.

## Risks

- The first CSV record supplies column headers. Files without headers
use that same rule.
- The table preview is bounded. Download and raw view retain the
complete file.
- No API, database, or dependency changes.

## Model Used

- OpenAI Codex, GPT-6. The runtime does not expose a more precise model
ID or context window. Used code editing, shell tools, and browser
checks.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-07 09:17:18 -05:00
DottaandPaperclip 1640c5b6ab feat(ui): render HTML artifacts in a secure sandbox (#15447)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents deliver reports as task attachments and workspace files.
> - The board already renders Markdown, but it shows HTML reports as
source or rejects their preview.
> - Reports can need inline scripts to build charts and tables.
> - Artifact scripts must not read board cookies, storage, or the parent
page.
> - This pull request adds an opaque-origin HTML preview and view
controls beside Download.
> - Users can explore a report, inspect its source, and download the
original file.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

The task attachment panel and workspace file viewers.

**Subsystem affected**

ui/ and the server workspace file preview service.

**Current behavior**

HTML attachments show source text. Workspace HTML files cannot be
previewed.

**Proposed behavior**

Render self-contained HTML reports in a sandboxed iframe. Place eye and
code controls beside Download. Preserve the original file for raw view
and download.

**Reason and benefit**

Users can read and filter agent reports in Paperclip without giving
artifact scripts access to the board session.

**Breaking changes**

HTML previews now open rendered. Scripts and styles must be embedded in
the report. Remote resources, API requests, forms, popups, and host
navigation are blocked. Workspace HTML content uses the existing bounded
UTF-8 JSON response.

**Additional context**

Related closed proposals: #3293 and #4857. This change uses the current
attachment and workspace viewers and adds no report-serving endpoint.
The Artifacts and Work Products roadmap item is complete; this improves
its existing preview behavior.

## What Changed

- Add a shared HTML iframe renderer with `sandbox="allow-scripts"` and
no `allow-same-origin`.
- Install a restrictive CSP before artifact markup. Keep inline report
scripts and styles.
- Add shared rendered/raw icon controls beside Download in the
attachment panel, workspace panel, and file sheet.
- Return workspace HTML as bounded text inside JSON. Keep download and
path-access protections.
- Add Storybooks for reports, security probes, workspace viewers, a task
journey, and mobile layouts.
- Document the security boundary and browser test command.

## Verification

- `pnpm -r typecheck` and `pnpm build` pass.
- Token gates and the static Storybook build pass.
- Four browser tests pass. They cover report filters, raw mode, task
entry, workspace viewers, cookies, storage, host DOM, resource requests,
forms, popups, and host navigation.
- The focused UI suite passes 26 tests. The file-resource server suite
passes 36 tests.
- Local full-stack acceptance passes with an actual 52 KB HTML report.
Its chart and filter render. Raw mode preserves the source. Download
bytes match the original file.
- The complete local UI suite passes: 698 files and 7,754 tests.
- `pnpm test:run` was attempted. Its server group recorded two unrelated
timeouts and eight connector failures. All failed cases pass in isolated
reruns, including the complete 388-test connector suite. The remaining
local run was stopped after the full CI test coverage passed, to release
test database resources. The full local command did not complete
cleanly.
- The separate local CLI group passes 511 tests. Three worktree database
cases fail to start embedded PostgreSQL because this Mac has exhausted
its shared-memory allocation. One isolated rerun fails at the same
database startup step. No CLI code was changed. The corresponding CI
test jobs pass.
- All 56 PR checks are clean on
`c771255f1a2e86bd8825f22eafb5d14a5e342cce`. Greptile gives 5/5 with no
actionable findings. The security scan passes. The branch has no merge
conflict.
- Review the **HTML artifacts** section in Storybook. Run `pnpm exec
playwright test --config tests/html-preview/playwright.config.ts` for
the browser checks.

## Risks

- Never add `allow-same-origin` to this iframe. Its opaque origin is the
cookie, storage, and host-DOM boundary.
- Reports that depend on CDN assets need to embed their dependencies.
- The frame can navigate itself. Sandbox restrictions remain after
navigation. This is not full network isolation.
- Inline scripts can consume browser resources. This change does not
isolate CPU or memory use.
- No database migrations or API shape changes are required.

## Model Used

OpenAI Codex, based on GPT-6, with reasoning, repository editing, code
execution, and browser tool use. The runtime does not expose a more
specific model ID or context window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass — feature and UI checks
pass; the full local run has the infrastructure limits recorded above.
All CI test jobs pass.
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-07 09:01:05 -05:00
DottaandPaperclip 8cfedd7df8 fix(runner): restore task monitors and durable timed waits (#15446)
## Thinking Path

> - Paperclip manages AI agents and their task execution.
> - Agents need a durable way to return to work after a delayed check.
> - The issue monitor scheduler already provides a one-shot wake for an
assignee.
> - Native runners reject generic execution-policy writes and had no
bound monitor tool.
> - Scheduling alone is insufficient because native completion also
needs to accept a timed wait.
> - This pull request adds an authorized monitor tool and connects it to
completion and the existing scheduler.
> - An agent can now schedule its next check, end the run, and resume on
the same task.

## Linked Issues or Issue Description

**What happened?**

A native runner could not set its own task monitor. `call_api` correctly
rejected execution-policy writes, while `schedule_wake` had no
production binding. `paperclip_finish` also rejected monitor waits.

**Expected behavior**

A standard native run can set a one-shot monitor on its current task or
another accessible task assigned to the same agent. After a confirmed
schedule on the current task, it can yield. The scheduler later delivers
`issue_monitor_due`.

**Steps to reproduce**

1. Start a standard native task.
2. Ask the agent to check the task again later and end its current run.
3. Inspect available tools and try the generic issue execution-policy
update.
4. Observe the missing native tool and the lifecycle-write denial.

Related PRs: #14680 concerns monitor notes in the shared wake prompt.
#11919 changes attempt-limit scope. This PR adds native scheduling and
completion authority and retains the existing cumulative attempt bounds.
It does not depend on either PR.

## What Changed

- Add provider-neutral `set_task_monitor` with a default current-task
target, future timestamp, required notes, existing bounds, and explicit
clearing.
- Check company, task visibility, ownership, runtime permissions, work
mode, and active-run authority. Preserve review-only restrictions.
Reject the reserved server-owned quota-recovery name before saving or
accepting a native wait.
- Commit the monitor, audit event, and retry receipt together. Retry
receipts survive a successor run without re-arming cleared or consumed
timers.
- Permit `paperclip_finish` to yield to a persisted monitor. Recheck
ownership and the schedule when committing final disposition. Release
execution without an immediate continuation.
- Preserve due monitors during native execution. Fence wake admission
and consumption against replacement, clearing, reassignment, and
completion. Preserve unrelated review policy.
- Expose scheduled and consumed monitor instructions in task context.
Update provider schemas, Rust validation, generated contracts, and
execution documentation.
- Add an opt-in live Codex smoke script with isolated data and explicit
run/session/runner/process evidence.

## Verification

- Repository `pnpm -r typecheck` and `pnpm build` passed after rebase.
Server typecheck passed again after review fixes. All CI test shards
pass on `3def77b1b`, including runner TypeScript/Rust, server,
serialized server, workspace, and browser tests. All CI gates are green,
including the canary dry run. Greptile is 5/5 on the same commit with
zero unresolved threads.
- The local monolithic `pnpm test:run`, started before the rebase, was
interrupted after current-head CI test coverage passed. It is not
counted as a standalone full-suite pass; the focused local regression
suites passed.
- Targeted server tests cover scheduling, replacement, clearing, policy
preservation, cumulative bounds, cross-run retries, permissions,
provider-neutral discovery, review restrictions, completion authority,
and scheduler/finalizer races.
- Runner contract/catalog/semantic tests and Rust terminal-tool tests
cover the new operation and monitor completion.
- Live Codex test passed twice (latest live run on `e0bcd63e6`) in a
temporary database and workspace, with a 300,000 ms warm window. First
run `67bd7709-c089-4d4a-9d2b-0d6b618a34b0` yielded at
`2026-10-07T13:15:03.274Z`. Second run
`97c323f9-595a-4cc5-a007-db5a2fbb937c` started at `13:15:30.952Z`,
received `issue_monitor_due`, and completed the same task. Exactly one
monitor wake was recorded.
- Both live runs used native session
`7b1dd753-1c9b-4e7a-b22f-a125dbc3748c`, runner
`e90d9a1b-3502-4ee3-b15e-edc024c555d4`, provider session
`01a11680-6d01-70c0-9a55-db7246ed66c3`, and PID `64218` with the same
process start time. This proves warm reuse for that local Codex test,
not only successful scheduling.
- Reproduce the paid live test with `node --import
./server/node_modules/tsx/dist/loader.mjs
server/scripts/smoke-native-task-monitor.ts --run`, with the installed
Codex binary on `PATH` and a valid local login.

## Risks

- The scheduler now defers monitor dispatch while the task has an active
native run. A stuck run still depends on the existing recovery
lifecycle.
- Idempotency uses the existing run ledger; no table or migration is
added.
- Other providers share the tested tool and completion contracts. Only
Codex received a live model test.
- Existing `call_api` lifecycle restrictions remain enforced. Monitor
waits do not bypass task blockers, reviews, or approvals.

## Model Used

OpenAI Codex, GPT-6 family, with tool use, code execution, and
TypeScript/Rust editing. The session does not expose the exact deployed
model ID or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.1007.0-canary.9
2026-10-07 08:41:24 -05:00
DottaandPaperclip b315580640 feat: add private task creation and sharing controls (#14718)
Add private task creation and sharing controls, effective access explanations, private project management, and locked task references.

Constrain the mobile composer plus menu to the viewport and explain named inherited privacy on hover, focus, or tap. Include 125 production-component Storybook stories and thirteen interactive user journeys.

Final head passes all CI checks and Greptile Apex 5/5 with no findings.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
canary/v2026.1007.0-canary.8
2026-10-07 07:27:52 -05:00
DottaandPaperclip abbd88007f fix(hermes): keep managed instructions out of resumed user turns (#15439)
Deliver managed instructions through Hermes's native system overlay while keeping current wake and runtime identity in user turns. Add fresh/resumed regression coverage and include Hermes in the default CI test roster.

Fixes #15385

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-07 07:07:32 -05:00
DottaandPaperclip 3a726e676f feat: enforce private task permissions across execution and data (#10633)
Enforce private task and project access across direct reads, search, execution, files, plugins, live delivery, and sharing mutations. Preserve downward-only sharing, current responsible-user authorization, and audited emergency access. Bind historical draft assets with migration 0314.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-07 07:06:40 -05:00
DottaandPaperclip b67db12d90 feat: add durable storage for private tasks (#14717)
Add private task and project ownership, downward access grants, and immutable run/workspace provenance. Apply migration 0313 with bounded batch commits and concurrent indexes. Keep the enforcement and sharing changes in their dependent PRs.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
canary/v2026.1007.0-canary.7
2026-10-07 06:53:20 -05:00