Commit Graph
1623 Commits
Author SHA1 Message Date
DottaandPaperclip 1bd40b69e2 fix(hermes): release assigned skill copies after failed state save
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 12:27:37 -05:00
DottaandPaperclip a2643eb538 fix(runner): preserve Dot v6 while versioning Hermes input as v7
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 12:07:09 -05:00
DottaandPaperclip 8cf6f8335d fix(hermes): publish and enforce native answer limits
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 12:07:09 -05:00
DottaandPaperclip fd054630f1 fix(runner): bound encoded Hermes attachment messages
Count serialized message, attachments, and metadata before turn admission. Keep TypeScript and Rust on a 7 MiB encoded content allowance within the existing encrypted frame bound. Validate before active-turn state changes, and prove accepted images and escaped documents traverse the secure PRP command frame.

Validation: 28 attachment/permission tests, two encrypted-frame regressions, the Rust encoded-admission regression, and Runner TypeScript compilation passed.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 12:07:09 -05:00
DottaandPaperclip 6d68b96bc5 fix(hermes): preserve planning authority and native transcript results
Qualification exposed denied plan workflows, missing structured tool output, and cancellation replay in subsequent turns. Admit assigned control-plane operations through their existing semantic authority, project denied native calls by identity, and keep managed history restoration separate from transcript replay.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 12:07:09 -05:00
DottaandPaperclip 63ec34ee79 fix(hermes): require the resolved image lock digest
Require the caller-verified target lock checksum instead of an incompatible checked-in default. Document lock-bot ownership and the companion browser qualification PR.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 12:07:09 -05:00
DottaandPaperclip 7d89e304f0 fix(acpx): separate restored history from active turn events
Keep history reconstruction without publishing restored text and tools as live output. Version the changed Cursor contract as profile 16 and retain replay compatibility.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 12:07:09 -05:00
DottaandPaperclip 68b1a39459 fix(hermes): keep native candidate independently restorable
Reject replaced learned-state parents, include authenticated run attachment and the advertised native transport fixtures in this candidate. Keep paid browser campaign definitions in the qualification PR.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 12:07:09 -05:00
Dotta c834f9dd8a fix(hermes): release credentials and reproduce relocatable runtime assets 2026-10-08 12:07:09 -05:00
DottaandPaperclip 9e07748e91 feat(runner): add native Hermes ACP qualification candidate
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 12:07:09 -05:00
Dotta 79e5d38139 fix(routines): remap annotations and verify atomic schedule edits 2026-10-08 02:27:42 -05:00
DottaandPaperclip 49fdc5899e feat(runner): bind routine management to the live service
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 02:27:42 -05:00
Devin FoleyandPaperclip e4b39da6f6 Isolate ACPX admission deadline tests from lease ports (#15527)
Apply the reviewed change for Isolate ACPX admission deadline tests from lease ports.

Validation: required local typecheck and tests, passing CI, Greptile 5/5, and independent review of the exact source head.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-07 19:20:08 -07:00
Devin FoleyandPaperclip ed0c6958f8 Explain refused local Hermes gateway connections (#15520)
Apply the reviewed change for Explain refused local Hermes gateway connections.

Validation: required local typecheck and tests, passing CI, Greptile 5/5, and independent review of the exact source head.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-07 18:55:55 -07:00
DottaandPaperclip dd777f4b73 feat(dot): complete onboarding and expand governed Runner capabilities (#15414)
## Thinking Path

> - Paperclip manages AI agents, their work, and their permissions.
> - Paperclip Runner supplies the same admitted tool authority to each
provider.
> - The Dot provider in #15402 needs reliable onboarding and useful
agent capabilities.
> - An idle Dot could not start work, assign a human, read skills, or
produce a workspace artifact.
> - Pairing also relied on a second configuration save before ordinary
admission could work.
> - This pull request adds governed idle admission and shared Runner
tools, and completes pairing atomically.
> - Operators can test Dot with fresh data while keeping normal company,
approval, budget, and run ownership checks.

## Linked Issues or Issue Description

**Subsystem affected**

Paperclip Runner, dedicated Dot MCP access, OAuth onboarding,
experimental settings, and empty worktree startup. This PR builds on
merged provider PR #15402. It reuses the merged MCP gateway from #14846
and assistant connection work from #14933 and #15380.

**Problem or motivation**

An idle Dot could see assignments but could not act on a conversation
request until someone created a task first. Its Runner catalog could not
assign tasks to humans or read pinned skills and workspace files.
First-time OAuth discovery and pairing also needed browser fixes, and a
completed pairing did not persist its binding reference on the agent.

**Proposed solution**

Keep Dot within the existing Runner. Admit a visible agent-authored
intake task for idle requests. Add people, human assignment, cross-task,
skill, and optional sandbox workspace tools through shared authority.
Relay assigned app calls through the configured MCP gateway. Add lease
renewal and follow-up references. Save pairing and its configuration
revision atomically. Show prerequisites and provide a complete copy
prompt.

**Roadmap alignment**

This extends the experimental provider in #15402. It uses the existing
governed gateway, task model, skills, artifact path, and native Runner.
It adds no separate execution subsystem.

## What Changed

- Add a standalone OpenAI Dot agent choice with an independent
experimental opt-in. It works with the general Runner option off.
Require Assistant connections (MCP), authenticated sign-in, and public
HTTPS for pairing.
- Use Dot’s own option for company package import. Allow an unpaired Dot
configuration to save after external billing acknowledgement; task
admission still requires pairing. Prepare the shared dev binary for
Dot-only opt-in.
- Label saved Dot agents as OpenAI Dot. Use prerequisite-check copy in
setup and runtime configuration.
- Route Dot creation directly to pairing after explicit external billing
acknowledgement. Hide local CLI, model, and harness setup for Dot.
Preserve the shared Runner implementation and canonical API type. Other
Runner providers still require their general opt-in.
- Add an empty worktree option with fresh signing keys and no production
data copy.
- Fix public-client OAuth negotiation, discovery compatibility, and
optional separate browser authorization origin.
- Add one-use pairing consent preview, clear copied setup instructions,
and atomic binding persistence and cleanup.
- Add idle request admission with stable request IDs and normal
scheduling, permissions, budgets, and task ownership.
- Add identity and people discovery, human task assignment and
reassignment, and authorized cross-task comments and documents.
- Read assigned skill files from pinned manifests. Relay assigned app
calls through the merged gateway without exposing credentials.
- Add an off-by-default workspace bridge. Constrain paths and writes.
Run commands in a deny-by-default OS sandbox with no network or injected
credentials. Reserve mutations before effects and never blindly repeat
uncertain work.
- Add rolling lease renewal, task pagination, bounded operation limits,
and deduplicated follow-up references without comment bodies in
webhooks.
- Add an off-by-default attachment reading setting. Restrict reads to
files on the current assigned task. Verify size and hash, cache bounded
verified copies per run, paginate text or binary bytes, and recheck live
authority before returning.
- Keep file grants operator-owned. Reject agent self-grants across
configuration routes. Preserve attachment consent in create/import
forms. Close generic API file bypasses while retaining current-run
response snapshots and permitted uploads.
- Keep provider limits explicit. Do not inherit a Dot binding,
attachment permission, or workspace permission when hiring another
agent.
- Include the required Markdown format in cross-task document writes and
validate the API title limit. Verify real creation and revision
persistence.
- Restrict command execution to Linux bubblewrap with descendant
containment. macOS retains workspace file tools and artifact publishing,
while refusing command calls. Explain the platform limit in setup.
- Exclude Paperclip instance state from workspace files, uploads,
artifact publication, and sandbox commands. Protect nested directories
and case variants. Fail closed when the directory protection scan
exceeds 4,096 directories.
- Add static UI compression for slow public tunnels. Document setup, the
complete tool inventory, and qualification limits.

## Verification

- Merge preparation on `00ca2c75b` integrates merged base #15402 and
master `fc6304dfe`. The ancestry commit preserves the reviewed follow-up
source tree. The subsequent security fix excludes instance state from
file tools, uploads, artifact publication, and sandbox commands. It
preserves private task authorization and task monitors. It regenerates
the combined tool catalog, seeded catalog digest, and protocol manifest.
Workspace typecheck, full build, and UI token gates pass. All sixteen
real Dot broker cases and 93 company import cases pass after the review
fixes and creator-attribution test correction. The exported
human-assignment catalog and Unicode page boundaries are also fixed. All
ten catalog tests and nine workspace/skill bridge tests pass. All 35
workspace bridge and authority tests passed after the instance-state
fix, including real macOS commands in that intermediate version. The
subsequent document and descendant-containment fixes pass 44 focused
tests across bridge, authority, and setup UI, with six Linux command
cases skipped on macOS. Real cross-task documents pass the route
validator and persist two revisions. macOS refuses command execution and
does not advertise the tool. Full workspace typecheck and build, changed
server/UI typechecks, and token gates pass again on the final commit.
Greptile rates final head 00ca2c75b 5/5 and its completed check
concludes success; all review threads are resolved. All 58 current-head
checks are complete: 54 passed and 4 intentionally skipped. This
includes full tests, typecheck, build, Runner, browser E2E, release
verification, and Canary Dry Run. The older full local root attempt
finished with 16,602 passes, 93 skips, and ten failures across two
suites: it began before source edits and retained earlier imported
implementations while loading later tests. Both suites pass a clean
final-head rerun (25 passed, six Linux-only command cases skipped on
macOS). This mixed-source full attempt is not a final-head full-suite
pass; use the fresh CI evidence below.
- Full workspace typecheck and build pass for the standalone Dot change
on head `392d54f80`. UI token gates are clean. The full workspace
typecheck passed before the final importer UI edit; the changed UI
typecheck, build, and token gates pass again on the final commit. The
server build passes for its unchanged final source.
- Standalone selection, creation, settings, harness visibility,
inventory, and runtime admission checks pass. OAuth onboarding and the
real Rust/PostgreSQL broker suite pass 27 cases with the general Runner
option disabled. Dot and MCP revocation still block pairing; other
Runner providers remain disabled. Create, hire, conversion, and
inherited-hire route regressions pass 68 cases. The runtime selection
suite passes 20 cases. The setup UI suite passes 50 cases. Company
import and dev binary checks pass 97 cases. Import UI checks pass 29
cases, including Dot-only selection, preservation of imported Dot
agents, generic Runner fallback, canonical configuration serialization,
and required billing acknowledgement. A read of the synthetic instance
reports the native binary required with Dot enabled, the general rollout
disabled, and no persisted native work.
- The latest attachment, operator-consent, generic API file-guard, and
bridge checks pass 80 tests. Cache and native lifecycle checks pass 550
tests; create/import builders pass 45 tests; attachment setup UI checks
pass 16 tests. The real Rust/PostgreSQL Dot broker suite passes 15
cases.
- Earlier authority, human assignment, pinned skill, confinement,
cancellation, inheritance, onboarding, replay, OpenAPI, and gateway
regressions pass their focused rechecks. Workspace writes reject
concurrent stale hashes.
- In the live empty test-drive, turn the general Runner option off and
leave Dot and MCP on. Add agent shows a separate OpenAI Dot choice.
Create rejects missing billing acknowledgement, saves the canonical Dot
configuration, and opens pairing. The existing paired Dot remains ready
and passes its prerequisite checks. No new plugin pairing was needed for
this UI change. The final browser import preview keeps a Dot source
agent as OpenAI Dot, falls back a generic Runner agent while its switch
is off, and offers Dot independently.
- Real Dot completed OAuth, signed MCP Events readiness, and an
event-only assigned document task in an empty synthetic instance.
Pairing saved its binding without a second configuration save.
- Real Dot verified human assignment, cross-task comments and documents,
pinned skill reading, sandbox commands, lease renewal, follow-up input,
and a downloaded artifact whose bytes and hash matched its receipt. An
idle conversation request created an intake, created a task for its
human owner, continued after a definite missing-file read error, and
finalized Done with exit code 0.
- After operator approval, real Dot read a synthetic assigned-task
attachment and wrote its file-only random proof into an agent-authored
document. Its first test required accepting review because the existing
document tool removed the final newline.
- On head `2626c8f9b`, real Dot read a 13,849-byte synthetic attachment
in two pages, used `expectedSha256` on page two, saved exactly its final
random marker, verified readback, and finalized Done with exit code 0. A
direct inbox check was required for this attachment qualification; it
does not claim event-only delivery.
- Previous head `392d54f80` passes all 55 checks (53 passed, two
intentionally skipped), including the full test, typecheck, build,
native Runner, browser, and release verification gates. Greptile rates
this head 5/5; all review threads are resolved.
- Head `2626c8f9b` passed all 56 checks: 54 passed and two intentionally
skipped. This includes full typecheck, build, native Runner, server
tests, serialized suites, browser E2E, and Canary Dry Run. The unchanged
Cursor managed-runtime test exceeded its five-second deadline on the
first attempt; its focused local suite passed all eight cases, and the
single CI rerun plus dependent verify gate passed.
- A prior full local serial test attempt reported 19 failures (16,141
passed, 88 skipped) and stopped before later wrapper groups. It began
before the final source edits. Every failed suite has a passing fresh
recheck; the macOS snapshot stress case passes alone in 227 seconds. The
previous head `000fb9261` passed all 56 CI checks. PRs leave
`pnpm-lock.yaml` unchanged; CI and the refresh bot own dependency
resolution.

## Risks

- Dot remains off by default. Its own opt-in does not enable other
Runner providers. It still requires the MCP option; disabling Dot blocks
new work while keeping existing recovery and saved bindings.
- Dot requires a stable public HTTPS origin. Its base provider in #15402
is merged. Temporary tunnels are useful for testing but are not
permanent deployments.
- Idle intake creates a visible agent-authored task. It does not
fabricate a human message or bypass ordinary admission.
- Workspace commands require Linux bubblewrap with a private PID
namespace; deployment qualification is still needed. macOS file tools
and artifact publishing remain available, but commands are disabled
because sandbox-exec does not contain detached descendants. There is no
unrestricted fallback. Command protection fails closed above 4,096
workspace directories. The historical macOS command walkthrough does not
qualify the current Linux implementation.
- Provider model choice, token usage, cost, native thread control, and
global external stopping remain unavailable. Known Paperclip budget
gates still apply.
- The workspace bridge and attachment reader are separate opt-in
settings. Reading assigned task files sends their contents to OpenAI.
Revocation prevents future reads but cannot withdraw bytes already sent.
Files are read on request; automatic inbound attachment staging remains
disabled.
- An existing event registration can retain earlier instructions that
prohibit extra tasks. Dot asked for permission before a second intake
after the final event-only bootstrap; that extra no-nudge continuation
is not live-qualified under that earlier registration. The later
attachment qualification submitted normally through the composer. The
revised onboarding prompt describes the new idle entry point, and its
protocol polling path is tested.
- OpenAI event delivery can be delayed; a webhook acknowledgement is not
proof that Dot has begun work.
- A running Dot can retain an old plugin catalog after tool refresh.
Refresh and actual tool exposure must be checked before using new
top-level actions.
- Hosted and remote controller modes are not qualified.

## Model Used

OpenAI Codex, based on GPT-6. The exact deployment ID and context window
size are not exposed in this session. Capabilities used: reasoning,
repository editing, code execution, test inspection, and browser
control.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either linked existing issues or described the issue in-PR
following the relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused suites and all
fresh failure rechecks pass; the earlier full attempt is disclosed
above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-07 19:44:17 -05:00
DottaandPaperclip fc6304dfe5 feat(runner): add experimental OpenAI Dot provider over MCP Events (#15402)
## Thinking Path

> - Paperclip manages AI agents, tasks, permissions, and execution
budgets.
> - Paperclip Runner gives each provider the same admitted task and tool
authority.
> - OpenAI Dot runs outside the local process tree and needs
asynchronous work delivery.
> - The merged MCP gateway supplies OAuth consent and signed event
delivery.
> - A personal assistant grant cannot safely stand in for an assigned
agent.
> - This pull request adds a separate Dot agent connection and a durable
Rust Runner bridge.
> - The operator can assign work to Dot and inspect its accepted work,
tool receipts, and result.

## Linked Issues or Issue Description

**Agent or provider**

OpenAI Dot, as an experimental provider of the existing Paperclip Runner
adapter.

**Why this adapter is useful**

An operator can assign normal Paperclip tasks to an existing Dot. Dot
can read its mailbox, request work on an assigned task, use admitted
task tools, and submit a result. Paperclip keeps company scope,
checkout, approvals, known budget limits, and activity attribution.

**How the agent is invoked**

A dedicated `/mcp/runner` OAuth resource pairs one Dot grant with one
agent. A signed MCP mailbox event wakes Dot. Dot explicitly accepts the
assignment. The Rust Runner owns the durable turn and operation
receipts. The first release supports self-hosted instances with a local
Runner controller.

**Additional context**

This extends the merged public MCP gateway from #14846 and the assistant
invitation and device-consent work from #14933. This also integrates the
merged assistant tool and configuration expansion in #15380. Dot retains
its dedicated agent resource and cannot receive personal configuration
permission. The public assistant connection remains a personal
connection.

## What Changed

- Add a durable Rust Dot provider and its TypeScript Runner driver.
- Add closed PRP v3 external-provider operations and native execution
input v6.
- Add company-scoped pairing, mailbox, assignment, and operation
records.
- Reuse merged browser/device consent, client metadata verification,
webhook admissions, refresh, secret rotation, and warm-standby gates.
- Keep Dot scopes, issuer, grants, event workers, and tool access
separate from personal assistant access.
- Add Dot configuration, pairing, readiness, and consent UI. Keep agent
grants out of the personal Connections entry.
- Regenerate the Dot-only migration after master. Preserve published
gateway migrations. Make the new migration safe to reapply.
- Document setup, recovery, accounting limits, evidence, and remaining
account qualification.
- Reverify reconnect callbacks and wake outstanding work with a fresh
mailbox reference; preserve the existing assignment and operation
receipts.
- Clean up Dot bindings and waiting runs on OAuth revoke and
refresh-token replay. Old grants cannot revoke replacement bindings.
- Restore the pairing reference when an unsaved agent form is reopened;
document board-only pairing routes in OpenAPI.
- Accept a clean Rust exit after the acknowledged shutdown receipt.
Unexpected exits still require recovery.
- Clear the cached binding after a successful revoke so a failed
connection refresh cannot restore it.
- Add production-component Storybook states and screenshots for pairing
and connection review. All preview account data is synthetic.
- Persist normalized completion, serialize Dot turns and durable work
admission, and poll subscription readiness.
- Serialize mailbox writes and cursor reads; retain paused fence
acknowledgement without task authority.
- Authorize admitted review runs without changing the worker assignee.
Include the fenced assignment ID in production stop notices.

## Verification

- This PR integrates master `4a8178e9c`. Dot migration
`0317_messy_famine.sql` follows the published history and is safe to
reapply. The merge preserves the reserved migration connection,
batch-commit handling, private task checks, task monitors, and native
accounting.
- Local workspace typecheck, full build, and UI token gates pass. The
server typecheck passes after the review fixes. Database and native
executor regressions pass.
- All twelve real Rust/PostgreSQL Dot integration tests and twelve Dot
driver tests pass. The tests cover native document writing and
finalization, durable replay, queue admission, mailbox ordering,
admitted reviews, stale authority, production stop references, and
paused acknowledgements.
- Current head `d0e7e0626` passes all 57 checks: 53 pass and four are
intentionally skipped. This includes full typecheck, build, tests, Rust
Runner verification, browser E2E, release verification, and Canary Dry
Run. Greptile rates this exact head 5/5. All review threads are
resolved.
- The full local root test run is slower than the sharded CI run and has
not completed. The full CI test gates pass on the current commit.
Focused local regressions pass.
- Real-account pairing and event delivery on this base commit remain
unqualified. Live account and setup proof are recorded in the follow-up
#15414.

The following screenshots use synthetic preview data. They show the
production pairing component and do not qualify a real account or the
full agent setup journey.

![Synthetic pairing
preview](https://github.com/paperclipai/paperclip/blob/codex/dot-events-prototype/doc/screenshots/openai-dot-runner/pairing.jpg?raw=true)

![Synthetic connected
preview](https://github.com/paperclipai/paperclip/blob/codex/dot-events-prototype/doc/screenshots/openai-dot-runner/connected.jpg?raw=true)

## Risks

- This base adapter uses `PAPERCLIP_ENABLE_OPENAI_DOT=1` plus Public MCP
and Paperclip Runner. The separate experimental-settings follow-up in
#15414 replaces this environment flag with saved operator settings.
- Dot does not expose provider token usage or cost. The operator must
acknowledge external billing. Known Paperclip budget gates still apply.
- Cancellation fences Paperclip authority. It does not confirm that Dot
stopped all external activity.
- Assigned skill files and third-party MCP bindings are unsupported and
reject admission. There is no mounted workspace, model selector, or
provider thread identifier.
- Hosted agent-broker and remote controller deployments are not
qualified.
- The new migration follows the merged master history. Existing
prototype databases still need the normal master migration history
before this Dot-only migration.

## Model Used

OpenAI Codex, based on GPT-6. The exact deployment ID and context window
size are not exposed in this session. Capabilities used: reasoning,
repository editing, code execution, and test inspection.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #123` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub references)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-07 19:11:50 -05:00
DottaandPaperclip a7a244ab33 feat: add company decision models with permission and cost controls (#15473)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Optional product features need small, typed model decisions.
> - Each company needs to choose which shared API connection pays for
those decisions.
> - Calls must retain user and task permissions, budget limits, and cost
attribution.
> - This pull request adds a managed decision service, setup UI, and
request history.
> - Features can check availability cheaply and keep their existing
behavior when decisions are unavailable.

## Linked Issues or Issue Description

**Subsystem affected**

Server services, shared contracts, database accounting, Company
Settings, and Costs.

**Problem or motivation**

Paperclip has no common decision-model service. Adding provider calls
within each feature would duplicate credential access, permission
checks, and billing rules.

**Proposed solution**

Let a connection manager configure one company decision model. Support
OpenAI Decisions and Jev through OpenRouter. Provide a fixed setup test
and metadata-only history. Default company-sponsored background
decisions to on during setup, and preserve a saved off setting.

**Alternatives considered**

Per-feature credentials would duplicate existing connection management.
Personal overrides and provider fallback chains add permission and
billing complexity; they remain deferred.

**Roadmap alignment**

Reviewed ROADMAP.md and searched open PRs. This extends existing
connection access and budget accounting. Product features that call the
service remain outside this change. No matching decision-model service
PR was found.

## What Changed

- Add company settings, an internal `decisionModelService`, local
availability checks, and trusted human, agent/run, and system contexts.
- Pin Vercel AI SDK provider dependencies and adapt boolean, choice, and
ordered-score decisions for both providers. Bound requests and time;
disable paid retries.
- Add durable invocation metadata and agentless decision ledger charges.
Preserve fractional cents, pricing evidence, dispatch identity, and
unresolved billing holds.
- Share the company accounting lock and apply company, agent, and
project budgets. Settle charges once, retain unknown holds, and recover
interrupted calls without resubmission.
- Reuse connection setup and management UI. Add a Decisions view under
Costs, production-component Storybook coverage, database migration, and
service documentation.

## Verification

- Passed 166 current-code tests covering the decision service/provider,
setup component, Costs, OpenAPI, and every failure from the earlier
broad run. Coverage includes native SDK wire formats, refusals, billed
malformed responses, permission and secret-rotation races, identity
changes, concurrent budget admission, unresolved holds, agent/task
deletion, and stale setup feedback.
- Passed 170 existing connection, cost, budget, heartbeat-accounting,
and profile regression tests.
- Passed repository typecheck, production build, and design token gates
after integrating master. Verified the generated migration on a fresh
test database and upgraded the populated preview database from the
branch's earlier migration without losing settings or usage.
- Ran the required full `pnpm test:run`: its general phase completed
with 16,354 passed and 10 failures across five files while this branch
was still being updated. Every reported failure passes in the
current-code rerun; the serialized phase did not run after that failure.
The full GitHub CI suite passed on `b1b856a88`: general and serialized
tests, browser shards, runner checks, typecheck, production build,
packaging/canary, and policy gates. [CI
evidence](https://github.com/paperclipai/paperclip/actions/runs/37673106360).
- Passed the full-shell Storybook setup-to-history interaction test
again after integrating master. Greptile rates the final revision 5/5
with zero unresolved threads.
- Walked through the running app: empty setup, add each provider, save,
reload, run all three sample questions, inspect fractional charges in
history, switch provider while sponsorship is off, disable, and
reconnect. Checked mobile settings. These tests used the actual UI,
vault, server, SDKs, and database with simulated upstream responses.
- Live paid setup tests remain unverified: this environment has no
authorized OpenAI/OpenRouter credentials available. No mocked test is
presented as live provider evidence.

Reviewer journey: Company Settings → General → Decision model.
Add/select a shared API connection, save, run the billed sample, open
View usage, then disable decisions and verify Run test is disabled after
reload.

## Risks

- The SDK decision interface is experimental. Pinned versions and
wire-format tests limit upgrade drift.
- The migration allows agentless service charges and reservations.
Existing agent cost-reporting APIs still require an agent, and decision
receipts stay separate from run reconciliation.
- Timeouts can have unknown provider charges. Holds remain until an
audited accounting correction resolves them.
- OpenAI prices use a versioned Decisions rate snapshot; OpenRouter
costs use provider receipts. Unknown pricing is retained as unknown.
- A configured company authorizes background spending by default. Setup
explains this, and managers can turn it off.
- Live provider account/model availability still needs the two
credentialed acceptance checks.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository inspection, code
execution, and browser tools. The exact serving revision and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-07 15:53:13 -05:00
Devin FoleyandPaperclip f1c44b7b56 Preserve remote work when required workspace restore fails (#15479)
Persist exact source-retention obligations before run finalization and protect them across cancellation, restart and task changes. Require board-authorized repair evidence without replaying old work or changing current task ownership, state or locks. Hide an unavailable retry action and document operator recovery and retention costs.

Validated with 674 scoped regressions, 224 combined integration tests, full typecheck/build, independent safety reviews and green CI with Greptile 5/5. No historical file recovery is claimed.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-07 13:29:44 -07:00
Devin FoleyandPaperclip cd40ebb95b Preserve ACP bridge terminal evidence before disconnect (#15478)
Drain terminal frames through backpressure and record bounded terminal evidence without waiting for log persistence. Guard later events and input failures against writes after socket end.

Validated with 217 focused tests, full typecheck/build, independent review and green CI with Greptile 5/5.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-07 13:01:22 -07:00
DottaandPaperclip fd8c6b920a fix(native): resume connection tasks after approval decisions (#15471)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Native tasks can pause while a human decides whether to allow a
connection action.
> - The next turn needs both the saved decision and complete accounting
for the previous turn.
> - Cancellation could discard final usage, and complete direct Claude
API receipts could remain unpriced.
> - A stale blocked or review report could also request approval again
after the original card was declined.
> - This pull request retains shutdown accounting and rejects approval
waits bound to an already resolved action.
> - The benefit is reliable continuation with the existing budget and
approval controls.

## Linked Issues or Issue Description

Refs #15420. Related: #15312 addresses requester ownership during
dispatch. This change addresses receipt capture and final-response
validation.

## What Changed

- Retain usage events after native cancellation. Continue to reject late
provider messages and work.
- Drain same-turn accounting and terminal events for at most fifteen
seconds after a durable governed wait. Keep incomplete accounting
blocked.
- Estimate complete, unpriced, direct Anthropic API receipts for the
exact `claude-sonnet-5` model. Record the rate version and assumptions.
Use the one-hour cache-write rate when the receipt lacks cache TTL.
- Bind stale approval reports to exact interaction, action-request, or
invocation IDs in the same company, task, agent, and run. Cover blocked,
review, and response-wake reports. Keep independent reviews valid.
- Fence checkpoint and result writes after a controller detaches for
restart, including operations waiting for a database lock. Reject stale
successful returns before certifying accounting.
- Allow bounded subscription teardown only after retaining an actual
provider terminal.
- Journal the exact governed-wait trigger and disposition before
provider interruption. Recover that wait independently of a later saved
answer, replay retained accounting, and reject mismatched or unproven
terminal evidence.
- Give settling governed turns a bounded window before shutdown detaches
their controller.
- Keep fuzzy external app matches alongside installed capability matches
instead of forcing an unrelated provider question for a generic query.
- Return up to twenty exact active catalog tool names after an invalid
request, after eligibility checks; still reject the request without
granting access or creating an approval.
- Require retained provider terminal proof before settling a governed
wait, including when complete usage arrives before stream
closure/error/timeout. Retain harmless numbered cancellation events so
restart replay stays contiguous.
- Isolate accounting-test OpenCode config from the host plugin
directory.
- Add regression coverage and document the accounting, restart and
connection-search behavior.

## Verification

Current PR source: `0cf08efd75f8fb23f7989beda6dbda92587088bf`. The live
matrix below measured frozen `09a776bcb1f77422156f8be11e13f8c29f43e7f7`;
later review fixes are verified separately and do not relabel those
runs.

- Repository typecheck and build pass.
- Current runner runtime and cancellation suites: 191 tests pass. Six
new regressions cover stream end/error/timeout without provider-stop
proof and contiguous cancellation acknowledgement/request replay; all
six failed before the fix. The existing bounded cleanup case now
explicitly supplies terminal proof. All 31 adapter accounting tests pass
with isolated fixture config.
- New checkpoint-rebinding and approval-criterion suites: 65 tests pass;
runner HTTP integration: 29 tests pass. Unchanged executor/control-plane
suites: 640 tests pass; database-backed connection suites: 68 tests
pass.
- Current retained evidence verification covers 222 file hashes across
all fifteen original result artifacts. The three-case [published
report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37651896525-1/index.html)
and all eight screenshot hashes verify. The [three-case recovery
report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37662587356-1/index.html)
and all six screenshots also verify; the nine-result campaign did not
publish.
- Evaluation support: 1,805 Vitest tests pass, one skipped; 128 Node
tests pass. Eval typecheck and catalog discovery pass.
- Full local repository unit run was interrupted before the follow-up
edits after three tool-access failures and one runner HTTP failure.
Those failures pass in isolation; the 09a full run was interrupted after
one rapid Slack callback-ordering failure and seven skill-service
failures. All eight pass both isolated and with full-runner environment
settings, and all 74 skill-service tests pass together; the subsequent
full run reported two 15-second OpenCode accounting timeouts and was
stopped with exit 130 to apply review fixes. The timeouts reproduce
while copying this host’s 61 MB OpenCode config. All 31 tests pass after
isolating config inside each fixture without increasing timeouts or
changing assertions. A complete local full-suite pass is not claimed.
Repository-wide CI also passes on the final review-fix head in [run
37669185952](https://github.com/paperclipai/paperclip/actions/runs/37669185952).
Earlier database-skipped diagnostics and the older ENFILE run are
retained and are not full-suite passing evidence.
- Three-case live campaign
[37651896525](https://github.com/paperclipai/paperclip/actions/runs/37651896525)
passes all three original grades on frozen source 09a: 55/55 checks,
eight succeeded run records, complete accounting receipts, matching
checkpoint identities and no pending approvals. Campaign
[37653533353](https://github.com/paperclipai/paperclip/actions/runs/37653533353)
adds nine original passes (173/173 checks, eighteen succeeded records)
on the identical source. Its other three jobs failed before runner
assignment or any step while GitHub could not load the paid environment;
those original infrastructure failures are retained. Campaign
[37662587356](https://github.com/paperclipai/paperclip/actions/runs/37662587356)
completes only those unstarted cells: all three original grades pass
(57/57 checks, six succeeded run records). All fifteen exact cases now
pass on source 09a: eight FAIL → PASS, seven PASS → PASS, zero new
overall failures and zero pending pairs. Total current evidence: 285/285
checks and thirty-two succeeded run records, complete accounting
receipts, matching checkpoint identities, no pending approvals or retry
records. Earlier campaigns retain forty-seven additional run records and
two known same-run recovery attempts; actual provider-call counts and
invoices remain unknown. The nine-result campaign skipped publication
and its public URL returns 403; original artifacts remain retained.
Previous ba9 campaign
[37645656840](https://github.com/paperclipai/paperclip/actions/runs/37645656840)
completed 2 PASS / 1 FAIL: Claude restart/approval and Codex decline
pass, while OpenCode resumes but times out searching for exact tool
names and never creates the access card. That failure and incomplete
cancelled-run accounting remain preserved; the new catalog error
guidance targets this observed dead end. Campaign
[37642957312](https://github.com/paperclipai/paperclip/actions/runs/37642957312)
remains 0 PASS / 3 FAIL and exposed the now-corrected cross-run marker
leak and unknown-criterion approval gap. The original baseline remains 7
PASS / 8 FAIL, first repair 2 PASS / 3 FAIL, and second repair 0 PASS /
3 FAIL. All fifteen selected cases are qualified by their original
grades in this bounded trial. These live grades belong to 09a. Its
twelve governed-wait checkpoints retain matching same-turn terminal
fingerprints, but passing artifacts omit detailed event journals; the
later six adversarial regressions qualify the new terminal-proof and
replay guards separately. Final-head repository CI passes. [Fresh
Greptile
review](https://github.com/paperclipai/paperclip/pull/15471#issuecomment-6044369780)
is 5/5, confirms both findings are fixed, and reports no new actionable
issues. All review threads are resolved and the PR has no merge
conflicts.

## Risks

- Governed cancellation drains accounting for up to fifteen seconds.
Restart detachment gives a settling batch up to twenty seconds to
finish. An incomplete receipt or unproven provider terminal still
prevents successful qualification.
- Claude prices are estimates, not invoices. The estimate assumes
standard global API pricing and uses a conservative cache-write rate.
Unsupported models, billers, and billing modes remain unpriced.
- Approval identity matching must remain scoped to the current run and
the requested approval. It does not authorize execution of a declined
call.
- Original baseline and final candidate have different merged master
context. Exact-case outcomes are before/after observations, not isolated
causal attribution to this repair.
- No schema migration, fixture, oracle or grader change. Search-result
guidance now treats fuzzy external matches as suggestions. Existing app
authorization and provider-consent checks remain required.

## Model Used

- OpenAI GPT-6 through Codex, with code editing, terminal tools, and
test execution. The exact deployment identifier and context window are
not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (191 current runtime tests
and 31 accounting tests; interrupted full-suite history and CI coverage
are disclosed above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-07 14:43:07 -05:00
Devin FoleyandPaperclip 0514e8abd5 Keep Cursor managed-runtime test installation offline (#15484)
Intercept the Cursor installer in its local shell fixture and guard against accidental curl downloads. Preserve real managed-home archive operations and the default timeout; change no production code.

Verified 46 related tests, safe negative mutation, independent review and all exact-head CI. Greptile 5/5 with no unresolved comments.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-07 12:38:06 -07:00
Devin FoleyandPaperclip 465140596f Preserve bounded workspace sync diagnostics across RPC (#15481)
Preserve bounded error codes and HTTP/exit statuses across the environmentSyncOut worker RPC boundary, and revalidate that method-scoped envelope before attaching host restore diagnostics. Keep the original error and all recovery policy unchanged; do not transmit provider payloads or credentials.

Verified real RPC roundtrip and privacy regressions, 95 focused tests, 68 independent tests, full typecheck/build and all exact-head CI. Greptile 5/5 with no unresolved comments.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-07 12:22:53 -07:00
DottaandPaperclip ae6f95ed7a feat: add personal primary agents (#15470)
Add a personal primary agent per company and user. Initialize it from the first human-created agent, expose profile-only switching with confirmation, and use it after recent choices for task and Chat defaults.

Persist authenticated preferences, preserve lifecycle and membership rules, keep selections out of shared audit events, and document the API contract. Include the reviewed Storybook surfaces and regression coverage for concurrent choices, onboarding, cross-device updates, and browser journeys.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-07 13:53:46 -05:00
Devin FoleyandPaperclip 06484b3c41 fix: preserve conversation retries through execution cleanup (#15463)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Recovery schedules bounded retries after a provider disconnects.
> - A stopped run can still hold its environment lease while cleanup
runs.
> - Retrying before that lease is released cancels the new run before it
starts and spends another retry.
> - Restoring the task can also send the worker repair instructions from
an already resolved recovery action.
> - This pull request preserves the waiting retry and removes settled
recovery instructions from later wakes.
> - The task can continue after cleanup without an operator repairing
the same incident again.

## Linked Issues or Issue Description

**What happened?**

A legacy conversation run disconnected while it was doing ordinary work.
Its environment cleanup took longer than the retry delay. Two retries
were cancelled before dispatch with `execution_reconciliation_required`.
Those cancellations exhausted the failure budget. After an operator
restored the task, the wake still told the original worker to repair the
runtime and hand the task back to itself.

**Expected behavior**

Cleanup waits preserve the pending attempt. The same retry can continue
after ownership is released, subject to all current gates. Once a
recovery action is resolved or cancelled, subsequent task wakes omit its
repair instructions.

**Steps to reproduce**

1. Fail a legacy conversation run while its environment lease remains in
`pending_cleanup`.
2. Schedule a bounded retry and run promotion before cleanup releases
that lease.
3. Repeat the scheduler sweep. Before this fix, retries promote and then
cancel without starting.
4. Resolve a stranded-task recovery action and build the restored task
wake with that action ID. Before this fix, the wake still includes the
settled repair instructions.

**Paperclip version or commit**

Reproduced with database regressions against `ceabc3bc880` on master.

Related: #15019 restores a skipped assignment handoff after lease
release. #15235 filters stale handoff evidence in the recovery sweep.
This change preserves an existing scheduled conversation retry and
corrects restored wake content. It does not create a new handoff wake.

## What Changed

- Keep an unstarted legacy conversation retry on the same durable row
while prior execution ownership remains active. Recheck after 30 seconds
without increasing retry accounting.
- Return a queued retry to scheduled state if it encounters that hold at
the claim gate. Retain its issue claim and publish the status change.
- Record one local lifecycle diagnostic per blocking run. Remove that
wait marker on promotion.
- Include recovery action metadata only while the referenced action is
active or escalated.
- Return an explicit `waiting` response and the saved schedule when
Retry now meets cleanup. Show the wait inline without a false success or
disabled button.
- Add database, rendered-prompt, route, and UI regressions. Document the
execution and run-log contracts.

## Verification

- Five cleanup and restored-wake regressions fail against the original
production code. The Retry now route and UI regressions also fail before
their correction.
- Related retry, dispatch, stale-queue, and recovery suites: 321 tests
pass across seven files. All ten focused cleanup/restored-wake cases
pass after rebase. The final dispatch adjustment passes all 46 adapter
tests.
- Retry now routes and affected UI suites: all 46 tests pass. The tests
cover repeated clicks, the saved schedule, unchanged accounting,
promotion after release, and no false success or error state.
- `pnpm -r typecheck`, `pnpm build`, `pnpm build-storybook`, and `pnpm
check:token-gates` pass.
- Full local `pnpm test:run` was started and then stopped after the
final commit passed all GitHub CI test shards. No complete local
full-suite result is claimed; CI supplies the complete test result for
the final commit.
- Final head `6f1058a332c039e33c4f002b296d20a5554e760e`: all 55 GitHub
checks are green or intentionally skipped. Apex review is 5/5 after two
reviews, with no unresolved threads. The PR has no merge conflicts.

## Risks

The wait applies only to unstarted legacy conversation retries. Native
runs and non-conversation execution keep their existing recovery rules.
Cleanup must actually release ownership before execution can resume. The
wait does not fix a cleanup service that never finishes. Promotion and
dispatch still enforce cancellation, reassignment, pause, budget, and
reconciliation gates. No schema change is required. The Retry now
response adds a `waiting` outcome; the shared contract and all three UI
controls handle it.

## Model Used

OpenAI Codex based on GPT-6, with repository analysis, tool use, and
local code execution. The runtime does not expose the precise serving
model ID or context window.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused suites; complete
final-head test coverage in CI)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-07 11:15:49 -07:00
9fb955e9a2 fix(codex): gate models on the Codex CLI floor before the ChatGPT backend rejects them (#15399)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The Codex adapter and the native runner run Codex in local hosts and
in managed sandboxes, and the model catalog lists the models an operator
can select
> - The Codex backend accepts a new model with ChatGPT sign-in only from
a recent enough Codex CLI. An older CLI fails every turn with `The
'<model>' model is not supported when using Codex with a ChatGPT
account.`
> - After #14942 added `gpt-6.1-sol`, an operator selected it for an
agent in a managed Daytona sandbox. The sandbox image still shipped
Codex 0.156.0. The connection test and runs failed with that sentence,
which reads like an account problem
> - Paperclip only checked the Codex compatibility window (`>=0.149.0
<0.161.0`), so 0.156.0 passed and the failure surfaced from the backend
without a cause
> - This pull request records the verified Codex CLI floor for each
gated model, compares the installed `codex --version` with that floor in
the environment Test and in the remote runner, and names the backend
rejection when it still happens
> - The benefit is a precise, actionable message ("gpt-6.1-sol requires
Codex CLI 0.159.0 or newer; detected 0.156.0; promote a sandbox image
with Codex 0.160.0") instead of an opaque 400, and no doomed hello probe

## Linked Issues or Issue Description

Refs #14942

**Bug: selecting `gpt-6.1-sol` in a managed sandbox fails with a ChatGPT
account error.**

**Steps to reproduce**
1. Promote a sandbox image that ships Codex CLI 0.156.0.
2. Create a Codex agent with a ChatGPT sign-in connection, select
`gpt-6.1-sol`, and run the environment Test or a task in that sandbox.

**Expected behavior**
The Test or the run tells the operator that the Codex CLI in the sandbox
is older than the model needs.

**Actual behavior**
The Test reports `codex_hello_probe_failed` with
`{"type":"error","status":400,"error":{"type":"invalid_request_error","message":"The
'gpt-6.1-sol' model is not supported when using Codex with a ChatGPT
account."}}`. A run fails with "The selected model is not supported by
the current ChatGPT connection."

**Evidence**
- Codex 0.156.1 and older are rejected for `gpt-6.1-sol` with ChatGPT
sign-in: https://github.com/openai/codex/issues/49396
- Codex 0.159.0 and 0.159.2 are accepted:
https://github.com/openai/codex/issues/49464 and
https://github.com/decolua/9router/issues/4471 (same error, fixed by
raising the client version identity from 0.155.0 to 0.159.0)
- Codex 0.157.0 release notes: "Add GPT-6 Sol and Luna to the model
catalog": https://github.com/openai/codex/releases/tag/rust-v0.157.0
- GPT-6.1 Sol is included for Plus, Pro, Business, Enterprise, and Edu
with ChatGPT sign-in: https://learn.chatgpt.com/docs/models

## What Changed

- `packages/adapters/codex-local/src/index.ts`: add
`minimumCodexCliVersionForModel` (`gpt-6.1-sol` → 0.159.0, `gpt-6-sol`
and `gpt-6-luna` → 0.157.0, with the sources above),
`parseCodexCliVersionOutput`, `codexCliVersionAtLeast`, and
`CODEX_CHATGPT_MODEL_REJECTION_RE`. Models without a verified floor
return `null`, so nothing probes them. The agent configuration doc
describes the floors.
- `packages/adapters/codex-local/src/server/cli-version.ts` (new): run
`codex --version` where the run would execute it (local, SSH, or
sandbox), compare it with the model floor, and build the
`codex_cli_version_compatible` / `codex_cli_version_incompatible`
checks. Map the backend rejection sentence to
`codex_hello_probe_model_rejected` with the detected CLI version and a
hint that separates a stale CLI from a plan that does not include the
model.
- `packages/adapters/codex-local/src/server/test.ts` (CLI lane Test):
run the version check before the hello probe when the model has a floor.
Skip the hello probe when the CLI is too old
(`codex_hello_probe_skipped_cli_version`). When the probe still fails
with the backend sentence, report `codex_hello_probe_model_rejected`
instead of the generic `codex_hello_probe_failed`. The new codes do not
match the auth-failure patterns, so a managed connection is not
invalidated.
- `packages/adapters/codex-local/src/server/acp.ts` (ACP lane Test): for
remote targets, run the same version check against the shared `codex`
the ACP server spawns.
- `server/src/services/native-runtime/native-session-executor.ts`: after
the compatibility-window check, compare the remote Codex with the
configured model's floor and fail with
`runner_remote_provider_artifact_incompatible: <model> requires Codex
<floor> or newer with ChatGPT sign-in, received <version> from the
sandbox image; promote a sandbox image with Codex 0.160.0 or configure
PAPERCLIP_RUNNER_REMOTE_CODEX_NPM_SPEC=...`. When a preinstalled Codex
fails this check and an npm spec is configured, the existing fallback
installs the pinned release.
- Tests: adapter metadata, CLI-lane Test (too old, compatible, no floor,
backend rejection), ACP-lane Test (too old, compatible, no floor), and
the remote runner floor (nine version/model cases).
- `doc/adapter-model-audit-2026-10-02.md`: record the floors and the
reason.

No catalog entry, default model, saved agent configuration, pin,
lockfile, or image definition changes.

## Verification

Local (Node 25.9, pnpm 9.15.4 via corepack):

```sh
pnpm --filter @paperclipai/adapter-codex-local typecheck
pnpm --filter @paperclipai/adapter-codex-local exec vitest run            # 81 passed
pnpm --filter @paperclipai/server exec vitest run \
  src/services/native-runtime/native-session-executor.test.ts \
  src/services/native-runtime/codex-runtime-compatibility.test.ts -t Codex   # 84 passed
```

Also run locally: the full `native-session-executor.test.ts` file (533
passed) and the full Codex adapter suite (81 passed).

Follow-up commit `9fd4d3b25` (Greptile P2: the ACP-lane version probe
ignored the agent's configured env): the probe now receives the
adapter's string-valued `env` entries, the same ones
`buildCodexAcpConfig` hands to remote ACP runs, so a `PATH` override
selects the same `codex` for the Test as for the run. New test covers a
`PATH` + `CODEX_HOME` override and a dropped non-string entry. Re-run on
that head: `pnpm --filter @paperclipai/adapter-codex-local typecheck`
clean; full Codex adapter suite 503 passed (31 files).

Not run here: `pnpm -r typecheck` for `server` (the direct `tsc
--noEmit` was killed by the sandbox memory cap; the adapter package
typecheck passes and CI covers the server), `pnpm build`, browser
suites, and a live sandbox probe. No UI files changed, so token gates do
not apply.

Manual check for a reviewer: set an agent's Codex model to
`gpt-6.1-sol`, point its environment at a sandbox whose `codex
--version` prints `codex-cli 0.156.0`, and run the environment Test. The
result contains `codex_cli_version_incompatible` with `Detected Codex
CLI 0.156.0.` and no hello probe check. With Codex 0.160.0 the Test
contains `codex_cli_version_compatible` and runs the hello probe.

## Risks

- A floor is enforced for all authentication modes, because the runner
does not know the credential method at verification time. An OpenAI API
key on Codex 0.157.0 or 0.158.0 with `gpt-6.1-sol` is now rejected
before launch. Paperclip pins Codex 0.160.0 everywhere, so this only
affects installs that lag the pin, and the message names the fix.
- The 0.159.0 floor for `gpt-6.1-sol` is the oldest stable release
verified to work; 0.157.0 and 0.158.0 were not verified either way. If
OpenAI accepts an older client, the floor can be lowered in one table.
- The `codex --version` probe adds one short process run to the Test
only for models with a floor.
- Rollback: revert this pull request. No data or configuration migrates.

## Model Used

Claude Fable 5.1 (`claude-fable-5-1`), Anthropic, with extended
thinking, tool use, and web research, operating as a Paperclip agent
through Claude Code.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Bender (Fable) <noreply@paperclip.ing>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-07 10:37:55 -07:00
DottaandPaperclip f669194298 fix(runner): propagate configured environment to future turns (#15451)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Native runners start provider processes and enforce a separate tool
environment policy.
> - The server resolves task environment bindings at each run boundary.
> - Fixed launch allowlists dropped custom variables after resolution.
> - Warm process reuse and a blanket fingerprint exclusion also hid
configuration changes.
> - This pull request carries a bounded, server-selected list of task
variable names through each process boundary.
> - The benefit is that later turns and tool commands receive configured
values while ambient host secrets stay excluded.

## Linked Issues or Issue Description

**What happened?**

A native Codex agent could not use configured task credentials. Adding
the variables in Settings did not fix the next turn. The provider
sanitizer, runner launch, Rust provider launch, and shell policy each
used fixed allowlists. Effective config fingerprints also excluded all
variables with the `PAPERCLIP_` prefix.

**Expected behavior**

Explicitly bound task variables must reach provider processes and tool
commands. Added, changed, and removed values must take effect at the
next run boundary. A running turn keeps its original configuration.

**Steps to reproduce**

1. Start a native Codex task without a custom environment binding.
2. Add a fake `PAPERCLIP_PAGE_BUCKET` value and a fake Pages credential
binding in Settings.
3. Continue the task and inspect the tool environment.
4. Before this fix, those values are absent even from a fresh runner
launch.

**Paperclip version or commit**

Reproduced at `b31558064`. The fix is rebased on current master.

**Deployment mode**

Self-hosted server with a native process runner.

Related changes: [the legacy Codex MCP environment
fix](https://github.com/paperclipai/paperclip/pull/13321) and [ambient
server-secret
exclusion](https://github.com/paperclipai/paperclip/pull/12870). These
affect different launch paths. This change preserves their credential
boundaries.

## What Changed

- Capture scoped task bindings after resolution, before managed provider
credential injection. Mint and validate the names-only projection at
native dispatch before host inheritance. Legacy adapters retain their
previous environment limits.
- Strip user-supplied projection markers from agent, environment,
project, and routine config.
- Carry selected values through the Codex, ACPX, OpenCode, runnerd, and
Rust subprocess launch boundaries.
- Add selected names to native and ACPX Codex shell include lists. Keep
selected values out of command arguments, including selected bootstrap
values.
- Reject malformed projections, reserved authority and loader names,
missing values, null bytes, and oversized input.
- Replace a retained native process when projected values change. Keep
unchanged processes reusable.
- Fingerprint custom namespaced variables while excluding known
generated runtime variables.
- Document next-run behavior and add regression coverage.

## Verification

- Red: the permanent reproduction failed at four launch/tool boundaries
and the namespaced fingerprint check. Two control checks passed.
- Green: the initial regression plus existing Codex environment and
shell tests passed (39 tests).
- Server config resolution, fingerprints, and native-session suites
passed (653 tests), including addition, rotation, removal, and unchanged
warm-session reuse.
- Rust regression tests passed. A real shell child received added and
rotated values, then lost them after removal. Unselected host variables
stayed absent.
- `pnpm -r typecheck` and `pnpm build` passed after rebase. The full
`pnpm test:run` was attempted but could not complete: fresh embedded
PostgreSQL databases fail during bootstrap on this macOS host. An
isolated suite and a disposable native `initdb` probe reproduced the
failure before test execution. `shmget` reports `No space left on
device` because the host has exhausted shared-memory IDs. This is not
disk exhaustion. The run was stopped after confirming the external setup
failure. CI results will be recorded separately.
- Runner boundary suites: 361 tests passed. Two process-launch errors
during concurrent binary staging passed on isolated rerun.
- Rust Codex provider and process supervisor integration suites: 97
passed, 2 intentionally ignored.
- ACPX shell and selected-bootstrap argv regressions: 3 failed before
the fix, then all 55 relevant tests passed.
- Review compatibility regression: 129-variable and large-value legacy
configurations failed before the correction and passed after moving
native-only validation to dispatch.
- Final review head `3f11e8d25`: repository typechecks and production
build passed. CI completed its implementation checks; one general-server
shard hit SQL `40P01` in `heartbeat-runtime-skills.test.ts` during its
`beforeEach` table truncate (1,199 tests passed in that shard). The
shard passed on its single rerun. All CI checks for this head are green.
Greptile reviewed this head at 5/5 with no new actionable findings; the
compatibility thread is resolved.
- No live provider credentials or model calls are required by these
tests.

## Risks

- Configured task credentials now reach the tools they were configured
for. The controller selects names only after existing scope and
secret-binding authorization.
- TypeScript and Rust validate the same bounded projection. Their
reserved-name rules must stay aligned.
- A changed projection replaces an idle provider process. Unchanged
values preserve reuse. Active turns retain their original environment.
- No database migration or API schema change.

> This fixes existing runner configuration behavior. It does not add a
new roadmap capability.

## Model Used

- OpenAI GPT-6 (Codex), with reasoning, local code execution, and
repository tools. The session does not expose a more specific API model
identifier or context window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-07 10:16:49 -05:00
DottaandPaperclip 8cfedd7df8 fix(runner): restore task monitors and durable timed waits (#15446)
## Thinking Path

> - Paperclip manages AI agents and their task execution.
> - Agents need a durable way to return to work after a delayed check.
> - The issue monitor scheduler already provides a one-shot wake for an
assignee.
> - Native runners reject generic execution-policy writes and had no
bound monitor tool.
> - Scheduling alone is insufficient because native completion also
needs to accept a timed wait.
> - This pull request adds an authorized monitor tool and connects it to
completion and the existing scheduler.
> - An agent can now schedule its next check, end the run, and resume on
the same task.

## Linked Issues or Issue Description

**What happened?**

A native runner could not set its own task monitor. `call_api` correctly
rejected execution-policy writes, while `schedule_wake` had no
production binding. `paperclip_finish` also rejected monitor waits.

**Expected behavior**

A standard native run can set a one-shot monitor on its current task or
another accessible task assigned to the same agent. After a confirmed
schedule on the current task, it can yield. The scheduler later delivers
`issue_monitor_due`.

**Steps to reproduce**

1. Start a standard native task.
2. Ask the agent to check the task again later and end its current run.
3. Inspect available tools and try the generic issue execution-policy
update.
4. Observe the missing native tool and the lifecycle-write denial.

Related PRs: #14680 concerns monitor notes in the shared wake prompt.
#11919 changes attempt-limit scope. This PR adds native scheduling and
completion authority and retains the existing cumulative attempt bounds.
It does not depend on either PR.

## What Changed

- Add provider-neutral `set_task_monitor` with a default current-task
target, future timestamp, required notes, existing bounds, and explicit
clearing.
- Check company, task visibility, ownership, runtime permissions, work
mode, and active-run authority. Preserve review-only restrictions.
Reject the reserved server-owned quota-recovery name before saving or
accepting a native wait.
- Commit the monitor, audit event, and retry receipt together. Retry
receipts survive a successor run without re-arming cleared or consumed
timers.
- Permit `paperclip_finish` to yield to a persisted monitor. Recheck
ownership and the schedule when committing final disposition. Release
execution without an immediate continuation.
- Preserve due monitors during native execution. Fence wake admission
and consumption against replacement, clearing, reassignment, and
completion. Preserve unrelated review policy.
- Expose scheduled and consumed monitor instructions in task context.
Update provider schemas, Rust validation, generated contracts, and
execution documentation.
- Add an opt-in live Codex smoke script with isolated data and explicit
run/session/runner/process evidence.

## Verification

- Repository `pnpm -r typecheck` and `pnpm build` passed after rebase.
Server typecheck passed again after review fixes. All CI test shards
pass on `3def77b1b`, including runner TypeScript/Rust, server,
serialized server, workspace, and browser tests. All CI gates are green,
including the canary dry run. Greptile is 5/5 on the same commit with
zero unresolved threads.
- The local monolithic `pnpm test:run`, started before the rebase, was
interrupted after current-head CI test coverage passed. It is not
counted as a standalone full-suite pass; the focused local regression
suites passed.
- Targeted server tests cover scheduling, replacement, clearing, policy
preservation, cumulative bounds, cross-run retries, permissions,
provider-neutral discovery, review restrictions, completion authority,
and scheduler/finalizer races.
- Runner contract/catalog/semantic tests and Rust terminal-tool tests
cover the new operation and monitor completion.
- Live Codex test passed twice (latest live run on `e0bcd63e6`) in a
temporary database and workspace, with a 300,000 ms warm window. First
run `67bd7709-c089-4d4a-9d2b-0d6b618a34b0` yielded at
`2026-10-07T13:15:03.274Z`. Second run
`97c323f9-595a-4cc5-a007-db5a2fbb937c` started at `13:15:30.952Z`,
received `issue_monitor_due`, and completed the same task. Exactly one
monitor wake was recorded.
- Both live runs used native session
`7b1dd753-1c9b-4e7a-b22f-a125dbc3748c`, runner
`e90d9a1b-3502-4ee3-b15e-edc024c555d4`, provider session
`01a11680-6d01-70c0-9a55-db7246ed66c3`, and PID `64218` with the same
process start time. This proves warm reuse for that local Codex test,
not only successful scheduling.
- Reproduce the paid live test with `node --import
./server/node_modules/tsx/dist/loader.mjs
server/scripts/smoke-native-task-monitor.ts --run`, with the installed
Codex binary on `PATH` and a valid local login.

## Risks

- The scheduler now defers monitor dispatch while the task has an active
native run. A stuck run still depends on the existing recovery
lifecycle.
- Idempotency uses the existing run ledger; no table or migration is
added.
- Other providers share the tested tool and completion contracts. Only
Codex received a live model test.
- Existing `call_api` lifecycle restrictions remain enforced. Monitor
waits do not bypass task blockers, reviews, or approvals.

## Model Used

OpenAI Codex, GPT-6 family, with tool use, code execution, and
TypeScript/Rust editing. The session does not expose the exact deployed
model ID or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-07 08:41:24 -05:00
DottaandPaperclip abbd88007f fix(hermes): keep managed instructions out of resumed user turns (#15439)
Deliver managed instructions through Hermes's native system overlay while keeping current wake and runtime identity in user turns. Add fresh/resumed regression coverage and include Hermes in the default CI test roster.

Fixes #15385

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-07 07:07:32 -05:00
DottaandPaperclip 3a726e676f feat: enforce private task permissions across execution and data (#10633)
Enforce private task and project access across direct reads, search, execution, files, plugins, live delivery, and sharing mutations. Preserve downward-only sharing, current responsible-user authorization, and audited emergency access. Bind historical draft assets with migration 0314.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-07 07:06:40 -05:00
DottaandPaperclip b67db12d90 feat: add durable storage for private tasks (#14717)
Add private task and project ownership, downward access grants, and immutable run/workspace provenance. Apply migration 0313 with bounded batch commits and concurrent indexes. Keep the enforcement and sharing changes in their dependent PRs.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-07 06:53:20 -05:00
DottaandPaperclip 1ead554bd1 fix(claude-local): resume sessions across agent file working copies (#15437)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The Claude CLI adapter resumes task sessions between runs.
> - Each run now receives a private copy of the agent files.
> - The adapter put that copy's path in the cached system instructions.
> - The new run ID changed the prompt bundle even when all instructions
and skills stayed the same.
> - This pull request sends the current file location with each run's
prompt and keeps it out of the stable bundle.
> - Agents can resume unchanged task sessions and use the current
working copy.

## Linked Issues or Issue Description

Fixes: #15373

Refs #14420, which introduced the per-run agent directory copies.
Related PRs checked: #5699 changes the fingerprint algorithm, and #12034
handles saved sessions without a bundle key. Neither fixes this
working-copy path regression.

## What Changed

- Keep instruction and skill contents in the cached system prompt.
Supply the current instruction path and relative-file base in every run
prompt, including resumed turns and fresh retries.
- Report a cwd or execution-target mismatch only when that value
differs. A bundle mismatch no longer produces a false cwd warning.
- Add a four-run regression: initial run, relocated copy, changed
instructions, and changed skill contents. Check the CLI arguments,
bundle keys, current file guidance, and reset logs.
- Cover and explain remote-to-local execution resets, even when the
working directory matches.
- Update the agent-file documentation and existing resume/fallback
assertions.

## Verification

- **Red:** With only the new regression test added, the second run fails
because the CLI arguments do not contain `--resume`.
- **Green:** All 64 tests pass in the command below. This uses a fake
Claude subprocess and real adapter execution, file caching, and session
serialization; it does not call a paid model.

```sh
pnpm exec vitest run server/src/__tests__/claude-local-execute.test.ts packages/adapters/claude-local/src/server/execute.remote.test.ts packages/adapters/claude-local/src/server/execute.acp-fallback.test.ts server/src/__tests__/adapter-session-codecs.test.ts
```

- Repository-wide `pnpm -r typecheck` and `pnpm build`: passed. The
adapter typecheck, build, and all 64 focused tests also pass after the
review fix.
- Local `pnpm test:run` was stopped after it reported a failure in the
unchanged native-session recovery database orchestration test. That test
passes in isolation with PostgreSQL enabled (1 passed, 47 filtered out).
The full local run did not complete; this is not a clean local
full-suite result. All CI checks pass on
`a71d23a39f1cc874d23a8715cf29bdea6edbb8ff`, including the full test
shards.

## Risks

- A session saved with the old path-bearing bundle starts fresh once
after upgrade. Later runs resume when instruction and skill contents
stay unchanged.
- The current location now travels in the run prompt. Stable system
guidance directs relative file references to that location, and each
turn explicitly replaces earlier locations.
- No database, API, authentication, permission, or UI change.

## Model Used

OpenAI Codex (GPT-6), with reasoning, repository inspection, code
execution, and test tools. The exact deployment model ID and context
window are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-07 06:48:19 -05:00
DottaandPaperclip d0db8820db fix(connections): repair native baseline and approval continuations (#15420)
## Thinking Path

> - Paperclip manages AI agents and the tools they may use.
> - Connection setup separates provider preference from permission to
use a tool.
> - The first native connection baseline could not exercise its intended
decisions.
> - The browser used mutable task titles, and the provider fixture
already granted access.
> - Native provider-choice instructions also disagreed with the
preferred question format. Schema rejection gave no field guidance.
> - This pull request repairs those test preconditions and native
guidance, then fixes restart/approval defects exposed by the corrected
baseline. It also restores missing OpenCode tool-error evidence.
> - The benefit is an inspectable baseline before any further
instruction reduction.

## Linked Issues or Issue Description

Refs #15407.

The original 15-cell baseline remains 0 PASS / 15 FAIL. Ten cells
stopped on stale titles, two Codex cells had schema denials, two
OpenCode cells used already-granted tools, and one Claude cell returned
no native result. No intended user decisions were submitted. The exact
invalid Codex field and underlying Claude failure cause remain unknown.

[Original
campaign](https://github.com/paperclipai/paperclip/actions/runs/37562577199)
· [Original
report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37562577199-1/index.html)

## What Changed

- Match the browser's task route and visible identifier instead of a
title the agent can change.
- Start native provider-choice fixtures with no agent tool access.
Verify the public effective-access records.
- In the positive case, select Arcade, then grant its exact HubSpot tool
through the real access card. Require both saved decisions and exactly
one observed call.
- Return a canonical `providerQuestionSet` for native input and retain
the equivalent legacy `providerQuestion`.
- Keep invalid input rejected. Return bounded schema locations and
required field names without submitted values.
- Preserve Claude's exact session/content identity while allowing
authenticated registered instruction-copy paths to rotate on a new run.
- Reject duplicate approval reports for an existing exact tool-action
card before they create another human review.
- Wait for a recorded service-approval continuation within the existing
deadline; retain missing or failed continuation grades.
- Forward OpenCode tool activity through the runner facade, preserving
bounded errors and execution-part identity without inventing host-call
joins or exposing arguments.
- Preserve original grades, costs, scope limits and diagnoses in the
dated repair report.

## Verification

- Eval typecheck passes. Support suite: 1,801 PASS, one intentional
skip; Node checks: 128 PASS.
- Connection/schema tests: 51 PASS. Real-server public fixture setup:
one PASS with zero providers.
- Browser support regression: five PASS, including renamed and wrong
tasks.
- Focused Rust safe-feedback test: one PASS.
- Repository typecheck and build pass before the latest master replay.
Post-replay connection/shared/real-server fixture checks: 52 PASS; eval
typecheck passes. The browser review fix additionally passes all five
browser checks and seven suite checks.
- The full local repository run was interrupted incomplete after about
45 minutes, with five integration failures retained. All five pass in a
separate targeted invocation (1,250 unrelated tests skipped). No full
local-suite pass or root cause for the initial local failures is
claimed.
- Corrected frozen source `162cc90fdabe7f505b88ae095044531b82784c92`:
**10 PASS / 5 FAIL** across the [passing Codex
canary](https://github.com/paperclipai/paperclip/actions/runs/37575158761)
and [remaining 14
cells](https://github.com/paperclipai/paperclip/actions/runs/37576261807).
The canary passes all 17 checks. Claude's two provider-choice
continuations fail on restart, Claude service approval exposes an early
evaluator rejection, Codex service approval creates a duplicate
approval, and OpenCode provider-second times out after both decisions
with no HubSpot call. No original result is regraded.
- Final ledgers count 31 actual runs: 27 succeeded, two failed, two
cancelled during cleanup. All 15 cleanup/budget checks pass. The late
Claude continuation is absent from its earlier workflow snapshot; it
remains in the result/API/final ledger. Original evidence retains 279
hashes. Recorded LLM subtotal $0.04553787 is incomplete billing, not
actual total cost; local runtime is unmetered.
- New repair regressions reproduce the Claude attach failure, duplicate
approval acceptance and dropped OpenCode tool events before their
respective fixes. Nine Rust attachment checks, 127 ACPX host/adapter
tests, 33 completion/control-plane checks, nine eval deadline tests, 59
OpenCode proxy/driver tests, one Rust tool-error/redaction check, and
TypeScript/Rust composer parity pass. Eval typecheck, repository
typecheck and build pass. Existing support coverage is 1,802 PASS plus
128 Node PASS, one intentional support skip; two additional deadline
tests also pass.
- New-source full CI/review and live canaries are pending. The next
bounded selection is Claude provider-decline, Codex service-approve and
one OpenCode provider-second diagnostic with repaired event evidence. No
broader campaign or instruction-reduction qualification is claimed.
- Initial corrected campaign
[37574251834](https://github.com/paperclipai/paperclip/actions/runs/37574251834)
was cancelled during shared build after review found the breadcrumb
whitespace assumption. Its matrix job has zero steps and no provider
execution. The real adjacent-span browser regression now reproduces the
old failure and passes after the fix.

## Risks

- The corrected baseline remains 10/15. The new restart/approval fixes
require live qualification; OpenCode evidence forwarding does not itself
establish or fix its prior behavioral failure.
- The positive provider case now expects three runs, including separate
access approval. Its new results are distinct from the original invalid
fixture.
- The old Claude missing-result cause and rejected Codex field are
unknown. These repairs do not retroactively explain or erase either
failure.
- Path rotation must preserve prompt, custom instruction, skill/content
identity and protected provider settings; regression checks reject stale
or changed content. No connection authorization, JSON schema, budget,
cleanup, or final-result requirement is relaxed. Historical Everyday
prompts and gateway setup remain unchanged.

## Model Used

OpenAI Codex, GPT-6. The exact deployment variant and context window are
not exposed in this session. Used repository inspection, code editing,
test execution and retained-evidence analysis.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-07 06:18:51 -05:00
Devin FoleyandPaperclip 799e4d556f fix: make accounting durable and synchronize cost reporting (#14997)
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-06 22:26:48 -07:00
DottaandPaperclip 99a9de9940 fix(mcp): personalize assistant connections and hide revoked grants (#15411)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Assistant connections let people use their organization from another
assistant.
> - Setup begins with an invitation and returns to a list of connected
assistants.
> - Client-only labels obscure the authorizing person, while revoked
rows clutter that list.
> - This pull request shows the person’s avatar, names connections by
owner and client, and hides revoked rows.
> - It also shortens the copied invitation while keeping the approval
instructions.

## Linked Issues or Issue Description

Related: #14933 and #15380.

**What existing behavior does this improve?**

Invitation copy and assistant connection management, in the
organization’s Connections screen and the account-wide management page.

**Subsystem affected**

Cross-cutting: shared MCP connection types, server profile projection,
UI and Storybook.

**Current behavior**

Connections are labeled only with a client name such as “Codex.” Revoked
connections remain visible. The invitation includes an extra sentence
about agent identity.

**Proposed behavior**

Show the authorizing person’s avatar and use names such as “Dotta’s
Codex connection.” Hide revoked rows after successful revocation and
when loading retained revoked grants. Failed revocation leaves the
connection visible. Remove the extra identity sentence from invitation
copy.

**Reason and benefit**

Make connection identity clear and keep the list focused on usable
connections.

**Breaking changes**

The connection response adds optional `user` metadata with name and
image. Older servers remain usable. Names and revoked-row visibility
change in the UI; OAuth client identity, authorization and audit
retention remain unchanged.

## What Changed

- Shorten the shared invitation text.
- Project the authorizing person’s name and avatar through the
user-scoped connection endpoint, without returning email or credentials.
- Reuse the existing Identity component and owner naming conventions
across the connection page, catalog card and account-wide list.
- Hide revoked grants and remove a successfully revoked row from the
shared cache, even if the subsequent refresh fails.
- Update documentation, regression tests and production-page Storybook
fixtures and revocation journeys.

## Verification

- `pnpm -r typecheck`, `pnpm build`, `pnpm build-storybook` and `pnpm
check:token-gates` pass. Final UI type checks also pass.
- Focused consent and connection UI tests: 31 pass, including company
filtering, legacy metadata, custom client names, user-initial fallbacks,
failed revocation, retained revoked rows and refresh failure after
successful revocation.
- Existing MCP regression suite: 6 pass; 71 database checks are skipped
locally because embedded PostgreSQL cannot start on this machine. The
added database check verifies user-profile isolation and retained
revocation history; CI runs these checks.
- Browser verification with Storybook fixtures: owner avatar and name
render; revocation removes the selected row in both production pages,
leaves other connections visible, and restores the empty state after the
last revocation.
- All 54 current-head CI checks pass, with two optional Storybook jobs
skipped. CI includes database, browser, runner, typecheck, build and
clean-install canary coverage.
- Greptile reviewed commit `1368d79e1066b418712224378d89d64c2b11cb86`:
5/5, no actionable findings or unresolved threads.
- The full local `pnpm test:run` was stopped after complete CI passed.
Local database coverage remains unavailable because embedded PostgreSQL
cannot start; no full local-suite pass is claimed.

## Risks

Low risk. The additive profile field is optional for compatibility.
Revoked grants are filtered only from management UI and retained for
audit. Revocation failure does not hide an active connection. No
authorization scopes, token handling or schema changes.

## Model Used

OpenAI GPT-6 through Codex, with code editing, command execution and
browser verification. The exact deployment ID and context-window size
are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 22:09:41 -05:00
DottaandPaperclip a9a20fb5c6 feat(security): add read-only customer-success inspection APIs (#15405)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents need a persistent identity and a verified active run for
governed access.
> - Customer-success inspection needs broad reads without tenant writes
or secret access.
> - Ordinary board login and database credentials give more authority
than this task needs.
> - This pull request adds a dedicated inspection API and strict
managed-run authority.
> - Cloud owns short grants, human approval, replay protection, and
audit records.
> - The benefit is inspectable access that an operator can disable
immediately.

## Linked Issues or Issue Description

Refs: #15352. This change reuses the persistent Ed25519 identity from
that PR.

**Subsystem affected**

Server authentication, pure resource readers, and the shared wire
contract.

**Problem or motivation**

One internal Paperclip agent must inspect customer onboarding work. It
must not receive owner login or database credentials. Reads must not
create customer sessions, memberships, activity, or read receipts.

**Proposed solution**

Add disabled-by-default run authority and versioned tenant inspection
endpoints. Require strict instance-bound managed-run JWTs on the home
instance. Require exact-operation, single-use Cloud permits on tenants.
Execute a reviewed company-scoped catalog in read-only transactions.
Cloud applies seven-day stack-age eligibility and human exceptions.

**Roadmap alignment**

This is access support for Cloud deployments and governed agent
identities. Bot creation, scheduling, scoring, and reports are separate
work. The maintainer requested this implementation.

## What Changed

- Reuse existing public identity reads and managed private-key
injection. Reject unprovisioned keys, paused agents, ended runs, legacy
signatures, and wrong instances.
- Mount `/api/customer-success/v1` before actor/session synchronization.
Verify Cloud permits and consume them centrally before reading.
- Add explicit company-scoped database readers and bounded instruction,
skill snapshot, run log, workspace, and asset reads. Preserve existing
redactions and file protections.
- Add protocol, security, database immutability, and managed-agent
qualification tests. Add deployment and rollback documentation.

## Verification

- Full `pnpm -r typecheck` and `pnpm build` passed. Server typecheck
passed after review fixes.
- The broad local `pnpm test:run` recorded 14,277 passes and four
failures in unchanged suites: two timeouts and two PR-metadata mock
assertions. All three affected suites passed on isolated reruns (36
tests). The complete CI matrix passes at the final head, including every
test lane, typecheck, build, runner checks, canary dry run, and the
security scan.
- Focused inspection, JWT, and existing identity tests pass. The catalog
test compares every public database table before and after reads.
- Inspection and route-contract tests: 22 passed. The coordinated test
runs a real managed process agent against separate home/customer
PostgreSQL databases and a PostgreSQL broker over HTTP. It proves wake
through the existing controller, bounded binary file reads, single
challenge consumption across replicas, concurrent grants with a
two-connection pool, scoped SQL audits, append-only runtime auditing,
one-year retention, and unchanged tenant data/files.
- Run the coordinated test with `PAPERCLIP_INSPECTION_CLOUD_DIST`
pointing at the sibling Cloud build. Normal unit runs skip that optional
private integration.
- Final-head Greptile is 5/5 with no unresolved findings.
- No production deployment or customer inspection occurred.

## Risks

- This adds an authentication boundary. Keep both feature flags disabled
until coordinated staging and canary qualification.
- Cloud support must deploy after this API. Unsupported tenants fail
closed. There is no owner-login or database fallback.
- Existing redactions remain the content boundary. Arbitrary pasted
secrets in readable prose or files may remain.
- Remote files and suppressed provider traces remain unavailable. Wake
can cause normal startup/background writes; test those separately.
- Disable Cloud policy first during rollback. Preserve existing identity
material and Cloud audit history.

## Model Used

OpenAI GPT-6 (Codex), with reasoning, code execution, and browser
testing. The session does not expose a more specific deployment ID or
context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused checks and
isolated reruns; broad-run flakes are documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 21:38:15 -05:00
DottaandPaperclip a6306ba606 feat(runner): consolidate Cursor production integration (#15075)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The native Runner keeps provider sessions under company authority,
approvals, budgets and durable recovery.
> - Cursor work was spread across candidate branches. The published
branch lacked later plan, permission and cleanup fixes.
> - Production also needs public installation and matching runtime
assets for local and Daytona execution.
> - This pull request consolidates Cursor onto current mainline recovery
behavior and completes that installation path.
> - The installed v11 release passed focused local and Daytona
qualification after the generic mode and lifecycle cleanup. The later
model-selection correction and current mainline merge produce v14
artifacts that need matching release qualification.
> - Cursor admission is enabled in source; publish only an artifact
combination with matching qualification. Native AskQuestion and complete
per-run dollar accounting remain excluded.

## Linked Issues or Issue Description

Refs: #14435, #14631, #14669, #14699, #14724.

This completes the Cursor implementation by @cryppadotta from combined
source `22c78242a4e0c2369fecf0c2dc4e7600fbad6706`. It preserves newer
mainline recovery, completion and warm-directory behavior. Pi and
Copilot remain gated.

## What Changed

- Generate named Rust and TypeScript ACPX release profiles from one
manifest. Share runtime pins with packaging and server verification.
Preserve vendor runtime versions; bind the updated ACPX patch to Cursor
profile v14 and reject stale generated declarations at build/typecheck.
- Remove ACPX model allowlists, including the former Codex and Pi
restrictions and the duplicate developer test-drive gate. Send any
explicit model ID unchanged to its provider and verify the effective
selection before prompting. The bundled ACPX package forwards unlisted
IDs, rejects mismatched acknowledgements, and restores the exact
selection after session load. It does not expand Cursor model aliases.
Provider rejection, mismatch, or missing model controls fails without a
fallback. Model examples live in evaluation fixtures, outside runtime
declarations.

- Add pinned Cursor execution, contained instructions, exact model
verification and Agent/Plan/Ask modes.
- Carry an opaque generic `mode` identifier in shared native execution,
sidecar, Rust and recovery contracts. The provider adapter owns
supported modes, defaults, native translation and acknowledgement.
- Keep native RPC recognition, accepted-plan interpretation and
permission evidence behind provider adapters. Shared settlement and
recovery verify normalized facts and their committed evidence.
- Replace the Cursor-only warm-attachment branch with a runner-owned
capability. Only Cursor opts into it. Move profile compatibility and
optional usage parsing into provider metadata and adapters.
- Write generic plan-wait receipts. Read exact historical Cursor
receipts through a separate compatibility decoder. Reject mixed formats
and preserve existing authority checks.
- Carry native plans, semantic questions, todos, child activity,
permission identities and partial usage diagnostics through the Runner.
- Preserve durable response delivery, cancellation, warm ownership and
process retirement.
- Finish accepted planning runs successfully. Keep their tasks open for
explicit direction. Acceptance does not start implementation.
- Ship `paperclipai runtime setup cursor` and its provisioner through
the public package. npm installation does not download Cursor. Setup
uses the OS account's closure-keyed cache so system-wide npm packages
can remain read-only. Run it as the Paperclip service account.
- Include Cursor in normal provider packs and Daytona images for macOS
ARM64/x64 and Linux x64.
- Reject stale release packs by source revision and current ACPX/Cursor
pins before assembly writes files. Verify current Cursor
version/profile/closure again at runtime.
- Ship all three daemon targets and the expected Linux image-pack
identity. A macOS controller uses its packaged Linux daemon for Daytona.
Image mismatches fail before provider launch.
- Use the vendored Runner boundary for installed readiness probes.
Verify the actual installed Cursor probe.
- Verify compiled public Daytona plugins and their release versions in
installed smokes.
- Record exact artifacts, the acceptance matrix, retained failures,
supported capabilities and rollback behavior in the [readiness
report](https://github.com/paperclipai/paperclip/blob/codex/cursor-production-readiness/doc/plans/2026-10-03-cursor-production-readiness.md).

## Verification

- Current head `f7ec5cc1f0e30c62a829c016ff2013a2a9d79037` merges
mainline `faa8e452c73bae5e044dd6379179a00106abb131`. It keeps Cursor
plan and cancellation guards alongside mainline historical-question
filtering. The evaluation catalog includes both Cursor and expanded
adapter accounting cases (683 total). Recursive typecheck, full build,
696 lifecycle/recovery tests, 45 fixture tests and fixture typecheck
passed. Current-head CI passed: 56 successful checks, one neutral and
four skipped. [Complete
CI](https://github.com/paperclipai/paperclip/actions/runs/37557996535).
The fresh Base Greptile review is 5/5 on this exact head, with 304 files
reviewed, zero new comments and zero unresolved threads. The user
authorized overriding the CODEOWNER review gate after checks passed; no
failing checks are overridden. Prior results below retain their own head
identities.
- Corrective head `3d2b168366258036f6b6a6fccb382c49138cc601` fixes the
post-merge Apex finding. Automatic-review and new-evidence
reconciliation preserve pending child results and recheck delivery under
the status lock before completing. Account repair now excludes unrelated
secret consumers and requires the failed agent's identity. Regression
coverage includes the commit race, delivery statuses,
current-run/current-intent exclusions, repeated reconciliation, both
database reconciliation paths, and credential consumer boundaries. All
184 affected tests, server typecheck and server build passed.
Current-head Base Greptile review is 5/5, with 304 files reviewed, zero
new comments and zero unresolved threads. Current-head CI passed: 56
successful checks, one neutral and four skipped. [Complete
CI](https://github.com/paperclipai/paperclip/actions/runs/37535994724).
This Base review is distinct from the earlier Apex review.
- Merge head `5957c257a` reconciles mainline `b508a05c4`. It preserves
both accepted-plan waits and pending-child-completion checks, current
provider selectors, task-creation response identities, and mainline ACPX
missing-file handling. The combined patch is bound to Cursor profile
v14; historical records keep their original identities.
- Merge head `5957c257a` passed recursive typecheck, full build, 43
installed ACPX/package contracts, 107 provider UI and plan/recovery
tests, 593 database-backed lifecycle tests, 49 profile/native contract
tests, 45 Product E2E fixture tests, fixture typecheck, token gates,
three provider-free browser task-creation cases, and Runner
conformance/replay checks. Its complete CI passed (55 successful checks,
one neutral and four skipped), while Apex returned 2/5 with a
child-delivery finding addressed below.
- The local full-suite attempt again failed the unchanged Git streaming
test (360-second timeout) and was stopped. The concurrent local Rust
attempt failed four unchanged Codex process/deadline tests; all four
passed serially without code changes in 7.29 seconds after removing the
competing test load. These failed commands are retained and are not
reported as full-suite passes; the fresh Linux CI runs are tracked
separately.
- The previous head `907bdb2a2778c7ffeb4a662a91460c9d1ddfc9c5` earned
Apex 5/5 with zero comments after fixing all three findings: per-user
install cache, stale release-pack rejection, and public Linux smoke
account/home handling. Its real built installer passed from read-only
public packages on macOS ARM64 and Linux x64. All 137 release-registry
checks and 64 ACPX package contracts passed. That review does not cover
this mainline reconciliation.
- Prior `beadd3654` passed the full CI matrix; its one unchanged chat
test failure and successful single retry remain in the [CI
history](https://github.com/paperclipai/paperclip/actions/runs/37521449327).
Historical results below remain attributed to their original builds.

- Fixture follow-up `dd59d7e82b103a88b7cbd7d2c38b612c0fbbff7a` removes
provider-specific model choices from generic offline ACPX tests. The
fake sidecar preserves the model and session identity selected at open
through suspension. Affected verification passed: 106 Rust tests and 73
TypeScript tests. This commit changes test code only; the
production-code checks below retain their recorded identities. Its CI
and Greptile review later passed; those results belong to that
historical head.
- Model-selection cleanup `9a070808b48960a41fdfd369ae0636b95af82459`:
252 focused Runner tests passed (six platform skips), covering all six
ACPX agents, native model acknowledgement, rejected selections,
installation integrity and recovery identity. The merged branch passed
recursive typecheck, full build, token gates, server admission (19
tests), and the Product E2E catalog (45 tests). The acceptance catalog
passed all four tests. The full Rust suite passed: 643 tests, 2 ignored.
It verifies sidecar acknowledgement of unlisted models and rejection of
model mismatches. The final commits only update Rust tests; production
sources match the verified build at
`65ec3279ac50185e3cda109b5cfd9b4f56105de0`. No new paid provider calls
were made.
- The merge preserves both Cursor and the new mainline public-MCP
fixture cases. Auto-merge remains disabled; the latest follow-up status
is recorded above. The local `pnpm test:run` attempt hit the unchanged
Git streaming test's 300-second timeout and was interrupted before
merging mainline. The broad Runner attempt found obsolete single-model
assertions plus three macOS fixture-path failures caused by a
`/private/tmp` override. The assertions are corrected; affected
TypeScript checks passed with the standard macOS temporary directory,
and the complete Rust suite passed. Neither interrupted command is a
full-suite pass.
- Earlier declaration-cleanup head `6f4a5e9e2` passed recursive
typecheck, build, Rust and focused tests. Its CI later exposed a test
expecting duplicated Grok digest literals. The current source fixes that
assertion to compare launcher bytes with the shared manifest. Historical
successes and failed attempts are retained; no new live provider
qualification is claimed.
- Previous head `e75fde6098b0ddd8cec765bfb6ecaeecb88a26a6` passed
complete CI (56 successful checks, one neutral, four skipped) and
Greptile 5/5. [Historical complete
CI](https://github.com/paperclipai/paperclip/actions/runs/37489112305).
Those results are not claimed for the cleanup head.
- Frozen live application: `d7b696f9b8f79095233e9e3d56d23e6a6018dd48`.
Public package version: `0.0.0-cursor-verify.3d0c9b7761c6`. The
declaration cleanup preserves release pins and does not relabel that
tested artifact as a build of the new source. Mainline through
`e34abee670` was reconciled while preserving accepted-plan waits,
provider-capacity handling, and both Cursor and public-MCP fixtures.
- Clean normal installation, explicit Cursor setup and daemon resolution
passed on macOS ARM64, macOS x64 under Rosetta, and Linux x64. npm
lifecycle hooks ran without silently downloading Cursor.
- Historical v11 live matrix: **18/18 passed with cleanup** (nine local,
nine Daytona) after the generic mode and lifecycle cleanup. The campaign
has 23 attempts; all five failures and their diagnoses remain recorded.
Exact case identities, hashes and limits are in the readiness report.
All provider calls are real, use the explicit Luna model and
company-bound credentials, and run without qualification or
runtime-asset overrides.
- The immutable Daytona image is
`ghcr.io/paperclipai/paperclip-daytona-runner@sha256:d6259b6bba094702c13fc2283bd85550849c1c53145b656fb2746778f9fa1747`.
The public Daytona plugin is installed independently and its version is
checked.
- Recursive typecheck, full build, token gates and Runner
contract/conformance/replay checks passed on the frozen application. Its
complete Linux CI suite passed. The duplicate local full-suite command
was incomplete after timing failures; affected repeats passed, but that
command is not reported as a clean pass.
- Qualification fixtures passed typecheck, 1,675 Vitest tests (one
skip), 128 Node checks, three provider-free browser tests, and 150
focused lifecycle tests after the final diagnostic correction. The
affected legacy Cursor command file also passed all five tests after
removing its shorter 10-second override; it now inherits the suite’s
standard 15-second timeout. Greptile is 5/5 on `e75fde609` with no
unresolved review threads. CI results above are recorded separately from
historical build results.

## Risks

- Cursor v14 includes the updated ACPX dependency patch and release
identity. The v11 live matrix and image below remain historical
evidence. They do not certify new v14 package/image artifacts.

- ACPX accepts models beyond the qualification fixtures. Availability
and entitlement depend on the provider. Successful configuration is not
a claim of live qualification for every model.
- Shared mode is an opaque identifier. Provider adapters own its
meaning. Incompatible historical sessions remain fenced; exact committed
plan waits and task history remain inspectable.
- Native AskQuestion is excluded. Paperclip semantic questions are
supported. Authoritative per-run dollar accounting is unavailable;
partial counters remain diagnostics and unknown cost is not zero.
- Image input, detailed native diffs, deeper child transcripts and
native plan-file export remain follow-ups.
- macOS x64 has clean-install and daemon-startup proof under Rosetta,
not a separate live campaign on Intel hardware.
- Release only the tested package/image combination. Merging this PR
does not publish npm packages or deploy that image. Later builds need
their own release verification. Rollback disables new Cursor admission
while preserving records and recovery inspection.
- A model can fail an exact instruction: one cancelled-plan attempt
returned the wrong summary marker despite correct cancellation. The
unchanged repeat passed; both results remain in the report.

> ROADMAP.md was checked. This completes existing native Runner/Cursor
work; it does not add an independent core feature proposal.

## Model Used

OpenAI Codex, GPT-6. The exact serving variant and context window are
not exposed in this session. The agent used reasoning, repository
inspection, code execution, protocol tests and browser-backed Product
E2E tools. Cursor acceptance uses the explicit
`gpt-5.6-luna[context=272k,reasoning=medium,fast=false]` model. That is
the evaluated provider model.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass — affected suites passed;
full CI and the retained local failed attempts are recorded separately
above.
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green — 56 successful checks, one
neutral and four skipped on `f7ec5cc1f0e30c62a829c016ff2013a2a9d79037`
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups —
fresh Base review passed on `f7ec5cc1f0e30c62a829c016ff2013a2a9d79037`;
zero new comments and no unresolved threads. The earlier Apex finding
remains fixed.
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 20:48:15 -05:00
DottaandPaperclip faa8e452c7 fix(tasks): stop repeated reminders for historical questions (#15392)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Task questions remain saved so a person can answer them later.
> - The composer moves an old question into history after a newer human
message.
> - Native completion still treated every pending question as a required
response.
> - This caused agents to demand an old answer after the person moved
work forward.
> - This pull request shares the historical-question rule across context
and task execution.
> - Agents can finish verified work while the original question remains
answerable.

## Linked Issues or Issue Description

Refs #15229, Refs #14613, Refs #13130.

**What happened?**

An agent repeatedly asked a person to answer a question that had moved
into feed history. Completion feedback explicitly told the agent to
request a response. The saved pending row also blocked task completion.

**Expected behavior**

A question before newer human direction remains answerable in the feed.
Its pending state alone must not require another reminder or stop
completed work. A new input blocker, an approval, or a configured review
stage must keep its gate.

**Steps to reproduce**

1. Let a task agent create an ordinary question.
2. Dismiss the question and send a newer task message.
3. Let the agent finish the requested work and submit its completion
report.
4. Observe a demand to answer the old question and a retained completion
gate.

## What Changed

- Add one company-scoped predicate for historical questions. Only later
human comments count. Exclude agent attribution, run attribution, system
notices, and untrusted source data.
- Apply the predicate to completion feedback, native waits,
finalization, commit validation, retry validation, blocked routing, and
successful-run handoff.
- Include question classification and guidance in heartbeat context and
both native task-context tools. Add the guidance to fresh and resumed
task prompts.
- Replace automatic reminders for current ordinary questions with
instructions to assess the real blocker, continue independent work, and
withdraw obsolete questions through the existing API.
- Preserve historical question rows during completion while cancelling
their live native source runs through the existing post-commit and
recovery paths. Let an authorized human answer them after completion
without reopening work or creating a response wake. Preserve
cancellation, current-input, approval, permission, credential,
connection, and review gates. Add no dismissal storage or migration.
- Document the rule in the execution contract and agent skill. Refresh
generated capability source anchors. Add database-backed status,
context, attribution, and governance regression tests.

## Verification

- `pnpm build` passed. The server rebuild also passed after the
lifecycle fix. Generated capability contract and inventory checks passed
after the agent documentation update.
- `pnpm -r typecheck` passed. Final `pnpm --filter @paperclipai/server
exec tsc --noEmit` also passed after the last test additions.
- Lifecycle and interaction regressions passed: 224 tests in 3 suites.
Context and prompt tests also passed. Final historical-question cases
passed (33 tests), native cancellation/recovery cases passed (6 tests),
and the existing interaction/confirmation suites passed (72 tests). The
full local `pnpm test:run` was attempted and stopped after more than two
hours with unrelated fixture/hook timeout failures; it did not pass. All
52 successful GitHub checks are green on the latest commit, including
the complete test matrix; no checks are pending or failing. Greptile is
5/5 and both review threads are resolved.
- Regression cases cover the old-question/new-human-message sequence,
final task status, answering after completion with no wake, live-run
cancellation and crash recovery, current input blockers, both
task-context tools, API context, timestamp precision, attribution
boundaries, and protected gates.

## Risks

- A later human task message makes an earlier ordinary question
historical even if its input is still missing. The agent must identify
the current blocker and ask only for information that still prevents
work.
- Browser dismissal remains a local preference. Dismissal without a
later human message is not recorded by this change.
- No schema change or data migration. Completion retains ordinary
historical questions; cancellation still expires them. A completed task
accepts historical answers only from an authorized human and creates no
response-delivery outbox row. Governed requests retain their gates.

## Model Used

OpenAI Codex, GPT-6, with repository editing, code execution, and
browser diagnostics. The exact deployment model ID and context-window
size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 20:33:13 -05:00
DottaandPaperclip 2d0c138122 Expand direct assistant MCP tools for work and configuration (#15380)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - People also use assistants in Codex, Claude, and other MCP clients.
> - The existing assistant connection can read work and create tasks or
comments.
> - It cannot edit tasks, exchange files, or manage normal agent and
project settings.
> - These operations must retain the person's permissions and
Paperclip's execution rules.
> - This pull request adds an explicit operation registry and separately
consented configuration access.
> - Assistants can manage work without receiving credentials or runner
authority.

## Linked Issues or Issue Description

**Subsystem affected**

Cross-cutting: assistant MCP, domain routes, consent UI, storage, and
Product E2E.

**Problem or motivation**

A connected assistant cannot update tasks, maintain documents, attach
files, or configure existing agents, projects, and skills. Users must
leave the assistant for these routine actions.

**Proposed solution**

Add named tools and a restricted API registry to direct connections.
Require separate configuration consent. Reuse domain routes and retry
receipts. File uploads save the attachment when the byte transfer
succeeds.

**Alternatives considered**

Arbitrary REST forwarding would expose administration and credential
operations. Runner impersonation would bypass execution ownership.
Separate upload completion calls add unnecessary client state.

**Roadmap alignment**

Checked ROADMAP.md and related MCP pull requests. This extends the
human-authorized connection from #14933. It does not replace the runner
or introduce agent impersonation.

Companion Cloud routing and directory isolation:
https://github.com/paperclipai/paperclip-cloud/pull/678.

## What Changed

- Add task editing, finish/block, documents/revisions, deliverables,
agent settings/instructions, projects/repositories, and skills/files.
- Add an allowlisted API search/call registry with identical field
restrictions, scopes, and retry identities.
- Add unchecked configuration consent. Existing write grants retain
their current authority.
- Add hashed, expiring file transfer tickets and atomic upload receipts.
No completion call is required.
- Preserve company boundaries, human attribution, active native
execution ownership, and execution review gates.
- Add protocol/domain tests, consent stories, and eight paid Product E2E
workflows.
- Repair two CI fixture races: await cold route setup before assertions,
and wait for asynchronously loaded connection copy. Both fixture suites
pass (24 + 48 tests).

## Verification

- Consent revision: one write-access checkbox controls requested work
and configuration permissions in browser and device flows. All 16
consent tests, UI typecheck/build and token gates pass. Updated
interactive stories cover default approval, opt-out and viewer
restrictions. The paid browser helper uses the new exact label. Real
GPT-5.4 Mini Product E2E passes 2/2 at
`64f96373118eb190f8cba1c2ab17cb979555f3ad` (configuration + permission
denial), campaign `local-2026-10-07T00-51-14-337Z`, no automatic
retries, cleanup passed; $0.04149375 estimated assistant cost plus
unpriced worker usage. Raw results, usage and source fingerprints are
retained in the worktree. UI and Product E2E typechecks pass.

- Prior head `2f246d4b74f1f98c75ebcb37ae6753a748237fac`: all 52 checks
pass; two optional Storybook checks skip. Greptile 5/5 on that head, no
unresolved review threads. Final consent head
`64f96373118eb190f8cba1c2ab17cb979555f3ad` also has all checks passing
and Greptile 5/5 with no unresolved threads. The unchanged Cursor
sandbox test had one 10-second timeout, passed in local isolation, and
passed its single CI rerun; the failed attempt remains in [the CI
run](https://github.com/paperclipai/paperclip/actions/runs/37554106934).
The existing chat retry-denial browser test had one visibility failure;
its single rerun passes, and the failed attempt remains in [the CI
run](https://github.com/paperclipai/paperclip/actions/runs/37542735691).

- Full workspace `pnpm -r typecheck` and `pnpm build` pass at final
runtime source `b2196fae1`. UI token gates pass.
- 139 MCP/OAuth/transfer/privacy tests and 76 grader calibration tests
pass, including one-connection PostgreSQL OAuth and concurrent upload
retries.
- Paid Product E2E: all eight expanded cases qualified across Mini,
Haiku and Sonnet. A merged-source repeat passed 23/24; one Haiku cell
timed out before application startup. Final affected-case qualification
passes 9/9 on all three models with grader v16, including the failed
cell. Automatic retries disabled; failures, costs, source hashes and
independent durable-state/file assertions are retained in [the
verification
record](doc/plans/2026-10-06-expanded-assistant-mcp-verification.md).
- Actual Codex CLI, Claude Code and OpenCode clients completed local
reads/mutations. Codex wrote a report, Claude updated it in a later
conversation, and OpenCode uploaded/downloaded a file with matching
SHA-256 and registered the attachment. Revoking the CLI grant rejects
subsequent bridge initialization.
- Butter staging is verified on final runtime `b2196fae1`
([deployment](https://github.com/paperclipai/paperclip-cloud/actions/runs/37538432138)).
A fresh OpenCode workspace fetched the copied invitation, configured
remote MCP, started OAuth and reached real consent with configuration
unchecked. Invalid transfer tickets return 403 through Cloud. Human
approval for the new persistent staging grant is pending; hosted
task/file success is not yet claimed. The final transaction fix is
deployed.
- Full local `pnpm test:run` passed 15,614 general-server tests but
stopped on two macOS timeouts. The heartbeat test passed in isolation;
the existing 40,000-file Git stress fixture timed out again. Its Linux
CI lane passes. Later local full-suite phases did not run after the
timeout; this is not an all-green local full-suite claim.
- Instructions and security limits are in `doc/public-mcp.md`; the saved
plan is `doc/plans/2026-10-06-expanded-assistant-mcp-tools.md`.

## Risks

- This expands the experimental direct MCP surface. Explicit schemas and
domain permissions must stay synchronized.
- Migration 0311 adds transfer tickets and upload receipts. Expired
orphan cleanup must not remove committed attachments.
- Configuration requires a new consent request containing that scope;
the single write-access choice controls it alongside work mutations.
Refreshing an old grant does not add it.
- The public directory keeps its original ten tools through the
companion Cloud change.
- Hosted consent/work proof remains the final delivery gate. The PR
stays draft while approval of the new staging grant is pending; code
checks and review are green. Merging is a separate action.

## Model Used

OpenAI Codex (GPT-6, tool use and code execution). The exact serving
model ID and context window are not exposed in this session. Paid
evaluation models: gpt-5.4-mini, claude-haiku-4-5-20251001;
claude-sonnet-4-6.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused suites; full-suite
macOS limitation disclosed above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 20:29:58 -05:00
Devin FoleyandPaperclip 892b0b3606 fix: scope quota reports to authorized subscription accounts (#14994)
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-06 18:20:37 -07:00
Devin FoleyandPaperclip eab93fd4a0 fix: checkpoint adapter usage and preserve unknown prices (#14991)
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-06 18:16:18 -07:00
Devin FoleyandPaperclip 3ebd7bc0c9 test: isolate accounting fixtures and include all adapter suites (#14989)
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-06 18:13:46 -07:00
f77fcbf4bf feat(apps): add Telem.AI web search connection (#15379)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Research agents need current web information.
> - The Apps catalog connects agents to remote MCP tools through the
normal access rules.
> - Telem.AI supplies web search and page reading through one API key.
> - This PR adds its catalog entry, optional search settings, artwork,
and setup guide.
> - The connection keeps an operator's saved header policy when they
reconnect.

## Linked Issues or Issue Description

Refs #15302. Related catalog work: #13881.

This PR continues #15302 by Yifei Ai (@aiwen324). Thank you for the
connector and its review fixes. All seven original commits are
preserved. GitHub denied the attempt to push to the contributor's fork,
so this branch retains the repair commit and merges current master.
Master now includes the same test fix.

For a squash merge, keep the original author in the final commit
message:

```text
Co-Authored-By: Yifei Ai <aiwen324@users.noreply.github.com>
Co-Authored-By: Paperclip <noreply@paperclip.ing>
```

A search found no separate public Telem issue or competing Telem PR. The
existing Apps path matches the roadmap.

**Agent or provider**

Telem.AI provides web search and page reading through a hosted MCP
server.

**Why this adapter is useful**

Agents can use multiple search providers through one governed
connection. Operators can set the search tier, auto routing, and
provider lists.

**How the agent is invoked**

The remote MCP server uses Streamable HTTP at
`https://mcp.telem.ai/mcp`. The API key uses an `Authorization: Bearer`
header. See the [official MCP
guide](https://docs.telem.ai/integrations/mcp/).

## What Changed

- Add the Telem.AI definition, research entry, permission review, and
generated registry entry.
- Add four optional settings. Unset settings send no request header.
- Add official light and dark artwork, source records, and a setup
guide.
- Forward company, issue, agent, run, project, and correlation IDs by
default. Preserve a saved policy, including disabled forwarding, on
reconnect.
- Add catalog and connection tests.

## Verification

Current head: `b4164477fb1b312a504789bf17b51c963244c0fd`.
Merged master: `228f0e2807c5b59d2aa129cf2d80b9777ebabf07`.

- Resolved five shared catalog conflicts after the Superagent connection
merged.
- Keep both providers in the research ledger, generated registry,
generator, branding manifest, and connection guide index.
- Correct the combined catalog totals: 52 self-serve candidates, 55
research entries, and 68 Apps entries.
- The published Git tree exactly matches the tested local resolution.
- Catalog and Apps UI suites: **295 tests pass** after the catalog count
fixes.
- Connection service suite: **387 tests pass** in the full run. Its only
failure was the old catalog count. That test passes on a focused rerun
after the fix. This gives **388 passing service tests** across the two
runs.
- Total focused coverage: **683 passing tests**. The first runs exposed
four fixed-count assertions that needed the combined totals.
- Shared package build and plugin SDK compile pass. Token gates and
whitespace checks pass.
- Generation with `--definitions-only` reproduces the Telem definition
and registry. The unrelated AgentMail and Linear drift remains excluded.
- Local UI typecheck ended with exit 137 at the container memory limit.
Full local typecheck, test, and build are not claimed. Earlier runs also
recorded missing Cargo and Node development headers.
- GitHub reports a clean merge state against master `228f0e280`.
- All 54 checks are complete: **52 passed and two Storybook checks
skipped**. No check failed or remains pending.
-
[CI](https://github.com/paperclipai/paperclip/actions/runs/37545412406)
passes on this head. This includes typecheck, build, tests, browser
shards, Runner checks, and Canary Dry Run.
-
[Greptile](https://github.com/paperclipai/paperclip/pull/15379#issuecomment-6024519537)
is **5/5 on this head**. There are no review threads, open P2s,
recommendations, or follow-ups.
-
[Superagent](https://github.com/paperclipai/paperclip/runs/112548142930)
passes.
- Final recovery checks confirm all seven original commits and current
master remain in history. Token gates and whitespace checks pass.
- The final recovery run makes no source change. It verifies the
published repair and retains the local check limits below.
-
[Commitperclip](https://github.com/paperclipai/paperclip/actions/runs/37545407943)
passes with no failures. Its only informational note asks the merger to
keep the author trailer above.
- All seven original contribution commits remain in history. The diff
against master contains the same 14 Telem files. It adds no dependency,
lockfile, schema, or workflow change.
- The managed GitHub CLI capability was missing in this run. The
installed GitHub connection applied the base files, merged master, then
restored the tested combined catalog. No history was rewritten.

The previous head `2d32a0094` passed all remote gates and had Greptile
5/5. Those results do not verify this new head.

The original PR reports live setup, discovery, settings headers, gateway
calls, and context-header forwarding. This repair does not repeat those
account-bound checks. The permission record still marks maintainer live
qualification as outstanding.

## Risks

- Telem.AI receives the six context IDs by default. The saved header
policy controls forwarding. Search use is billed to the account that
owns the key.
- All agents on a connection share its search settings.
- The merge uses master's route-test setup unchanged. The
company-boundary assertions remain intact.
- No schema, dependency, or workflow change is included.
- Live provider evidence is attributed to the original contributor.
Maintainer live qualification remains outside this CI repair.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- Original contribution: Anthropic Claude Opus 5.5 (`claude-opus-5-5`),
1M-token context, through Claude Code with shell, editing, and test
tools, as disclosed in #15302.
- CI repair and review: OpenAI `gpt-6-astra`, through Codex with
reasoning, shell, editing, and GitHub tools. The runtime does not expose
the context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Yifei Ai <aiwen324@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: devinfoley <139239+devinfoley@users.noreply.github.com>
Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 16:33:18 -07:00
Devin Foley 8cbd21b3e7 feat(apps): add Superagent connection (#15394)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents reach external services through the Apps catalog. Each
catalog entry is a reviewed `AppDefinition` that connects a provider's
hosted MCP server to Paperclip's shared vault, grants, policies,
gateway, and audit trail.
> - Superagent (`superagent.sh`) is a security platform. Its hosted MCP
server lets agents read and triage security findings, start red-team
reports, run Contributor Trust and dependency update jobs, and score web
pages, email, files, skills, MCP repositories, and packages before an
agent trusts them.
> - Superagent is not in the catalog. Operators must use the generic
"Connect your own MCP server" flow. That flow has no branding, no key
guidance, and no warning about billable or destructive tools.
> - The connector playbook supports this provider with existing
definition fields. The server accepts an organization API key as a
bearer header. It publishes no OAuth authorization-server metadata, so
browser sign-in is not possible.
> - This pull request adds the Superagent definition, its artwork, the
research and permission-review ledger rows, documentation, deterministic
tests, and one reviewed risk-classification rule.
> - The benefit is a branded, governed Superagent connection with clear
key guidance and a warning about tools that cost credits or delete data.

## Linked Issues or Issue Description

**Problem or motivation**

Teams that use Superagent for PR security, red teaming, and agent
guardrails want their Paperclip agents to read findings, return
structured reports, and score content before they use it. Superagent is
not in the Apps catalog. Operators must paste the MCP URL and an
`Authorization` header into the generic remote-MCP flow. That flow gives
no branding and no provider guidance. It also does not tell the operator
that the key reaches the whole organization, or that some tools consume
credits or permanently delete findings.

**Proposed solution**

Add a catalog-only Superagent connection that follows the connector
playbook. It has one method: a customer organization API key
(`sk_live_...`), sent as an `Authorization: Bearer` header to
`https://www.superagent.sh/mcp`. The field helper text explains that
Superagent keys are not scoped. The method warning tells operators to
set billable and destructive actions to Ask first before agents run
unattended. All discovered tools stay governed by the normal per-action
policies.

**Alternatives considered**

A browser sign-in method was not added. The server's protected-resource
metadata names `https://superagent.sh` as its authorization server, but
that origin publishes no `oauth-authorization-server` or
`openid-configuration` document, so Paperclip cannot discover OAuth
endpoints. A plugin was not needed because the connection needs no
custom UI, tables, workers, or webhooks. Relying on the generic risk
classifier was not enough. Several Superagent mutations
(`triage_finding`, `scan_*`, `restore_agent_builtin_rule`) use names
that it reads as reads, so a narrow reviewed Superagent rule was added
instead.

**Roadmap alignment**

This extends the existing self-serve remote-MCP connection catalog. It
does not overlap planned core work.

## What Changed

- Added the `superagent` row to
`packages/shared/src/self-serve-mcp-research.json` (API-key auth, risk
tier S4).
- Added the `superagent` provider to
`scripts/ingest-app-definitions.mjs` (category, key placement and
placeholder, console links, guidance, description). Regenerated
`packages/shared/src/app-definitions/superagent.json` and the generated
registry.
- Added the `superagent/mcp-api-key` permission review to
`doc/connections/tool-method-permission-reviews.json`, with
key-permission text and evidence links.
- Added Superagent's official mark
(`ui/public/brands/apps/superagent.png`, the 460×460 avatar of the
official `superagent-ai` GitHub organization) and the brand manifest
entry.
- Added gallery copy for the Superagent card.
- Added a reviewed Superagent rule to `classifyRisk` in
`server/src/services/tool-access.ts`. Only `list_*` and `get_*` tools,
and tools that Superagent marks read-only, are reads. `delete_*` and
`revoke_agent_client` are destructive. All other tools are writes, so
billable and rule-changing tools can be set to Ask first.
- Put the Ask-first advice in the API-key helper text, because the key
form shows helper text and not method warnings.
- Added `doc/connections/SUPERAGENT.md` (transport and auth, why there
is no OAuth, administrator setup, capabilities and policy, manifest,
brand provenance, validation hook). Linked it from the connections
README and the permission audit.
- Tests: definition shape, store visibility and artwork, URL
recognition, the bearer header on discovery with the key kept out of
connection config, read/write/destructive classification of fixture
tools, the Superagent risk rule (including `triage_finding` and
`restore_agent_builtin_rule`), the visible Ask-first advice and API-key
gating of the connect form, and the pinned catalog counts.

## Verification

- `pnpm exec vitest run packages/shared/src/app-definitions.test.ts
packages/shared/src/app-definitions-url.test.ts
ui/src/pages/apps/AppsConnect.test.tsx ui/src/pages/apps/Browse.test.tsx
ui/src/lib/app-brand-assets.test.ts
ui/src/pages/apps/AppLogo.brand-assets.test.tsx`: 317 passed.
- `pnpm exec vitest run
server/src/__tests__/tool-access-service.test.ts`: 385 passed.
- `node scripts/check-app-brand-assets.mjs` and `node --test
scripts/app-brand-validation.test.mjs`: passed.
- `pnpm --filter @paperclipai/shared typecheck`, `pnpm --filter
@paperclipai/server typecheck`, and `pnpm --filter @paperclipai/ui
typecheck`: clean.
- Manual: in a local instance, open Apps → Browse and confirm the
Superagent card and icon. Open `/apps/connect?source=superagent`.
Confirm the single API-key method, and confirm that Connect enables only
after a key is entered.
- Live metadata probe on 2026-10-06: an unauthenticated `initialize` on
`https://www.superagent.sh/mcp` returns 401 with
`resource_metadata="https://www.superagent.sh/.well-known/oauth-protected-resource"`.
That document returns 200. No authorization-server metadata exists at
the named issuer.

## Risks

- Low risk to existing providers. The change is additive catalog data
plus tests. The generated registry only gains one import. The new risk
rule runs only for Superagent connections.
- A Superagent key reaches its whole organization. Some tools consume
credits (`create_*_report`, `triage_finding`) or delete data permanently
(`delete_finding`). Every action starts Allowed under the current
product default. The key helper text tells operators to set these
actions to Ask first.
- The permission-review ledger records live proof as not run. No
Superagent account was used. The lifecycle checklist in
`doc/connections/SUPERAGENT.md` needs a documented pass before the entry
is fully qualified.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- Claude Opus 5.5 (`claude-opus-5-5`, 1M context) in Claude Code, with
extended thinking and tool use (shell, file editing, web fetch, browser
checks).

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-10-06 15:57:16 -07:00
DottaandPaperclip b508a05c43 feat: add internal agent complaints and suggestions (#15367)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents use legacy skills or native runner tools to work on tasks.
> - Those agents can encounter friction that does not belong in the task
thread.
> - A complaint should preserve the raw reaction. A suggestion should
describe an improvement.
> - This pull request adds attributed local storage and both submission
paths.
> - Agents can submit feedback once and continue their primary work.

## Linked Issues or Issue Description

**Subsystem affected**

Server, database, shared contracts, runtime skills, and native runner
tools.

**Problem or motivation**

Agents have no default internal channel for incidental complaints and
suggestions. Sending this feedback through task comments adds noise and
can alter task workflows.

**Proposed solution**

Store free-form feedback in the current instance database. Derive agent,
run, company, and task attribution from active authority. Provide
default legacy skills and provider-neutral native actions. Keep the
instructions close to Warp's MIT-licensed originals.

**Alternatives considered**

Task comments and external Slack delivery add unwanted side effects.
Mandatory suggestion fields and short editorial limits would discard
useful feedback. This release has no listing API, UI, read tool,
automatic triage, or external forwarding.

**Roadmap alignment**

This is a maintainer-requested addition to the existing runtime skills
and runner tool paths. It does not duplicate a listed roadmap milestone.
Searches for complaint tooling, suggestion-box, and agent commentary
found no overlapping public PR or issue.

## What Changed

- Add the company-scoped `agent_commentary` table, shared validation,
and idempotent migration `0310`.
- Add one transactional service and the agent-only POST route. Validate
active authority before writes or replay. Redact known credentials.
Commit a content-free audit with each new record.
- Add `submit_complaint` and `submit_suggestion` to standard, ask, and
planning modes. Keep review, revocation, and completion restrictions.
Store replay identity on the commentary row.
- Mount `complain` and `suggestion-box` by default for legacy agents.
Bundle a dependency-free Node.js stdin helper in the operational skill
and allow its POST through the sandbox bridge.
- Preserve Warp's complaint voice and suggestion guidance, with
attribution and local transport adaptations. Keep source attribution and
MIT notices in each skill's LICENSE, outside runtime instructions.
- Document custom-runtime HTTP use and database inspection. Add
real-database tests and a repeatable live Codex smoke for local and
Daytona execution.
- Pin the lagging-source migration fixture before the identity-repair
migration so later migrations preserve its regression coverage.

## Verification

- Personally ran real Codex submissions in all four environments on
2026-10-06. Local runs passed at 20:35 UTC. Daytona native passed at
20:31 UTC; Daytona legacy passed at 20:33 UTC. Each stored exactly two
rows with company, agent, run, and task attribution, wrote the
continuation marker, exited zero, created no task comments, and left
task status unchanged. Each recorded two content-free activity entries.
- Daytona used production provider hooks, real remote execution and file
transfer, the legacy queue callback bridge, and native private WebSocket
ingress. The current Linux runner was built from `abf47b595`, staged,
and verified against controller contracts. Both sandboxes were confirmed
deleted. This is a focused feedback transport smoke; it does not claim
full Runner E2E catalog or browser qualification.
- The immutable base image and Linux binary digest are recorded in [the
verification
documentation](https://github.com/paperclipai/paperclip/blob/codex/agent-commentary/doc/agent-commentary.md#verification).
The smoke script can save content-free JSON evidence. No credentials or
feedback bodies are in these reports.

| Environment | Runner | Complaint row | Suggestion row |
| --- | --- | --- | --- |
| local | legacy Codex | `59413a00-1de2-4bb1-bcc6-9c4b54c64aa6` |
`3db2364d-3e15-4f47-846f-875d3902999d` |
| local | native Codex | `5da22b5f-41df-4de5-8ba0-d9345ab01267` |
`2d5abe17-dd41-403c-a5ee-4729f2d58921` |
| daytona | legacy Codex | `27c9d0aa-8477-409f-9da0-e8ffa48dee50` |
`209681c9-d1e9-4ce1-999e-48fa07692389` |
| daytona | native Codex | `6eb001bb-4bcf-43f7-8717-f662f53dc7c3` |
`77c383d8-a997-49e5-a33e-25c70e15c0b2` |

- Run the local check with `node cli/node_modules/tsx/dist/cli.mjs
server/scripts/verify-agent-commentary-live.ts`. The documentation gives
the Daytona invocation. Both use disposable instance databases and
normal Codex provider usage.
- Repository `pnpm -r typecheck` and `pnpm build` passed after the test
extension. The build includes runner generation, contracts, and replay
checks. The smoke scripts also passed a separate TypeScript check. The
lagging-source migration regression passed. All equivalent current-head
Vitest CI shards passed. The local monolithic `pnpm test:run` invocation
was stopped after CI supplied that coverage; it did not complete
locally.
- Focused tests cover company isolation, spoofing, revoked credentials,
stale ownership, post-finish rejection, concurrent replay, conflicting
keys, atomic rollback, and deletion through existing services. Boundary
tests cover empty text, Unicode, text beyond 8,000 characters, and the
524,288-character ceiling without truncation. Mounting tests cover
Codex, Claude, and sandbox staging. Helper tests cover standalone Node
execution, stdin, invalid UTF-8, redirects, HTTP failure, and its
deadline. Privacy and bridge tests cover successful and rejected
requests.
- Instructions were compared with Warp's originals. MIT notices and
source credits live only in LICENSE files. Native tools preserve
truthful disclosure when asked, without routine announcements.
- [Full
CI](https://github.com/paperclipai/paperclip/actions/runs/37508559190)
and Greptile 5/5 passed on the earlier feature commit `5209c3501`. The
later head found the migration-fixture assumption fixed in this update.
On `8a4965164`, all 55 check contexts passed after one browser shard
rerun. Its initial reviewer signoff failure also passed an isolated
local browser run (1 test). Greptile scored that head 5/5 and identified
one smoke cleanup gap. `6ecbafb0b` fixes failed-acquisition cleanup with
four passing tests and a passing smoke-script typecheck. Fresh CI is
pending for this final test-only fix. No commentary production code
changed during verification.

## Risks

- Feedback is internally attributed. It is not anonymous. Existing
redaction removes known credentials, but agents must still omit
sensitive content. Normal provider transcripts can include their
submitted arguments.
- Default skill availability changes for existing legacy agents. Runtime
policy filtering still applies. The helper uses the existing Node.js
runtime with no extra dependencies; custom runtimes can call the HTTP
endpoint.
- Feedback is removed with its run, agent, or company. Task deletion
clears only the issue pointer. Normal database backups include the
table.
- The migration is additive and has no backfill. Writes serialize on the
active run for replay consistency. No server suggestion quota is
imposed.

## Model Used

OpenAI `gpt-6-astra` through Codex, with `xhigh` reasoning effort and a
reported 258,400-token context window. Capabilities used: repository
inspection, code execution, and live runtime verification. No subagents
were used.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 16:14:06 -05:00
ea8e686197 fix(tools): distinguish requested from granted OAuth scopes (#14059)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Apps use OAuth connection grants to let agents reach external
services.
> - Paperclip stored a requested scope as a granted scope when a
provider omitted `scope` from its token response.
> - A live MCP connection showed why that matters: the provider reported
scopes beyond the requested scope. Scope names alone do not establish
write capability.
> - This PR records whether a scope came from the provider or from
Paperclip's request, and flags asserted extra scopes.
> - The result is an honest scope record across authorization and
refresh, including for self-hosted and Cloud instances.

**Review order:** shared scope resolver and OAuth write paths in
`server/src/services/tool-access.ts` → optional schema/types/validator
fields → focused tests in
`server/src/__tests__/tool-access-service.test.ts`. This PR changes the
shared OAuth record. The Enterpret catalog definition is separate in
**#13906**.

## Linked Issues or Issue Description

No public issue covers this bug. Related: **#13906** adds the official
read-only Enterpret connector with OAuth and organization tokens.

**What happened?**

The OAuth callback stored `normalizeOauthScopes(token.scope ??
requestedScopes)`. When the provider omitted `scope`, Paperclip recorded
the request as though it were a verified grant. In the live case:

```text
Paperclip requested     mcp:read
Provider token reply    no scope field
Paperclip recorded      ["mcp:read"]
Token introspection    openid email profile mcp:read mcp:write
```

The access token is opaque. A negative introspection control returned
`active: false` with no scope, confirming the wider scope belonged to
the live token. The grants API could therefore present requested scopes
as though the provider had asserted them. This observation did not
establish access to write tools.

**Expected behavior**

Store the provider's asserted scope when present. When absent, identify
the value as an inference from the request. Preserve that provenance
through refresh and report extra scopes only when the provider actually
asserts them.

**Steps to reproduce**

1. Use an OAuth provider that omits `scope` from its token response.
2. Complete consent after requesting `mcp:read`.
3. Read the connection or grant: before this fix, the stored `scopes`
looked like an asserted read-only grant.
4. The focused fixture tests reproduce the callback and refresh record
without needing a live provider.

**Deployment mode**

The shared OAuth path affects self-hosted and Cloud. The live provider
evidence came from an isolated self-hosted runtime.

## What Changed

- Added `resolveGrantedOauthScopes`. It records `scopeSource: provider`
when the token response asserts scope, or `requested_fallback` when it
does not. `unrequestedScopes` contains only provider-asserted scopes
outside the request.
- Applied the resolver at initial authorization and refresh. A refresh
without `scope` retains a previous provider assertion and warning; a
fresh assertion can replace them.
- Stored the actual authorization request per grant as
`providerTenant.oauth.requestedScopes`. This keeps a multi-user
connection's refresh baseline tied to the right grant. Generic MCP OAuth
also records discovered scopes when it sends them; curated apps retain
their reviewed request.
- Carried provenance onto the connection and default organization grant.
The organization grant does not receive token expiry, so its existing
refresh/reconnect behavior remains intact.
- Added optional fields to database JSONB schema types, shared types,
and validation. Existing grants need no migration or backfill.

This PR **does not change authorization decisions**, reject tokens,
introspect providers, or show a new UI warning. If a provider hides an
over-grant by omitting `scope`, Paperclip still cannot discover it.
`unrequestedScopes: []` with `requested_fallback` means **unknown**, not
least privilege.

## Verification

**Current head:** `7fde922a8`. Reconciled with master `88ff98b83`; the
final diff is six OAuth implementation, contract, and test files.
Preserved master's generic `offline_access` consent and legacy callback
baseline, and verified requested versus provider-asserted scopes on each
grant. Removed four duplicated GitHub token-method fixture properties
after fresh review.

All 504 focused tests in seven suites pass on the final head. Local
workspace build, recursive typecheck, build-gap typecheck, module
boundaries, node-version, token gates, and runtime push-policy checks
pass. The full local Vitest run completed with 15,610 passing tests, 91
skipped, and six failures in heartbeat/workspace suites; all six
failures reproduced on unchanged master `88ff98b83` in an isolated
baseline worktree. These suites and their runtime code are unchanged by
this PR. Greptile is 5/5 on this exact head with no actionable findings.
All current-head GitHub checks are green: 53 successful check runs, two
intentional Storybook skips, and the successful Snyk status. The
initially failed adapter-access and signoff-heartbeat shards both passed
their single rerun. [CI
run](https://github.com/paperclipai/paperclip/actions/runs/37521588330).

The regression coverage includes omitted and explicit scopes, asserted
extra scopes, refresh preservation, per-user request baselines, generic
MCP discovery, and organization grants. No live provider token is needed
to reproduce these recordkeeping cases.

## Risks

- Existing records keep their scope list but have no source marker.
Historical provenance cannot be recovered from them.
- A refresh that omits `scope` preserves the last known assertion. A
provider that silently narrows a grant will not update the record until
it asserts a new scope.
- A provider that omits `scope` leaves the actual scope unasserted. This
PR cannot discover scopes omitted from the response or establish
endpoint capabilities. Enterpret’s provider fixes are separate
follow-ups.
- Scope provenance is additive metadata. No policy path uses these new
fields to allow or deny an action.

## Model Used

Implementation and tests: Claude Opus 5 (`claude-opus-5`) through Claude
Code with extended thinking, tools, and code execution. PR text cleanup:
Codex GPT-6 (exact host model ID and context window were not exposed;
tool use).

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] Focused tests pass locally; full-suite baseline failures are
documented above
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All current-head CI gates are green
- [x] Greptile is 5/5 on the current head with no actionable follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

This shared scope-recording fix can be reviewed and merged independently
of the Enterpret connector.

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 13:25:20 -07:00
DottaandPaperclip e38d6d16b6 feat(connections): add advanced provider setup and live browser qualification (#15341)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Users connect accounts and choose an agent harness and model.
> - The runtime change in #14970 supports custom providers on those
connections.
> - Normal setup must stay simple while advanced users can choose a
compatible gateway.
> - Shared connector rows and access controls keep these choices
consistent.
> - This pull request refines the agent setup UI and adds review stories
and repeatable browser qualification.
> - The qualification checks real tools and downloaded outputs, not only
a successful run status.

## Linked Issues or Issue Description

Refs #14970, #37, #13083, #14104, #14565, #12692.

The core implementation in #14970 is merged. This branch incorporates
its squash commit and targets `master`. Both PRs contain our
implementation. #14016 is a reference only and is not a dependency. This
PR has 96 changed files.

## What Changed

- Complete model-provider connector presentation beside other
connectors. Each row uses the existing Connect action and connection
list. Tags are stored without category UI. The base PR includes the
provider forms and routes.
- Show persistent Subscription, API Key, and Advanced choices. Label
Advanced as Custom Gateway. Reuse provider logos, connection lists, and
permissions controls. Default access to the organization and all agents
when permitted; keep narrowing controls under Advanced.
- Keep Configure reachable before subscription sign-in, so users can
select a supported environment when the default cannot sign in. Testing
and saving still require a connection. Show the execution environment in
Configure. Preserve the confirmed Connect choice. Editing a method,
credential, saved account, or advanced choice requires that current
choice to connect before testing or saving. Use matching model and
thinking-effort dropdowns and retain connection icons in selected
values.
- Preserve the new harness model default when switching an existing
OpenCode agent to Codex or Claude, and resolve user-selected model names
with the effective harness.
- Load popular OpenRouter models through the shared connection-model
discovery path. Keep explicit model lists and manual model entry
available.
- Group onboarding, connection setup, agent runtime, management,
recovery, and production-component stories under AI Connections /
Provider routing.
- Add an explicit-only provider-connections browser suite for managed
local or existing local/staging targets. Use private browser profiles
and credential handoffs. Support human-assisted subscription sign-in
without sharing passwords or tokens in reports.
- Verify persisted connection identity, runtime probes, tool execution,
exact artifact bytes, completion, and context-dependent follow-up.
Retain source/model provenance, cost bounds, closed error diagnostics,
original failures, and cleanup evidence.
- Add Gemini startup-model and skill-root fixes, Grok private-history
detection, ACP filesystem regression fixtures, selected-workspace
handling for local Hermes, and artifact-helper workspace fallback.
- Keep managed Grok runtime homes disposable. Remove host-side
transcript retention/restoration because private file modes do not
isolate same-user agent processes. Ignore earlier development archives
and use a fresh task handoff when history is unavailable. Verify the
absence of restored transcripts with a separate same-user process.
- Capture stopped-run diagnostics before deleting an attached-company
fixture agent. Track creation and owned sign-in receipts; revoke only
this attempt's accounts and never adopt a concurrent campaign's newly
created account. Preserve failure signals and final status through
cleanup.
- Require the requested environment in the saved agent and every run,
including follow-ups. Reject a forced incompatible target. Keep one
cancellation state through startup, every cell, reporting, and teardown
for SIGINT, SIGTERM, and SIGHUP. Stop further paid cells after
interruption. Document qualification limits.

## Verification

- Current head `b3bb3e94d577d43d9965a6b9daba039f599b2e49` includes
master `d9f600043`. The security fix in `a758fde31` passes full
workspace typecheck, production build, and 119 connection/Grok
regressions. The unchanged UI passes all 126
configuration/model-discovery tests and token gates. The final
published-guide correction passes Grok adapter typecheck. Earlier head
`eebd8225c` passed the complete deterministic runner suite (1,404 Vitest
tests and 128 Node tests) and all CI jobs. Current-head CI run
`37520147514` passed all 47 jobs, including the full sharded Vitest and
browser matrix, production build, and canary dry run. All 55 checks
completed: 53 successes and two expected skips. The current-head
security scan passed, Greptile is 5/5, and no review threads remain
open.
- A separate same-user process reproduced reading a restored Grok
transcript before the security fix. The regression now finds no
transcript. Existing fresh-session fallback and ordinary session
metadata behavior pass.
- The final account-choice and cleanup fixes pass 85 setup tests and 26
qualification-harness tests. Regressions verify that editing a
connection invalidates confirmation, Configure remains reachable before
sign-in, diagnostics are captured before fixture deletion, and
concurrent campaigns cannot adopt or revoke each other's accounts. UI
and E2E typechecks pass.
- The Storybook build and actual Chromium production-component stories
passed during this change. Review the neighboring AI Connections /
Provider routing stories, regular connector rows, three connection
modes, model discovery, and the single execution-environment control in
Configure.
- Cancellation smoke verified authenticated cleanup before browser close
for SIGINT, SIGTERM, and SIGHUP. Regressions cover interruption during
startup and reporting, missing-file ACP resource errors, and preserved
permission denials. Both ACP runtime versions and 54 ACPX/Grok
regressions passed. The deterministic connection-intent browser suite
passed two tests.
- Historical local qualification retained 43 passing API/gateway cells
out of 46, with downloaded outputs and follow-up receipts. These
attempts span earlier builds; they do not qualify this exact commit or
staging. Subscription combinations, Gemini overloads, and the unresolved
follow-up failure remain recorded rather than counted as passing.
- Use `pnpm test:e2e:runner -- --list --suite provider-connections` to
inspect the matrix. Follow `tests/runner-e2e/PROVIDER-CONNECTIONS.md`
for credentials, target URL, sign-in assistance, budget, evidence, and
cleanup. Paid live tests remain opt-in.

## Risks

- The core implementation in #14970 is merged. This PR adds no database
migration of its own.
- Subscription login needs an interactive provider session. Dedicated
accounts and staging qualification remain follow-up work; this PR does
not certify every login combination for production.
- Managed Grok transcript resume is deferred until provider history has
an OS isolation or authorized broker solution. Follow-ups start fresh
with Paperclip task context; earlier live Grok results do not qualify
this behavior.
- Gemini CLI 0.58.0 has an upstream ACP new-file error conversion
defect. Live overloads and one unresolved follow-up timeout remain
recorded. The stock CLI is unchanged, and those cases are not marked as
passing.
- Real-provider tests spend credits and use private credential/evidence
directories. The launcher requires explicit selection and checks target
ownership. It must not attach to a developer's database by accident.
- OpenClaw Gateway, Hermes Gateway, Claude Managed, AWS AgentCore,
Process, HTTP, and legacy ACPX local remain outside custom provider
setup.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, code
execution, and browser testing. The exact deployment model ID and
context window size were not exposed in this session.

## Checklist


- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 15:21:22 -05:00
88ff98b83d feat(apps): add read-only Enterpret MCP with OAuth and org tokens (#13906)
## Thinking Path

> - Paperclip agents need governed access to Enterpret's hosted MCP
tools for customer feedback.
> - The official Enterpret MCP is read-only. Enterpret Agent’s beta
write MCP is a separate service and is outside this PR.
> - This PR supports both organization auth tokens and personal browser
OAuth for the official MCP.
> - Enterpret committed to fixing OAuth scope behavior and reducing its
token-revocation cache window. The provider fixes will follow after
connector deployment and do not gate this release. Paperclip functional
validation remains required. New tools stay quarantined on refresh.
> - This PR combines the original connector and Vivek's token-path work
on current master, with live self-hosted QA recorded in the validation
guide.

## Linked Issues or Issue Description

No public issue covers this connector. This PR incorporates the work
from #14737. The independent OAuth scope-provenance fix is #14059; it is
not required for the organization-token path. #13890 is a merged
connector precedent. #13692 provides the connector-authoring skill.

**Problem or motivation**

The generic MCP setup can reach Enterpret, but it has no reviewed
Enterpret entry, method-specific setup, artwork, or policy defaults. The
original connector branch had also fallen behind master.

**Proposed solution**

Publish the Enterpret catalog entry with an organization auth token as
the selectable method. Support browser OAuth through instance-local DCR
on the same read-only endpoint. Default `run_graph_query` to Ask first
and quarantine tools newly found on refresh.

## What Changed

- Added the Enterpret app definition, generated registry, official icon,
Browse card, setup copy, tests, and Storybook states.
- Combined #14737 and resolved the current-master conflicts, including
the connector permission audit.
- Preserved catalog quarantine on token reconnect and updated expiry
guidance to follow the token's dashboard expiry. The QA token displayed
three months, contradicting the former six-month copy.
- Enabled personal OAuth for the official read-only MCP. Removed the
write-capability inference and OAuth draft-only setup restriction. The
beta Agent endpoint is excluded.
- Preserved tool quarantine during token reconnect, with a regression
test for existing quarantined and newly discovered tools.
- Added OAuth scope and revocation-cache disclosures and an interactive
Storybook retry check.
- Merged master `d9f600043` into the existing PR branch without
rewriting history. Resolved four conflicts, regenerated the combined
registry, and retained the model-provider assertions with a catalog
count of 66.
- Recorded live QA and its limits in
[doc/connections/ENTERPRET.md](https://github.com/paperclipai/paperclip/blob/950bacc52/doc/connections/ENTERPRET.md).

## Verification

Current head: `6c73d3084202ccbb25080e7a4e9c3467bc1341dd`. All 513
focused connector, discovery, OAuth, policy, shared-definition, and
UI-handoff tests passed across 11 suites. Full repository typecheck,
build, Storybook build, token gates, module boundaries, Node policy,
no-git-push policy, and typecheck-build-gaps passed. Release-registry
tests passed 129/129 with modern Bash. The four failures with macOS Bash
3.2 also reproduced on untouched master.

The full test suite passed in CI on this exact head, including every
general-test and serialized-server shard and all eight E2E shards. The
duplicate local `pnpm test:run` was stopped after these remote test
lanes passed; it did not complete locally. The combined registry was
regenerated. Regeneration also reproduces existing Linear and AgentMail
definition drift on untouched master; those unrelated changes were
excluded.

On an isolated self-hosted Paperclip instance, an organization token
connected and discovered eight actions. `get_organization_details`
returned the intended organization before and after reconnect. Catalog
refresh, invalid-token failure and valid-token recovery, activity
logging, local disable, and local removal were exercised. No feedback
quotes or graph query were requested. Paperclip revoked both QA
connection secrets and archived the QA connections.

The token-path agent access preview correctly showed an action Off, but
the board Test endpoint still invoked it as the board user. This shared
Test-path mismatch is documented; it is not evidence that a real agent
can bypass policy. An actual token-path agent-session denial, a real
agent process, expiry recovery, server/VPS, and Cloud remain unverified.

Enterpret's dashboard confirmed the QA token was revoked, but the same
token immediately reconnected and completed a safe read. Vivek explained
this as a 24-hour validity cache and committed to reducing the window.
The new maximum delay and deployment remain unverified. The earlier
OAuth check reported `mcp:write` and `email` for an `mcp:read` request.
This is a scope mismatch; it did not prove access to write tools. On
2026-10-01 the account holder relayed Vivek’s clarification that the
official MCP is read-only and the beta Agent write MCP is separate.

## Risks

The organization token can read the organization's customer feedback,
and provider-side revocation did not immediately stop access. Operators
should check each token's displayed expiry and remove Paperclip's local
connection when retiring access. The account holder confirmed the
provider fixes are fast follows after deployment. Release can proceed
with the current scope reporting and up-to-24-hour revocation delay
documented, once fresh OAuth authorization/refresh, real-agent
allow/deny, CI, and review pass. Retest the reduced revocation window
after Enterpret deploys it. Scope labels alone must not be used to claim
write access.

Greptile is 5/5 on head `6c73d3084` with no actionable findings. All
current-head CI checks pass, including the canary dry run. Live
agent-session denial and fresh OAuth refresh remain explicit deployment
QA items. The board Test panel does not prove agent authorization.
Provider scope-reporting and revocation-cache fixes are accepted fast
follows after deployment.

## Model Used

Original connector and validation record: Claude Opus 5 through Claude
Code. Integration, current-master reconciliation, and token-path QA:
OpenAI Codex, GPT-6 family (exact host variant not exposed), using
repository tools and an isolated runtime. Contributor commits retain
their authors.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either linked existing issues or described the issue in-PR
- [x] I have not referenced internal or instance-local Paperclip issues
or links
- [x] My branch name describes the change and contains no internal issue
ID
- [x] I have run relevant local tests and checks
- [x] I have updated relevant documentation
- [x] I have considered and documented the risks above
- [x] All current-head CI gates are green
- [x] Greptile is 5/5 with no actionable follow-ups

If squash-merging, preserve contributor credit in the squash message:

Co-Authored-By: Vivek Kaushal <kaushalvivek@users.noreply.github.com>

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Vivek Kaushal <vivek@enterpret.com>
2026-10-06 14:35:51 -05:00
Nicky LeachandPaperclip 1cebdd4c4f Add plugin lifecycle delivery and an initial resource baseline (#15343)
## Thinking Path

> - Paperclip manages agents and projects for work.
> - Their lifecycle changes already commit to a durable journal.
> - Plugins need to read these records and track completed work.
> - Existing resources need a one-time baseline before delivery is
enabled.
> - Worker crashes must leave unfinished records available for retry.
> - Each plugin needs its own progress and company access checks.
> - This PR adds a pull inbox through the existing plugin SDK and job
system.

## Linked Issues or Issue Description

**Problem or motivation**

Plugins cannot consume the durable resource lifecycle journal.
In-process notifications can disappear during a restart and cannot
record successful completion. Resource plugins need retries, company
boundaries, and ordered transitions.

**Proposed solution**

Add `ctx.events.listLifecycle(companyId, limit?, afterId?)` and
`ctx.events.acknowledgeLifecycle(companyId, eventId)` under
`events.subscribe`. Seed a one-time baseline from current resource
state. Deliver creation first, then the remaining transitions in ID
order. Store acknowledgments for each plugin. Use existing plugin jobs
to poll configured companies.

**Alternatives considered**

A global sequence cursor can skip lower IDs that commit later.
Fire-and-forget subscriptions cannot record completion. A new dispatcher
is unnecessary because plugin jobs support polling. Consumers serialize
polling and use provider idempotency keys.

**Roadmap alignment**

This extends the existing plugin system and builds on #15280 and #15306.
Searches found no duplicate resource inbox work. Related #13306 exposes
decision events on the in-process bus. This PR includes the initial
journal baseline. Provider provisioning remains separate work.

## What Changed

- Add company-scoped lifecycle reads and acknowledgments to the SDK and
worker RPC host.
- Gate both methods by capability, invocation or proactive company
scope, plugin readiness, and company enablement.
- Seed hired agents and all projects once. Preserve paused/terminated
agent state and archived project state. Keep pending hires behind
approval.
- Preserve existing history and deliver backfilled creation before
partial transition histories.
- Deliver project archive events from #15371 through the SDK. Seed
archive intents for archived projects and update intents for active
projects with a partial archive history.
- Store acknowledgments per plugin and reject acknowledgments that skip
earlier resource events.
- Page past failed resources while retaining their pending records.
Reset the page cursor each sweep to include late commits.
- Add acknowledgment storage, an index for resource ordering, tests, and
authoring guidance.
- Preserve native identity definitions and sequence progress in
JavaScript backups. Repair journal ID generators lost by older backups
before seeding the baseline, without changing existing IDs.

## Verification

- `pnpm -r typecheck` passed after the final origin/master rebase.
- `pnpm build` passed before the final metadata rebase. The final CI
build also passed.
- All 46 focused lifecycle, SDK harness/RPC, migration snapshot, and
legacy restore checks passed again with migration 0309.
- Checks cover retries, per-plugin progress, resource order,
capability/company boundaries, baseline idempotency, approval gates,
archive delivery, restored projects with partial histories, and atomic
migration rollback.
- Full CI passed on final commit `0e674b87fc`. One unrelated Cursor
fixture hit a 10-second timeout; it passed locally in under one second,
and the failed server shard passed on retry.
- Greptile scored the final commit 5/5 with no unresolved review
threads.
- The branch is rebased onto origin/master and is conflict-free.
- Earlier full local runs were stopped as scope changed. CI ran the
complete repository suite.
- `git diff --check` and a local secret/PII scan passed.

## Risks


- Apply migration `0309_loving_the_hood.sql` before starting the new
server. It creates acknowledgment storage and an index, then seeds the
baseline in the same transaction. Agent/project writes wait for the
migration to commit. Keep these writes quiesced through the migration
and activation of the new capture-capable server; do not resume an older
runtime that lacks capture after the baseline.
- The baseline records current desired state, not historical
transitions. It includes archived projects and terminated-agent cleanup
intents. It runs once with the delivery migration; no later or runtime
journal backfill is planned.
- Delivery is at least once. Concurrent reads can repeat an event.
Consumers must serialize polling, use stable company/event idempotency
keys, and acknowledge successful operations only.
- Reset `afterId` at the start of every polling sweep. It is a page
cursor, not a persisted high-water mark.
- Deleting a plugin removes its acknowledgment records. A new
installation may replay existing journal events.
- Records contain identity and action. Consumers must load current
authorized data before acting. Cleanup and retention policy belong to
the provider plugin.
- Provider calls, VM/volume provisioning, and journal retention are
outside this PR.

## Model Used

OpenAI Codex, based on GPT-6, with reasoning, repository inspection,
code execution, and tool use. The exact deployment model ID and context
window are not exposed in this session.

## Checklist


- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 12:27:09 -07:00
854af7df19 build(deps-dev): bump vitest from 4.1.11 to 5.0.3 (#12969)
Bumps
[vitest](https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest)
from 4.1.11 to 5.0.3.
<details>
<summary>Release notes</summary>
<p><em>Sourced from <a
href="https://github.com/vitest-dev/vitest/releases">vitest's
releases</a>.</em></p>
<blockquote>
<h2>v5.0.3</h2>
<h3>   🐞 Bug Fixes</h3>
<ul>
<li>Isolate <code>result.status</code> between <code>repeats</code> runs
 -  by <a
href="https://github.com/hi-ogawa"><code>@​hi-ogawa</code></a>,
<strong>Hiroshi Ogawa</strong> and <strong>Codex (GPT-6)</strong> in <a
href="https://redirect.github.com/vitest-dev/vitest/issues/11218">vitest-dev/vitest#11218</a>
<a href="https://github.com/vitest-dev/vitest/commit/5dbebe9e3"><!-- raw
HTML omitted -->(5dbeb)<!-- raw HTML omitted --></a></li>
<li>Don't print an interceptor warning in browser mode  -  by <a
href="https://github.com/sheremet-va"><code>@​sheremet-va</code></a> in
<a
href="https://redirect.github.com/vitest-dev/vitest/issues/11377">vitest-dev/vitest#11377</a>
<a href="https://github.com/vitest-dev/vitest/commit/15cc006aa"><!-- raw
HTML omitted -->(15cc0)<!-- raw HTML omitted --></a></li>
<li>Don't retry when <code>test.fails</code> expectedly failed  -  by <a
href="https://github.com/hi-ogawa"><code>@​hi-ogawa</code></a>,
<strong>Hiroshi Ogawa</strong> and <strong>Codex (GPT-6)</strong> in <a
href="https://redirect.github.com/vitest-dev/vitest/issues/11219">vitest-dev/vitest#11219</a>
<a href="https://github.com/vitest-dev/vitest/commit/b24585f08"><!-- raw
HTML omitted -->(b2458)<!-- raw HTML omitted --></a></li>
<li>Scope cache key generators to projects  -  by <a
href="https://github.com/ecoyoung"><code>@​ecoyoung</code></a> in <a
href="https://redirect.github.com/vitest-dev/vitest/issues/11281">vitest-dev/vitest#11281</a>
and <a
href="https://redirect.github.com/vitest-dev/vitest/issues/11301">vitest-dev/vitest#11301</a>
<a href="https://github.com/vitest-dev/vitest/commit/92ba7fc1d"><!-- raw
HTML omitted -->(92ba7)<!-- raw HTML omitted --></a></li>
<li><strong>browser</strong>:
<ul>
<li>Delay server <code>listen</code> until tests start running  -  by <a
href="https://github.com/sheremet-va"><code>@​sheremet-va</code></a> in
<a
href="https://redirect.github.com/vitest-dev/vitest/issues/11366">vitest-dev/vitest#11366</a>
<a href="https://github.com/vitest-dev/vitest/commit/7d8ed3e9b"><!-- raw
HTML omitted -->(7d8ed)<!-- raw HTML omitted --></a></li>
<li>Check mock path boundaries  -  by <a
href="https://github.com/saryn17"><code>@​saryn17</code></a>,
<strong>Ryosei Sato</strong> and <a
href="https://github.com/sheremet-va"><code>@​sheremet-va</code></a> in
<a
href="https://redirect.github.com/vitest-dev/vitest/issues/11361">vitest-dev/vitest#11361</a>
and <a
href="https://redirect.github.com/vitest-dev/vitest/issues/11362">vitest-dev/vitest#11362</a>
<a href="https://github.com/vitest-dev/vitest/commit/1c3888bce"><!-- raw
HTML omitted -->(1c388)<!-- raw HTML omitted --></a></li>
<li>Keep config of browser-consumed environments  -  by <a
href="https://github.com/kasperpeulen"><code>@​kasperpeulen</code></a>
and <strong>Claude</strong> in <a
href="https://redirect.github.com/vitest-dev/vitest/issues/11378">vitest-dev/vitest#11378</a>
<a href="https://github.com/vitest-dev/vitest/commit/aafc0996f"><!-- raw
HTML omitted -->(aafc0)<!-- raw HTML omitted --></a></li>
<li>Ignore page crash while cancelling  -  by <a
href="https://github.com/sheremet-va"><code>@​sheremet-va</code></a> in
<a
href="https://redirect.github.com/vitest-dev/vitest/issues/11386">vitest-dev/vitest#11386</a>
<a href="https://github.com/vitest-dev/vitest/commit/7c36748fa"><!-- raw
HTML omitted -->(7c367)<!-- raw HTML omitted --></a></li>
<li><code>toMatchScreenshot</code> uses wrong reference on retried tests
 -  by <a href="https://github.com/macarie"><code>@​macarie</code></a>
in <a
href="https://redirect.github.com/vitest-dev/vitest/issues/11393">vitest-dev/vitest#11393</a>
<a href="https://github.com/vitest-dev/vitest/commit/c22aba992"><!-- raw
HTML omitted -->(c22ab)<!-- raw HTML omitted --></a></li>
</ul>
</li>
<li><strong>cache</strong>:
<ul>
<li>Revalidate imports of cached modules  -  by <a
href="https://github.com/sheremet-va"><code>@​sheremet-va</code></a> in
<a
href="https://redirect.github.com/vitest-dev/vitest/issues/11381">vitest-dev/vitest#11381</a>
<a href="https://github.com/vitest-dev/vitest/commit/38f98855f"><!-- raw
HTML omitted -->(38f98)<!-- raw HTML omitted --></a></li>
</ul>
</li>
<li><strong>deps</strong>:
<ul>
<li>Pin <code>why-is-node-running</code> to <code>3.2.1</code> to avoid
users running into <code>ERR_PNPM_TRUST_DOWNGRADE</code>  -  by <a
href="https://github.com/AriPerkkio"><code>@​AriPerkkio</code></a> in <a
href="https://redirect.github.com/vitest-dev/vitest/issues/11403">vitest-dev/vitest#11403</a>
<a href="https://github.com/vitest-dev/vitest/commit/f6c9a4977"><!-- raw
HTML omitted -->(f6c9a)<!-- raw HTML omitted --></a></li>
</ul>
</li>
<li><strong>expect</strong>:
<ul>
<li>Pass current equality testers to <code>expect.extend</code>
asymmetric matchers  -  by <a
href="https://github.com/hi-ogawa"><code>@​hi-ogawa</code></a>,
<strong>Hiroshi Ogawa</strong> and <strong>Claude</strong> in <a
href="https://redirect.github.com/vitest-dev/vitest/issues/11401">vitest-dev/vitest#11401</a>
<a href="https://github.com/vitest-dev/vitest/commit/3e794a96b"><!-- raw
HTML omitted -->(3e794)<!-- raw HTML omitted --></a></li>
</ul>
</li>
<li><strong>jsdom</strong>:
<ul>
<li>Support Blob on jsdom 30.1  -  by <a
href="https://github.com/sheremet-va"><code>@​sheremet-va</code></a> in
<a
href="https://redirect.github.com/vitest-dev/vitest/issues/11379">vitest-dev/vitest#11379</a>
<a href="https://github.com/vitest-dev/vitest/commit/6c49b7197"><!-- raw
HTML omitted -->(6c49b)<!-- raw HTML omitted --></a></li>
</ul>
</li>
<li><strong>pool</strong>:
<ul>
<li>Preserve unique pool ids when <code>groupOrder</code> is set  -  by
<a href="https://github.com/mtorp"><code>@​mtorp</code></a> in <a
href="https://redirect.github.com/vitest-dev/vitest/issues/11392">vitest-dev/vitest#11392</a>
<a href="https://github.com/vitest-dev/vitest/commit/50312ebb4"><!-- raw
HTML omitted -->(50312)<!-- raw HTML omitted --></a></li>
</ul>
</li>
<li><strong>ui</strong>:
<ul>
<li>Split-pane handle overlapping iframe  -  by <a
href="https://github.com/macarie"><code>@​macarie</code></a> in <a
href="https://redirect.github.com/vitest-dev/vitest/issues/11221">vitest-dev/vitest#11221</a>
<a href="https://github.com/vitest-dev/vitest/commit/f91db0dfd"><!-- raw
HTML omitted -->(f91db)<!-- raw HTML omitted --></a></li>
</ul>
</li>
<li><strong>vitest</strong>:
<ul>
<li>Remove root temp dir on close  -  by <a
href="https://github.com/abhinav-phi"><code>@​abhinav-phi</code></a> in
<a
href="https://redirect.github.com/vitest-dev/vitest/issues/11248">vitest-dev/vitest#11248</a>
<a href="https://github.com/vitest-dev/vitest/commit/7c7119cf7"><!-- raw
HTML omitted -->(7c711)<!-- raw HTML omitted --></a></li>
</ul>
</li>
<li><strong>vm</strong>:
<ul>
<li>Do not optimize deps from index.html  -  by <a
href="https://github.com/ezefernandezyf"><code>@​ezefernandezyf</code></a>,
<strong>Hiroshi Ogawa</strong> and <strong>Codex (GPT-6)</strong> in <a
href="https://redirect.github.com/vitest-dev/vitest/issues/11329">vitest-dev/vitest#11329</a>
and <a
href="https://redirect.github.com/vitest-dev/vitest/issues/11360">vitest-dev/vitest#11360</a>
<a href="https://github.com/vitest-dev/vitest/commit/caf2887de"><!-- raw
HTML omitted -->(caf28)<!-- raw HTML omitted --></a></li>
<li>Don't reuse scripts across vite environments  -  by <a
href="https://github.com/MO2k4"><code>@​MO2k4</code></a>, <strong>Martin
Oehlert</strong> and <strong>Claude</strong> in <a
href="https://redirect.github.com/vitest-dev/vitest/issues/11395">vitest-dev/vitest#11395</a>
<a href="https://github.com/vitest-dev/vitest/commit/346d3896b"><!-- raw
HTML omitted -->(346d3)<!-- raw HTML omitted --></a></li>
</ul>
</li>
</ul>
<h5>    <a
href="https://github.com/vitest-dev/vitest/compare/v5.0.2...v5.0.3">View
changes on GitHub</a></h5>
<h2>v5.0.2</h2>
<h3>   🐞 Bug Fixes</h3>
<ul>
<li>Bind <code>process</code> in case global is overwritten  -  by <a
href="https://github.com/AriPerkkio"><code>@​AriPerkkio</code></a> in <a
href="https://redirect.github.com/vitest-dev/vitest/issues/11343">vitest-dev/vitest#11343</a>
<a href="https://github.com/vitest-dev/vitest/commit/0b79231ad"><!-- raw
HTML omitted -->(0b792)<!-- raw HTML omitted --></a></li>
<li><strong>detect-async-leaks</strong>:
<ul>
<li>Ignore <code>process.stdio</code> handles  -  by <a
href="https://github.com/AriPerkkio"><code>@​AriPerkkio</code></a> in <a
href="https://redirect.github.com/vitest-dev/vitest/issues/11333">vitest-dev/vitest#11333</a>
<a href="https://github.com/vitest-dev/vitest/commit/0fd6b9790"><!-- raw
HTML omitted -->(0fd6b)<!-- raw HTML omitted --></a></li>
</ul>
</li>
<li><strong>expect</strong>:
<ul>
<li>Fix <code>toMatchObject</code> with asymmetric matchers  -  by <a
href="https://github.com/ShreeBohara"><code>@​ShreeBohara</code></a>,
<strong>Claude Opus 5</strong>, <a
href="https://github.com/hi-ogawa"><code>@​hi-ogawa</code></a>,
<strong>Hiroshi Ogawa</strong> and <strong>Codex (GPT-5)</strong> in <a
href="https://redirect.github.com/vitest-dev/vitest/issues/11100">vitest-dev/vitest#11100</a>
<a href="https://github.com/vitest-dev/vitest/commit/42523289e"><!-- raw
HTML omitted -->(42523)<!-- raw HTML omitted --></a></li>
</ul>
</li>
<li><strong>jsdom</strong>:
<ul>
<li>Fix <code>Request</code> with <code>Blob</code> body on jsdom 28+
 -  by <a
href="https://github.com/harshit-d3v"><code>@​harshit-d3v</code></a> in
<a
href="https://redirect.github.com/vitest-dev/vitest/issues/11295">vitest-dev/vitest#11295</a>
<a href="https://github.com/vitest-dev/vitest/commit/d1c3ecc93"><!-- raw
HTML omitted -->(d1c3e)<!-- raw HTML omitted --></a></li>
</ul>
</li>
<li><strong>reporter</strong>:
<ul>
<li><code>agent</code> to respect <code>--silent</code>  -  by <a
href="https://github.com/Raj4478"><code>@​Raj4478</code></a> and <a
href="https://github.com/AriPerkkio"><code>@​AriPerkkio</code></a> in <a
href="https://redirect.github.com/vitest-dev/vitest/issues/11271">vitest-dev/vitest#11271</a>
<a href="https://github.com/vitest-dev/vitest/commit/5b95efb6d"><!-- raw
HTML omitted -->(5b95e)<!-- raw HTML omitted --></a></li>
</ul>
</li>
<li><strong>reporters</strong>:
<ul>
<li>Handle concurrent <code>createReport</code> calls  -  by <a
href="https://github.com/7rulnik"><code>@​7rulnik</code></a> in <a
href="https://redirect.github.com/vitest-dev/vitest/issues/11278">vitest-dev/vitest#11278</a>
<a href="https://github.com/vitest-dev/vitest/commit/e8e556ff7"><!-- raw
HTML omitted -->(e8e55)<!-- raw HTML omitted --></a></li>
<li><code>hanging-process</code> to use ESM entrypoint  -  by <a
href="https://github.com/AriPerkkio"><code>@​AriPerkkio</code></a> in <a
href="https://redirect.github.com/vitest-dev/vitest/issues/11316">vitest-dev/vitest#11316</a>
<a href="https://github.com/vitest-dev/vitest/commit/4e91e5668"><!-- raw
HTML omitted -->(4e91e)<!-- raw HTML omitted --></a></li>
</ul>
</li>
<li><strong>spy</strong>:
<ul>
<li>Fix stack overflow when spying <code>Set.prototype.add</code>  -  by
<a href="https://github.com/fengmk2"><code>@​fengmk2</code></a> in <a
href="https://redirect.github.com/vitest-dev/vitest/issues/11299">vitest-dev/vitest#11299</a>
<a href="https://github.com/vitest-dev/vitest/commit/a0a939653"><!-- raw
HTML omitted -->(a0a93)<!-- raw HTML omitted --></a></li>
</ul>
</li>
</ul>
<!-- raw HTML omitted -->
</blockquote>
<p>... (truncated)</p>
</details>
<details>
<summary>Commits</summary>
<ul>
<li><a
href="https://github.com/vitest-dev/vitest/commit/33cadea62e8763c455c7fca38d9ab1dda87c5f75"><code>33cadea</code></a>
chore: release v5.0.3 (<a
href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11409">#11409</a>)</li>
<li><a
href="https://github.com/vitest-dev/vitest/commit/346d3896b65c3c907174447a035807342799f346"><code>346d389</code></a>
fix(vm): don't reuse scripts across vite environments (<a
href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11395">#11395</a>)</li>
<li><a
href="https://github.com/vitest-dev/vitest/commit/f6c9a4977ad3363f796a737834572e54c6ad5c18"><code>f6c9a49</code></a>
fix(deps): pin <code>why-is-node-running</code> to <code>3.2.1</code> to
avoid users running into `...</li>
<li><a
href="https://github.com/vitest-dev/vitest/commit/062c75d8b63519211d951d8293ea81b5a9e3c124"><code>062c75d</code></a>
chore: fix standalone docs build, update exports maps (<a
href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11394">#11394</a>)</li>
<li><a
href="https://github.com/vitest-dev/vitest/commit/caf2887dee8987a60118d53933f6e9cabd6b3e2a"><code>caf2887</code></a>
fix(vm): do not optimize deps from index.html (fix <a
href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11329">#11329</a>)
(<a
href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11360">#11360</a>)</li>
<li><a
href="https://github.com/vitest-dev/vitest/commit/50312ebb4eca6a98f6d0b2b61d5d9d38cbbabcef"><code>50312eb</code></a>
fix(pool): preserve unique pool ids when <code>groupOrder</code> is set
(<a
href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11392">#11392</a>)</li>
<li><a
href="https://github.com/vitest-dev/vitest/commit/7c36748fad1eae9687312f2f7ceadce6ec88b5df"><code>7c36748</code></a>
fix(browser): ignore page crash while cancelling (<a
href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11386">#11386</a>)</li>
<li><a
href="https://github.com/vitest-dev/vitest/commit/92ba7fc1df16a4fa5bbee3f198c582fbd56689d8"><code>92ba7fc</code></a>
fix: scope cache key generators to projects (fix <a
href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11281">#11281</a>)
(<a
href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11301">#11301</a>)</li>
<li><a
href="https://github.com/vitest-dev/vitest/commit/38f98855fa9cd7fd376afb84094eba0fda256a74"><code>38f9885</code></a>
fix(cache): revalidate imports of cached modules (<a
href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11381">#11381</a>)</li>
<li><a
href="https://github.com/vitest-dev/vitest/commit/b24585f08f2ea267746a2d6ca0e43edcbb29726f"><code>b24585f</code></a>
fix: don't retry when <code>test.fails</code> expectedly failed (<a
href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11219">#11219</a>)</li>
<li>Additional commits viewable in <a
href="https://github.com/vitest-dev/vitest/commits/v5.0.3/packages/vitest">compare
view</a></li>
</ul>
</details>
<br />

---------

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Priya Raman <priya.raman@paperclip.ing>
Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-10-06 12:13:46 -07:00
DottaandPaperclip a590ab769d Give managed agents persistent cryptographic identities (#15352)
Give agents persistent Ed25519 identities encrypted with the existing instance master key. Create keys transactionally for new agents and lazily before supported managed runs, expose public identities in the API and agent UI, and protect private material during runtime delivery and output persistence.

Preserve identities in recovery backups while giving imported and development-cloned agents fresh keys.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-06 13:57:03 -05:00