Commit Graph
849 Commits
Author SHA1 Message Date
DottaandPaperclip b0bbaffec6 Clarify the Hermes native question fixture objective
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 12:28:13 -05:00
DottaandPaperclip 3c7e0ef3a1 Test Hermes native question batches through the browser
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 12:28:13 -05:00
DottaandPaperclip 1500900ece Test Hermes native question batches through runnerd
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 12:28:13 -05:00
DottaandPaperclip a8cce899fc Fix recovery question submission labels
Use the canonical saved question presentation when submitting recovery answers and record current Mac qualification evidence.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 12:28:13 -05:00
DottaandPaperclip 6788c30cf2 fix(hermes): make Mac runtime signing reproducible across hosts
Specify 16 KiB code-signing pages for the normalized Python library. The 4 KiB default reproduces the exact rejected cloud hash; 16 KiB reproduces the existing reviewed bytes. Keep closure pins and dependency locks unchanged.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 12:28:13 -05:00
DottaandPaperclip a60b6badb8 ci(hermes): retain rejected Mac closure manifest for byte comparison
Preserve file-level provisioning evidence before cleanup without weakening reviewed closure admission. Record the original cloud failure and the passing current-source native fixtures on macOS 26.5.2.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 12:28:13 -05:00
DottaandPaperclip 40f6ddac33 ci(hermes): qualify Mac arm64 and export tested cloud runner
Build both native targets on standard cloud runners. Keep fixtures credential-free, validate host and binary architecture, and publish exact source, runtime, tool and binary provenance for acceptance without local Rust builds.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 12:28:13 -05:00
DottaandPaperclip 1b31d77c17 fix(hermes): admit verified generated assets without weakening source checks
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 12:28:13 -05:00
DottaandPaperclip 83f4a1a136 fix(hermes): retain v7 permissions after master synchronization
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 12:28:13 -05:00
DottaandPaperclip efd5f5d74b fix(hermes): settle native usage before governed confirmations
Extend the managed human-input boundary to confirmations and checkbox
confirmations. Preserve native receipts before controller parking and
cover immediate shutdown for all three canonical input kinds.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 12:28:13 -05:00
DottaandPaperclip f29f7d89b9 fix(hermes): delay bridge question completion until native receipt
Return the committed question to Hermes before publishing the tool bridge's
own completion fact. Hold that fact until native terminal delivery, after
final prompt usage. Keep other harnesses on their existing order and treat a
typed already-terminal passive interrupt as settled cancellation.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 12:28:12 -05:00
DottaandPaperclip 23919a8787 fix(hermes): settle native usage before question yield
Stop native Hermes at a committed assigned question. Preserve final prompt
usage before its completed tool result reaches the controller cancellation
boundary. Keep ordinary tools streaming and add an immediate-shutdown native
regression with pinned runtime closure updates.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 12:28:12 -05:00
DottaandPaperclip e3be06f25e fix(hermes): retain usage after input-yield shutdown
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 12:28:12 -05:00
DottaandPaperclip 842925d08c fix(hermes): settle closed retry work with unknown usage
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 12:28:12 -05:00
DottaandPaperclip e5afbd7107 fix(hermes): retain per-turn reported OpenRouter billing
Capture pinned SDK wire attempts, carry closed versioned receipts through ACPX and PRP, and bind reported subtotals to native turn accounting. Keep incomplete attempts unpriced and require settled cost plus budget health in the live OpenRouter oracle.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 12:28:12 -05:00
DottaandPaperclip 01855199ad fix(hermes): verify approved qualification lockfile
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 12:28:12 -05:00
DottaandPaperclip 6d81d5c0dd fix(hermes): record checked-out qualification source before credentials
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 12:28:12 -05:00
DottaandPaperclip 29f1c087c5 fix(hermes): preserve pinned setup configuration and verify runtime files
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 12:28:12 -05:00
DottaandPaperclip 51700ff8c6 feat(hermes): ship explicit public runtime setup
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 12:28:12 -05:00
DottaandPaperclip af50ff553d test(hermes): add bounded managed Bedrock qualification
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 12:28:12 -05:00
DottaandPaperclip 08d9803b8f test(hermes): verify the expected account owner independently
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 12:28:12 -05:00
DottaandPaperclip 6f5748bb95 docs(hermes): record API completion and answer-limit evidence
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 12:28:12 -05:00
DottaandPaperclip 49e99eca47 test(hermes): use the current Google qualification model
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 12:28:12 -05:00
DottaandPaperclip 6f75f7a59c ci(hermes): supply the selected Google qualification key
Map GEMINI_API_KEY only for the selected paid matrix credential. Extend the trusted-workflow credential checks to prevent ambient Google secrets in unapproved workflows. This fixes the reviewed Google cells startup failure.

Validation: all 15 workflow security tests, actionlint, and diff checks pass.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 12:28:11 -05:00
DottaandPaperclip 4313a55dde test(hermes): bound API qualification before task dispatch
Require one attempt and public readback of scoped 200-cent company and agent budgets before the paid API task is created. Reject unlimited or foreign budget records. Preserve unknown usage as unknown.

Record all five successful paid Linux browser workflows and verified temporary AWS host/access cleanup. The complete support suite passes 1,833 Vitest tests and 128 Node assertions, with one skipped Vitest test. Product E2E TypeScript and diff checks pass.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 12:28:11 -05:00
DottaandPaperclip b10746b4e6 test(hermes): qualify managed API accounts through the runner
Add ten explicit pending API-account browser cells and independently grade the selected account, native harness, and exact model. Keep fixture credential staging on the shared Connections path. Carry candidate metadata into CI so every selected Hermes profile receives verified runtime assets before secrets.

Record the private AWS image build and three actual Linux browser passes without promoting Hermes. Product E2E support typecheck and all 102 affected tests pass; the broader support suite passed 1,826 Vitest tests and 128 node assertions before the final oracle refinements.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 12:28:11 -05:00
DottaandPaperclip 6f87ea523b ci(hermes): cover shared native fixture inputs
Trigger credential-free Linux native fixtures for shared Runner sources, Rust and build manifests, dependency patches, and workspace inputs. Verify actual glob matching for attachment, backend, transport and dependency changes. Paid workflow authority is unchanged.

Validation: 15 workflow security tests and actionlint passed. Also record 26 passing real-database authority tests, the expected old-lock negative proof, and 143 passing ACPX lifecycle regressions.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 12:28:11 -05:00
DottaandPaperclip d8453bf775 docs(hermes): record review fixes and replayed Linux proof
Record the encoded attachment and independent planning fixes, the routine contention regression awaiting a healthy database host, and the passing replayed Linux transport evidence. Preserve all paid qualification failures and the pending release gate.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 12:28:11 -05:00
DottaandPaperclip 3dfd3da7f9 docs(hermes): record Linux transport proof and stack replay
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 12:28:11 -05:00
DottaandPaperclip 3561e1edca ci(hermes): qualify the fixture Node interpreter
Remove group/world write bits from the dedicated CI job's Node interpreter,
matching the protected paid workflow setup. Preserve the Rust launch verifier
and record the first Linux fixture result and interrupted PR checks.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 12:28:11 -05:00
DottaandPaperclip 6cf6cef34a fix(hermes): bound cold native admission separately
Give Hermes native session opening a 60-second bound for verified private
runtime copying and ACP initialization. Allow controller startup and per-turn
restoration overhead while preserving ordinary command and stop deadlines.

Add delayed-admission regressions and credential-free Linux native fixture CI.
Record the failed final-head paid campaign without promoting qualification.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 12:28:11 -05:00
DottaandPaperclip 42f91e33ac docs(hermes): record paid acceptance and qualification limits
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 12:28:11 -05:00
DottaandPaperclip 4a0b327727 test(hermes): add paid browser qualification campaigns
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 12:28:11 -05:00
Dotta c834f9dd8a fix(hermes): release credentials and reproduce relocatable runtime assets 2026-10-08 12:07:09 -05:00
DottaandPaperclip 9e07748e91 feat(runner): add native Hermes ACP qualification candidate
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-08 12:07:09 -05:00
Devin FoleyandPaperclip 5717523b9e fix: retain the original repository for reused task workspaces (#15528)
Apply the reviewed change for fix: retain the original repository for reused task workspaces.

Validation: required local typecheck and tests, passing CI, Greptile 5/5, and independent review of the exact source head.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-07 19:52:30 -07:00
Devin FoleyandPaperclip 1d19f9b562 Distinguish failed Git inspection from missing worktree registration (#15525)
Apply the reviewed change for Distinguish failed Git inspection from missing worktree registration.

Validation: required local typecheck and tests, passing CI, Greptile 5/5, and independent review of the exact source head.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-07 19:05:04 -07:00
Devin FoleyandPaperclip ed0c6958f8 Explain refused local Hermes gateway connections (#15520)
Apply the reviewed change for Explain refused local Hermes gateway connections.

Validation: required local typecheck and tests, passing CI, Greptile 5/5, and independent review of the exact source head.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-07 18:55:55 -07:00
Devin FoleyandPaperclip 6b171c9616 Keep proven local workspace configuration conflicts out of Sentry (#15521)
Apply the reviewed change for Keep proven local workspace configuration conflicts out of Sentry.

Validation: required local typecheck and tests, passing CI, Greptile 5/5, and independent review of the exact source head.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-07 18:54:35 -07:00
DottaandPaperclip dd777f4b73 feat(dot): complete onboarding and expand governed Runner capabilities (#15414)
## Thinking Path

> - Paperclip manages AI agents, their work, and their permissions.
> - Paperclip Runner supplies the same admitted tool authority to each
provider.
> - The Dot provider in #15402 needs reliable onboarding and useful
agent capabilities.
> - An idle Dot could not start work, assign a human, read skills, or
produce a workspace artifact.
> - Pairing also relied on a second configuration save before ordinary
admission could work.
> - This pull request adds governed idle admission and shared Runner
tools, and completes pairing atomically.
> - Operators can test Dot with fresh data while keeping normal company,
approval, budget, and run ownership checks.

## Linked Issues or Issue Description

**Subsystem affected**

Paperclip Runner, dedicated Dot MCP access, OAuth onboarding,
experimental settings, and empty worktree startup. This PR builds on
merged provider PR #15402. It reuses the merged MCP gateway from #14846
and assistant connection work from #14933 and #15380.

**Problem or motivation**

An idle Dot could see assignments but could not act on a conversation
request until someone created a task first. Its Runner catalog could not
assign tasks to humans or read pinned skills and workspace files.
First-time OAuth discovery and pairing also needed browser fixes, and a
completed pairing did not persist its binding reference on the agent.

**Proposed solution**

Keep Dot within the existing Runner. Admit a visible agent-authored
intake task for idle requests. Add people, human assignment, cross-task,
skill, and optional sandbox workspace tools through shared authority.
Relay assigned app calls through the configured MCP gateway. Add lease
renewal and follow-up references. Save pairing and its configuration
revision atomically. Show prerequisites and provide a complete copy
prompt.

**Roadmap alignment**

This extends the experimental provider in #15402. It uses the existing
governed gateway, task model, skills, artifact path, and native Runner.
It adds no separate execution subsystem.

## What Changed

- Add a standalone OpenAI Dot agent choice with an independent
experimental opt-in. It works with the general Runner option off.
Require Assistant connections (MCP), authenticated sign-in, and public
HTTPS for pairing.
- Use Dot’s own option for company package import. Allow an unpaired Dot
configuration to save after external billing acknowledgement; task
admission still requires pairing. Prepare the shared dev binary for
Dot-only opt-in.
- Label saved Dot agents as OpenAI Dot. Use prerequisite-check copy in
setup and runtime configuration.
- Route Dot creation directly to pairing after explicit external billing
acknowledgement. Hide local CLI, model, and harness setup for Dot.
Preserve the shared Runner implementation and canonical API type. Other
Runner providers still require their general opt-in.
- Add an empty worktree option with fresh signing keys and no production
data copy.
- Fix public-client OAuth negotiation, discovery compatibility, and
optional separate browser authorization origin.
- Add one-use pairing consent preview, clear copied setup instructions,
and atomic binding persistence and cleanup.
- Add idle request admission with stable request IDs and normal
scheduling, permissions, budgets, and task ownership.
- Add identity and people discovery, human task assignment and
reassignment, and authorized cross-task comments and documents.
- Read assigned skill files from pinned manifests. Relay assigned app
calls through the merged gateway without exposing credentials.
- Add an off-by-default workspace bridge. Constrain paths and writes.
Run commands in a deny-by-default OS sandbox with no network or injected
credentials. Reserve mutations before effects and never blindly repeat
uncertain work.
- Add rolling lease renewal, task pagination, bounded operation limits,
and deduplicated follow-up references without comment bodies in
webhooks.
- Add an off-by-default attachment reading setting. Restrict reads to
files on the current assigned task. Verify size and hash, cache bounded
verified copies per run, paginate text or binary bytes, and recheck live
authority before returning.
- Keep file grants operator-owned. Reject agent self-grants across
configuration routes. Preserve attachment consent in create/import
forms. Close generic API file bypasses while retaining current-run
response snapshots and permitted uploads.
- Keep provider limits explicit. Do not inherit a Dot binding,
attachment permission, or workspace permission when hiring another
agent.
- Include the required Markdown format in cross-task document writes and
validate the API title limit. Verify real creation and revision
persistence.
- Restrict command execution to Linux bubblewrap with descendant
containment. macOS retains workspace file tools and artifact publishing,
while refusing command calls. Explain the platform limit in setup.
- Exclude Paperclip instance state from workspace files, uploads,
artifact publication, and sandbox commands. Protect nested directories
and case variants. Fail closed when the directory protection scan
exceeds 4,096 directories.
- Add static UI compression for slow public tunnels. Document setup, the
complete tool inventory, and qualification limits.

## Verification

- Merge preparation on `00ca2c75b` integrates merged base #15402 and
master `fc6304dfe`. The ancestry commit preserves the reviewed follow-up
source tree. The subsequent security fix excludes instance state from
file tools, uploads, artifact publication, and sandbox commands. It
preserves private task authorization and task monitors. It regenerates
the combined tool catalog, seeded catalog digest, and protocol manifest.
Workspace typecheck, full build, and UI token gates pass. All sixteen
real Dot broker cases and 93 company import cases pass after the review
fixes and creator-attribution test correction. The exported
human-assignment catalog and Unicode page boundaries are also fixed. All
ten catalog tests and nine workspace/skill bridge tests pass. All 35
workspace bridge and authority tests passed after the instance-state
fix, including real macOS commands in that intermediate version. The
subsequent document and descendant-containment fixes pass 44 focused
tests across bridge, authority, and setup UI, with six Linux command
cases skipped on macOS. Real cross-task documents pass the route
validator and persist two revisions. macOS refuses command execution and
does not advertise the tool. Full workspace typecheck and build, changed
server/UI typechecks, and token gates pass again on the final commit.
Greptile rates final head 00ca2c75b 5/5 and its completed check
concludes success; all review threads are resolved. All 58 current-head
checks are complete: 54 passed and 4 intentionally skipped. This
includes full tests, typecheck, build, Runner, browser E2E, release
verification, and Canary Dry Run. The older full local root attempt
finished with 16,602 passes, 93 skips, and ten failures across two
suites: it began before source edits and retained earlier imported
implementations while loading later tests. Both suites pass a clean
final-head rerun (25 passed, six Linux-only command cases skipped on
macOS). This mixed-source full attempt is not a final-head full-suite
pass; use the fresh CI evidence below.
- Full workspace typecheck and build pass for the standalone Dot change
on head `392d54f80`. UI token gates are clean. The full workspace
typecheck passed before the final importer UI edit; the changed UI
typecheck, build, and token gates pass again on the final commit. The
server build passes for its unchanged final source.
- Standalone selection, creation, settings, harness visibility,
inventory, and runtime admission checks pass. OAuth onboarding and the
real Rust/PostgreSQL broker suite pass 27 cases with the general Runner
option disabled. Dot and MCP revocation still block pairing; other
Runner providers remain disabled. Create, hire, conversion, and
inherited-hire route regressions pass 68 cases. The runtime selection
suite passes 20 cases. The setup UI suite passes 50 cases. Company
import and dev binary checks pass 97 cases. Import UI checks pass 29
cases, including Dot-only selection, preservation of imported Dot
agents, generic Runner fallback, canonical configuration serialization,
and required billing acknowledgement. A read of the synthetic instance
reports the native binary required with Dot enabled, the general rollout
disabled, and no persisted native work.
- The latest attachment, operator-consent, generic API file-guard, and
bridge checks pass 80 tests. Cache and native lifecycle checks pass 550
tests; create/import builders pass 45 tests; attachment setup UI checks
pass 16 tests. The real Rust/PostgreSQL Dot broker suite passes 15
cases.
- Earlier authority, human assignment, pinned skill, confinement,
cancellation, inheritance, onboarding, replay, OpenAPI, and gateway
regressions pass their focused rechecks. Workspace writes reject
concurrent stale hashes.
- In the live empty test-drive, turn the general Runner option off and
leave Dot and MCP on. Add agent shows a separate OpenAI Dot choice.
Create rejects missing billing acknowledgement, saves the canonical Dot
configuration, and opens pairing. The existing paired Dot remains ready
and passes its prerequisite checks. No new plugin pairing was needed for
this UI change. The final browser import preview keeps a Dot source
agent as OpenAI Dot, falls back a generic Runner agent while its switch
is off, and offers Dot independently.
- Real Dot completed OAuth, signed MCP Events readiness, and an
event-only assigned document task in an empty synthetic instance.
Pairing saved its binding without a second configuration save.
- Real Dot verified human assignment, cross-task comments and documents,
pinned skill reading, sandbox commands, lease renewal, follow-up input,
and a downloaded artifact whose bytes and hash matched its receipt. An
idle conversation request created an intake, created a task for its
human owner, continued after a definite missing-file read error, and
finalized Done with exit code 0.
- After operator approval, real Dot read a synthetic assigned-task
attachment and wrote its file-only random proof into an agent-authored
document. Its first test required accepting review because the existing
document tool removed the final newline.
- On head `2626c8f9b`, real Dot read a 13,849-byte synthetic attachment
in two pages, used `expectedSha256` on page two, saved exactly its final
random marker, verified readback, and finalized Done with exit code 0. A
direct inbox check was required for this attachment qualification; it
does not claim event-only delivery.
- Previous head `392d54f80` passes all 55 checks (53 passed, two
intentionally skipped), including the full test, typecheck, build,
native Runner, browser, and release verification gates. Greptile rates
this head 5/5; all review threads are resolved.
- Head `2626c8f9b` passed all 56 checks: 54 passed and two intentionally
skipped. This includes full typecheck, build, native Runner, server
tests, serialized suites, browser E2E, and Canary Dry Run. The unchanged
Cursor managed-runtime test exceeded its five-second deadline on the
first attempt; its focused local suite passed all eight cases, and the
single CI rerun plus dependent verify gate passed.
- A prior full local serial test attempt reported 19 failures (16,141
passed, 88 skipped) and stopped before later wrapper groups. It began
before the final source edits. Every failed suite has a passing fresh
recheck; the macOS snapshot stress case passes alone in 227 seconds. The
previous head `000fb9261` passed all 56 CI checks. PRs leave
`pnpm-lock.yaml` unchanged; CI and the refresh bot own dependency
resolution.

## Risks

- Dot remains off by default. Its own opt-in does not enable other
Runner providers. It still requires the MCP option; disabling Dot blocks
new work while keeping existing recovery and saved bindings.
- Dot requires a stable public HTTPS origin. Its base provider in #15402
is merged. Temporary tunnels are useful for testing but are not
permanent deployments.
- Idle intake creates a visible agent-authored task. It does not
fabricate a human message or bypass ordinary admission.
- Workspace commands require Linux bubblewrap with a private PID
namespace; deployment qualification is still needed. macOS file tools
and artifact publishing remain available, but commands are disabled
because sandbox-exec does not contain detached descendants. There is no
unrestricted fallback. Command protection fails closed above 4,096
workspace directories. The historical macOS command walkthrough does not
qualify the current Linux implementation.
- Provider model choice, token usage, cost, native thread control, and
global external stopping remain unavailable. Known Paperclip budget
gates still apply.
- The workspace bridge and attachment reader are separate opt-in
settings. Reading assigned task files sends their contents to OpenAI.
Revocation prevents future reads but cannot withdraw bytes already sent.
Files are read on request; automatic inbound attachment staging remains
disabled.
- An existing event registration can retain earlier instructions that
prohibit extra tasks. Dot asked for permission before a second intake
after the final event-only bootstrap; that extra no-nudge continuation
is not live-qualified under that earlier registration. The later
attachment qualification submitted normally through the composer. The
revised onboarding prompt describes the new idle entry point, and its
protocol polling path is tested.
- OpenAI event delivery can be delayed; a webhook acknowledgement is not
proof that Dot has begun work.
- A running Dot can retain an old plugin catalog after tool refresh.
Refresh and actual tool exposure must be checked before using new
top-level actions.
- Hosted and remote controller modes are not qualified.

## Model Used

OpenAI Codex, based on GPT-6. The exact deployment ID and context window
size are not exposed in this session. Capabilities used: reasoning,
repository editing, code execution, test inspection, and browser
control.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either linked existing issues or described the issue in-PR
following the relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused suites and all
fresh failure rechecks pass; the earlier full attempt is disclosed
above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-07 19:44:17 -05:00
DottaandPaperclip fc6304dfe5 feat(runner): add experimental OpenAI Dot provider over MCP Events (#15402)
## Thinking Path

> - Paperclip manages AI agents, tasks, permissions, and execution
budgets.
> - Paperclip Runner gives each provider the same admitted task and tool
authority.
> - OpenAI Dot runs outside the local process tree and needs
asynchronous work delivery.
> - The merged MCP gateway supplies OAuth consent and signed event
delivery.
> - A personal assistant grant cannot safely stand in for an assigned
agent.
> - This pull request adds a separate Dot agent connection and a durable
Rust Runner bridge.
> - The operator can assign work to Dot and inspect its accepted work,
tool receipts, and result.

## Linked Issues or Issue Description

**Agent or provider**

OpenAI Dot, as an experimental provider of the existing Paperclip Runner
adapter.

**Why this adapter is useful**

An operator can assign normal Paperclip tasks to an existing Dot. Dot
can read its mailbox, request work on an assigned task, use admitted
task tools, and submit a result. Paperclip keeps company scope,
checkout, approvals, known budget limits, and activity attribution.

**How the agent is invoked**

A dedicated `/mcp/runner` OAuth resource pairs one Dot grant with one
agent. A signed MCP mailbox event wakes Dot. Dot explicitly accepts the
assignment. The Rust Runner owns the durable turn and operation
receipts. The first release supports self-hosted instances with a local
Runner controller.

**Additional context**

This extends the merged public MCP gateway from #14846 and the assistant
invitation and device-consent work from #14933. This also integrates the
merged assistant tool and configuration expansion in #15380. Dot retains
its dedicated agent resource and cannot receive personal configuration
permission. The public assistant connection remains a personal
connection.

## What Changed

- Add a durable Rust Dot provider and its TypeScript Runner driver.
- Add closed PRP v3 external-provider operations and native execution
input v6.
- Add company-scoped pairing, mailbox, assignment, and operation
records.
- Reuse merged browser/device consent, client metadata verification,
webhook admissions, refresh, secret rotation, and warm-standby gates.
- Keep Dot scopes, issuer, grants, event workers, and tool access
separate from personal assistant access.
- Add Dot configuration, pairing, readiness, and consent UI. Keep agent
grants out of the personal Connections entry.
- Regenerate the Dot-only migration after master. Preserve published
gateway migrations. Make the new migration safe to reapply.
- Document setup, recovery, accounting limits, evidence, and remaining
account qualification.
- Reverify reconnect callbacks and wake outstanding work with a fresh
mailbox reference; preserve the existing assignment and operation
receipts.
- Clean up Dot bindings and waiting runs on OAuth revoke and
refresh-token replay. Old grants cannot revoke replacement bindings.
- Restore the pairing reference when an unsaved agent form is reopened;
document board-only pairing routes in OpenAPI.
- Accept a clean Rust exit after the acknowledged shutdown receipt.
Unexpected exits still require recovery.
- Clear the cached binding after a successful revoke so a failed
connection refresh cannot restore it.
- Add production-component Storybook states and screenshots for pairing
and connection review. All preview account data is synthetic.
- Persist normalized completion, serialize Dot turns and durable work
admission, and poll subscription readiness.
- Serialize mailbox writes and cursor reads; retain paused fence
acknowledgement without task authority.
- Authorize admitted review runs without changing the worker assignee.
Include the fenced assignment ID in production stop notices.

## Verification

- This PR integrates master `4a8178e9c`. Dot migration
`0317_messy_famine.sql` follows the published history and is safe to
reapply. The merge preserves the reserved migration connection,
batch-commit handling, private task checks, task monitors, and native
accounting.
- Local workspace typecheck, full build, and UI token gates pass. The
server typecheck passes after the review fixes. Database and native
executor regressions pass.
- All twelve real Rust/PostgreSQL Dot integration tests and twelve Dot
driver tests pass. The tests cover native document writing and
finalization, durable replay, queue admission, mailbox ordering,
admitted reviews, stale authority, production stop references, and
paused acknowledgements.
- Current head `d0e7e0626` passes all 57 checks: 53 pass and four are
intentionally skipped. This includes full typecheck, build, tests, Rust
Runner verification, browser E2E, release verification, and Canary Dry
Run. Greptile rates this exact head 5/5. All review threads are
resolved.
- The full local root test run is slower than the sharded CI run and has
not completed. The full CI test gates pass on the current commit.
Focused local regressions pass.
- Real-account pairing and event delivery on this base commit remain
unqualified. Live account and setup proof are recorded in the follow-up
#15414.

The following screenshots use synthetic preview data. They show the
production pairing component and do not qualify a real account or the
full agent setup journey.

![Synthetic pairing
preview](https://github.com/paperclipai/paperclip/blob/codex/dot-events-prototype/doc/screenshots/openai-dot-runner/pairing.jpg?raw=true)

![Synthetic connected
preview](https://github.com/paperclipai/paperclip/blob/codex/dot-events-prototype/doc/screenshots/openai-dot-runner/connected.jpg?raw=true)

## Risks

- This base adapter uses `PAPERCLIP_ENABLE_OPENAI_DOT=1` plus Public MCP
and Paperclip Runner. The separate experimental-settings follow-up in
#15414 replaces this environment flag with saved operator settings.
- Dot does not expose provider token usage or cost. The operator must
acknowledge external billing. Known Paperclip budget gates still apply.
- Cancellation fences Paperclip authority. It does not confirm that Dot
stopped all external activity.
- Assigned skill files and third-party MCP bindings are unsupported and
reject admission. There is no mounted workspace, model selector, or
provider thread identifier.
- Hosted agent-broker and remote controller deployments are not
qualified.
- The new migration follows the merged master history. Existing
prototype databases still need the normal master migration history
before this Dot-only migration.

## Model Used

OpenAI Codex, based on GPT-6. The exact deployment ID and context window
size are not exposed in this session. Capabilities used: reasoning,
repository editing, code execution, and test inspection.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #123` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub references)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-07 19:11:50 -05:00
DottaandPaperclip a7a244ab33 feat: add company decision models with permission and cost controls (#15473)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Optional product features need small, typed model decisions.
> - Each company needs to choose which shared API connection pays for
those decisions.
> - Calls must retain user and task permissions, budget limits, and cost
attribution.
> - This pull request adds a managed decision service, setup UI, and
request history.
> - Features can check availability cheaply and keep their existing
behavior when decisions are unavailable.

## Linked Issues or Issue Description

**Subsystem affected**

Server services, shared contracts, database accounting, Company
Settings, and Costs.

**Problem or motivation**

Paperclip has no common decision-model service. Adding provider calls
within each feature would duplicate credential access, permission
checks, and billing rules.

**Proposed solution**

Let a connection manager configure one company decision model. Support
OpenAI Decisions and Jev through OpenRouter. Provide a fixed setup test
and metadata-only history. Default company-sponsored background
decisions to on during setup, and preserve a saved off setting.

**Alternatives considered**

Per-feature credentials would duplicate existing connection management.
Personal overrides and provider fallback chains add permission and
billing complexity; they remain deferred.

**Roadmap alignment**

Reviewed ROADMAP.md and searched open PRs. This extends existing
connection access and budget accounting. Product features that call the
service remain outside this change. No matching decision-model service
PR was found.

## What Changed

- Add company settings, an internal `decisionModelService`, local
availability checks, and trusted human, agent/run, and system contexts.
- Pin Vercel AI SDK provider dependencies and adapt boolean, choice, and
ordered-score decisions for both providers. Bound requests and time;
disable paid retries.
- Add durable invocation metadata and agentless decision ledger charges.
Preserve fractional cents, pricing evidence, dispatch identity, and
unresolved billing holds.
- Share the company accounting lock and apply company, agent, and
project budgets. Settle charges once, retain unknown holds, and recover
interrupted calls without resubmission.
- Reuse connection setup and management UI. Add a Decisions view under
Costs, production-component Storybook coverage, database migration, and
service documentation.

## Verification

- Passed 166 current-code tests covering the decision service/provider,
setup component, Costs, OpenAPI, and every failure from the earlier
broad run. Coverage includes native SDK wire formats, refusals, billed
malformed responses, permission and secret-rotation races, identity
changes, concurrent budget admission, unresolved holds, agent/task
deletion, and stale setup feedback.
- Passed 170 existing connection, cost, budget, heartbeat-accounting,
and profile regression tests.
- Passed repository typecheck, production build, and design token gates
after integrating master. Verified the generated migration on a fresh
test database and upgraded the populated preview database from the
branch's earlier migration without losing settings or usage.
- Ran the required full `pnpm test:run`: its general phase completed
with 16,354 passed and 10 failures across five files while this branch
was still being updated. Every reported failure passes in the
current-code rerun; the serialized phase did not run after that failure.
The full GitHub CI suite passed on `b1b856a88`: general and serialized
tests, browser shards, runner checks, typecheck, production build,
packaging/canary, and policy gates. [CI
evidence](https://github.com/paperclipai/paperclip/actions/runs/37673106360).
- Passed the full-shell Storybook setup-to-history interaction test
again after integrating master. Greptile rates the final revision 5/5
with zero unresolved threads.
- Walked through the running app: empty setup, add each provider, save,
reload, run all three sample questions, inspect fractional charges in
history, switch provider while sponsorship is off, disable, and
reconnect. Checked mobile settings. These tests used the actual UI,
vault, server, SDKs, and database with simulated upstream responses.
- Live paid setup tests remain unverified: this environment has no
authorized OpenAI/OpenRouter credentials available. No mocked test is
presented as live provider evidence.

Reviewer journey: Company Settings → General → Decision model.
Add/select a shared API connection, save, run the billed sample, open
View usage, then disable decisions and verify Run test is disabled after
reload.

## Risks

- The SDK decision interface is experimental. Pinned versions and
wire-format tests limit upgrade drift.
- The migration allows agentless service charges and reservations.
Existing agent cost-reporting APIs still require an agent, and decision
receipts stay separate from run reconciliation.
- Timeouts can have unknown provider charges. Holds remain until an
audited accounting correction resolves them.
- OpenAI prices use a versioned Decisions rate snapshot; OpenRouter
costs use provider receipts. Unknown pricing is retained as unknown.
- A configured company authorizes background spending by default. Setup
explains this, and managers can turn it off.
- Live provider account/model availability still needs the two
credentialed acceptance checks.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository inspection, code
execution, and browser tools. The exact serving revision and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-07 15:53:13 -05:00
Devin FoleyandPaperclip f1c44b7b56 Preserve remote work when required workspace restore fails (#15479)
Persist exact source-retention obligations before run finalization and protect them across cancellation, restart and task changes. Require board-authorized repair evidence without replaying old work or changing current task ownership, state or locks. Hide an unavailable retry action and document operator recovery and retention costs.

Validated with 674 scoped regressions, 224 combined integration tests, full typecheck/build, independent safety reviews and green CI with Greptile 5/5. No historical file recovery is claimed.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-07 13:29:44 -07:00
Devin FoleyandPaperclip cd40ebb95b Preserve ACP bridge terminal evidence before disconnect (#15478)
Drain terminal frames through backpressure and record bounded terminal evidence without waiting for log persistence. Guard later events and input failures against writes after socket end.

Validated with 217 focused tests, full typecheck/build, independent review and green CI with Greptile 5/5.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-07 13:01:22 -07:00
DottaandPaperclip fd8c6b920a fix(native): resume connection tasks after approval decisions (#15471)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Native tasks can pause while a human decides whether to allow a
connection action.
> - The next turn needs both the saved decision and complete accounting
for the previous turn.
> - Cancellation could discard final usage, and complete direct Claude
API receipts could remain unpriced.
> - A stale blocked or review report could also request approval again
after the original card was declined.
> - This pull request retains shutdown accounting and rejects approval
waits bound to an already resolved action.
> - The benefit is reliable continuation with the existing budget and
approval controls.

## Linked Issues or Issue Description

Refs #15420. Related: #15312 addresses requester ownership during
dispatch. This change addresses receipt capture and final-response
validation.

## What Changed

- Retain usage events after native cancellation. Continue to reject late
provider messages and work.
- Drain same-turn accounting and terminal events for at most fifteen
seconds after a durable governed wait. Keep incomplete accounting
blocked.
- Estimate complete, unpriced, direct Anthropic API receipts for the
exact `claude-sonnet-5` model. Record the rate version and assumptions.
Use the one-hour cache-write rate when the receipt lacks cache TTL.
- Bind stale approval reports to exact interaction, action-request, or
invocation IDs in the same company, task, agent, and run. Cover blocked,
review, and response-wake reports. Keep independent reviews valid.
- Fence checkpoint and result writes after a controller detaches for
restart, including operations waiting for a database lock. Reject stale
successful returns before certifying accounting.
- Allow bounded subscription teardown only after retaining an actual
provider terminal.
- Journal the exact governed-wait trigger and disposition before
provider interruption. Recover that wait independently of a later saved
answer, replay retained accounting, and reject mismatched or unproven
terminal evidence.
- Give settling governed turns a bounded window before shutdown detaches
their controller.
- Keep fuzzy external app matches alongside installed capability matches
instead of forcing an unrelated provider question for a generic query.
- Return up to twenty exact active catalog tool names after an invalid
request, after eligibility checks; still reject the request without
granting access or creating an approval.
- Require retained provider terminal proof before settling a governed
wait, including when complete usage arrives before stream
closure/error/timeout. Retain harmless numbered cancellation events so
restart replay stays contiguous.
- Isolate accounting-test OpenCode config from the host plugin
directory.
- Add regression coverage and document the accounting, restart and
connection-search behavior.

## Verification

Current PR source: `0cf08efd75f8fb23f7989beda6dbda92587088bf`. The live
matrix below measured frozen `09a776bcb1f77422156f8be11e13f8c29f43e7f7`;
later review fixes are verified separately and do not relabel those
runs.

- Repository typecheck and build pass.
- Current runner runtime and cancellation suites: 191 tests pass. Six
new regressions cover stream end/error/timeout without provider-stop
proof and contiguous cancellation acknowledgement/request replay; all
six failed before the fix. The existing bounded cleanup case now
explicitly supplies terminal proof. All 31 adapter accounting tests pass
with isolated fixture config.
- New checkpoint-rebinding and approval-criterion suites: 65 tests pass;
runner HTTP integration: 29 tests pass. Unchanged executor/control-plane
suites: 640 tests pass; database-backed connection suites: 68 tests
pass.
- Current retained evidence verification covers 222 file hashes across
all fifteen original result artifacts. The three-case [published
report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37651896525-1/index.html)
and all eight screenshot hashes verify. The [three-case recovery
report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37662587356-1/index.html)
and all six screenshots also verify; the nine-result campaign did not
publish.
- Evaluation support: 1,805 Vitest tests pass, one skipped; 128 Node
tests pass. Eval typecheck and catalog discovery pass.
- Full local repository unit run was interrupted before the follow-up
edits after three tool-access failures and one runner HTTP failure.
Those failures pass in isolation; the 09a full run was interrupted after
one rapid Slack callback-ordering failure and seven skill-service
failures. All eight pass both isolated and with full-runner environment
settings, and all 74 skill-service tests pass together; the subsequent
full run reported two 15-second OpenCode accounting timeouts and was
stopped with exit 130 to apply review fixes. The timeouts reproduce
while copying this host’s 61 MB OpenCode config. All 31 tests pass after
isolating config inside each fixture without increasing timeouts or
changing assertions. A complete local full-suite pass is not claimed.
Repository-wide CI also passes on the final review-fix head in [run
37669185952](https://github.com/paperclipai/paperclip/actions/runs/37669185952).
Earlier database-skipped diagnostics and the older ENFILE run are
retained and are not full-suite passing evidence.
- Three-case live campaign
[37651896525](https://github.com/paperclipai/paperclip/actions/runs/37651896525)
passes all three original grades on frozen source 09a: 55/55 checks,
eight succeeded run records, complete accounting receipts, matching
checkpoint identities and no pending approvals. Campaign
[37653533353](https://github.com/paperclipai/paperclip/actions/runs/37653533353)
adds nine original passes (173/173 checks, eighteen succeeded records)
on the identical source. Its other three jobs failed before runner
assignment or any step while GitHub could not load the paid environment;
those original infrastructure failures are retained. Campaign
[37662587356](https://github.com/paperclipai/paperclip/actions/runs/37662587356)
completes only those unstarted cells: all three original grades pass
(57/57 checks, six succeeded run records). All fifteen exact cases now
pass on source 09a: eight FAIL → PASS, seven PASS → PASS, zero new
overall failures and zero pending pairs. Total current evidence: 285/285
checks and thirty-two succeeded run records, complete accounting
receipts, matching checkpoint identities, no pending approvals or retry
records. Earlier campaigns retain forty-seven additional run records and
two known same-run recovery attempts; actual provider-call counts and
invoices remain unknown. The nine-result campaign skipped publication
and its public URL returns 403; original artifacts remain retained.
Previous ba9 campaign
[37645656840](https://github.com/paperclipai/paperclip/actions/runs/37645656840)
completed 2 PASS / 1 FAIL: Claude restart/approval and Codex decline
pass, while OpenCode resumes but times out searching for exact tool
names and never creates the access card. That failure and incomplete
cancelled-run accounting remain preserved; the new catalog error
guidance targets this observed dead end. Campaign
[37642957312](https://github.com/paperclipai/paperclip/actions/runs/37642957312)
remains 0 PASS / 3 FAIL and exposed the now-corrected cross-run marker
leak and unknown-criterion approval gap. The original baseline remains 7
PASS / 8 FAIL, first repair 2 PASS / 3 FAIL, and second repair 0 PASS /
3 FAIL. All fifteen selected cases are qualified by their original
grades in this bounded trial. These live grades belong to 09a. Its
twelve governed-wait checkpoints retain matching same-turn terminal
fingerprints, but passing artifacts omit detailed event journals; the
later six adversarial regressions qualify the new terminal-proof and
replay guards separately. Final-head repository CI passes. [Fresh
Greptile
review](https://github.com/paperclipai/paperclip/pull/15471#issuecomment-6044369780)
is 5/5, confirms both findings are fixed, and reports no new actionable
issues. All review threads are resolved and the PR has no merge
conflicts.

## Risks

- Governed cancellation drains accounting for up to fifteen seconds.
Restart detachment gives a settling batch up to twenty seconds to
finish. An incomplete receipt or unproven provider terminal still
prevents successful qualification.
- Claude prices are estimates, not invoices. The estimate assumes
standard global API pricing and uses a conservative cache-write rate.
Unsupported models, billers, and billing modes remain unpriced.
- Approval identity matching must remain scoped to the current run and
the requested approval. It does not authorize execution of a declined
call.
- Original baseline and final candidate have different merged master
context. Exact-case outcomes are before/after observations, not isolated
causal attribution to this repair.
- No schema migration, fixture, oracle or grader change. Search-result
guidance now treats fuzzy external matches as suggestions. Existing app
authorization and provider-consent checks remain required.

## Model Used

- OpenAI GPT-6 through Codex, with code editing, terminal tools, and
test execution. The exact deployment identifier and context window are
not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (191 current runtime tests
and 31 accounting tests; interrupted full-suite history and CI coverage
are disclosed above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-07 14:43:07 -05:00
Devin FoleyandPaperclip 465140596f Preserve bounded workspace sync diagnostics across RPC (#15481)
Preserve bounded error codes and HTTP/exit statuses across the environmentSyncOut worker RPC boundary, and revalidate that method-scoped envelope before attaching host restore diagnostics. Keep the original error and all recovery policy unchanged; do not transmit provider payloads or credentials.

Verified real RPC roundtrip and privacy regressions, 95 focused tests, 68 independent tests, full typecheck/build and all exact-head CI. Greptile 5/5 with no unresolved comments.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-07 12:22:53 -07:00
Devin FoleyandPaperclip 128c95837b Retain bounded orphan process-loss diagnostics (#15475)
Record bounded observer uptime, run and output ages, process-check observations and retry eligibility before orphan cleanup changes the evidence. Preserve existing recovery and reporting behavior and omit process IDs, raw paths and credentials.

Verified focused helper, actual reaper and real SDK regressions, full typecheck/build, exact-head CI and independent review. An unchanged Cursor timeout passed its isolated retry; unrelated local Slack timing failure is documented separately. Greptile 5/5 with no unresolved comments.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-07 12:05:21 -07:00
Devin FoleyandPaperclip 43b0aff54b Keep pre-dispatch secret configuration blockers out of Sentry (#15472)
Keep confirmed pre-dispatch missing-secret setup blockers out of Sentry while preserving failed runs and owner recovery actions. Runtime, provider and ambiguous failures remain reportable.

Verified full exact-head CI, focused local regressions, workspace typecheck and independent review; Greptile 5/5 with no unresolved comments.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-07 12:02:29 -07:00
DottaandPaperclip ae6f95ed7a feat: add personal primary agents (#15470)
Add a personal primary agent per company and user. Initialize it from the first human-created agent, expose profile-only switching with confirmation, and use it after recent choices for task and Chat defaults.

Persist authenticated preferences, preserve lifecycle and membership rules, keep selections out of shared audit events, and document the API contract. Include the reviewed Storybook surfaces and regression coverage for concurrent choices, onboarding, cross-device updates, and browser journeys.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-07 13:53:46 -05:00
Devin FoleyandPaperclip 06484b3c41 fix: preserve conversation retries through execution cleanup (#15463)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Recovery schedules bounded retries after a provider disconnects.
> - A stopped run can still hold its environment lease while cleanup
runs.
> - Retrying before that lease is released cancels the new run before it
starts and spends another retry.
> - Restoring the task can also send the worker repair instructions from
an already resolved recovery action.
> - This pull request preserves the waiting retry and removes settled
recovery instructions from later wakes.
> - The task can continue after cleanup without an operator repairing
the same incident again.

## Linked Issues or Issue Description

**What happened?**

A legacy conversation run disconnected while it was doing ordinary work.
Its environment cleanup took longer than the retry delay. Two retries
were cancelled before dispatch with `execution_reconciliation_required`.
Those cancellations exhausted the failure budget. After an operator
restored the task, the wake still told the original worker to repair the
runtime and hand the task back to itself.

**Expected behavior**

Cleanup waits preserve the pending attempt. The same retry can continue
after ownership is released, subject to all current gates. Once a
recovery action is resolved or cancelled, subsequent task wakes omit its
repair instructions.

**Steps to reproduce**

1. Fail a legacy conversation run while its environment lease remains in
`pending_cleanup`.
2. Schedule a bounded retry and run promotion before cleanup releases
that lease.
3. Repeat the scheduler sweep. Before this fix, retries promote and then
cancel without starting.
4. Resolve a stranded-task recovery action and build the restored task
wake with that action ID. Before this fix, the wake still includes the
settled repair instructions.

**Paperclip version or commit**

Reproduced with database regressions against `ceabc3bc880` on master.

Related: #15019 restores a skipped assignment handoff after lease
release. #15235 filters stale handoff evidence in the recovery sweep.
This change preserves an existing scheduled conversation retry and
corrects restored wake content. It does not create a new handoff wake.

## What Changed

- Keep an unstarted legacy conversation retry on the same durable row
while prior execution ownership remains active. Recheck after 30 seconds
without increasing retry accounting.
- Return a queued retry to scheduled state if it encounters that hold at
the claim gate. Retain its issue claim and publish the status change.
- Record one local lifecycle diagnostic per blocking run. Remove that
wait marker on promotion.
- Include recovery action metadata only while the referenced action is
active or escalated.
- Return an explicit `waiting` response and the saved schedule when
Retry now meets cleanup. Show the wait inline without a false success or
disabled button.
- Add database, rendered-prompt, route, and UI regressions. Document the
execution and run-log contracts.

## Verification

- Five cleanup and restored-wake regressions fail against the original
production code. The Retry now route and UI regressions also fail before
their correction.
- Related retry, dispatch, stale-queue, and recovery suites: 321 tests
pass across seven files. All ten focused cleanup/restored-wake cases
pass after rebase. The final dispatch adjustment passes all 46 adapter
tests.
- Retry now routes and affected UI suites: all 46 tests pass. The tests
cover repeated clicks, the saved schedule, unchanged accounting,
promotion after release, and no false success or error state.
- `pnpm -r typecheck`, `pnpm build`, `pnpm build-storybook`, and `pnpm
check:token-gates` pass.
- Full local `pnpm test:run` was started and then stopped after the
final commit passed all GitHub CI test shards. No complete local
full-suite result is claimed; CI supplies the complete test result for
the final commit.
- Final head `6f1058a332c039e33c4f002b296d20a5554e760e`: all 55 GitHub
checks are green or intentionally skipped. Apex review is 5/5 after two
reviews, with no unresolved threads. The PR has no merge conflicts.

## Risks

The wait applies only to unstarted legacy conversation retries. Native
runs and non-conversation execution keep their existing recovery rules.
Cleanup must actually release ownership before execution can resume. The
wait does not fix a cleanup service that never finishes. Promotion and
dispatch still enforce cancellation, reassignment, pause, budget, and
reconciliation gates. No schema change is required. The Retry now
response adds a `waiting` outcome; the shared contract and all three UI
controls handle it.

## Model Used

OpenAI Codex based on GPT-6, with repository analysis, tool use, and
local code execution. The runtime does not expose the precise serving
model ID or context window.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused suites; complete
final-head test coverage in CI)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-07 11:15:49 -07:00