Commit Graph
4854 Commits
Author SHA1 Message Date
dependabot[bot] bf3eecdd2d build(deps-dev): bump vite from 6.4.3 to 8.3.2
Bumps [vite](https://github.com/vitejs/vite/tree/HEAD/packages/vite) from 6.4.3 to 8.3.2.
- [Release notes](https://github.com/vitejs/vite/releases)
- [Changelog](https://github.com/vitejs/vite/blob/main/packages/vite/CHANGELOG.md)
- [Commits](https://github.com/vitejs/vite/commits/v8.3.2/packages/vite)

---
updated-dependencies:
- dependency-name: vite
  dependency-version: 8.3.2
  dependency-type: direct:development
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
2026-10-07 05:47:15 +00:00
03cf6a6ecb chore(lockfile): refresh pnpm-lock.yaml (#15359)
Co-Authored-By: lockfile-bot <lockfile-bot@users.noreply.github.com>
Co-Authored-By: Paperclip <noreply@paperclip.ing>
nightly/v2026.1007.0-nightly.0 canary/v2026.1007.0-canary.5
2026-10-06 22:44:24 -07:00
Devin FoleyandPaperclip 799e4d556f fix: make accounting durable and synchronize cost reporting (#14997)
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-06 22:26:48 -07:00
DottaandPaperclip 99a9de9940 fix(mcp): personalize assistant connections and hide revoked grants (#15411)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Assistant connections let people use their organization from another
assistant.
> - Setup begins with an invitation and returns to a list of connected
assistants.
> - Client-only labels obscure the authorizing person, while revoked
rows clutter that list.
> - This pull request shows the person’s avatar, names connections by
owner and client, and hides revoked rows.
> - It also shortens the copied invitation while keeping the approval
instructions.

## Linked Issues or Issue Description

Related: #14933 and #15380.

**What existing behavior does this improve?**

Invitation copy and assistant connection management, in the
organization’s Connections screen and the account-wide management page.

**Subsystem affected**

Cross-cutting: shared MCP connection types, server profile projection,
UI and Storybook.

**Current behavior**

Connections are labeled only with a client name such as “Codex.” Revoked
connections remain visible. The invitation includes an extra sentence
about agent identity.

**Proposed behavior**

Show the authorizing person’s avatar and use names such as “Dotta’s
Codex connection.” Hide revoked rows after successful revocation and
when loading retained revoked grants. Failed revocation leaves the
connection visible. Remove the extra identity sentence from invitation
copy.

**Reason and benefit**

Make connection identity clear and keep the list focused on usable
connections.

**Breaking changes**

The connection response adds optional `user` metadata with name and
image. Older servers remain usable. Names and revoked-row visibility
change in the UI; OAuth client identity, authorization and audit
retention remain unchanged.

## What Changed

- Shorten the shared invitation text.
- Project the authorizing person’s name and avatar through the
user-scoped connection endpoint, without returning email or credentials.
- Reuse the existing Identity component and owner naming conventions
across the connection page, catalog card and account-wide list.
- Hide revoked grants and remove a successfully revoked row from the
shared cache, even if the subsequent refresh fails.
- Update documentation, regression tests and production-page Storybook
fixtures and revocation journeys.

## Verification

- `pnpm -r typecheck`, `pnpm build`, `pnpm build-storybook` and `pnpm
check:token-gates` pass. Final UI type checks also pass.
- Focused consent and connection UI tests: 31 pass, including company
filtering, legacy metadata, custom client names, user-initial fallbacks,
failed revocation, retained revoked rows and refresh failure after
successful revocation.
- Existing MCP regression suite: 6 pass; 71 database checks are skipped
locally because embedded PostgreSQL cannot start on this machine. The
added database check verifies user-profile isolation and retained
revocation history; CI runs these checks.
- Browser verification with Storybook fixtures: owner avatar and name
render; revocation removes the selected row in both production pages,
leaves other connections visible, and restores the empty state after the
last revocation.
- All 54 current-head CI checks pass, with two optional Storybook jobs
skipped. CI includes database, browser, runner, typecheck, build and
clean-install canary coverage.
- Greptile reviewed commit `1368d79e1066b418712224378d89d64c2b11cb86`:
5/5, no actionable findings or unresolved threads.
- The full local `pnpm test:run` was stopped after complete CI passed.
Local database coverage remains unavailable because embedded PostgreSQL
cannot start; no full local-suite pass is claimed.

## Risks

Low risk. The additive profile field is optional for compatibility.
Revoked grants are filtered only from management UI and retained for
audit. Revocation failure does not hide an active connection. No
authorization scopes, token handling or schema changes.

## Model Used

OpenAI GPT-6 through Codex, with code editing, command execution and
browser verification. The exact deployment ID and context-window size
are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.1007.0-canary.4
2026-10-06 22:09:41 -05:00
DottaandPaperclip a9a20fb5c6 feat(security): add read-only customer-success inspection APIs (#15405)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents need a persistent identity and a verified active run for
governed access.
> - Customer-success inspection needs broad reads without tenant writes
or secret access.
> - Ordinary board login and database credentials give more authority
than this task needs.
> - This pull request adds a dedicated inspection API and strict
managed-run authority.
> - Cloud owns short grants, human approval, replay protection, and
audit records.
> - The benefit is inspectable access that an operator can disable
immediately.

## Linked Issues or Issue Description

Refs: #15352. This change reuses the persistent Ed25519 identity from
that PR.

**Subsystem affected**

Server authentication, pure resource readers, and the shared wire
contract.

**Problem or motivation**

One internal Paperclip agent must inspect customer onboarding work. It
must not receive owner login or database credentials. Reads must not
create customer sessions, memberships, activity, or read receipts.

**Proposed solution**

Add disabled-by-default run authority and versioned tenant inspection
endpoints. Require strict instance-bound managed-run JWTs on the home
instance. Require exact-operation, single-use Cloud permits on tenants.
Execute a reviewed company-scoped catalog in read-only transactions.
Cloud applies seven-day stack-age eligibility and human exceptions.

**Roadmap alignment**

This is access support for Cloud deployments and governed agent
identities. Bot creation, scheduling, scoring, and reports are separate
work. The maintainer requested this implementation.

## What Changed

- Reuse existing public identity reads and managed private-key
injection. Reject unprovisioned keys, paused agents, ended runs, legacy
signatures, and wrong instances.
- Mount `/api/customer-success/v1` before actor/session synchronization.
Verify Cloud permits and consume them centrally before reading.
- Add explicit company-scoped database readers and bounded instruction,
skill snapshot, run log, workspace, and asset reads. Preserve existing
redactions and file protections.
- Add protocol, security, database immutability, and managed-agent
qualification tests. Add deployment and rollback documentation.

## Verification

- Full `pnpm -r typecheck` and `pnpm build` passed. Server typecheck
passed after review fixes.
- The broad local `pnpm test:run` recorded 14,277 passes and four
failures in unchanged suites: two timeouts and two PR-metadata mock
assertions. All three affected suites passed on isolated reruns (36
tests). The complete CI matrix passes at the final head, including every
test lane, typecheck, build, runner checks, canary dry run, and the
security scan.
- Focused inspection, JWT, and existing identity tests pass. The catalog
test compares every public database table before and after reads.
- Inspection and route-contract tests: 22 passed. The coordinated test
runs a real managed process agent against separate home/customer
PostgreSQL databases and a PostgreSQL broker over HTTP. It proves wake
through the existing controller, bounded binary file reads, single
challenge consumption across replicas, concurrent grants with a
two-connection pool, scoped SQL audits, append-only runtime auditing,
one-year retention, and unchanged tenant data/files.
- Run the coordinated test with `PAPERCLIP_INSPECTION_CLOUD_DIST`
pointing at the sibling Cloud build. Normal unit runs skip that optional
private integration.
- Final-head Greptile is 5/5 with no unresolved findings.
- No production deployment or customer inspection occurred.

## Risks

- This adds an authentication boundary. Keep both feature flags disabled
until coordinated staging and canary qualification.
- Cloud support must deploy after this API. Unsupported tenants fail
closed. There is no owner-login or database fallback.
- Existing redactions remain the content boundary. Arbitrary pasted
secrets in readable prose or files may remain.
- Remote files and suppressed provider traces remain unavailable. Wake
can cause normal startup/background writes; test those separately.
- Disable Cloud policy first during rollback. Preserve existing identity
material and Cloud audit history.

## Model Used

OpenAI GPT-6 (Codex), with reasoning, code execution, and browser
testing. The session does not expose a more specific deployment ID or
context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused checks and
isolated reruns; broad-run flakes are documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.1007.0-canary.3
2026-10-06 21:38:15 -05:00
DottaandPaperclip caf120105c test: prepare neutral native connection guidance evals (#15407)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents need to discover connections, obtain consent, and continue
from saved decisions.
> - We want to reduce repeated instructions only when measured behavior
supports the change.
> - The existing decline tasks tell the model not to retry. One
provider-decline check can pass without an explanation or an observed
service counter.
> - This PR adds neutral tasks and stricter saved-evidence checks before
any connection instruction reduction.
> - Production instructions remain unchanged. The new cells are
configured, not live-qualified.

## Linked Issues or Issue Description

Refs #15218. Refs #15389.

**What existing behavior does this improve?**

The Product E2E connection workflow evaluation and its instruction
measurement provenance.

**Current behavior**

Some decline prompts supply the policy they intend to test. The
provider-decline workflow does not require a saved post-decision
explanation. Its old no-call check can use a missing fixture counter as
zero. OpenCode has no connection cases in the original Everyday matrix.

**Proposed behavior**

Add an explicit-only suite with five connection stories on native Codex,
ACPX Claude, and OpenCode. Require an explanation attributed by exact
run ID after a saved decline. Observe the provider fixture counter.
Preserve the original cases and grades.

## What Changed

- Add fifteen configured cells with one attempt, twelve-minute
deadlines, and verified 1,000-cent company and agent budget stops.
- Remove procedure hints from the three new decline prompts. Keep a
user-permitted explanation fallback and the existing positive controls.
- Require saved decline state, one decision, unchanged connections,
observed zero service calls where applicable, and a post-decision
explanation from a successful run on the same task.
- Add negative grader calibration and test the actual fixture budget
payloads. Exclude the suite from default and generic selection.
- Extend the existing full-catalog measurement source manifest with
connection descriptions and schemas. Add an audit of fixed text, tool
descriptions, returned instructions, and unqualified behavior.
- Rebase on master `a6306ba606eb87c89b9ef0344e9fe8e0025580f9` and
preserve its new Cursor suites. No production, credential, workflow, or
lockfile change.

## Verification

- Before rebase: Product E2E support passed 1,424 TypeScript tests and
128 Node checks. Six catalog measurement tests, repository
typecheck/build, Product E2E typecheck, and exact fifteen-cell discovery
passed.
- The full pre-rebase repository test run was stopped when master
advanced. Its partial result is not a pass.
- After rebase and the review correction: repository build/typecheck,
Product E2E typecheck, 1,799 TypeScript support tests (one skipped), 128
Node checks, six measurement tests, and exact fifteen-cell discovery
pass. The duplicate local full-suite run was stopped incomplete after
about 20 minutes once complete CI passed; no local full-suite pass is
claimed.
- Review found that the initial grader read `runId` instead of public
`createdByRunId`. A regression calibration reproduced both rejection of
valid public comments and acceptance of the wrong alias. The fix uses
the actual field and binds the evidence type to the shared
`IssueComment` contract. A subsequent type-only import path correction
passes Product E2E typecheck.
- Final source `0de306b9664bfbdebb6709ddb54c95152740d1ad` passes
[complete
CI](https://github.com/paperclipai/paperclip/actions/runs/37560250545):
51 successful checks and two intentional Storybook skips, plus separate
Snyk success. Fresh Greptile review is 5/5 with the single review thread
resolved and no new findings. The PR is clean and mergeable.
- Local commands: `pnpm build`, `pnpm -r typecheck`, `pnpm
test:e2e:runner:unit`, `pnpm test:e2e:runner:typecheck`, and `pnpm
test:e2e:runner -- --list --suite native-connection-guidance`. The
measurement uses
`PAPERCLIP_NATIVE_PROCEDURE_MEASUREMENT=/tmp/connection-measurement.json
pnpm exec vitest run --project @paperclipai/server
server/src/__tests__/native-procedure-measurement.test.ts`. Validation
used pinned pnpm 9.15.4.
- No paid provider campaign was started. There is no baseline/candidate
behavior result for these new cells.
- The audit records 654 UTF-8 bytes of fixed connection guidance. A
clean capture at `a04b8c6a452315625014888335d45670a2094fb6` confirms 41
supplied tools, 53,341 normalized bytes at start/resume, 50,949 at
compact continuation, and a 48,195-byte authenticated OpenCode MCP
catalog. These are byte counts, not tokens, bills, vendor-private prompt
sizes, or savings from this PR.

## Risks

- This is eval preparation. Passing support tests do not establish live
model behavior or qualify an instruction reduction.
- The explanation oracle checks attributed saved output. It does not
prove cognition or arbitrary prose truthfulness. One saved interaction
also does not prove the absence of repeated idempotent tool calls.
- Successful new authentication and tool refresh, existing-connection
agent grants, independent work while waiting, explicit retry after
decline, and blocking when mandatory work remains still need separate
coverage.
- Notion setup decline does not execute a real Notion service. Positive
service approval uses an already installed deterministic service; it
does not qualify new connection creation.
- The original historical failures remain unchanged. Future comparisons
must freeze source, fixture, model, input, and grading controls and
retain every actual attempt.

## Model Used

OpenAI Codex, GPT-6 family, with reasoning, repository tools, and code
execution. The exact serving model ID and context-window size were not
exposed in this session; they are not inferred. No model provider was
invoked by the eval suite in this PR preparation.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.1007.0-canary.2
2026-10-06 21:28:16 -05:00
DottaandPaperclip a6306ba606 feat(runner): consolidate Cursor production integration (#15075)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The native Runner keeps provider sessions under company authority,
approvals, budgets and durable recovery.
> - Cursor work was spread across candidate branches. The published
branch lacked later plan, permission and cleanup fixes.
> - Production also needs public installation and matching runtime
assets for local and Daytona execution.
> - This pull request consolidates Cursor onto current mainline recovery
behavior and completes that installation path.
> - The installed v11 release passed focused local and Daytona
qualification after the generic mode and lifecycle cleanup. The later
model-selection correction and current mainline merge produce v14
artifacts that need matching release qualification.
> - Cursor admission is enabled in source; publish only an artifact
combination with matching qualification. Native AskQuestion and complete
per-run dollar accounting remain excluded.

## Linked Issues or Issue Description

Refs: #14435, #14631, #14669, #14699, #14724.

This completes the Cursor implementation by @cryppadotta from combined
source `22c78242a4e0c2369fecf0c2dc4e7600fbad6706`. It preserves newer
mainline recovery, completion and warm-directory behavior. Pi and
Copilot remain gated.

## What Changed

- Generate named Rust and TypeScript ACPX release profiles from one
manifest. Share runtime pins with packaging and server verification.
Preserve vendor runtime versions; bind the updated ACPX patch to Cursor
profile v14 and reject stale generated declarations at build/typecheck.
- Remove ACPX model allowlists, including the former Codex and Pi
restrictions and the duplicate developer test-drive gate. Send any
explicit model ID unchanged to its provider and verify the effective
selection before prompting. The bundled ACPX package forwards unlisted
IDs, rejects mismatched acknowledgements, and restores the exact
selection after session load. It does not expand Cursor model aliases.
Provider rejection, mismatch, or missing model controls fails without a
fallback. Model examples live in evaluation fixtures, outside runtime
declarations.

- Add pinned Cursor execution, contained instructions, exact model
verification and Agent/Plan/Ask modes.
- Carry an opaque generic `mode` identifier in shared native execution,
sidecar, Rust and recovery contracts. The provider adapter owns
supported modes, defaults, native translation and acknowledgement.
- Keep native RPC recognition, accepted-plan interpretation and
permission evidence behind provider adapters. Shared settlement and
recovery verify normalized facts and their committed evidence.
- Replace the Cursor-only warm-attachment branch with a runner-owned
capability. Only Cursor opts into it. Move profile compatibility and
optional usage parsing into provider metadata and adapters.
- Write generic plan-wait receipts. Read exact historical Cursor
receipts through a separate compatibility decoder. Reject mixed formats
and preserve existing authority checks.
- Carry native plans, semantic questions, todos, child activity,
permission identities and partial usage diagnostics through the Runner.
- Preserve durable response delivery, cancellation, warm ownership and
process retirement.
- Finish accepted planning runs successfully. Keep their tasks open for
explicit direction. Acceptance does not start implementation.
- Ship `paperclipai runtime setup cursor` and its provisioner through
the public package. npm installation does not download Cursor. Setup
uses the OS account's closure-keyed cache so system-wide npm packages
can remain read-only. Run it as the Paperclip service account.
- Include Cursor in normal provider packs and Daytona images for macOS
ARM64/x64 and Linux x64.
- Reject stale release packs by source revision and current ACPX/Cursor
pins before assembly writes files. Verify current Cursor
version/profile/closure again at runtime.
- Ship all three daemon targets and the expected Linux image-pack
identity. A macOS controller uses its packaged Linux daemon for Daytona.
Image mismatches fail before provider launch.
- Use the vendored Runner boundary for installed readiness probes.
Verify the actual installed Cursor probe.
- Verify compiled public Daytona plugins and their release versions in
installed smokes.
- Record exact artifacts, the acceptance matrix, retained failures,
supported capabilities and rollback behavior in the [readiness
report](https://github.com/paperclipai/paperclip/blob/codex/cursor-production-readiness/doc/plans/2026-10-03-cursor-production-readiness.md).

## Verification

- Current head `f7ec5cc1f0e30c62a829c016ff2013a2a9d79037` merges
mainline `faa8e452c73bae5e044dd6379179a00106abb131`. It keeps Cursor
plan and cancellation guards alongside mainline historical-question
filtering. The evaluation catalog includes both Cursor and expanded
adapter accounting cases (683 total). Recursive typecheck, full build,
696 lifecycle/recovery tests, 45 fixture tests and fixture typecheck
passed. Current-head CI passed: 56 successful checks, one neutral and
four skipped. [Complete
CI](https://github.com/paperclipai/paperclip/actions/runs/37557996535).
The fresh Base Greptile review is 5/5 on this exact head, with 304 files
reviewed, zero new comments and zero unresolved threads. The user
authorized overriding the CODEOWNER review gate after checks passed; no
failing checks are overridden. Prior results below retain their own head
identities.
- Corrective head `3d2b168366258036f6b6a6fccb382c49138cc601` fixes the
post-merge Apex finding. Automatic-review and new-evidence
reconciliation preserve pending child results and recheck delivery under
the status lock before completing. Account repair now excludes unrelated
secret consumers and requires the failed agent's identity. Regression
coverage includes the commit race, delivery statuses,
current-run/current-intent exclusions, repeated reconciliation, both
database reconciliation paths, and credential consumer boundaries. All
184 affected tests, server typecheck and server build passed.
Current-head Base Greptile review is 5/5, with 304 files reviewed, zero
new comments and zero unresolved threads. Current-head CI passed: 56
successful checks, one neutral and four skipped. [Complete
CI](https://github.com/paperclipai/paperclip/actions/runs/37535994724).
This Base review is distinct from the earlier Apex review.
- Merge head `5957c257a` reconciles mainline `b508a05c4`. It preserves
both accepted-plan waits and pending-child-completion checks, current
provider selectors, task-creation response identities, and mainline ACPX
missing-file handling. The combined patch is bound to Cursor profile
v14; historical records keep their original identities.
- Merge head `5957c257a` passed recursive typecheck, full build, 43
installed ACPX/package contracts, 107 provider UI and plan/recovery
tests, 593 database-backed lifecycle tests, 49 profile/native contract
tests, 45 Product E2E fixture tests, fixture typecheck, token gates,
three provider-free browser task-creation cases, and Runner
conformance/replay checks. Its complete CI passed (55 successful checks,
one neutral and four skipped), while Apex returned 2/5 with a
child-delivery finding addressed below.
- The local full-suite attempt again failed the unchanged Git streaming
test (360-second timeout) and was stopped. The concurrent local Rust
attempt failed four unchanged Codex process/deadline tests; all four
passed serially without code changes in 7.29 seconds after removing the
competing test load. These failed commands are retained and are not
reported as full-suite passes; the fresh Linux CI runs are tracked
separately.
- The previous head `907bdb2a2778c7ffeb4a662a91460c9d1ddfc9c5` earned
Apex 5/5 with zero comments after fixing all three findings: per-user
install cache, stale release-pack rejection, and public Linux smoke
account/home handling. Its real built installer passed from read-only
public packages on macOS ARM64 and Linux x64. All 137 release-registry
checks and 64 ACPX package contracts passed. That review does not cover
this mainline reconciliation.
- Prior `beadd3654` passed the full CI matrix; its one unchanged chat
test failure and successful single retry remain in the [CI
history](https://github.com/paperclipai/paperclip/actions/runs/37521449327).
Historical results below remain attributed to their original builds.

- Fixture follow-up `dd59d7e82b103a88b7cbd7d2c38b612c0fbbff7a` removes
provider-specific model choices from generic offline ACPX tests. The
fake sidecar preserves the model and session identity selected at open
through suspension. Affected verification passed: 106 Rust tests and 73
TypeScript tests. This commit changes test code only; the
production-code checks below retain their recorded identities. Its CI
and Greptile review later passed; those results belong to that
historical head.
- Model-selection cleanup `9a070808b48960a41fdfd369ae0636b95af82459`:
252 focused Runner tests passed (six platform skips), covering all six
ACPX agents, native model acknowledgement, rejected selections,
installation integrity and recovery identity. The merged branch passed
recursive typecheck, full build, token gates, server admission (19
tests), and the Product E2E catalog (45 tests). The acceptance catalog
passed all four tests. The full Rust suite passed: 643 tests, 2 ignored.
It verifies sidecar acknowledgement of unlisted models and rejection of
model mismatches. The final commits only update Rust tests; production
sources match the verified build at
`65ec3279ac50185e3cda109b5cfd9b4f56105de0`. No new paid provider calls
were made.
- The merge preserves both Cursor and the new mainline public-MCP
fixture cases. Auto-merge remains disabled; the latest follow-up status
is recorded above. The local `pnpm test:run` attempt hit the unchanged
Git streaming test's 300-second timeout and was interrupted before
merging mainline. The broad Runner attempt found obsolete single-model
assertions plus three macOS fixture-path failures caused by a
`/private/tmp` override. The assertions are corrected; affected
TypeScript checks passed with the standard macOS temporary directory,
and the complete Rust suite passed. Neither interrupted command is a
full-suite pass.
- Earlier declaration-cleanup head `6f4a5e9e2` passed recursive
typecheck, build, Rust and focused tests. Its CI later exposed a test
expecting duplicated Grok digest literals. The current source fixes that
assertion to compare launcher bytes with the shared manifest. Historical
successes and failed attempts are retained; no new live provider
qualification is claimed.
- Previous head `e75fde6098b0ddd8cec765bfb6ecaeecb88a26a6` passed
complete CI (56 successful checks, one neutral, four skipped) and
Greptile 5/5. [Historical complete
CI](https://github.com/paperclipai/paperclip/actions/runs/37489112305).
Those results are not claimed for the cleanup head.
- Frozen live application: `d7b696f9b8f79095233e9e3d56d23e6a6018dd48`.
Public package version: `0.0.0-cursor-verify.3d0c9b7761c6`. The
declaration cleanup preserves release pins and does not relabel that
tested artifact as a build of the new source. Mainline through
`e34abee670` was reconciled while preserving accepted-plan waits,
provider-capacity handling, and both Cursor and public-MCP fixtures.
- Clean normal installation, explicit Cursor setup and daemon resolution
passed on macOS ARM64, macOS x64 under Rosetta, and Linux x64. npm
lifecycle hooks ran without silently downloading Cursor.
- Historical v11 live matrix: **18/18 passed with cleanup** (nine local,
nine Daytona) after the generic mode and lifecycle cleanup. The campaign
has 23 attempts; all five failures and their diagnoses remain recorded.
Exact case identities, hashes and limits are in the readiness report.
All provider calls are real, use the explicit Luna model and
company-bound credentials, and run without qualification or
runtime-asset overrides.
- The immutable Daytona image is
`ghcr.io/paperclipai/paperclip-daytona-runner@sha256:d6259b6bba094702c13fc2283bd85550849c1c53145b656fb2746778f9fa1747`.
The public Daytona plugin is installed independently and its version is
checked.
- Recursive typecheck, full build, token gates and Runner
contract/conformance/replay checks passed on the frozen application. Its
complete Linux CI suite passed. The duplicate local full-suite command
was incomplete after timing failures; affected repeats passed, but that
command is not reported as a clean pass.
- Qualification fixtures passed typecheck, 1,675 Vitest tests (one
skip), 128 Node checks, three provider-free browser tests, and 150
focused lifecycle tests after the final diagnostic correction. The
affected legacy Cursor command file also passed all five tests after
removing its shorter 10-second override; it now inherits the suite’s
standard 15-second timeout. Greptile is 5/5 on `e75fde609` with no
unresolved review threads. CI results above are recorded separately from
historical build results.

## Risks

- Cursor v14 includes the updated ACPX dependency patch and release
identity. The v11 live matrix and image below remain historical
evidence. They do not certify new v14 package/image artifacts.

- ACPX accepts models beyond the qualification fixtures. Availability
and entitlement depend on the provider. Successful configuration is not
a claim of live qualification for every model.
- Shared mode is an opaque identifier. Provider adapters own its
meaning. Incompatible historical sessions remain fenced; exact committed
plan waits and task history remain inspectable.
- Native AskQuestion is excluded. Paperclip semantic questions are
supported. Authoritative per-run dollar accounting is unavailable;
partial counters remain diagnostics and unknown cost is not zero.
- Image input, detailed native diffs, deeper child transcripts and
native plan-file export remain follow-ups.
- macOS x64 has clean-install and daemon-startup proof under Rosetta,
not a separate live campaign on Intel hardware.
- Release only the tested package/image combination. Merging this PR
does not publish npm packages or deploy that image. Later builds need
their own release verification. Rollback disables new Cursor admission
while preserving records and recovery inspection.
- A model can fail an exact instruction: one cancelled-plan attempt
returned the wrong summary marker despite correct cancellation. The
unchanged repeat passed; both results remain in the report.

> ROADMAP.md was checked. This completes existing native Runner/Cursor
work; it does not add an independent core feature proposal.

## Model Used

OpenAI Codex, GPT-6. The exact serving variant and context window are
not exposed in this session. The agent used reasoning, repository
inspection, code execution, protocol tests and browser-backed Product
E2E tools. Cursor acceptance uses the explicit
`gpt-5.6-luna[context=272k,reasoning=medium,fast=false]` model. That is
the evaluated provider model.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass — affected suites passed;
full CI and the retained local failed attempts are recorded separately
above.
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green — 56 successful checks, one
neutral and four skipped on `f7ec5cc1f0e30c62a829c016ff2013a2a9d79037`
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups —
fresh Base review passed on `f7ec5cc1f0e30c62a829c016ff2013a2a9d79037`;
zero new comments and no unresolved threads. The earlier Apex finding
remains fixed.
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.1007.0-canary.1
2026-10-06 20:48:15 -05:00
DottaandPaperclip faa8e452c7 fix(tasks): stop repeated reminders for historical questions (#15392)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Task questions remain saved so a person can answer them later.
> - The composer moves an old question into history after a newer human
message.
> - Native completion still treated every pending question as a required
response.
> - This caused agents to demand an old answer after the person moved
work forward.
> - This pull request shares the historical-question rule across context
and task execution.
> - Agents can finish verified work while the original question remains
answerable.

## Linked Issues or Issue Description

Refs #15229, Refs #14613, Refs #13130.

**What happened?**

An agent repeatedly asked a person to answer a question that had moved
into feed history. Completion feedback explicitly told the agent to
request a response. The saved pending row also blocked task completion.

**Expected behavior**

A question before newer human direction remains answerable in the feed.
Its pending state alone must not require another reminder or stop
completed work. A new input blocker, an approval, or a configured review
stage must keep its gate.

**Steps to reproduce**

1. Let a task agent create an ordinary question.
2. Dismiss the question and send a newer task message.
3. Let the agent finish the requested work and submit its completion
report.
4. Observe a demand to answer the old question and a retained completion
gate.

## What Changed

- Add one company-scoped predicate for historical questions. Only later
human comments count. Exclude agent attribution, run attribution, system
notices, and untrusted source data.
- Apply the predicate to completion feedback, native waits,
finalization, commit validation, retry validation, blocked routing, and
successful-run handoff.
- Include question classification and guidance in heartbeat context and
both native task-context tools. Add the guidance to fresh and resumed
task prompts.
- Replace automatic reminders for current ordinary questions with
instructions to assess the real blocker, continue independent work, and
withdraw obsolete questions through the existing API.
- Preserve historical question rows during completion while cancelling
their live native source runs through the existing post-commit and
recovery paths. Let an authorized human answer them after completion
without reopening work or creating a response wake. Preserve
cancellation, current-input, approval, permission, credential,
connection, and review gates. Add no dismissal storage or migration.
- Document the rule in the execution contract and agent skill. Refresh
generated capability source anchors. Add database-backed status,
context, attribution, and governance regression tests.

## Verification

- `pnpm build` passed. The server rebuild also passed after the
lifecycle fix. Generated capability contract and inventory checks passed
after the agent documentation update.
- `pnpm -r typecheck` passed. Final `pnpm --filter @paperclipai/server
exec tsc --noEmit` also passed after the last test additions.
- Lifecycle and interaction regressions passed: 224 tests in 3 suites.
Context and prompt tests also passed. Final historical-question cases
passed (33 tests), native cancellation/recovery cases passed (6 tests),
and the existing interaction/confirmation suites passed (72 tests). The
full local `pnpm test:run` was attempted and stopped after more than two
hours with unrelated fixture/hook timeout failures; it did not pass. All
52 successful GitHub checks are green on the latest commit, including
the complete test matrix; no checks are pending or failing. Greptile is
5/5 and both review threads are resolved.
- Regression cases cover the old-question/new-human-message sequence,
final task status, answering after completion with no wake, live-run
cancellation and crash recovery, current input blockers, both
task-context tools, API context, timestamp precision, attribution
boundaries, and protected gates.

## Risks

- A later human task message makes an earlier ordinary question
historical even if its input is still missing. The agent must identify
the current blocker and ask only for information that still prevents
work.
- Browser dismissal remains a local preference. Dismissal without a
later human message is not recorded by this change.
- No schema change or data migration. Completion retains ordinary
historical questions; cancellation still expires them. A completed task
accepts historical answers only from an authorized human and creates no
response-delivery outbox row. Governed requests retain their gates.

## Model Used

OpenAI Codex, GPT-6, with repository editing, code execution, and
browser diagnostics. The exact deployment model ID and context-window
size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 20:33:13 -05:00
DottaandPaperclip 2d0c138122 Expand direct assistant MCP tools for work and configuration (#15380)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - People also use assistants in Codex, Claude, and other MCP clients.
> - The existing assistant connection can read work and create tasks or
comments.
> - It cannot edit tasks, exchange files, or manage normal agent and
project settings.
> - These operations must retain the person's permissions and
Paperclip's execution rules.
> - This pull request adds an explicit operation registry and separately
consented configuration access.
> - Assistants can manage work without receiving credentials or runner
authority.

## Linked Issues or Issue Description

**Subsystem affected**

Cross-cutting: assistant MCP, domain routes, consent UI, storage, and
Product E2E.

**Problem or motivation**

A connected assistant cannot update tasks, maintain documents, attach
files, or configure existing agents, projects, and skills. Users must
leave the assistant for these routine actions.

**Proposed solution**

Add named tools and a restricted API registry to direct connections.
Require separate configuration consent. Reuse domain routes and retry
receipts. File uploads save the attachment when the byte transfer
succeeds.

**Alternatives considered**

Arbitrary REST forwarding would expose administration and credential
operations. Runner impersonation would bypass execution ownership.
Separate upload completion calls add unnecessary client state.

**Roadmap alignment**

Checked ROADMAP.md and related MCP pull requests. This extends the
human-authorized connection from #14933. It does not replace the runner
or introduce agent impersonation.

Companion Cloud routing and directory isolation:
https://github.com/paperclipai/paperclip-cloud/pull/678.

## What Changed

- Add task editing, finish/block, documents/revisions, deliverables,
agent settings/instructions, projects/repositories, and skills/files.
- Add an allowlisted API search/call registry with identical field
restrictions, scopes, and retry identities.
- Add unchecked configuration consent. Existing write grants retain
their current authority.
- Add hashed, expiring file transfer tickets and atomic upload receipts.
No completion call is required.
- Preserve company boundaries, human attribution, active native
execution ownership, and execution review gates.
- Add protocol/domain tests, consent stories, and eight paid Product E2E
workflows.
- Repair two CI fixture races: await cold route setup before assertions,
and wait for asynchronously loaded connection copy. Both fixture suites
pass (24 + 48 tests).

## Verification

- Consent revision: one write-access checkbox controls requested work
and configuration permissions in browser and device flows. All 16
consent tests, UI typecheck/build and token gates pass. Updated
interactive stories cover default approval, opt-out and viewer
restrictions. The paid browser helper uses the new exact label. Real
GPT-5.4 Mini Product E2E passes 2/2 at
`64f96373118eb190f8cba1c2ab17cb979555f3ad` (configuration + permission
denial), campaign `local-2026-10-07T00-51-14-337Z`, no automatic
retries, cleanup passed; $0.04149375 estimated assistant cost plus
unpriced worker usage. Raw results, usage and source fingerprints are
retained in the worktree. UI and Product E2E typechecks pass.

- Prior head `2f246d4b74f1f98c75ebcb37ae6753a748237fac`: all 52 checks
pass; two optional Storybook checks skip. Greptile 5/5 on that head, no
unresolved review threads. Final consent head
`64f96373118eb190f8cba1c2ab17cb979555f3ad` also has all checks passing
and Greptile 5/5 with no unresolved threads. The unchanged Cursor
sandbox test had one 10-second timeout, passed in local isolation, and
passed its single CI rerun; the failed attempt remains in [the CI
run](https://github.com/paperclipai/paperclip/actions/runs/37554106934).
The existing chat retry-denial browser test had one visibility failure;
its single rerun passes, and the failed attempt remains in [the CI
run](https://github.com/paperclipai/paperclip/actions/runs/37542735691).

- Full workspace `pnpm -r typecheck` and `pnpm build` pass at final
runtime source `b2196fae1`. UI token gates pass.
- 139 MCP/OAuth/transfer/privacy tests and 76 grader calibration tests
pass, including one-connection PostgreSQL OAuth and concurrent upload
retries.
- Paid Product E2E: all eight expanded cases qualified across Mini,
Haiku and Sonnet. A merged-source repeat passed 23/24; one Haiku cell
timed out before application startup. Final affected-case qualification
passes 9/9 on all three models with grader v16, including the failed
cell. Automatic retries disabled; failures, costs, source hashes and
independent durable-state/file assertions are retained in [the
verification
record](doc/plans/2026-10-06-expanded-assistant-mcp-verification.md).
- Actual Codex CLI, Claude Code and OpenCode clients completed local
reads/mutations. Codex wrote a report, Claude updated it in a later
conversation, and OpenCode uploaded/downloaded a file with matching
SHA-256 and registered the attachment. Revoking the CLI grant rejects
subsequent bridge initialization.
- Butter staging is verified on final runtime `b2196fae1`
([deployment](https://github.com/paperclipai/paperclip-cloud/actions/runs/37538432138)).
A fresh OpenCode workspace fetched the copied invitation, configured
remote MCP, started OAuth and reached real consent with configuration
unchecked. Invalid transfer tickets return 403 through Cloud. Human
approval for the new persistent staging grant is pending; hosted
task/file success is not yet claimed. The final transaction fix is
deployed.
- Full local `pnpm test:run` passed 15,614 general-server tests but
stopped on two macOS timeouts. The heartbeat test passed in isolation;
the existing 40,000-file Git stress fixture timed out again. Its Linux
CI lane passes. Later local full-suite phases did not run after the
timeout; this is not an all-green local full-suite claim.
- Instructions and security limits are in `doc/public-mcp.md`; the saved
plan is `doc/plans/2026-10-06-expanded-assistant-mcp-tools.md`.

## Risks

- This expands the experimental direct MCP surface. Explicit schemas and
domain permissions must stay synchronized.
- Migration 0311 adds transfer tickets and upload receipts. Expired
orphan cleanup must not remove committed attachments.
- Configuration requires a new consent request containing that scope;
the single write-access choice controls it alongside work mutations.
Refreshing an old grant does not add it.
- The public directory keeps its original ten tools through the
companion Cloud change.
- Hosted consent/work proof remains the final delivery gate. The PR
stays draft while approval of the new staging grant is pending; code
checks and review are green. Merging is a separate action.

## Model Used

OpenAI Codex (GPT-6, tool use and code execution). The exact serving
model ID and context window are not exposed in this session. Paid
evaluation models: gpt-5.4-mini, claude-haiku-4-5-20251001;
claude-sonnet-4-6.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused suites; full-suite
macOS limitation disclosed above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 20:29:58 -05:00
Devin FoleyandPaperclip 892b0b3606 fix: scope quota reports to authorized subscription accounts (#14994)
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-06 18:20:37 -07:00
Devin FoleyandPaperclip eab93fd4a0 fix: checkpoint adapter usage and preserve unknown prices (#14991)
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-06 18:16:18 -07:00
Devin FoleyandPaperclip 3ebd7bc0c9 test: isolate accounting fixtures and include all adapter suites (#14989)
Co-Authored-By: Paperclip <noreply@paperclip.ing>
canary/v2026.1007.0-canary.0
2026-10-06 18:13:46 -07:00
ac8f3eb143 feat(ui): remove star and leave buttons from agent index (#15400)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The Agents index page (`/agents/all`) shows one row per agent in
both the streamlined and production sides of the UI
> - Each row had a star toggle and a leave/join button that write
per-user state through the resource-memberships API (`agent_memberships`
table)
> - The default streamlined sidebar no longer lists agents, so the star
and leave actions on the index no longer feed a visible pinned list
> - This pull request removes the star toggle and the leave/join button
from agent index rows and org-tree nodes
> - The backend resource-memberships endpoints and the
`agent_memberships` table stay live because the agent detail page and
the legacy sidebar still use them
> - The benefit is a cleaner index page with no redundant actions for
state that is not surfaced on the page itself

## Linked Issues or Issue Description

No public GitHub issue exists for this change; the description follows
the feature request issue template inline.

**Subsystem affected**
ui/ — React + Vite board UI. The change is presentational only; no
backend or schema code is touched.

**Problem or motivation**
The agent index rows show a star toggle and a leave/join button. Those
two buttons write per-user membership state that the default streamlined
UI no longer surfaces anywhere on the page itself. The default sidebar
shows a plain "Agents" nav item and does not list agents. The buttons
are now redundant on the index.

**Proposed solution**
Remove both buttons from the agent index list rows and org-tree nodes in
both UI variants. Keep the data layer. The agent detail page still
renders the header star toggle and the join/leave banner, and the legacy
sidebar still renders star and leave rows for instances that opt out of
the streamlined UI.

**Alternatives considered**
- Remove the backend endpoints too. Rejected: the agent detail page and
the legacy sidebar still read and write the same state.
- Keep the buttons hidden behind a setting. Rejected: no product need,
and the plan approved the straightforward removal.

**Roadmap alignment**
This is a small presentational change. It does not overlap any section
in `ROADMAP.md`.

**Additional context**
Rows whose membership state is `left` keep their dimmed styling on the
index.

## What Changed

- Removed the `StarToggle` component from agent index list rows in
`ui/src/pages/Agents.tsx` and `ui/src/pages/Agents.production.tsx`
- Removed the `MembershipAction` (Leave/Join) component from agent index
list rows in both files
- Removed both actions from the org-tree nodes in both files
- Dropped now-unused imports and derived values (`StarToggle`,
`MembershipAction`, `isStarred`, `useResourceMembershipMutation`, and
the per-row pending/starred variables)
- Kept `useResourceMemberships` and the dim-on-left row styling so rows
whose membership is `left` remain visually identified
- Updated `ui/src/pages/Agents.test.tsx`: a new test asserts both modes
omit star and leave/join actions from list and org-chart views; the
dim-on-left tests are kept

## Verification

- `pnpm --filter @paperclipai/ui typecheck` passes
- `pnpm --filter @paperclipai/ui build` passes
- UI Vitest suite for `Agents.test.tsx` passes (21 tests)
- `pnpm check:token-gates` is clean
- Manual check on `/agents/all`: no star and no Leave/Join buttons in
the list view or the org-chart view; star/leave still work on the agent
detail page

## Risks

Low risk. This is a presentational UI change with no backend or schema
changes. Star/leave remain available on the agent detail page and the
legacy sidebar. No migration is involved.

## Model Used

- Provider: DeepSeek (via the Paperclip opencode_local adapter, model
`openrouter/~deepseek/deepseek-v4-flash-latest`)
- Model: deepseek-v4-flash (OpenRouter
`openrouter/~deepseek/deepseek-v4-flash-latest`)
- Context window: 128K; used with tool use in the repository
- Assisted with the code change and this PR body; the plan and scope
came from the tracking work item

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [ ] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: devinfoley <139239+devinfoley@users.noreply.github.com>
Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 17:25:41 -07:00
f77fcbf4bf feat(apps): add Telem.AI web search connection (#15379)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Research agents need current web information.
> - The Apps catalog connects agents to remote MCP tools through the
normal access rules.
> - Telem.AI supplies web search and page reading through one API key.
> - This PR adds its catalog entry, optional search settings, artwork,
and setup guide.
> - The connection keeps an operator's saved header policy when they
reconnect.

## Linked Issues or Issue Description

Refs #15302. Related catalog work: #13881.

This PR continues #15302 by Yifei Ai (@aiwen324). Thank you for the
connector and its review fixes. All seven original commits are
preserved. GitHub denied the attempt to push to the contributor's fork,
so this branch retains the repair commit and merges current master.
Master now includes the same test fix.

For a squash merge, keep the original author in the final commit
message:

```text
Co-Authored-By: Yifei Ai <aiwen324@users.noreply.github.com>
Co-Authored-By: Paperclip <noreply@paperclip.ing>
```

A search found no separate public Telem issue or competing Telem PR. The
existing Apps path matches the roadmap.

**Agent or provider**

Telem.AI provides web search and page reading through a hosted MCP
server.

**Why this adapter is useful**

Agents can use multiple search providers through one governed
connection. Operators can set the search tier, auto routing, and
provider lists.

**How the agent is invoked**

The remote MCP server uses Streamable HTTP at
`https://mcp.telem.ai/mcp`. The API key uses an `Authorization: Bearer`
header. See the [official MCP
guide](https://docs.telem.ai/integrations/mcp/).

## What Changed

- Add the Telem.AI definition, research entry, permission review, and
generated registry entry.
- Add four optional settings. Unset settings send no request header.
- Add official light and dark artwork, source records, and a setup
guide.
- Forward company, issue, agent, run, project, and correlation IDs by
default. Preserve a saved policy, including disabled forwarding, on
reconnect.
- Add catalog and connection tests.

## Verification

Current head: `b4164477fb1b312a504789bf17b51c963244c0fd`.
Merged master: `228f0e2807c5b59d2aa129cf2d80b9777ebabf07`.

- Resolved five shared catalog conflicts after the Superagent connection
merged.
- Keep both providers in the research ledger, generated registry,
generator, branding manifest, and connection guide index.
- Correct the combined catalog totals: 52 self-serve candidates, 55
research entries, and 68 Apps entries.
- The published Git tree exactly matches the tested local resolution.
- Catalog and Apps UI suites: **295 tests pass** after the catalog count
fixes.
- Connection service suite: **387 tests pass** in the full run. Its only
failure was the old catalog count. That test passes on a focused rerun
after the fix. This gives **388 passing service tests** across the two
runs.
- Total focused coverage: **683 passing tests**. The first runs exposed
four fixed-count assertions that needed the combined totals.
- Shared package build and plugin SDK compile pass. Token gates and
whitespace checks pass.
- Generation with `--definitions-only` reproduces the Telem definition
and registry. The unrelated AgentMail and Linear drift remains excluded.
- Local UI typecheck ended with exit 137 at the container memory limit.
Full local typecheck, test, and build are not claimed. Earlier runs also
recorded missing Cargo and Node development headers.
- GitHub reports a clean merge state against master `228f0e280`.
- All 54 checks are complete: **52 passed and two Storybook checks
skipped**. No check failed or remains pending.
-
[CI](https://github.com/paperclipai/paperclip/actions/runs/37545412406)
passes on this head. This includes typecheck, build, tests, browser
shards, Runner checks, and Canary Dry Run.
-
[Greptile](https://github.com/paperclipai/paperclip/pull/15379#issuecomment-6024519537)
is **5/5 on this head**. There are no review threads, open P2s,
recommendations, or follow-ups.
-
[Superagent](https://github.com/paperclipai/paperclip/runs/112548142930)
passes.
- Final recovery checks confirm all seven original commits and current
master remain in history. Token gates and whitespace checks pass.
- The final recovery run makes no source change. It verifies the
published repair and retains the local check limits below.
-
[Commitperclip](https://github.com/paperclipai/paperclip/actions/runs/37545407943)
passes with no failures. Its only informational note asks the merger to
keep the author trailer above.
- All seven original contribution commits remain in history. The diff
against master contains the same 14 Telem files. It adds no dependency,
lockfile, schema, or workflow change.
- The managed GitHub CLI capability was missing in this run. The
installed GitHub connection applied the base files, merged master, then
restored the tested combined catalog. No history was rewritten.

The previous head `2d32a0094` passed all remote gates and had Greptile
5/5. Those results do not verify this new head.

The original PR reports live setup, discovery, settings headers, gateway
calls, and context-header forwarding. This repair does not repeat those
account-bound checks. The permission record still marks maintainer live
qualification as outstanding.

## Risks

- Telem.AI receives the six context IDs by default. The saved header
policy controls forwarding. Search use is billed to the account that
owns the key.
- All agents on a connection share its search settings.
- The merge uses master's route-test setup unchanged. The
company-boundary assertions remain intact.
- No schema, dependency, or workflow change is included.
- Live provider evidence is attributed to the original contributor.
Maintainer live qualification remains outside this CI repair.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- Original contribution: Anthropic Claude Opus 5.5 (`claude-opus-5-5`),
1M-token context, through Claude Code with shell, editing, and test
tools, as disclosed in #15302.
- CI repair and review: OpenAI `gpt-6-astra`, through Codex with
reasoning, shell, editing, and GitHub tools. The runtime does not expose
the context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Yifei Ai <aiwen324@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: devinfoley <139239+devinfoley@users.noreply.github.com>
Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.1006.0-canary.25
2026-10-06 16:33:18 -07:00
Devin FoleyandPaperclip 228f0e2807 fix(ui): make pull request review waits visible (#15375)
## Thinking Path

> - Paperclip lets people supervise agent work and review its outputs.
> - A task can wait for a human to review a pull request while an agent
schedules checks.
> - The PR work product and detected external links use separate display
paths.
> - A private PR can disappear from the useful sidebar view when its
provider lookup fails.
> - The monitor countdown says when the next check runs but omits the
saved review request.
> - This change shows the saved PR and requested review in the task's
existing surfaces.

## Linked Issues or Issue Description

**What happened?**

A task kept checking whether a private GitHub PR had merged. Its work
product requested board review, but the Properties sidebar used detected
external objects instead. The PR URL was stored only in metadata, which
the PR refresh path did not read. The composer showed a generic monitor
countdown with no PR link or review action.

**Expected behavior**

Show saved PR links even if provider access fails. During a GitHub
monitor wait, make outstanding review requests visible with the next
check time and a status-check action. Stop requesting review when a PR
merges or closes.

**Steps to reproduce**

1. Register a PR work product with `reviewState: needs_board_review` and
its URL in `metadata.url`.
2. Schedule an external-service monitor for GitHub.
3. Open the task with external-object lookup unavailable or unable to
access the private repository.
4. Inspect Properties and the composer wait strip.

Related work: #14469 added rich artifact cards. #8759 addresses
attention on monitored task blockers. This change uses the existing work
products and monitor action; it adds no blocker or approval mechanism.

## What Changed

- Show saved PRs in Properties, including metadata-only links,
independent of external-object availability.
- Deduplicate equivalent GitHub PR URLs while preserving the saved
navigation link, and put explicit review requests first.
- Keep provider status and freshness on the combined PR row; keep
private PR links usable when lookup fails.
- Show the saved PR review request above the composer and in the
existing monitor banner during a GitHub monitor wait.
- Label the existing monitor action `Check status` and show check
failures inline.
- Suppress review prompts for merged, closed, or archived PRs even if
their review flag is stale.
- Refresh PR metadata using `metadata.url` and the existing `repository`
alias.
- Share saved work-product reads across the thread, Properties, and
Artifacts. Refresh GitHub in a separate query and enrich only matching
PR versions.
- Start a fresh saved-row request on live invalidation so late provider
responses cannot hide new artifacts or changed review requests.
- Cover stalled GitHub lookups in both panels so refresh latency cannot
hide saved work.
- Scope monitor-check mutation state to the task so failures and late
responses do not leak across navigation.
- Refresh GitHub status when the displayed run finishes, including when
saved PR rows have not changed.
- Clarify that PR review and external release handoffs need a saved
human-input interaction with an agent assigned for continuation; a flag,
monitor, or handoff comment alone does not create that card.

## Verification

- 386 tests pass across eleven affected UI and server suites on
`184c7d8746`. Regressions cover saved links, provider status, cold-cache
loading, panel reopening, monitor errors during navigation, and live
updates during provider refresh.
- Three live-update regressions and two run-completion regressions fail
before their fixes and pass afterward. These tests use the real API
client's GET coalescing and abort handling.
- Full repository typecheck and build passed on the merged parent
`1a49112dd2`. The latest UI changes pass UI typecheck and token gates;
CI also verifies the latest build. Capability contract/inventory checks
pass.
- Storybook build and browser checks passed before the query race fix.
Browser checks cover desktop and 390px mobile review waits, plus video,
mixed-file, and empty artifact galleries after background refresh.
- The earlier full local `pnpm test:run` was stopped under disk
pressure. It reported failures outside the changed suites; an isolated
skill-cache run reproduced three existing macOS permission failures. The
full local suite was not rerun for this follow-up.
- Latest-head CI passes on `184c7d8746`, including the browser shards
and canary dry run. Apex is 5/5 with no actionable findings and no
unresolved threads. All eight reported findings are addressed.

## Risks

- This uses saved PR review state. When GitHub access fails, the saved
state can remain stale until the agent updates it. The link remains
visible and the status check remains available.
- `Check status` wakes the existing monitor owner. It does not merge a
PR, accept an approval, or mark the task done.
- No schema, permissions, or scheduler behavior changes.

## Model Used

OpenAI GPT-6 via Codex, with reasoning, repository analysis, code
editing, and test execution. The exact deployment model ID and context
window were not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass — 277 affected tests
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 16:03:13 -07:00
Devin Foley 8cbd21b3e7 feat(apps): add Superagent connection (#15394)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents reach external services through the Apps catalog. Each
catalog entry is a reviewed `AppDefinition` that connects a provider's
hosted MCP server to Paperclip's shared vault, grants, policies,
gateway, and audit trail.
> - Superagent (`superagent.sh`) is a security platform. Its hosted MCP
server lets agents read and triage security findings, start red-team
reports, run Contributor Trust and dependency update jobs, and score web
pages, email, files, skills, MCP repositories, and packages before an
agent trusts them.
> - Superagent is not in the catalog. Operators must use the generic
"Connect your own MCP server" flow. That flow has no branding, no key
guidance, and no warning about billable or destructive tools.
> - The connector playbook supports this provider with existing
definition fields. The server accepts an organization API key as a
bearer header. It publishes no OAuth authorization-server metadata, so
browser sign-in is not possible.
> - This pull request adds the Superagent definition, its artwork, the
research and permission-review ledger rows, documentation, deterministic
tests, and one reviewed risk-classification rule.
> - The benefit is a branded, governed Superagent connection with clear
key guidance and a warning about tools that cost credits or delete data.

## Linked Issues or Issue Description

**Problem or motivation**

Teams that use Superagent for PR security, red teaming, and agent
guardrails want their Paperclip agents to read findings, return
structured reports, and score content before they use it. Superagent is
not in the Apps catalog. Operators must paste the MCP URL and an
`Authorization` header into the generic remote-MCP flow. That flow gives
no branding and no provider guidance. It also does not tell the operator
that the key reaches the whole organization, or that some tools consume
credits or permanently delete findings.

**Proposed solution**

Add a catalog-only Superagent connection that follows the connector
playbook. It has one method: a customer organization API key
(`sk_live_...`), sent as an `Authorization: Bearer` header to
`https://www.superagent.sh/mcp`. The field helper text explains that
Superagent keys are not scoped. The method warning tells operators to
set billable and destructive actions to Ask first before agents run
unattended. All discovered tools stay governed by the normal per-action
policies.

**Alternatives considered**

A browser sign-in method was not added. The server's protected-resource
metadata names `https://superagent.sh` as its authorization server, but
that origin publishes no `oauth-authorization-server` or
`openid-configuration` document, so Paperclip cannot discover OAuth
endpoints. A plugin was not needed because the connection needs no
custom UI, tables, workers, or webhooks. Relying on the generic risk
classifier was not enough. Several Superagent mutations
(`triage_finding`, `scan_*`, `restore_agent_builtin_rule`) use names
that it reads as reads, so a narrow reviewed Superagent rule was added
instead.

**Roadmap alignment**

This extends the existing self-serve remote-MCP connection catalog. It
does not overlap planned core work.

## What Changed

- Added the `superagent` row to
`packages/shared/src/self-serve-mcp-research.json` (API-key auth, risk
tier S4).
- Added the `superagent` provider to
`scripts/ingest-app-definitions.mjs` (category, key placement and
placeholder, console links, guidance, description). Regenerated
`packages/shared/src/app-definitions/superagent.json` and the generated
registry.
- Added the `superagent/mcp-api-key` permission review to
`doc/connections/tool-method-permission-reviews.json`, with
key-permission text and evidence links.
- Added Superagent's official mark
(`ui/public/brands/apps/superagent.png`, the 460×460 avatar of the
official `superagent-ai` GitHub organization) and the brand manifest
entry.
- Added gallery copy for the Superagent card.
- Added a reviewed Superagent rule to `classifyRisk` in
`server/src/services/tool-access.ts`. Only `list_*` and `get_*` tools,
and tools that Superagent marks read-only, are reads. `delete_*` and
`revoke_agent_client` are destructive. All other tools are writes, so
billable and rule-changing tools can be set to Ask first.
- Put the Ask-first advice in the API-key helper text, because the key
form shows helper text and not method warnings.
- Added `doc/connections/SUPERAGENT.md` (transport and auth, why there
is no OAuth, administrator setup, capabilities and policy, manifest,
brand provenance, validation hook). Linked it from the connections
README and the permission audit.
- Tests: definition shape, store visibility and artwork, URL
recognition, the bearer header on discovery with the key kept out of
connection config, read/write/destructive classification of fixture
tools, the Superagent risk rule (including `triage_finding` and
`restore_agent_builtin_rule`), the visible Ask-first advice and API-key
gating of the connect form, and the pinned catalog counts.

## Verification

- `pnpm exec vitest run packages/shared/src/app-definitions.test.ts
packages/shared/src/app-definitions-url.test.ts
ui/src/pages/apps/AppsConnect.test.tsx ui/src/pages/apps/Browse.test.tsx
ui/src/lib/app-brand-assets.test.ts
ui/src/pages/apps/AppLogo.brand-assets.test.tsx`: 317 passed.
- `pnpm exec vitest run
server/src/__tests__/tool-access-service.test.ts`: 385 passed.
- `node scripts/check-app-brand-assets.mjs` and `node --test
scripts/app-brand-validation.test.mjs`: passed.
- `pnpm --filter @paperclipai/shared typecheck`, `pnpm --filter
@paperclipai/server typecheck`, and `pnpm --filter @paperclipai/ui
typecheck`: clean.
- Manual: in a local instance, open Apps → Browse and confirm the
Superagent card and icon. Open `/apps/connect?source=superagent`.
Confirm the single API-key method, and confirm that Connect enables only
after a key is entered.
- Live metadata probe on 2026-10-06: an unauthenticated `initialize` on
`https://www.superagent.sh/mcp` returns 401 with
`resource_metadata="https://www.superagent.sh/.well-known/oauth-protected-resource"`.
That document returns 200. No authorization-server metadata exists at
the named issuer.

## Risks

- Low risk to existing providers. The change is additive catalog data
plus tests. The generated registry only gains one import. The new risk
rule runs only for Superagent connections.
- A Superagent key reaches its whole organization. Some tools consume
credits (`create_*_report`, `triage_finding`) or delete data permanently
(`delete_finding`). Every action starts Allowed under the current
product default. The key helper text tells operators to set these
actions to Ask first.
- The permission-review ledger records live proof as not run. No
Superagent account was used. The lifecycle checklist in
`doc/connections/SUPERAGENT.md` needs a documented pass before the entry
is fully qualified.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- Claude Opus 5.5 (`claude-opus-5-5`, 1M context) in Claude Code, with
extended thinking and tool use (shell, file editing, web fetch, browser
checks).

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-10-06 15:57:16 -07:00
dependabot[bot] 5c15be3715 build(deps): bump actions/cache from 5.1.0 to 6.1.0 (#12962)
Bumps [actions/cache](https://github.com/actions/cache) from 5.1.0 to
6.1.0.
<details>
<summary>Release notes</summary>
<p><em>Sourced from <a
href="https://github.com/actions/cache/releases">actions/cache's
releases</a>.</em></p>
<blockquote>
<h2>v6.1.0</h2>
<h2>What's Changed</h2>
<ul>
<li>Bump <code>@​actions/cache</code> to v6.1.0 - handle read-only cache
access by <a
href="https://github.com/jasongin"><code>@​jasongin</code></a> in <a
href="https://redirect.github.com/actions/cache/pull/1768">actions/cache#1768</a></li>
</ul>
<p><strong>Full Changelog</strong>: <a
href="https://github.com/actions/cache/compare/v6...v6.1.0">https://github.com/actions/cache/compare/v6...v6.1.0</a></p>
<h2>v6.0.0</h2>
<h2>What's Changed</h2>
<ul>
<li>Update packages, migrate to ESM by <a
href="https://github.com/Samirat"><code>@​Samirat</code></a> in <a
href="https://redirect.github.com/actions/cache/pull/1760">actions/cache#1760</a></li>
</ul>
<p><strong>Full Changelog</strong>: <a
href="https://github.com/actions/cache/compare/v5...v6.0.0">https://github.com/actions/cache/compare/v5...v6.0.0</a></p>
</blockquote>
</details>
<details>
<summary>Commits</summary>
<ul>
<li><a
href="https://github.com/actions/cache/commit/55cc8345863c7cc4c66a329aec7e433d2d1c52a9"><code>55cc834</code></a>
Merge pull request <a
href="https://redirect.github.com/actions/cache/issues/1768">#1768</a>
from jasongin/readonly-cache</li>
<li><a
href="https://github.com/actions/cache/commit/d8cd72f230726cdf4457ebb61ec1b593a8d12337"><code>d8cd72f</code></a>
Bump <code>@​actions/cache</code> to v6.1.0 - handle cache write error
due to RO token</li>
<li><a
href="https://github.com/actions/cache/commit/2c8a9bd7457de244a408f35966fab2fb45fda9c8"><code>2c8a9bd</code></a>
Merge pull request <a
href="https://redirect.github.com/actions/cache/issues/1760">#1760</a>
from actions/samirat/esm_migration_and_package_update</li>
<li><a
href="https://github.com/actions/cache/commit/e9b91fdc3fea7d79165fceb79042ef45c2d51023"><code>e9b91fd</code></a>
Prettier fixes</li>
<li><a
href="https://github.com/actions/cache/commit/e4884b8ff7f92ef6b52c79eda480bbc86e685adb"><code>e4884b8</code></a>
Rebuild dist</li>
<li><a
href="https://github.com/actions/cache/commit/10baf0191a3c426ea0fa4a3253a5c04233b6e18f"><code>10baf01</code></a>
Fixed licenses</li>
<li><a
href="https://github.com/actions/cache/commit/e39b386c9004d72a15d864ade8c0b3a702d47a37"><code>e39b386</code></a>
Fix test mock return order</li>
<li><a
href="https://github.com/actions/cache/commit/b6928203372a8571ff984c0c883ef3a1adfb0c06"><code>b692820</code></a>
PR feedback</li>
<li><a
href="https://github.com/actions/cache/commit/60749128a44d25d3c520a489e576380cf00ff3f1"><code>6074912</code></a>
Rebuild dist bundles as ESM to match type:module</li>
<li><a
href="https://github.com/actions/cache/commit/5a912e8b4af820fa082a0e75cfd2c782f8fbfe0e"><code>5a912e8</code></a>
Fix lint and jest issues</li>
<li>Additional commits viewable in <a
href="https://github.com/actions/cache/compare/v5.1.0...v6.1.0">compare
view</a></li>
</ul>
</details>
<br />

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-10-06 15:36:53 -07:00
45ac4e945d docs(release): stable notes for the 2026.1006.0-beta.0 soak (v2026.1009.0) (#15388)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The release process promotes a soaked beta to stable. The stable
GitHub Release body comes from `releases/beta/v<beta-version>.md` on
`master`
> - The `draft_stable_notes` job pushed a commit-log skeleton to this
branch when `2026.1006.0-beta.0` published today (workflow run
37516507660)
> - The skeleton is grouped commit subjects. It is not in release-notes
voice and it does not tell a self-hoster what to do before an upgrade
> - This pull request replaces the skeleton with curated notes. Each
claim traces to a pull request in `v2026.1005.0..22a3ea34`
> - The benefit is that the earliest stable promotion (v2026.1009.0)
finds finished notes on `master` on time

## Linked Issues or Issue Description

**Issue type**

Missing content

**Where is the issue?**

`releases/beta/v2026.1006.0-beta.0.md` — the stable-notes file for the
`2026.1006.0-beta.0` soak. It exists only on this machine-owned branch
and only as the auto-generated skeleton.

**What's wrong?**

The stable promotion reads this file from `master` and publishes it as
the GitHub Release body. Until this branch merges, the stable preflight
has no notes to resolve. The skeleton lists raw commit subjects with
nested PR summaries. It does not call out the two changes that need
operator action before the upgrade.

**Suggested fix**

Merge the curated notes so the notes invariant holds for the
v2026.1009.0 promotion. Correct the `> Released:` date in a follow-up if
the promotion date slips.

## What Changed

- Replaced the skeleton in `releases/beta/v2026.1006.0-beta.0.md` with
curated stable notes for v2026.1009.0 in the same layout as
`releases/v2026.1005.0.md`: overview, Breaking Changes, Highlights,
Fixes, Improvements, Upgrade Guide, Contributors
- Breaking Changes lists the SQLite restore-lock protocol change
(#14869) and stricter stored tool-grant enforcement (#14915)
- Upgrade Guide enumerates migrations `0294` through `0305`, the three
new default-off experimental settings, the
`PAPERCLIP_CONNECTION_INSTRUCTIONS_FILE` contract for custom adapters,
and the harness pin moves
- Release and CI internals, smoke specs, and canary tooling are left
out. Features already described in the v2026.1005.0 notes are not
repeated

## Verification

- Beta publish is complete: npm dist-tag `beta` is `2026.1006.0-beta.0`,
tag `beta/v2026.1006.0-beta.0` points at
`22a3ea3414e9039a538fbb0374b3cfa7fc4371eb`, and the `publish_beta`,
`smoke_beta`, and `draft_stable_notes` jobs in run 37516507660 all
succeeded
- `git rev-list --count
v2026.1005.0..22a3ea3414e9039a538fbb0374b3cfa7fc4371eb` returns 130,
which matches the overview and Contributors section
- `git shortlog -sn --no-merges` over the same range shows 8 human
authors after the two bot accounts are excluded
- `./scripts/release.sh stable --date 2026-10-09 --print-version`
returns `2026.1009.0`
- Every `#NNNN` link in the file resolves to a pull request inside the
range. Migration numbers, setting keys, env var names, and defaults were
checked in the code on the source commit, not in commit subjects
- The file contains no internal ticket ids or instance-local links

## Risks

- Low risk: a single markdown file, no source changes. If the promotion
date slips past 2026-10-09, the `> Released:` line and H1 need a
one-line update before the stable dispatch. The beta-keyed filename
makes that re-date harmless

## Model Used

- Claude (Anthropic), model ID `claude-fable-5-1` (Claude Fable 5.1),
extended thinking enabled, tool use via Claude Code

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: devinfoley <139239+devinfoley@users.noreply.github.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
canary/v2026.1006.0-canary.24
2026-10-06 15:35:36 -07:00
dependabot[bot] c365a16e34 build(deps): bump actions/deploy-pages from 4.0.5 to 5.0.1 (#12963)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The repository runs a weekly full-stack end-to-end campaign for the
runner, and that workflow publishes its merged dashboard to GitHub Pages
> - The publish flow has two actions: `actions/upload-pages-artifact`
packages the dashboard, and `actions/deploy-pages` deploys it
> - #12964 moved `actions/upload-pages-artifact` to version 5, and
`actions/deploy-pages` is still pinned to a version 4 commit
> - Version 4 of `actions/deploy-pages` runs on the Node.js 20 runtime,
and GitHub-hosted runners are moving actions to Node.js 24
> - This pull request moves the pin to the commit that upstream tags as
`v5.0.1`
> - The benefit is that both Pages actions are on the same major
version, the deploy step runs on Node.js 24, and the pin stays a commit
SHA

## Linked Issues or Issue Description

No public issue exists for this change. Refs #12964 (the matching
`actions/upload-pages-artifact` update, now merged) and Refs #12962
(another pinned GitHub Action update). They are related, and they are
not duplicates. The description below follows the enhancement template.

**What existing behavior does this improve?**

The weekly runner end-to-end workflow deploys its merged dashboard to
GitHub Pages. This change updates the action that performs that
deployment.

**Subsystem affected**

CI and release automation —
`.github/workflows/runner-full-stack-e2e.yml`.

**Current behavior**

The `pages` job uses `actions/deploy-pages` at commit
`d6db90164ac5ed86f2b6aed7e0febac5b3c0c03e`, which runs on `node20`.

**Proposed behavior**

The `pages` job uses `actions/deploy-pages` at commit
`368f82528645a54fb793d4d04e342629a3f51346`, which upstream tags as
`v5.0.1` and which runs on `node24`. Version 5.0.1 also adds backoff and
jitter to the deployment status polling.

**Breaking changes**

None for this repository. The `action.yml` of both versions declares the
same six inputs (`token`, `timeout`, `error_count`,
`reporting_interval`, `artifact_name`, `preview`) with the same
defaults, and the same `page_url` output. The job passes `artifact_name`
only. The job runs on `ubuntu-latest`, which supports the `node24`
runtime.

## What Changed

- Change the pinned commit of `actions/deploy-pages` in
`.github/workflows/runner-full-stack-e2e.yml` from the version 4 commit
to `368f82528645a54fb793d4d04e342629a3f51346`, which upstream tags as
`v5.0.1`.

## Verification

- The pinned commit matches the upstream tag. `gh api
repos/actions/deploy-pages/git/ref/tags/v5.0.1 --jq .object.sha` returns
`368f82528645a54fb793d4d04e342629a3f51346`.
- The input and output contract stays the same. A comparison of
`action.yml` at the old pin and at the new pin shows identical inputs,
defaults and outputs. Only the runtime changes from `node20` to
`node24`.
- The repository CI suite passes on this branch after a rebase onto the
current base branch.
- One limit applies. The changed step runs only in the `Runner
Full-Stack E2E` workflow. That workflow starts on a weekly schedule and
on a manual dispatch, so no pull-request run exercises the step. A
maintainer can exercise it with a manual dispatch of that workflow, or
the next scheduled run exercises it.

## Risks

- Low risk, with one limit. The changed step does not run on a pull
request, so the pull-request checks do not prove the new action version
in this workflow.
- The runtime moves to Node.js 24. GitHub-hosted runners support it, and
this job uses a GitHub-hosted runner.
- The rollback is one commit. Restore the previous pinned commit of
`actions/deploy-pages`.

## Model Used

Dependabot generated this dependency update automatically, so no AI
model produced the code change. A maintainer wrote this description with
Claude Opus 5 (Anthropic, model id `claude-opus-5`, extended thinking,
tool use).

## Checklist

Three boxes stay unticked on purpose. This change pins one GitHub Action
version in one workflow file. No local test covers a pinned action
version, no new test applies, and no document refers to this pin.

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub #NNN / github.com/paperclipai/paperclip URLs)
- [x] My branch name describes the change (e.g. docs/..., fix/...) and
contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [ ] I have added or updated tests where applicable
- [ ] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-10-06 14:32:16 -07:00
DottaandPaperclip b508a05c43 feat: add internal agent complaints and suggestions (#15367)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents use legacy skills or native runner tools to work on tasks.
> - Those agents can encounter friction that does not belong in the task
thread.
> - A complaint should preserve the raw reaction. A suggestion should
describe an improvement.
> - This pull request adds attributed local storage and both submission
paths.
> - Agents can submit feedback once and continue their primary work.

## Linked Issues or Issue Description

**Subsystem affected**

Server, database, shared contracts, runtime skills, and native runner
tools.

**Problem or motivation**

Agents have no default internal channel for incidental complaints and
suggestions. Sending this feedback through task comments adds noise and
can alter task workflows.

**Proposed solution**

Store free-form feedback in the current instance database. Derive agent,
run, company, and task attribution from active authority. Provide
default legacy skills and provider-neutral native actions. Keep the
instructions close to Warp's MIT-licensed originals.

**Alternatives considered**

Task comments and external Slack delivery add unwanted side effects.
Mandatory suggestion fields and short editorial limits would discard
useful feedback. This release has no listing API, UI, read tool,
automatic triage, or external forwarding.

**Roadmap alignment**

This is a maintainer-requested addition to the existing runtime skills
and runner tool paths. It does not duplicate a listed roadmap milestone.
Searches for complaint tooling, suggestion-box, and agent commentary
found no overlapping public PR or issue.

## What Changed

- Add the company-scoped `agent_commentary` table, shared validation,
and idempotent migration `0310`.
- Add one transactional service and the agent-only POST route. Validate
active authority before writes or replay. Redact known credentials.
Commit a content-free audit with each new record.
- Add `submit_complaint` and `submit_suggestion` to standard, ask, and
planning modes. Keep review, revocation, and completion restrictions.
Store replay identity on the commentary row.
- Mount `complain` and `suggestion-box` by default for legacy agents.
Bundle a dependency-free Node.js stdin helper in the operational skill
and allow its POST through the sandbox bridge.
- Preserve Warp's complaint voice and suggestion guidance, with
attribution and local transport adaptations. Keep source attribution and
MIT notices in each skill's LICENSE, outside runtime instructions.
- Document custom-runtime HTTP use and database inspection. Add
real-database tests and a repeatable live Codex smoke for local and
Daytona execution.
- Pin the lagging-source migration fixture before the identity-repair
migration so later migrations preserve its regression coverage.

## Verification

- Personally ran real Codex submissions in all four environments on
2026-10-06. Local runs passed at 20:35 UTC. Daytona native passed at
20:31 UTC; Daytona legacy passed at 20:33 UTC. Each stored exactly two
rows with company, agent, run, and task attribution, wrote the
continuation marker, exited zero, created no task comments, and left
task status unchanged. Each recorded two content-free activity entries.
- Daytona used production provider hooks, real remote execution and file
transfer, the legacy queue callback bridge, and native private WebSocket
ingress. The current Linux runner was built from `abf47b595`, staged,
and verified against controller contracts. Both sandboxes were confirmed
deleted. This is a focused feedback transport smoke; it does not claim
full Runner E2E catalog or browser qualification.
- The immutable base image and Linux binary digest are recorded in [the
verification
documentation](https://github.com/paperclipai/paperclip/blob/codex/agent-commentary/doc/agent-commentary.md#verification).
The smoke script can save content-free JSON evidence. No credentials or
feedback bodies are in these reports.

| Environment | Runner | Complaint row | Suggestion row |
| --- | --- | --- | --- |
| local | legacy Codex | `59413a00-1de2-4bb1-bcc6-9c4b54c64aa6` |
`3db2364d-3e15-4f47-846f-875d3902999d` |
| local | native Codex | `5da22b5f-41df-4de5-8ba0-d9345ab01267` |
`2d5abe17-dd41-403c-a5ee-4729f2d58921` |
| daytona | legacy Codex | `27c9d0aa-8477-409f-9da0-e8ffa48dee50` |
`209681c9-d1e9-4ce1-999e-48fa07692389` |
| daytona | native Codex | `6eb001bb-4bcf-43f7-8717-f662f53dc7c3` |
`77c383d8-a997-49e5-a33e-25c70e15c0b2` |

- Run the local check with `node cli/node_modules/tsx/dist/cli.mjs
server/scripts/verify-agent-commentary-live.ts`. The documentation gives
the Daytona invocation. Both use disposable instance databases and
normal Codex provider usage.
- Repository `pnpm -r typecheck` and `pnpm build` passed after the test
extension. The build includes runner generation, contracts, and replay
checks. The smoke scripts also passed a separate TypeScript check. The
lagging-source migration regression passed. All equivalent current-head
Vitest CI shards passed. The local monolithic `pnpm test:run` invocation
was stopped after CI supplied that coverage; it did not complete
locally.
- Focused tests cover company isolation, spoofing, revoked credentials,
stale ownership, post-finish rejection, concurrent replay, conflicting
keys, atomic rollback, and deletion through existing services. Boundary
tests cover empty text, Unicode, text beyond 8,000 characters, and the
524,288-character ceiling without truncation. Mounting tests cover
Codex, Claude, and sandbox staging. Helper tests cover standalone Node
execution, stdin, invalid UTF-8, redirects, HTTP failure, and its
deadline. Privacy and bridge tests cover successful and rejected
requests.
- Instructions were compared with Warp's originals. MIT notices and
source credits live only in LICENSE files. Native tools preserve
truthful disclosure when asked, without routine announcements.
- [Full
CI](https://github.com/paperclipai/paperclip/actions/runs/37508559190)
and Greptile 5/5 passed on the earlier feature commit `5209c3501`. The
later head found the migration-fixture assumption fixed in this update.
On `8a4965164`, all 55 check contexts passed after one browser shard
rerun. Its initial reviewer signoff failure also passed an isolated
local browser run (1 test). Greptile scored that head 5/5 and identified
one smoke cleanup gap. `6ecbafb0b` fixes failed-acquisition cleanup with
four passing tests and a passing smoke-script typecheck. Fresh CI is
pending for this final test-only fix. No commentary production code
changed during verification.

## Risks

- Feedback is internally attributed. It is not anonymous. Existing
redaction removes known credentials, but agents must still omit
sensitive content. Normal provider transcripts can include their
submitted arguments.
- Default skill availability changes for existing legacy agents. Runtime
policy filtering still applies. The helper uses the existing Node.js
runtime with no extra dependencies; custom runtimes can call the HTTP
endpoint.
- Feedback is removed with its run, agent, or company. Task deletion
clears only the issue pointer. Normal database backups include the
table.
- The migration is additive and has no backfill. Writes serialize on the
active run for replay consistency. No server suggestion quota is
imposed.

## Model Used

OpenAI `gpt-6-astra` through Codex, with `xhigh` reasoning effort and a
reported 258,400-token context window. Capabilities used: repository
inspection, code execution, and live runtime verification. No subagents
were used.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 16:14:06 -05:00
DottaandPaperclip fe9d865250 fix(native): wake parents after revised child completions (#15372)
## Thinking Path

> - Paperclip manages agents and the tasks they perform.
> - A parent needs a new notification when a child finishes requested
revisions.
> - Native completion currently identifies that wake by the parent and
child only.
> - An already consumed notification can suppress the child's next
completion.
> - This PR binds ordinary child notifications to the committed status
decision.
> - A replay keeps the same identity, while a new completion can wake
the parent.
> - Live evaluation then exposed new child results being swallowed by an
active parent.
> - The repair preserves that input for a fresh turn and prevents
premature Done from discarding it.

## Linked Issues or Issue Description

Refs: #15218. The procedure experiments remain unshipped. This fixes
native completion wake identity and durable delivery to a busy parent.

Related open work: #10559 and #4507 address legacy wake deduplication,
#11179 addresses delivery during an active parent run, and #13044
addresses watchdog signals. This PR changes native committed-decision
keys and their delivery to an active parent, while preserving exact
watchdog behavior.

**What happened?**

A native child completed, its parent consumed the notification, and
feedback reopened the child. The second completion found the old
completed wake and created no new notification.

**Expected behavior**

Each new ordinary child completion can notify the parent. Replaying the
same committed completion must not add a notification. If the parent is
already running, the new result must remain available for a fresh turn;
an unread result must survive a parent Done claim.

**Steps to reproduce**

Complete a native child, consume its parent wake, reopen and complete
that same child, then inspect the parent wakes. The regression fails on
unchanged master because only one wake exists after two completions.

**Paperclip version or commit**

Baseline: `0fe47882cfcb12082035113c59ca96091c46ebfc`.

**Deployment mode**

Native runner with the standard server and PostgreSQL control plane.

## What Changed

- Include the durable status-decision ID in ordinary child completion
wake keys.
- Defer new ordinary native child completions behind an active parent,
carrying the committed decision identity and revised summary.
- Keep a parent in progress while that result is still queued, claimed
or deferred; recheck before the status commit. Use the existing
continuation without reviving cancelled tasks.
- Lock the parent before ordinary child status writes, making the
notification/Done ordering explicit without upgrading an implicit
foreign-key lock.
- Cover exact sequential delivery, dispatcher replay, and
queued/claimed/deferred versus consumed/current-run completion
identities in database tests.
- Apply it when the child is also a dependency and when it is only a
child.
- Preserve the stable key for exact `task_watchdog` origins to avoid
repeated watchdog loops.
- Extend real database conformance coverage for both relationships,
ordinary and near-match origins, watchdogs, revised summaries, and
replay.
- Reconcile the working checklist with the merged guidance PRs and
record the bounded next step.
- Reuse the current composer helper for Everyday task creation: capture
the returned task ID, preserve the exact prompt and chosen
assignee/project, and cover the setup with paused-agent browser tests.
The same setup correction is present in both comparison variants.

## Verification

**Ready for review and merge at
`e2fc0c8e3ddb84dd9bc045704c3d1ecb23ee3447`: both original live cases
pass, zero new failures against the frozen baseline, all checks green,
CLEAN/MERGEABLE and out of draft. Not merged.**

| Profile | Original baseline | Initial candidate | Fixed candidate |
| --- | --- | --- | --- |
| native Codex / `gpt-5.6-sol` | PASS | FAIL | PASS |
| ACPX Claude / `claude-sonnet-5` | FAIL | PASS | PASS |

One new pass, one unchanged pass, zero new failures and no pending pairs
against the original baseline. The earlier failed candidate is
preserved; it was fixed and measured at a new source, not regraded or
rerolled unchanged.

### Current source and live evidence

- [Completed campaign
37523025407](https://github.com/paperclipai/paperclip/actions/runs/37523025407)
and [public report with original
screenshots](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37523025407-1/index.html).
Report and screenshots return HTTP 200; screenshot bytes match original
retained evidence.
- Measured candidate `e2fc0c8e3ddb84dd9bc045704c3d1ecb23ee3447`; reused
frozen baseline `7d700f43e93e89e4c196f8799b0aa3cef41d20da`. Trusted
workflow `88ff98b83d15eaa6640ee1848df4f0ad9bc82af3` is distinct from
measured sources; workflow blob
`0600886144d3e22ea2e4a38329a79177882f3948`. Prompts, fixture, oracle,
model/profile and input controls are unchanged. Both variants include
the same composer setup repair.
- Both original candidate grades PASS: all 31 checks per cell, parent
and child Done, independently tested revised ZIPs, successful cleanup.
Exactly one first attempt per profile on this source: 10 actual runs, no
retries. Reusing nine baseline runs gives 19 matched runs; the baseline
was not rerun.
- Claude directly exercised the repair: its revised child completed
while the parent was active, the parent's Done claim became InProgress
with `native_child_completion_pending`, and a fresh parent turn then
completed with the revised artifact.
- Both persisted final provider comments refer to the same attachment
whose parent-registration hash matches the revised ZIP tested by the
original oracle; native final evidence also references that attachment.
Codex provides a clickable download link. Claude describes the new
`--max-length` behavior but uses a backticked attachment ID, without a
clickable URL in the final prose. The artifact is registered on the
parent and independently downloaded/tested; prose-link usability remains
a presentation limit outside the original oracle. Artifact identity and
observed execution do not prove cognitive review.
- 68 focused scheduling/conformance/arbiter tests pass. Four regressions
fail against the original production files while four controls pass. The
database scheduler test proves one sequential continuation with the
revised summary and replay deduplication. It uses a mock adapter and is
separate from the live proof.
- Full local `pnpm -r typecheck` and `pnpm build` pass. [Exact-head full
CI
37522126627](https://github.com/paperclipai/paperclip/actions/runs/37522126627):
51 successful checks, two intentional skips, separate Snyk success.
Fresh exact-head Greptile
[5/5](https://github.com/paperclipai/paperclip/pull/15372#issuecomment-6022179172),
no new actionable findings; all review threads resolved.
- Ready-transition Contributor trust and Superagent Security Scan both
pass. The security scan completed at 2026-10-06T20:26:53Z with zero
annotations. Final total: 53 successful check runs, two intentional
skips and separate Snyk success; aggregate SUCCESS, source unchanged,
zero unresolved threads.
- The explicit parent lock precedes child writes. Its controlled
PostgreSQL ordering also passes on previous production through an
implicit foreign-key lock; this is hardening, not a reproduced
additional live failure. The four original failing regressions remain
the before/after proof of the active-parent repair.

### Preserved failures and accounting

- Initial candidate `a2ae2324ce7692e704fc43e99904b076d6246fee`: Codex
PASS → FAIL, Claude FAIL → PASS. Equal totals concealed a new failure
and did not qualify that source. Its Codex parent ended Blocked after
the revised child result coalesced into the active parent; there was no
retained later parent execution and the independent ZIP oracle was never
reached. A later server backstop log did not prove recovery.
- In the initial Claude cells, both parent finals referenced an earlier
parent ZIP while the oracle tested the revised child's ZIP. The original
candidate PASS did not establish latest-artifact delivery. Earlier
parent ZIP bytes are absent, so different hashes alone do not prove
missing functionality. Those original grades and content findings are
unchanged.
- Original baseline
[report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-37509893406-1/index.html);
original candidate Claude
[artifact](https://github.com/paperclipai/paperclip/actions/runs/37509916344/artifacts/11434353106);
failed candidate Codex
[campaign](https://github.com/paperclipai/paperclip/actions/runs/37513070706).
The original cohort contains 17 actual runs, no retries, successful
cleanup.
- Campaign 37509916344 was cancelled after its Codex job waited over 16
minutes without a runner or steps. Completed Claude evidence was
preserved; only the missing Codex first trial was dispatched in
37513070706. The trusted workflow revision differs but workflow bytes
are identical. The cancelled campaign has no published HTML. Earlier
campaigns 37506546098 / 37506550087 retain four composer setup failures,
zero agent runs; the identical fixture-only repair restored setup
without changing the prompt or outcome oracle.
- Intermediate `3c1cf6830ce9db4123b834ffb99c594c383a8986` [campaign
37518652522](https://github.com/paperclipai/paperclip/actions/runs/37518652522)
was cancelled after a concurrency review finding. Both paid-cell steps
started and both tasks were created. Only invocation policies survived;
no grade, run inventory or cleanup receipt. Provider activity and
charges are unknown: two incomplete attempts, not passes or
zero-provider setup failures.
- Cumulative accounting: **27 known actual runs** (17 original + 10
fixed-candidate), plus unknown activity in those two cancelled
intermediate attempts. The 19-run matched comparison reuses nine
baseline runs and is not additional execution. Reported LLM amounts are
zero with original billing `complete=true`; actual charges are unknown
and local/hosted runtime is unmetered. No free-run, speed or cost claim.
- Original and new results pass the canonical result validator. Retained
source, input, result/API/story/final-ledger/usage run identities and
artifact hashes were audited. Each downloaded package omits the
pre-upload-declared `playwright-output/.last-run.json`; primary result,
API, final ledger, story, screenshots and reached ZIP oracles are
retained. No full-package completeness claim.
- The initial redundant local full test invocation was stopped after
2,492.5 seconds once that head's CI passed; completed groups recorded
23,045 passes and 87 skips. That local invocation remains incomplete.
Current full CI is the repository-wide test evidence.

## Risks

A revised completion can schedule another sequential parent run and its
normal budget use. A parent Done claim remains non-terminal while a
newer native child result awaits delivery. Company scope, governance,
workspace-finalization, terminal cancellation and exact watchdog
behavior remain enforced. This changes native completion authority and
scheduling, so the durable identity, replay and concurrent-commit
controls matter.

The two live trials qualify the observed revised-child handoff, not
broad task quality or causal/general equivalence. Claude's final prose
still gives an attachment ID without a clickable URL, although the
registered revised artifact passes the original download oracle.
Completion before any new child result exists, removal of unfinished
dependencies and broader instruction reduction remain separate
questions. There is no schema or prompt change.

## Model Used

OpenAI Codex, GPT-6 family, with repository inspection, shell execution
and code editing. The exact serving model ID, reasoning setting and
context window are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 16:11:43 -05:00
DottaandPaperclip ee3c728b95 fix(ui): keep No project available in new composer (#15382)
Keep the no-project action first and available during project search, while preserving matching-project keyboard selection. Add focused regression coverage for ordering, projectless submission, and filtered Enter behavior.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
canary/v2026.1006.0-canary.23
2026-10-06 16:04:08 -05:00
DottaandPaperclip 582911ba74 fix(access): give Operators default company editing permissions (#15377)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Human company members receive permission grants from their role
preset.
> - Human invitations use the Operator role by default.
> - The old Operator preset only granted task assignment, so ordinary
members could not edit agents or manage connections.
> - This pull request gives Operators company editing and audit access
while keeping join approval and member-permission management separate.
> - Members can configure their agents and accounts without an Owner
role.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

The default permission grants for human Operators and the invitation
role description.

**Subsystem affected**

server/ permission presets and invitation authorization, plus ui/
invitation and tool access gates.

**Current behavior**

Operators receive only `tasks:assign` as an explicit default grant.
Agent configuration and managed account setup fail their permission
checks.

**Proposed behavior**

Operators receive agent creation/configuration, skills, environments,
invitations, task assignment, pipelines, connections, tool
management/use, and audit grants. The preset excludes `joins:approve`
and `users:manage_permissions`. Operators can invite Operators and
Viewers. Inviting a role with either excluded power requires that power,
so invitations cannot bypass the restriction.

**Reason and benefit**

Ordinary company members can edit company work and connect accounts
through the existing routes.

**Breaking changes**

The Operator preset grants more permissions. The existing startup and
Cloud sign-in seeding paths can insert these missing grants for existing
Operators. They retain custom grant scopes. Explicit invitation grants
still take precedence. Inviting Admins now requires join approval;
inviting Owners also requires member-permission management. This PR adds
no migration or new backfill path.

Related: #5945 describes missing role grants in another membership entry
point. This PR changes the preset and does not change that endpoint.

## What Changed

- Expanded the Operator preset to 14 company editing, invitation, tool,
and audit grants.
- Kept join approval and member-permission management out of the preset.
- Required those powers when an invitation's selected human role
includes them, preventing Operators from delegating the excluded powers
through Owner/Admin invitations.
- Updated the invitation role copy and the product specifications.
- Allowed active Operators (including legacy Member roles) through the
invite shortcut and advanced tool/profile UI gates. Kept inactive
memberships, Viewers, and other companies excluded; API grants remain
authoritative.
- Added tests for the default Operator grants, both excluded actions,
agent editing, company boundaries, and invitation role authorization.
- Simplified adapter-route test imports so authorization errors and the
error handler use one module graph; retained the upstream
mock-initialization fix.
- Kept Owner/Admin/Viewer presets and explicit invitation grants
unchanged. Added no EE code or database migration.

## Verification

- Permission and invitation tests: 38 passed across access service,
invitation defaults, and invitation creation routes.
- UI access tests: 61 passed across invitation shortcuts, Operator
tool/profile access, and invitation UI.
- Adapter route tests: 15 passed after resolving the upstream test setup
conflict.
- `pnpm -r typecheck`, `pnpm build`, and `pnpm check:token-gates`:
passed on the final branch.
- The initial local `pnpm test:run` general-server batch reported three
tool-access failures and eight upload-helper timeouts. Both complete
affected suites passed in isolation on the final branch: 403 tests. The
initial full local run was not green.
- Both complete tool-access and adapter route suites also passed on an
untouched snapshot of base `88ff98b83`: 387 tests. The earlier tool
failures did not reproduce there.
- [All CI gates
passed](https://github.com/paperclipai/paperclip/actions/runs/37527422346)
for `068be153b9c2a064f2aaccc0627fdf458e73747e`, including the previously
failing serialized adapter lane, general tests, E2E, typecheck, build,
Runner checks, and canary dry run.
- Greptile scored that exact commit 5/5. Superagent passed. Both
invitation review threads are resolved.

## Risks

- Operators gain broad company editing and invitation access by design.
- Existing default seeding can add the new missing grants during startup
or Cloud sign-in. It does not replace existing scopes.
- Company boundaries and the two excluded permission checks still apply.
- An Admin without member-permission management can no longer invite an
Owner. Operators retain invitation access for Operator/Viewer roles and
agent-only invites.
- This PR changes company permissions. Cloud workspace invitation rules
remain separate.

## Model Used

OpenAI Codex, GPT-6. The exact deployment model ID and context window
were not exposed in this session. Used repository inspection, code
editing, shell execution, and test tools.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting a merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 15:55:38 -05:00
dependabot[bot] da679f113a build(deps): bump actions/upload-pages-artifact from 4.0.0 to 5.0.0 (#12964)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The repository runs a weekly full-stack end-to-end campaign for the
runner, and that workflow publishes its merged dashboard to GitHub Pages
> - The publish step pins `actions/upload-pages-artifact` to a version 4
commit
> - Version 4 of that action embeds `actions/upload-artifact` at
`v4.6.2`, so the publish step keeps an old upload path
> - An old pinned action falls behind upstream fixes, and the gap grows
with each upstream release
> - This pull request moves the pin to the commit that upstream tags as
`v5.0.0`
> - The benefit is that the Pages publish step uses the current upload
path, and the pin stays a commit SHA

## Linked Issues or Issue Description

No public issue exists for this change. Refs #12963 and Refs #12962 —
two other pinned GitHub Action updates for the same repository. They are
related, and they are not duplicates. The description below follows the
enhancement template.

**What existing behavior does this improve?**

The weekly runner end-to-end workflow publishes its merged dashboard to
GitHub Pages. This change updates the action that packages that
dashboard.

**Subsystem affected**

CI and release automation —
`.github/workflows/runner-full-stack-e2e.yml`.

**Current behavior**

The publish step uses `actions/upload-pages-artifact` at the commit that
upstream tags as version 4. That version embeds
`actions/upload-artifact` at `v4.6.2`.

**Proposed behavior**

The publish step uses `actions/upload-pages-artifact` at commit
`fc324d3547104276b827a68afc52ff2a11cc49c9`, which upstream tags as
`v5.0.0`. That version embeds `actions/upload-artifact` at `v7.0.0`.

**Breaking changes**

None for this repository. Version 5 keeps the `name`, `path` and
`retention-days` inputs with the same defaults. Version 5 adds one
optional input, `include-hidden-files`, which defaults to `false`. The
step in this workflow passes `name` and `path` only. With the default
value of the new input, version 5 excludes hidden files, and that
matches version 4.

## What Changed

- Change the pinned commit of `actions/upload-pages-artifact` in
`.github/workflows/runner-full-stack-e2e.yml` from the version 4 commit
to `fc324d3547104276b827a68afc52ff2a11cc49c9`, which upstream tags as
`v5.0.0`.

## Verification

- The pinned commit matches the upstream tag. `gh api
repos/actions/upload-pages-artifact/git/ref/tags/v5.0.0 --jq
.object.sha` returns `fc324d3547104276b827a68afc52ff2a11cc49c9`.
- The input contract stays compatible. A comparison of `action.yml` at
the old pin and at the new pin shows the same `name`, `path` and
`retention-days` inputs with the same defaults, plus one new optional
input.
- The repository CI suite passes on this branch after a rebase onto the
current base branch.
- One limit applies. The changed step runs only in the `Runner
Full-Stack E2E` workflow. That workflow starts on a weekly schedule and
on a manual dispatch, so no pull-request run exercises the step. A
maintainer can exercise it with a manual dispatch of that workflow, or
the next scheduled run exercises it.

## Risks

- Low risk, with one limit. The changed step does not run on a pull
request, so the pull-request checks do not prove the new action version
in this workflow.
- Version 5 of the outer action embeds a newer `actions/upload-artifact`
version. A behaviour change in that inner action appears first in the
weekly campaign, and not in the pull-request checks.
- The rollback is one commit. Restore the previous pinned commit of
`actions/upload-pages-artifact`.

## Model Used

Dependabot generated this dependency update automatically, so no AI
model produced the code change. A maintainer wrote this description with
Claude Opus 5 (Anthropic, model id `claude-opus-5`, extended thinking,
tool use).

## Checklist

Three boxes stay unticked on purpose. This change pins one GitHub Action
version in one workflow file. No local test covers a pinned action
version, no new test applies, and no document refers to this pin.

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [ ] I have added or updated tests where applicable
- [ ] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-10-06 13:53:05 -07:00
DottaandPaperclip 2ca0d26a99 fix(connections): recover missing personal AI credentials in chat (#15376)
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-06 15:36:01 -05:00
ea8e686197 fix(tools): distinguish requested from granted OAuth scopes (#14059)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Apps use OAuth connection grants to let agents reach external
services.
> - Paperclip stored a requested scope as a granted scope when a
provider omitted `scope` from its token response.
> - A live MCP connection showed why that matters: the provider reported
scopes beyond the requested scope. Scope names alone do not establish
write capability.
> - This PR records whether a scope came from the provider or from
Paperclip's request, and flags asserted extra scopes.
> - The result is an honest scope record across authorization and
refresh, including for self-hosted and Cloud instances.

**Review order:** shared scope resolver and OAuth write paths in
`server/src/services/tool-access.ts` → optional schema/types/validator
fields → focused tests in
`server/src/__tests__/tool-access-service.test.ts`. This PR changes the
shared OAuth record. The Enterpret catalog definition is separate in
**#13906**.

## Linked Issues or Issue Description

No public issue covers this bug. Related: **#13906** adds the official
read-only Enterpret connector with OAuth and organization tokens.

**What happened?**

The OAuth callback stored `normalizeOauthScopes(token.scope ??
requestedScopes)`. When the provider omitted `scope`, Paperclip recorded
the request as though it were a verified grant. In the live case:

```text
Paperclip requested     mcp:read
Provider token reply    no scope field
Paperclip recorded      ["mcp:read"]
Token introspection    openid email profile mcp:read mcp:write
```

The access token is opaque. A negative introspection control returned
`active: false` with no scope, confirming the wider scope belonged to
the live token. The grants API could therefore present requested scopes
as though the provider had asserted them. This observation did not
establish access to write tools.

**Expected behavior**

Store the provider's asserted scope when present. When absent, identify
the value as an inference from the request. Preserve that provenance
through refresh and report extra scopes only when the provider actually
asserts them.

**Steps to reproduce**

1. Use an OAuth provider that omits `scope` from its token response.
2. Complete consent after requesting `mcp:read`.
3. Read the connection or grant: before this fix, the stored `scopes`
looked like an asserted read-only grant.
4. The focused fixture tests reproduce the callback and refresh record
without needing a live provider.

**Deployment mode**

The shared OAuth path affects self-hosted and Cloud. The live provider
evidence came from an isolated self-hosted runtime.

## What Changed

- Added `resolveGrantedOauthScopes`. It records `scopeSource: provider`
when the token response asserts scope, or `requested_fallback` when it
does not. `unrequestedScopes` contains only provider-asserted scopes
outside the request.
- Applied the resolver at initial authorization and refresh. A refresh
without `scope` retains a previous provider assertion and warning; a
fresh assertion can replace them.
- Stored the actual authorization request per grant as
`providerTenant.oauth.requestedScopes`. This keeps a multi-user
connection's refresh baseline tied to the right grant. Generic MCP OAuth
also records discovered scopes when it sends them; curated apps retain
their reviewed request.
- Carried provenance onto the connection and default organization grant.
The organization grant does not receive token expiry, so its existing
refresh/reconnect behavior remains intact.
- Added optional fields to database JSONB schema types, shared types,
and validation. Existing grants need no migration or backfill.

This PR **does not change authorization decisions**, reject tokens,
introspect providers, or show a new UI warning. If a provider hides an
over-grant by omitting `scope`, Paperclip still cannot discover it.
`unrequestedScopes: []` with `requested_fallback` means **unknown**, not
least privilege.

## Verification

**Current head:** `7fde922a8`. Reconciled with master `88ff98b83`; the
final diff is six OAuth implementation, contract, and test files.
Preserved master's generic `offline_access` consent and legacy callback
baseline, and verified requested versus provider-asserted scopes on each
grant. Removed four duplicated GitHub token-method fixture properties
after fresh review.

All 504 focused tests in seven suites pass on the final head. Local
workspace build, recursive typecheck, build-gap typecheck, module
boundaries, node-version, token gates, and runtime push-policy checks
pass. The full local Vitest run completed with 15,610 passing tests, 91
skipped, and six failures in heartbeat/workspace suites; all six
failures reproduced on unchanged master `88ff98b83` in an isolated
baseline worktree. These suites and their runtime code are unchanged by
this PR. Greptile is 5/5 on this exact head with no actionable findings.
All current-head GitHub checks are green: 53 successful check runs, two
intentional Storybook skips, and the successful Snyk status. The
initially failed adapter-access and signoff-heartbeat shards both passed
their single rerun. [CI
run](https://github.com/paperclipai/paperclip/actions/runs/37521588330).

The regression coverage includes omitted and explicit scopes, asserted
extra scopes, refresh preservation, per-user request baselines, generic
MCP discovery, and organization grants. No live provider token is needed
to reproduce these recordkeeping cases.

## Risks

- Existing records keep their scope list but have no source marker.
Historical provenance cannot be recovered from them.
- A refresh that omits `scope` preserves the last known assertion. A
provider that silently narrows a grant will not update the record until
it asserts a new scope.
- A provider that omits `scope` leaves the actual scope unasserted. This
PR cannot discover scopes omitted from the response or establish
endpoint capabilities. Enterpret’s provider fixes are separate
follow-ups.
- Scope provenance is additive metadata. No policy path uses these new
fields to allow or deny an action.

## Model Used

Implementation and tests: Claude Opus 5 (`claude-opus-5`) through Claude
Code with extended thinking, tools, and code execution. PR text cleanup:
Codex GPT-6 (exact host model ID and context window were not exposed;
tool use).

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] Focused tests pass locally; full-suite baseline failures are
documented above
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All current-head CI gates are green
- [x] Greptile is 5/5 on the current head with no actionable follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

This shared scope-recording fix can be reviewed and merged independently
of the Enterpret connector.

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 13:25:20 -07:00
DottaandPaperclip e38d6d16b6 feat(connections): add advanced provider setup and live browser qualification (#15341)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Users connect accounts and choose an agent harness and model.
> - The runtime change in #14970 supports custom providers on those
connections.
> - Normal setup must stay simple while advanced users can choose a
compatible gateway.
> - Shared connector rows and access controls keep these choices
consistent.
> - This pull request refines the agent setup UI and adds review stories
and repeatable browser qualification.
> - The qualification checks real tools and downloaded outputs, not only
a successful run status.

## Linked Issues or Issue Description

Refs #14970, #37, #13083, #14104, #14565, #12692.

The core implementation in #14970 is merged. This branch incorporates
its squash commit and targets `master`. Both PRs contain our
implementation. #14016 is a reference only and is not a dependency. This
PR has 96 changed files.

## What Changed

- Complete model-provider connector presentation beside other
connectors. Each row uses the existing Connect action and connection
list. Tags are stored without category UI. The base PR includes the
provider forms and routes.
- Show persistent Subscription, API Key, and Advanced choices. Label
Advanced as Custom Gateway. Reuse provider logos, connection lists, and
permissions controls. Default access to the organization and all agents
when permitted; keep narrowing controls under Advanced.
- Keep Configure reachable before subscription sign-in, so users can
select a supported environment when the default cannot sign in. Testing
and saving still require a connection. Show the execution environment in
Configure. Preserve the confirmed Connect choice. Editing a method,
credential, saved account, or advanced choice requires that current
choice to connect before testing or saving. Use matching model and
thinking-effort dropdowns and retain connection icons in selected
values.
- Preserve the new harness model default when switching an existing
OpenCode agent to Codex or Claude, and resolve user-selected model names
with the effective harness.
- Load popular OpenRouter models through the shared connection-model
discovery path. Keep explicit model lists and manual model entry
available.
- Group onboarding, connection setup, agent runtime, management,
recovery, and production-component stories under AI Connections /
Provider routing.
- Add an explicit-only provider-connections browser suite for managed
local or existing local/staging targets. Use private browser profiles
and credential handoffs. Support human-assisted subscription sign-in
without sharing passwords or tokens in reports.
- Verify persisted connection identity, runtime probes, tool execution,
exact artifact bytes, completion, and context-dependent follow-up.
Retain source/model provenance, cost bounds, closed error diagnostics,
original failures, and cleanup evidence.
- Add Gemini startup-model and skill-root fixes, Grok private-history
detection, ACP filesystem regression fixtures, selected-workspace
handling for local Hermes, and artifact-helper workspace fallback.
- Keep managed Grok runtime homes disposable. Remove host-side
transcript retention/restoration because private file modes do not
isolate same-user agent processes. Ignore earlier development archives
and use a fresh task handoff when history is unavailable. Verify the
absence of restored transcripts with a separate same-user process.
- Capture stopped-run diagnostics before deleting an attached-company
fixture agent. Track creation and owned sign-in receipts; revoke only
this attempt's accounts and never adopt a concurrent campaign's newly
created account. Preserve failure signals and final status through
cleanup.
- Require the requested environment in the saved agent and every run,
including follow-ups. Reject a forced incompatible target. Keep one
cancellation state through startup, every cell, reporting, and teardown
for SIGINT, SIGTERM, and SIGHUP. Stop further paid cells after
interruption. Document qualification limits.

## Verification

- Current head `b3bb3e94d577d43d9965a6b9daba039f599b2e49` includes
master `d9f600043`. The security fix in `a758fde31` passes full
workspace typecheck, production build, and 119 connection/Grok
regressions. The unchanged UI passes all 126
configuration/model-discovery tests and token gates. The final
published-guide correction passes Grok adapter typecheck. Earlier head
`eebd8225c` passed the complete deterministic runner suite (1,404 Vitest
tests and 128 Node tests) and all CI jobs. Current-head CI run
`37520147514` passed all 47 jobs, including the full sharded Vitest and
browser matrix, production build, and canary dry run. All 55 checks
completed: 53 successes and two expected skips. The current-head
security scan passed, Greptile is 5/5, and no review threads remain
open.
- A separate same-user process reproduced reading a restored Grok
transcript before the security fix. The regression now finds no
transcript. Existing fresh-session fallback and ordinary session
metadata behavior pass.
- The final account-choice and cleanup fixes pass 85 setup tests and 26
qualification-harness tests. Regressions verify that editing a
connection invalidates confirmation, Configure remains reachable before
sign-in, diagnostics are captured before fixture deletion, and
concurrent campaigns cannot adopt or revoke each other's accounts. UI
and E2E typechecks pass.
- The Storybook build and actual Chromium production-component stories
passed during this change. Review the neighboring AI Connections /
Provider routing stories, regular connector rows, three connection
modes, model discovery, and the single execution-environment control in
Configure.
- Cancellation smoke verified authenticated cleanup before browser close
for SIGINT, SIGTERM, and SIGHUP. Regressions cover interruption during
startup and reporting, missing-file ACP resource errors, and preserved
permission denials. Both ACP runtime versions and 54 ACPX/Grok
regressions passed. The deterministic connection-intent browser suite
passed two tests.
- Historical local qualification retained 43 passing API/gateway cells
out of 46, with downloaded outputs and follow-up receipts. These
attempts span earlier builds; they do not qualify this exact commit or
staging. Subscription combinations, Gemini overloads, and the unresolved
follow-up failure remain recorded rather than counted as passing.
- Use `pnpm test:e2e:runner -- --list --suite provider-connections` to
inspect the matrix. Follow `tests/runner-e2e/PROVIDER-CONNECTIONS.md`
for credentials, target URL, sign-in assistance, budget, evidence, and
cleanup. Paid live tests remain opt-in.

## Risks

- The core implementation in #14970 is merged. This PR adds no database
migration of its own.
- Subscription login needs an interactive provider session. Dedicated
accounts and staging qualification remain follow-up work; this PR does
not certify every login combination for production.
- Managed Grok transcript resume is deferred until provider history has
an OS isolation or authorized broker solution. Follow-ups start fresh
with Paperclip task context; earlier live Grok results do not qualify
this behavior.
- Gemini CLI 0.58.0 has an upstream ACP new-file error conversion
defect. Live overloads and one unresolved follow-up timeout remain
recorded. The stock CLI is unchanged, and those cases are not marked as
passing.
- Real-provider tests spend credits and use private credential/evidence
directories. The launcher requires explicit selection and checks target
ownership. It must not attach to a developer's database by accident.
- OpenClaw Gateway, Hermes Gateway, Claude Managed, AWS AgentCore,
Process, HTTP, and legacy ACPX local remain outside custom provider
setup.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, code
execution, and browser testing. The exact deployment model ID and
context window size were not exposed in this session.

## Checklist


- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 15:21:22 -05:00
DottaandPaperclip a7a2ed63a0 fix(ui): restore wheel and touch scrolling in task selectors (#15374)
Restore wheel and touch scrolling in nested task selectors and keep mobile picker controls reachable above the keyboard.

Stabilize the affected server route test harnesses for Vitest 5.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
canary/v2026.1006.0-canary.22
2026-10-06 15:18:07 -05:00
88ff98b83d feat(apps): add read-only Enterpret MCP with OAuth and org tokens (#13906)
## Thinking Path

> - Paperclip agents need governed access to Enterpret's hosted MCP
tools for customer feedback.
> - The official Enterpret MCP is read-only. Enterpret Agent’s beta
write MCP is a separate service and is outside this PR.
> - This PR supports both organization auth tokens and personal browser
OAuth for the official MCP.
> - Enterpret committed to fixing OAuth scope behavior and reducing its
token-revocation cache window. The provider fixes will follow after
connector deployment and do not gate this release. Paperclip functional
validation remains required. New tools stay quarantined on refresh.
> - This PR combines the original connector and Vivek's token-path work
on current master, with live self-hosted QA recorded in the validation
guide.

## Linked Issues or Issue Description

No public issue covers this connector. This PR incorporates the work
from #14737. The independent OAuth scope-provenance fix is #14059; it is
not required for the organization-token path. #13890 is a merged
connector precedent. #13692 provides the connector-authoring skill.

**Problem or motivation**

The generic MCP setup can reach Enterpret, but it has no reviewed
Enterpret entry, method-specific setup, artwork, or policy defaults. The
original connector branch had also fallen behind master.

**Proposed solution**

Publish the Enterpret catalog entry with an organization auth token as
the selectable method. Support browser OAuth through instance-local DCR
on the same read-only endpoint. Default `run_graph_query` to Ask first
and quarantine tools newly found on refresh.

## What Changed

- Added the Enterpret app definition, generated registry, official icon,
Browse card, setup copy, tests, and Storybook states.
- Combined #14737 and resolved the current-master conflicts, including
the connector permission audit.
- Preserved catalog quarantine on token reconnect and updated expiry
guidance to follow the token's dashboard expiry. The QA token displayed
three months, contradicting the former six-month copy.
- Enabled personal OAuth for the official read-only MCP. Removed the
write-capability inference and OAuth draft-only setup restriction. The
beta Agent endpoint is excluded.
- Preserved tool quarantine during token reconnect, with a regression
test for existing quarantined and newly discovered tools.
- Added OAuth scope and revocation-cache disclosures and an interactive
Storybook retry check.
- Merged master `d9f600043` into the existing PR branch without
rewriting history. Resolved four conflicts, regenerated the combined
registry, and retained the model-provider assertions with a catalog
count of 66.
- Recorded live QA and its limits in
[doc/connections/ENTERPRET.md](https://github.com/paperclipai/paperclip/blob/950bacc52/doc/connections/ENTERPRET.md).

## Verification

Current head: `6c73d3084202ccbb25080e7a4e9c3467bc1341dd`. All 513
focused connector, discovery, OAuth, policy, shared-definition, and
UI-handoff tests passed across 11 suites. Full repository typecheck,
build, Storybook build, token gates, module boundaries, Node policy,
no-git-push policy, and typecheck-build-gaps passed. Release-registry
tests passed 129/129 with modern Bash. The four failures with macOS Bash
3.2 also reproduced on untouched master.

The full test suite passed in CI on this exact head, including every
general-test and serialized-server shard and all eight E2E shards. The
duplicate local `pnpm test:run` was stopped after these remote test
lanes passed; it did not complete locally. The combined registry was
regenerated. Regeneration also reproduces existing Linear and AgentMail
definition drift on untouched master; those unrelated changes were
excluded.

On an isolated self-hosted Paperclip instance, an organization token
connected and discovered eight actions. `get_organization_details`
returned the intended organization before and after reconnect. Catalog
refresh, invalid-token failure and valid-token recovery, activity
logging, local disable, and local removal were exercised. No feedback
quotes or graph query were requested. Paperclip revoked both QA
connection secrets and archived the QA connections.

The token-path agent access preview correctly showed an action Off, but
the board Test endpoint still invoked it as the board user. This shared
Test-path mismatch is documented; it is not evidence that a real agent
can bypass policy. An actual token-path agent-session denial, a real
agent process, expiry recovery, server/VPS, and Cloud remain unverified.

Enterpret's dashboard confirmed the QA token was revoked, but the same
token immediately reconnected and completed a safe read. Vivek explained
this as a 24-hour validity cache and committed to reducing the window.
The new maximum delay and deployment remain unverified. The earlier
OAuth check reported `mcp:write` and `email` for an `mcp:read` request.
This is a scope mismatch; it did not prove access to write tools. On
2026-10-01 the account holder relayed Vivek’s clarification that the
official MCP is read-only and the beta Agent write MCP is separate.

## Risks

The organization token can read the organization's customer feedback,
and provider-side revocation did not immediately stop access. Operators
should check each token's displayed expiry and remove Paperclip's local
connection when retiring access. The account holder confirmed the
provider fixes are fast follows after deployment. Release can proceed
with the current scope reporting and up-to-24-hour revocation delay
documented, once fresh OAuth authorization/refresh, real-agent
allow/deny, CI, and review pass. Retest the reduced revocation window
after Enterpret deploys it. Scope labels alone must not be used to claim
write access.

Greptile is 5/5 on head `6c73d3084` with no actionable findings. All
current-head CI checks pass, including the canary dry run. Live
agent-session denial and fresh OAuth refresh remain explicit deployment
QA items. The board Test panel does not prove agent authorization.
Provider scope-reporting and revocation-cache fixes are accepted fast
follows after deployment.

## Model Used

Original connector and validation record: Claude Opus 5 through Claude
Code. Integration, current-master reconciliation, and token-path QA:
OpenAI Codex, GPT-6 family (exact host variant not exposed), using
repository tools and an isolated runtime. Contributor commits retain
their authors.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either linked existing issues or described the issue in-PR
- [x] I have not referenced internal or instance-local Paperclip issues
or links
- [x] My branch name describes the change and contains no internal issue
ID
- [x] I have run relevant local tests and checks
- [x] I have updated relevant documentation
- [x] I have considered and documented the risks above
- [x] All current-head CI gates are green
- [x] Greptile is 5/5 with no actionable follow-ups

If squash-merging, preserve contributor credit in the squash message:

Co-Authored-By: Vivek Kaushal <kaushalvivek@users.noreply.github.com>

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Vivek Kaushal <vivek@enterpret.com>
2026-10-06 14:35:51 -05:00
Nicky LeachandPaperclip 1cebdd4c4f Add plugin lifecycle delivery and an initial resource baseline (#15343)
## Thinking Path

> - Paperclip manages agents and projects for work.
> - Their lifecycle changes already commit to a durable journal.
> - Plugins need to read these records and track completed work.
> - Existing resources need a one-time baseline before delivery is
enabled.
> - Worker crashes must leave unfinished records available for retry.
> - Each plugin needs its own progress and company access checks.
> - This PR adds a pull inbox through the existing plugin SDK and job
system.

## Linked Issues or Issue Description

**Problem or motivation**

Plugins cannot consume the durable resource lifecycle journal.
In-process notifications can disappear during a restart and cannot
record successful completion. Resource plugins need retries, company
boundaries, and ordered transitions.

**Proposed solution**

Add `ctx.events.listLifecycle(companyId, limit?, afterId?)` and
`ctx.events.acknowledgeLifecycle(companyId, eventId)` under
`events.subscribe`. Seed a one-time baseline from current resource
state. Deliver creation first, then the remaining transitions in ID
order. Store acknowledgments for each plugin. Use existing plugin jobs
to poll configured companies.

**Alternatives considered**

A global sequence cursor can skip lower IDs that commit later.
Fire-and-forget subscriptions cannot record completion. A new dispatcher
is unnecessary because plugin jobs support polling. Consumers serialize
polling and use provider idempotency keys.

**Roadmap alignment**

This extends the existing plugin system and builds on #15280 and #15306.
Searches found no duplicate resource inbox work. Related #13306 exposes
decision events on the in-process bus. This PR includes the initial
journal baseline. Provider provisioning remains separate work.

## What Changed

- Add company-scoped lifecycle reads and acknowledgments to the SDK and
worker RPC host.
- Gate both methods by capability, invocation or proactive company
scope, plugin readiness, and company enablement.
- Seed hired agents and all projects once. Preserve paused/terminated
agent state and archived project state. Keep pending hires behind
approval.
- Preserve existing history and deliver backfilled creation before
partial transition histories.
- Deliver project archive events from #15371 through the SDK. Seed
archive intents for archived projects and update intents for active
projects with a partial archive history.
- Store acknowledgments per plugin and reject acknowledgments that skip
earlier resource events.
- Page past failed resources while retaining their pending records.
Reset the page cursor each sweep to include late commits.
- Add acknowledgment storage, an index for resource ordering, tests, and
authoring guidance.
- Preserve native identity definitions and sequence progress in
JavaScript backups. Repair journal ID generators lost by older backups
before seeding the baseline, without changing existing IDs.

## Verification

- `pnpm -r typecheck` passed after the final origin/master rebase.
- `pnpm build` passed before the final metadata rebase. The final CI
build also passed.
- All 46 focused lifecycle, SDK harness/RPC, migration snapshot, and
legacy restore checks passed again with migration 0309.
- Checks cover retries, per-plugin progress, resource order,
capability/company boundaries, baseline idempotency, approval gates,
archive delivery, restored projects with partial histories, and atomic
migration rollback.
- Full CI passed on final commit `0e674b87fc`. One unrelated Cursor
fixture hit a 10-second timeout; it passed locally in under one second,
and the failed server shard passed on retry.
- Greptile scored the final commit 5/5 with no unresolved review
threads.
- The branch is rebased onto origin/master and is conflict-free.
- Earlier full local runs were stopped as scope changed. CI ran the
complete repository suite.
- `git diff --check` and a local secret/PII scan passed.

## Risks


- Apply migration `0309_loving_the_hood.sql` before starting the new
server. It creates acknowledgment storage and an index, then seeds the
baseline in the same transaction. Agent/project writes wait for the
migration to commit. Keep these writes quiesced through the migration
and activation of the new capture-capable server; do not resume an older
runtime that lacks capture after the baseline.
- The baseline records current desired state, not historical
transitions. It includes archived projects and terminated-agent cleanup
intents. It runs once with the delivery migration; no later or runtime
journal backfill is planned.
- Delivery is at least once. Concurrent reads can repeat an event.
Consumers must serialize polling, use stable company/event idempotency
keys, and acknowledge successful operations only.
- Reset `afterId` at the start of every polling sweep. It is a page
cursor, not a persisted high-water mark.
- Deleting a plugin removes its acknowledgment records. A new
installation may replay existing journal events.
- Records contain identity and action. Consumers must load current
authorized data before acting. Cleanup and retention policy belong to
the provider plugin.
- Provider calls, VM/volume provisioning, and journal retention are
outside this PR.

## Model Used

OpenAI Codex, based on GPT-6, with reasoning, repository inspection,
code execution, and tool use. The exact deployment model ID and context
window are not exposed in this session.

## Checklist


- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 12:27:09 -07:00
854af7df19 build(deps-dev): bump vitest from 4.1.11 to 5.0.3 (#12969)
Bumps
[vitest](https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest)
from 4.1.11 to 5.0.3.
<details>
<summary>Release notes</summary>
<p><em>Sourced from <a
href="https://github.com/vitest-dev/vitest/releases">vitest's
releases</a>.</em></p>
<blockquote>
<h2>v5.0.3</h2>
<h3>   🐞 Bug Fixes</h3>
<ul>
<li>Isolate <code>result.status</code> between <code>repeats</code> runs
 -  by <a
href="https://github.com/hi-ogawa"><code>@​hi-ogawa</code></a>,
<strong>Hiroshi Ogawa</strong> and <strong>Codex (GPT-6)</strong> in <a
href="https://redirect.github.com/vitest-dev/vitest/issues/11218">vitest-dev/vitest#11218</a>
<a href="https://github.com/vitest-dev/vitest/commit/5dbebe9e3"><!-- raw
HTML omitted -->(5dbeb)<!-- raw HTML omitted --></a></li>
<li>Don't print an interceptor warning in browser mode  -  by <a
href="https://github.com/sheremet-va"><code>@​sheremet-va</code></a> in
<a
href="https://redirect.github.com/vitest-dev/vitest/issues/11377">vitest-dev/vitest#11377</a>
<a href="https://github.com/vitest-dev/vitest/commit/15cc006aa"><!-- raw
HTML omitted -->(15cc0)<!-- raw HTML omitted --></a></li>
<li>Don't retry when <code>test.fails</code> expectedly failed  -  by <a
href="https://github.com/hi-ogawa"><code>@​hi-ogawa</code></a>,
<strong>Hiroshi Ogawa</strong> and <strong>Codex (GPT-6)</strong> in <a
href="https://redirect.github.com/vitest-dev/vitest/issues/11219">vitest-dev/vitest#11219</a>
<a href="https://github.com/vitest-dev/vitest/commit/b24585f08"><!-- raw
HTML omitted -->(b2458)<!-- raw HTML omitted --></a></li>
<li>Scope cache key generators to projects  -  by <a
href="https://github.com/ecoyoung"><code>@​ecoyoung</code></a> in <a
href="https://redirect.github.com/vitest-dev/vitest/issues/11281">vitest-dev/vitest#11281</a>
and <a
href="https://redirect.github.com/vitest-dev/vitest/issues/11301">vitest-dev/vitest#11301</a>
<a href="https://github.com/vitest-dev/vitest/commit/92ba7fc1d"><!-- raw
HTML omitted -->(92ba7)<!-- raw HTML omitted --></a></li>
<li><strong>browser</strong>:
<ul>
<li>Delay server <code>listen</code> until tests start running  -  by <a
href="https://github.com/sheremet-va"><code>@​sheremet-va</code></a> in
<a
href="https://redirect.github.com/vitest-dev/vitest/issues/11366">vitest-dev/vitest#11366</a>
<a href="https://github.com/vitest-dev/vitest/commit/7d8ed3e9b"><!-- raw
HTML omitted -->(7d8ed)<!-- raw HTML omitted --></a></li>
<li>Check mock path boundaries  -  by <a
href="https://github.com/saryn17"><code>@​saryn17</code></a>,
<strong>Ryosei Sato</strong> and <a
href="https://github.com/sheremet-va"><code>@​sheremet-va</code></a> in
<a
href="https://redirect.github.com/vitest-dev/vitest/issues/11361">vitest-dev/vitest#11361</a>
and <a
href="https://redirect.github.com/vitest-dev/vitest/issues/11362">vitest-dev/vitest#11362</a>
<a href="https://github.com/vitest-dev/vitest/commit/1c3888bce"><!-- raw
HTML omitted -->(1c388)<!-- raw HTML omitted --></a></li>
<li>Keep config of browser-consumed environments  -  by <a
href="https://github.com/kasperpeulen"><code>@​kasperpeulen</code></a>
and <strong>Claude</strong> in <a
href="https://redirect.github.com/vitest-dev/vitest/issues/11378">vitest-dev/vitest#11378</a>
<a href="https://github.com/vitest-dev/vitest/commit/aafc0996f"><!-- raw
HTML omitted -->(aafc0)<!-- raw HTML omitted --></a></li>
<li>Ignore page crash while cancelling  -  by <a
href="https://github.com/sheremet-va"><code>@​sheremet-va</code></a> in
<a
href="https://redirect.github.com/vitest-dev/vitest/issues/11386">vitest-dev/vitest#11386</a>
<a href="https://github.com/vitest-dev/vitest/commit/7c36748fa"><!-- raw
HTML omitted -->(7c367)<!-- raw HTML omitted --></a></li>
<li><code>toMatchScreenshot</code> uses wrong reference on retried tests
 -  by <a href="https://github.com/macarie"><code>@​macarie</code></a>
in <a
href="https://redirect.github.com/vitest-dev/vitest/issues/11393">vitest-dev/vitest#11393</a>
<a href="https://github.com/vitest-dev/vitest/commit/c22aba992"><!-- raw
HTML omitted -->(c22ab)<!-- raw HTML omitted --></a></li>
</ul>
</li>
<li><strong>cache</strong>:
<ul>
<li>Revalidate imports of cached modules  -  by <a
href="https://github.com/sheremet-va"><code>@​sheremet-va</code></a> in
<a
href="https://redirect.github.com/vitest-dev/vitest/issues/11381">vitest-dev/vitest#11381</a>
<a href="https://github.com/vitest-dev/vitest/commit/38f98855f"><!-- raw
HTML omitted -->(38f98)<!-- raw HTML omitted --></a></li>
</ul>
</li>
<li><strong>deps</strong>:
<ul>
<li>Pin <code>why-is-node-running</code> to <code>3.2.1</code> to avoid
users running into <code>ERR_PNPM_TRUST_DOWNGRADE</code>  -  by <a
href="https://github.com/AriPerkkio"><code>@​AriPerkkio</code></a> in <a
href="https://redirect.github.com/vitest-dev/vitest/issues/11403">vitest-dev/vitest#11403</a>
<a href="https://github.com/vitest-dev/vitest/commit/f6c9a4977"><!-- raw
HTML omitted -->(f6c9a)<!-- raw HTML omitted --></a></li>
</ul>
</li>
<li><strong>expect</strong>:
<ul>
<li>Pass current equality testers to <code>expect.extend</code>
asymmetric matchers  -  by <a
href="https://github.com/hi-ogawa"><code>@​hi-ogawa</code></a>,
<strong>Hiroshi Ogawa</strong> and <strong>Claude</strong> in <a
href="https://redirect.github.com/vitest-dev/vitest/issues/11401">vitest-dev/vitest#11401</a>
<a href="https://github.com/vitest-dev/vitest/commit/3e794a96b"><!-- raw
HTML omitted -->(3e794)<!-- raw HTML omitted --></a></li>
</ul>
</li>
<li><strong>jsdom</strong>:
<ul>
<li>Support Blob on jsdom 30.1  -  by <a
href="https://github.com/sheremet-va"><code>@​sheremet-va</code></a> in
<a
href="https://redirect.github.com/vitest-dev/vitest/issues/11379">vitest-dev/vitest#11379</a>
<a href="https://github.com/vitest-dev/vitest/commit/6c49b7197"><!-- raw
HTML omitted -->(6c49b)<!-- raw HTML omitted --></a></li>
</ul>
</li>
<li><strong>pool</strong>:
<ul>
<li>Preserve unique pool ids when <code>groupOrder</code> is set  -  by
<a href="https://github.com/mtorp"><code>@​mtorp</code></a> in <a
href="https://redirect.github.com/vitest-dev/vitest/issues/11392">vitest-dev/vitest#11392</a>
<a href="https://github.com/vitest-dev/vitest/commit/50312ebb4"><!-- raw
HTML omitted -->(50312)<!-- raw HTML omitted --></a></li>
</ul>
</li>
<li><strong>ui</strong>:
<ul>
<li>Split-pane handle overlapping iframe  -  by <a
href="https://github.com/macarie"><code>@​macarie</code></a> in <a
href="https://redirect.github.com/vitest-dev/vitest/issues/11221">vitest-dev/vitest#11221</a>
<a href="https://github.com/vitest-dev/vitest/commit/f91db0dfd"><!-- raw
HTML omitted -->(f91db)<!-- raw HTML omitted --></a></li>
</ul>
</li>
<li><strong>vitest</strong>:
<ul>
<li>Remove root temp dir on close  -  by <a
href="https://github.com/abhinav-phi"><code>@​abhinav-phi</code></a> in
<a
href="https://redirect.github.com/vitest-dev/vitest/issues/11248">vitest-dev/vitest#11248</a>
<a href="https://github.com/vitest-dev/vitest/commit/7c7119cf7"><!-- raw
HTML omitted -->(7c711)<!-- raw HTML omitted --></a></li>
</ul>
</li>
<li><strong>vm</strong>:
<ul>
<li>Do not optimize deps from index.html  -  by <a
href="https://github.com/ezefernandezyf"><code>@​ezefernandezyf</code></a>,
<strong>Hiroshi Ogawa</strong> and <strong>Codex (GPT-6)</strong> in <a
href="https://redirect.github.com/vitest-dev/vitest/issues/11329">vitest-dev/vitest#11329</a>
and <a
href="https://redirect.github.com/vitest-dev/vitest/issues/11360">vitest-dev/vitest#11360</a>
<a href="https://github.com/vitest-dev/vitest/commit/caf2887de"><!-- raw
HTML omitted -->(caf28)<!-- raw HTML omitted --></a></li>
<li>Don't reuse scripts across vite environments  -  by <a
href="https://github.com/MO2k4"><code>@​MO2k4</code></a>, <strong>Martin
Oehlert</strong> and <strong>Claude</strong> in <a
href="https://redirect.github.com/vitest-dev/vitest/issues/11395">vitest-dev/vitest#11395</a>
<a href="https://github.com/vitest-dev/vitest/commit/346d3896b"><!-- raw
HTML omitted -->(346d3)<!-- raw HTML omitted --></a></li>
</ul>
</li>
</ul>
<h5>    <a
href="https://github.com/vitest-dev/vitest/compare/v5.0.2...v5.0.3">View
changes on GitHub</a></h5>
<h2>v5.0.2</h2>
<h3>   🐞 Bug Fixes</h3>
<ul>
<li>Bind <code>process</code> in case global is overwritten  -  by <a
href="https://github.com/AriPerkkio"><code>@​AriPerkkio</code></a> in <a
href="https://redirect.github.com/vitest-dev/vitest/issues/11343">vitest-dev/vitest#11343</a>
<a href="https://github.com/vitest-dev/vitest/commit/0b79231ad"><!-- raw
HTML omitted -->(0b792)<!-- raw HTML omitted --></a></li>
<li><strong>detect-async-leaks</strong>:
<ul>
<li>Ignore <code>process.stdio</code> handles  -  by <a
href="https://github.com/AriPerkkio"><code>@​AriPerkkio</code></a> in <a
href="https://redirect.github.com/vitest-dev/vitest/issues/11333">vitest-dev/vitest#11333</a>
<a href="https://github.com/vitest-dev/vitest/commit/0fd6b9790"><!-- raw
HTML omitted -->(0fd6b)<!-- raw HTML omitted --></a></li>
</ul>
</li>
<li><strong>expect</strong>:
<ul>
<li>Fix <code>toMatchObject</code> with asymmetric matchers  -  by <a
href="https://github.com/ShreeBohara"><code>@​ShreeBohara</code></a>,
<strong>Claude Opus 5</strong>, <a
href="https://github.com/hi-ogawa"><code>@​hi-ogawa</code></a>,
<strong>Hiroshi Ogawa</strong> and <strong>Codex (GPT-5)</strong> in <a
href="https://redirect.github.com/vitest-dev/vitest/issues/11100">vitest-dev/vitest#11100</a>
<a href="https://github.com/vitest-dev/vitest/commit/42523289e"><!-- raw
HTML omitted -->(42523)<!-- raw HTML omitted --></a></li>
</ul>
</li>
<li><strong>jsdom</strong>:
<ul>
<li>Fix <code>Request</code> with <code>Blob</code> body on jsdom 28+
 -  by <a
href="https://github.com/harshit-d3v"><code>@​harshit-d3v</code></a> in
<a
href="https://redirect.github.com/vitest-dev/vitest/issues/11295">vitest-dev/vitest#11295</a>
<a href="https://github.com/vitest-dev/vitest/commit/d1c3ecc93"><!-- raw
HTML omitted -->(d1c3e)<!-- raw HTML omitted --></a></li>
</ul>
</li>
<li><strong>reporter</strong>:
<ul>
<li><code>agent</code> to respect <code>--silent</code>  -  by <a
href="https://github.com/Raj4478"><code>@​Raj4478</code></a> and <a
href="https://github.com/AriPerkkio"><code>@​AriPerkkio</code></a> in <a
href="https://redirect.github.com/vitest-dev/vitest/issues/11271">vitest-dev/vitest#11271</a>
<a href="https://github.com/vitest-dev/vitest/commit/5b95efb6d"><!-- raw
HTML omitted -->(5b95e)<!-- raw HTML omitted --></a></li>
</ul>
</li>
<li><strong>reporters</strong>:
<ul>
<li>Handle concurrent <code>createReport</code> calls  -  by <a
href="https://github.com/7rulnik"><code>@​7rulnik</code></a> in <a
href="https://redirect.github.com/vitest-dev/vitest/issues/11278">vitest-dev/vitest#11278</a>
<a href="https://github.com/vitest-dev/vitest/commit/e8e556ff7"><!-- raw
HTML omitted -->(e8e55)<!-- raw HTML omitted --></a></li>
<li><code>hanging-process</code> to use ESM entrypoint  -  by <a
href="https://github.com/AriPerkkio"><code>@​AriPerkkio</code></a> in <a
href="https://redirect.github.com/vitest-dev/vitest/issues/11316">vitest-dev/vitest#11316</a>
<a href="https://github.com/vitest-dev/vitest/commit/4e91e5668"><!-- raw
HTML omitted -->(4e91e)<!-- raw HTML omitted --></a></li>
</ul>
</li>
<li><strong>spy</strong>:
<ul>
<li>Fix stack overflow when spying <code>Set.prototype.add</code>  -  by
<a href="https://github.com/fengmk2"><code>@​fengmk2</code></a> in <a
href="https://redirect.github.com/vitest-dev/vitest/issues/11299">vitest-dev/vitest#11299</a>
<a href="https://github.com/vitest-dev/vitest/commit/a0a939653"><!-- raw
HTML omitted -->(a0a93)<!-- raw HTML omitted --></a></li>
</ul>
</li>
</ul>
<!-- raw HTML omitted -->
</blockquote>
<p>... (truncated)</p>
</details>
<details>
<summary>Commits</summary>
<ul>
<li><a
href="https://github.com/vitest-dev/vitest/commit/33cadea62e8763c455c7fca38d9ab1dda87c5f75"><code>33cadea</code></a>
chore: release v5.0.3 (<a
href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11409">#11409</a>)</li>
<li><a
href="https://github.com/vitest-dev/vitest/commit/346d3896b65c3c907174447a035807342799f346"><code>346d389</code></a>
fix(vm): don't reuse scripts across vite environments (<a
href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11395">#11395</a>)</li>
<li><a
href="https://github.com/vitest-dev/vitest/commit/f6c9a4977ad3363f796a737834572e54c6ad5c18"><code>f6c9a49</code></a>
fix(deps): pin <code>why-is-node-running</code> to <code>3.2.1</code> to
avoid users running into `...</li>
<li><a
href="https://github.com/vitest-dev/vitest/commit/062c75d8b63519211d951d8293ea81b5a9e3c124"><code>062c75d</code></a>
chore: fix standalone docs build, update exports maps (<a
href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11394">#11394</a>)</li>
<li><a
href="https://github.com/vitest-dev/vitest/commit/caf2887dee8987a60118d53933f6e9cabd6b3e2a"><code>caf2887</code></a>
fix(vm): do not optimize deps from index.html (fix <a
href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11329">#11329</a>)
(<a
href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11360">#11360</a>)</li>
<li><a
href="https://github.com/vitest-dev/vitest/commit/50312ebb4eca6a98f6d0b2b61d5d9d38cbbabcef"><code>50312eb</code></a>
fix(pool): preserve unique pool ids when <code>groupOrder</code> is set
(<a
href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11392">#11392</a>)</li>
<li><a
href="https://github.com/vitest-dev/vitest/commit/7c36748fad1eae9687312f2f7ceadce6ec88b5df"><code>7c36748</code></a>
fix(browser): ignore page crash while cancelling (<a
href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11386">#11386</a>)</li>
<li><a
href="https://github.com/vitest-dev/vitest/commit/92ba7fc1df16a4fa5bbee3f198c582fbd56689d8"><code>92ba7fc</code></a>
fix: scope cache key generators to projects (fix <a
href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11281">#11281</a>)
(<a
href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11301">#11301</a>)</li>
<li><a
href="https://github.com/vitest-dev/vitest/commit/38f98855fa9cd7fd376afb84094eba0fda256a74"><code>38f9885</code></a>
fix(cache): revalidate imports of cached modules (<a
href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11381">#11381</a>)</li>
<li><a
href="https://github.com/vitest-dev/vitest/commit/b24585f08f2ea267746a2d6ca0e43edcbb29726f"><code>b24585f</code></a>
fix: don't retry when <code>test.fails</code> expectedly failed (<a
href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11219">#11219</a>)</li>
<li>Additional commits viewable in <a
href="https://github.com/vitest-dev/vitest/commits/v5.0.3/packages/vitest">compare
view</a></li>
</ul>
</details>
<br />

---------

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Priya Raman <priya.raman@paperclip.ing>
Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-10-06 12:13:46 -07:00
DottaandPaperclip a590ab769d Give managed agents persistent cryptographic identities (#15352)
Give agents persistent Ed25519 identities encrypted with the existing instance master key. Create keys transactionally for new agents and lazily before supported managed runs, expose public identities in the API and agent UI, and protect private material during runtime delivery and output persistence.

Preserve identities in recovery backups while giving imported and development-cloned agents fresh keys.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-06 13:57:03 -05:00
Nicky LeachandPaperclip d9f600043f Capture project archive and restore lifecycle events (#15371)
## Thinking Path

> - Paperclip manages projects and their repositories.
> - Resource changes commit to a durable lifecycle journal.
> - Project archive changes currently leave no journal record.
> - Consumers need to observe archive and restore transitions.
> - The record must commit with the project change.
> - This PR adds project archive capture before plugin delivery in
#15343.

## Linked Issues or Issue Description

**Problem or motivation**

Archiving a project does not record a lifecycle event. Restoring an
archived project also emits no event when it is the only change.
Consumers cannot observe these transitions through the journal.

**Proposed solution**

Record `archive` when an active project becomes archived. Record
`update` when an archived project is restored. Keep both writes in the
existing project transaction and row lock. Repeated status requests emit
no new status record.

**Alternatives considered**

An in-process notification can disappear on restart. A separate archive
service would duplicate the existing mutation path. Use the existing
journal and project transaction.

**Roadmap alignment**

This extends lifecycle capture from #15280 and #15306. It should merge
before the delivery PR #15343. The roadmap and related PR search showed
no duplicate archive lifecycle capture.

## What Changed

- Allow project `archive` records in the journal constraint and
TypeScript type.
- Capture archive and restore transitions under the existing project row
lock.
- Emit `update` then `archive` for an edit combined with archive.
- Preserve repository records and roll back the project change if
capture fails.
- Add migration `0307_cool_naoko.sql`, tests, and database
documentation.

## Verification

- `pnpm -r typecheck` passed after the final master rebase.
- `pnpm build` passed before the final master rebase; the final CI build
also passed.
- 20 targeted lifecycle, migration snapshot, and legacy restore tests
passed on the current migration.
- Tests cover concurrent archives, repeated archive/restore requests,
repository retention, combined edits, atomic rollback, and upgrades from
older JavaScript backups without the action constraint.
- All 53 applicable CI checks passed on final commit `1021953035`; the
two Storybook checks were skipped as expected.
- Greptile scored the final commit 5/5 with no unresolved review
threads.
- `git diff --check` and a local secret/PII scan passed.
- Prior CI found a missing-constraint restore failure; the migration now
handles it. A separate runtime readiness timeout passed locally. The
current full CI run passes both paths.

## Risks

- Older JavaScript restores may omit the action check. The migration
tolerates its absence and installs the complete check.
- Apply the constraint migration before running the new capture code. It
takes a short table lock and validates existing journal rows; it changes
no existing records.
- This PR captures future transitions only. The one-time baseline and
plugin consumption remain in #15343, which must be rebased after this PR
merges.
- Archive records authorize no provider cleanup. Provider behavior and
volume retention remain separate work.
- An edit combined with archive emits two ordered records in one
transaction.

## Model Used

OpenAI Codex, based on GPT-6, with reasoning, repository inspection,
code execution, and tool use. The exact deployment model ID and context
window are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.1006.0-canary.21
2026-10-06 11:43:16 -07:00
DottaandPaperclip 9f7057e122 feat(connections): configure custom model providers across agent harnesses (#14970)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents use a harness, a model, and a credential to run tasks.
> - Connections already store credentials and control who can use them.
> - Custom providers also need an endpoint and a supported API format.
> - A per-agent endpoint would duplicate credentials and access rules.
> - This pull request stores routing on the connection and projects it
into the harness.
> - Isolated credentials and protocol checks keep the selected
connection authoritative.

## Linked Issues or Issue Description

Refs #37, #13083, #14104, #14565, #12692.

#14016 is a reference only. This PR has its own schema, vault
persistence, routing validation, runtime projection, and tests. None of
the five commits in #14016 is an ancestor of this branch. We do not
depend on or plan to merge it. #14967 addresses task-pinned account
pools. #14422 addresses another provider integration.

This is the first of two linked PRs. Merge this connection change before
#15341, which refines agent setup and adds the qualification harness.
The split keeps each review below 100 changed files. Provider catalog
entries have usable setup forms in this PR. Local browser subscription
sign-in is included.

## What Changed

- Store non-secret routing metadata on AI connections. Vault provider
API keys, including Bedrock bearer API keys. Reject general AWS access
keys.
- Enforce company, owner, human audience, agent access, connection
status, and protocol checks before resolving credentials. Keep reconnect
destinations immutable and retain connection identity during key
rotation.
- Project OpenRouter and compatible custom endpoints into Codex, Claude,
OpenCode, and local Hermes. Carry these settings through both legacy and
native runner transports. Clear conflicting host credentials and redact
keys from diagnostics.
- Preserve older OpenRouter accounts and native personal defaults. Add
Google API-key accounts and migration `0306` for the two
provider-default constraints.
- Run local Claude and Codex subscription sign-in behind the existing
browser sign-in card. Use private attempt homes and owner-bound
completion instead of a copied terminal command.
- Seed isolated Gemini authentication and preserve OpenCode workspace
permissions. Keep the selected connection authoritative. The independent
Gemini and Grok workflow fixes are in #15341.
- Keep native OpenCode custom gateway keys in a runner-owned
selected-model proxy; the harness config contains only a session-scoped
capability. Honor runtime outgoing proxy and certificate settings.
Preserve streamed responses and revoke the proxy on close or startup
failure.
- Allow ordinary members to connect native personal accounts before an
agent exists.
- Repair routed accounts from task cards using the saved provider
destination, protocol, model aliases, and connection identity.
- Add provider catalog definitions, model discovery, pinned logos, and
complete native and routed setup forms. Allow a personal routed
connection before a new agent exists. Keep endpoint authentication keys
out of Hermes terminal children.
- Recover cancelled or restarted browser sign-in with a clear restart
action. Support no-auth endpoints without a vault credential. Add
isolation and recovery regressions and runtime documentation.

## Verification

- Updated with `origin/master` at `22a3ea341`. Migration `0306` follows
the new master migration and passes migration and snapshot checks.
- The integrated connection regressions passed 152 tests and 50 native
OpenCode driver tests, including key-free child-shell configuration
reads, authenticated/no-auth forwarding, streaming, model/path
restrictions, cancellation, outgoing proxy routing, and NO_PROXY bypass.
Provider setup has 14 passing tests. The pinned real OpenCode 1.18.34
executable also completed a turn through the proxy against a local
synthetic provider; the reusable key was absent from its config. A
second real-executable smoke passed with an HTTPS CONNECT proxy and
runtime-specific synthetic certificate trust. Certificate-file and
certificate-directory regressions pass.
- Task-card repair passed 48 tests, including OpenRouter, Bedrock, and
custom gateway reconnect cases. UI typecheck and token gates passed.
- The prior core regression set passed 133 tests across new-agent setup,
provider forms, browser sign-in, routing projection, and connection
authorization. Token gates and UI typecheck passed.
- Full workspace typecheck and production build passed again after the
latest integration and credential-proxy fix. The merged deterministic
runner E2E suite passed 1,400 Vitest tests and 128 Node tests.
- Full workspace typecheck passed on the prior linked combined
implementation. Production build, Storybook build, 1,316 browser-harness
Vitest tests, and 128 Node tests passed. Head `9d964c8d9` includes the
latest master integration and regenerated migration. This exact head
passed 54 remote checks with four expected skips and Greptile 5/5; no
review threads remain open. An unchanged server fixture had a random
six-character issue-prefix collision on its first attempt. All 245 tests
passed locally and the single CI retry passed.
- A provider-free terminal check used the cited supported Hermes source
and dummy keys. Gateway and OpenRouter terminal children could not read
the selected key.
- The broad local Vitest attempt passed 15,442 tests but was not green.
It had an embedded-Postgres startup failure, an HTTP logger timeout, an
origin socket error, and a browser cancellation wait timeout. The
cancellation wait was corrected. The relevant connection tests and the
full origin test file passed separately. Latest-head CI must pass before
merge.
- Prior credential-backed acceptance exercised task creation, tool use,
artifact delivery, completion, and context-dependent follow-up. Claude
legacy and native runners passed Bedrock with `us-east-1` and
`us.anthropic.claude-sonnet-4-6`.
- Historical local qualification retained 43 passing API/gateway cells
out of 46. Those attempts span earlier builds. They do not qualify this
exact commit or staging. All subscription combinations and staging
remain unqualified.
- Verify native subscription and API-key setup. Connect a regular
provider catalog row. Verify an incompatible harness and a changed
reconnect URL are rejected. Use #15341 for the complete browser
campaign.

## Risks

- Migration `0306` changes two check constraints. It preserves rows and
is safe to reapply. It takes normal constraint-change locks.
- Credential projection touches several harnesses. CLI upgrades can
change provider configuration and session behavior.
- The native OpenCode proxy adds a loopback hop, pins requests to the
selected model, limits request bodies to 16 MiB, rejects redirects, and
expires at session close. It prevents reusable keys in the child
configuration; it is not an OS isolation boundary against a process
debugger running as the same user.
- Custom endpoints must be reachable from the agent environment. Saving
a connection does not prove connectivity. Bedrock keys require rotation
before expiry.
- Gemini CLI 0.58.0 has an upstream ACP new-file error conversion
defect. Provider overloads and an unresolved follow-up timeout also
affect live Gemini qualification. We have not patched the installed CLI
or marked those cases as passing.
- OpenClaw Gateway, Hermes Gateway, Claude Managed, AWS AgentCore,
Process, HTTP, and legacy ACPX local are excluded. Vertex, ambient AWS
identity, arbitrary auth headers, and custom routing for other harnesses
are excluded.
- These PRs do not establish production or staging qualification for
every provider and login method.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, code
execution, and browser testing. The exact deployment model ID and
context window size were not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 13:29:03 -05:00
Devin Foleyandgithub-actions[bot] 7be51f30e1 docs(release): canonicalize stable notes for v2026.1005.0 (#15357)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The release process keeps stable notes under a beta-keyed path
during the soak and renames them to the versioned path after the stable
ships
> - Stable v2026.1005.0 is published. The `canonicalize_stable_notes`
job pushed the rename branch but it does not open a pull request
> - Until the rename merges, `releases/v2026.1005.0.md` does not exist
on master and the announcement links do not resolve
> - This pull request merges the workflow's rename commit. It is a pure
rename with no content changes
> - The benefit is that the repository returns to the canonical
release-notes layout and the notes link in the announcements resolves

## Linked Issues or Issue Description

**Issue type**

Missing content

**Where is the issue?**

`releases/` — the stable notes for v2026.1005.0 still live at the
beta-keyed path `releases/beta/v2026.1002.0-beta.0.md`.

**What's wrong?**

The `canonicalize_stable_notes` job in the stable release run pushed
branch `release-notes/v2026.1005.0-canonicalize` with the rename, but it
does not open a pull request. The versioned path
`releases/v2026.1005.0.md` does not exist on master until this merges.

**Suggested fix**

Merge the workflow's rename commit. The notes were already corrected
before the stable dispatch in #15251, so no content change is needed
here.

## What Changed

- Renamed `releases/beta/v2026.1002.0-beta.0.md` to
`releases/v2026.1005.0.md` (workflow commit `ecc10236`, authored by
`github-actions[bot]`)
- No content changed. This is a pure rename. The file already carries
the correct header, commit count, and Contributors section from #14928
and #15251

## Verification

- The compare view for this branch against master shows one commit and
one file with status `renamed`, zero line changes
- The GitHub Release body for v2026.1005.0 is identical to this file
(one trailing newline differs)
- `git rev-list --count v2026.1001.0..v2026.1005.0` returns 179, which
matches the Contributors section
- After merge,
https://github.com/paperclipai/paperclip/blob/master/releases/v2026.1005.0.md
returns 200

## Risks

- Low risk: a documentation-only rename with no content changes.

## Model Used

- Claude (Anthropic), model ID `claude-fable-5-1` (Claude Fable 5.1),
extended thinking enabled, tool use via Claude Code

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
canary/v2026.1006.0-canary.20
2026-10-06 11:10:54 -07:00
DottaandPaperclip 22a3ea3414 Invite assistants from Connections with scoped browser and device consent (#14933)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Public MCP lets people use their organization from an external
assistant.
> - Operators need a visible control for this experimental access.
> - Hosted users should select an organization once and then approve its
permissions.
> - This pull request adds the setting, invitation-first setup, and
browser or device consent.
> - Connections provides a copyable invitation with public instructions
that grant no access.
> - Users reach browser consent from their assistant and return to
inspect or revoke access.

## Linked Issues or Issue Description

Builds on merged foundation #14846. This PR now targets master. Related
settings convention: #13905.

**Current behavior**

The preview uses an environment variable to enable MCP. Hosted consent
repeats organization selection. Assistant access has no entry in
Connections, so users must already know the endpoint and how to reach
consent.

**Proposed behavior**

An administrator enables Settings → Experimental → Assistant connections
(MCP). A hosted connection shows the selected organization and its icon,
then asks for permissions. Requested write access starts checked when
the user’s role permits it; the user can opt out before connecting.
Direct instance connections show an organization picker with the first
available organization selected. The selection stays fixed across
refetches and still requires an explicit Connect action. Connections
includes Assistant Connection (MCP). Its setup page explains the
canonical endpoint, client configuration, browser authentication, and
connected access. It connects as the current person and does not select
or impersonate an agent.

**Reason and benefit**

Operators manage access with the other experiments. Users select one
organization, and both the UI and server enforce that choice.

**Breaking changes**

The old enable variable has no effect. Preview operators must enable the
setting once. Apply the additive consent-request migration before
deploying the tenant, then deploy the compatible Cloud broker. Existing
direct requests and grants keep their behavior.

## What Changed

- Simplify OAuth and device consent: show the Paperclip logo beside
“Connect {client} to Paperclip”, fall back to “your assistant”, and show
the identifying origin plus its favicon below, with the callback URL
also visible when different. Remove the hosted-organization creation
action. Default to the first available organization without silently
changing it on refetch; preserve company restrictions and write
opt-outs. Keep the button row contained on narrow screens.
- Make Copy invitation the primary action, using the shared animated
AgentSetupPrompt and a collapsed manual setup section with icon-labeled
line tabs. Remove redundant link actions, copy-status text, the extra
first-prompt well and revocation explanation from the setup page. Serve
shared version-aware HTML and Markdown instructions without private
organization data.
- Support guarded Client ID Metadata Documents alongside dynamic
registration, and include authorization response issuer identification.
- Add RFC 8628 device authorization with separately hashed codes,
expiry, shared request quotas, persistent polling backoff and atomic
redemption. Reuse human consent, role checks, scoped grants, audit and
revocation.
- Add CLI device login and a local stdio bridge. Store credentials
separately with private permissions and serialize rotating refreshes.
- Add device consent stories and five cold-start paid Product E2E cases
with independent grant, configuration and durable-work assertions.

- Add `enablePublicMcp` to the settings validator, normalizer, feature
catalog, and toggle UI. Check it live for OAuth, tools, subscriptions,
and event delivery. Keep connection management and revocation available
while disabled.
- Default the MCP origin to the existing auth public URL, with strict
validation and an explicit override.
- Persist the optional OAuth `company_id` restriction. Describe only
that company and reject approval for any other company, even if the
person belongs to both. Keep active-membership and role checks.
- Show the Paperclip icon and a large organization icon during consent.
Return the saved company logo through the company-scoped request
response and reuse the standard fallback icon. Use the requested concise
permission labels: “Read all of your Paperclip data” and “Allow write
access and creating tasks as me”. Use concise permission copy, retain a
compact client and callback-origin disclosure, and remove the footer
link.
- Default requested write access on for eligible roles. Preserve opt-out
across organization changes and refetch, reset defaults for a new
request, and submit read-only access when the request or role does not
allow writes. Align the shared checkbox with its label.
- Use organization wording in consent, management, settings, and
walkthroughs. Keep the organization fixed for hosted requests and retain
direct-instance choice.
- Add an Assistant Connection (MCP) card to the Connectors catalog, a
setup page in the app shell, and a return link from Experimental
settings. Include Codex, Claude Code, OpenCode, and generic remote MCP
instructions.
- Read the live gate and canonical server URL through authenticated
setup metadata. Show only the current person’s grants for the selected
organization, refresh after consent, and support revocation. Surface
catalog status failures with an explicit retry action; do not present
them as an empty connection list. Opening setup grants no authority.
- Start the eight guided chapters in Connections. Keep presenter notes
and chapter controls around real product pages in the app shell. Explain
the terminal, consent, delegation, retrieval, and revocation handoffs.
Mark conversation examples as illustrative. Cover first use, client
setup, connected, loading, and error states. Keep the existing consent
and management stories.
- Keep the paid-eval setup and browser helper aligned with the setting
and consent button.

## Verification

- Warm-standby integration fix `ec64ea05e`: public MCP ingress now
follows the Cloud claim guard; MCP and discovery paths return 503
instead of SPA HTML while unclaimed. Event polling checks the in-memory
claim before reading the persisted experimental setting. All 97 focused
OAuth/Cloud tests and server typecheck pass, including new request and
timer regressions for idle-before-claim and resume-after-claim behavior.
Fresh review is 5/5 with no unresolved threads, and all security scans
pass on this final head. All browser shards, typecheck, build, canary
installation and other test groups passed on the first attempt. The
unchanged Cursor sandbox default-command test timed out at 10 seconds;
the exact test passed locally without edits in 587 ms. The single
failed-job retry passed, with the original failure retained in workflow
37500711895. All 54 final-head checks pass on
`ec64ea05e9a03e2179d4e2f84c2de03761f7ce26` (two optional Storybook jobs
are intentionally skipped).

- Final master integration `8457828fc`: merged foundation #14846 and
current master, preserving the invitation changes and all 33 files from
the two newer upstream changes. No migration renumbering was required.
All 95 focused OAuth/Cloud integration tests, full recursive typecheck
and token gates pass. All CI gates passed on that integration head;
review identified the warm-standby issue fixed above.

- Security-review fix `8c1d0b696`: commit shared global/per-source
admission before outbound CIMD work, preserve failed-attempt receipts,
and validate resource/scope before fetching. Added migration
`0305_chubby_vin_gonzales.sql` and six concurrent/adversarial regression
cases. All 69 OAuth/metadata tests, 26 migration checks, full recursive
typecheck and production build pass. The security scanner passed that
commit. Follow-up `87f9658e7` limits only actual cache-miss fetches; 18
authorization requests sharing one proxy across two service instances
use just two fetches. All 70 OAuth/metadata tests and server typecheck
pass after that refinement. Final follow-up `8ebeae84c` reports
admission-storage failures as retryable HTTP 503 instead of invalid
client metadata. Its regression proves no outbound request before
admission and successful retry after storage recovers. All 71
OAuth/metadata tests and server typecheck pass. Final-head security
scanning passes; Greptile is 5/5 with no unresolved findings. CI passed
all browser shards, typecheck, build, token gates and canary
installation. One unchanged adapter-utils bridge test raced a
response-file write (expected a JSON error, received the safe
file-changed error). The exact test passed locally without edits. The
single failed-job retry passed; the original failure is retained in
workflow 37490609192. All 54 checks now pass on final head
`8ebeae84ca77c0cf7ac12c2006f0f8743fe50e0b`, with security scan and fresh
Greptile 5/5 and no unresolved threads. Foundation #14846 subsequently
merged as `e34abee670069cca84afb2efb86041bce7dccbec`; the final
integration above now targets master.

- Integration with current master: preserved the new Connections source
filters and pagination, kept all eval suites, and regenerated the
consent/device snapshots as migrations 0303/0304. All four MCP migration
SQL hashes are unchanged from the staging versions. Full recursive
typecheck and production build, 132 focused UI tests (including catalog
filtering), 89 server authorization/settings tests, 26 migration tests,
120 eval calibration tests and token gates pass. Review follow-up
`4820ce74c` also keeps active assistant grants in Installed, with
pending/error recovery and revocation/company-isolation coverage. All 76
setup/catalog tests, UI typecheck and token gates pass after that fix.
The unchanged signoff browser test timed out waiting for a heartbeat in
CI at `4820ce74c`; the exact test passed locally without code changes,
and the preceding CI head passed that shard. That same unchanged test
failed at the reviewer stage in the next CI run. All five signoff tests
passed three times locally (15/15), without test changes. All eight
browser shards pass at final head `8ebeae84c`; no browser-test edits or
failed-browser-job retries were needed.

- Setup-page refinement at `9ab009178`: all 17 focused setup/consent
tests pass, along with UI typecheck, production build, Storybook build
and token gates. Browser exercised the shared prompt preview and client
tab switching, and the updated InvitationCopied Storybook interaction
checks its clipboard fixture. All final-head CI checks pass at
`9ab009178`, with no unresolved review findings. Deployed successfully
to Butter in
https://github.com/paperclipai/paperclip-cloud/actions/runs/37475189524.
Verified the actual page, tab switching and line styling, removed
actions/copy, and successful native copy/paste of the complete Butter
invitation into a local-only test field. The existing Claude grant was
left intact.
- Consent follow-up at `dc8e9fd11`: all 10 consent tests and token gates
pass. UI typecheck and production build passed again at `4e4d5e4d9`;
Storybook build and eval-helper typecheck passed for `28101cf91`.
Follow-ups let the primary button wrap on narrow screens, preserve a
distinct callback URL, and use only bundled icons to avoid pre-consent
requests to client-selected sites. Browser-verified the real consent
component in desktop and 320px mobile stories, including default
selection, write access and preserved opt-out. Updated E2E
heading/default-selection helpers. All CI checks passed at `dc8e9fd11`,
with review 5/5 and no unresolved threads. The Butter preview
publication needed a retry because npm initially accepted the DB package
before making it visible; the retry succeeded and `dc8e9fd11` deployed.
Verified a fresh, unapproved native Codex CIMD request on Butter:
default organization/write selection, known-client heading and icon,
distinct callback origin, and removed creation action. No grant was
approved for this UI check. Prior paid runs below retain their exact
source provenance; this UI-only follow-up did not rerun paid
qualification.
- Source-pinned paid matrix at
`2992ef2710f47230e7f484c709c6ba02524f884c`: **15/15 passed**, five cases
each on GPT-5.4 Mini, Claude Haiku and Sonnet. Campaign
`local-2026-10-06T02-41-14-462Z`. Covers cold start, existing config,
unavailable host, denied consent and reconnect/later retrieval, with
independent configuration/grant/task/run/document assertions. Original
failures, transcripts, source fingerprints and billing remain retained.
- Final instruction follow-up `cda8178af`: **3/3 cold starts passed** on
Mini, Haiku and Sonnet. Campaign `local-2026-10-06T02-58-30-041Z`.
Latest `0637b9f1c` shares that same guidance across HTML, Markdown and
manual UI after review; generated Markdown is verified byte-identical to
the paid-evaluated version. Shared build, server/UI typechecks, token
gates and 63 auth/metadata tests passed again. Every CI gate passed at
prior HEAD `0637b9f1c`, with review 5/5 and no unresolved threads.
- Other focused checks: 11 CLI credential/refresh-lock tests, 120 eval
calibration tests, server/UI/eval typechecks, token gates and Storybook
build passed. Full recursive typecheck and production build passed
during implementation; CI also passed them at `2992ef271`.
- Local full-suite limitations: a large-file Git streaming test times
out on this Mac, and broader CLI/route runs hit DB hook timeouts. Fresh
MCP reruns passed, and the corresponding CI groups passed. No claim that
the local full suite is green.
- Actual clients: Codex 0.153.4 and Claude Code 2.1.245 reach CIMD
consent; device CLI reaches verification/consent. New grants await human
approval. Existing local OpenCode retrieved a saved result in a fresh
conversation through its previously approved grant.
- Fresh OpenCode 1.18.17 on Butter: started with no MCP config, received
the exact copied invitation, read public setup, configured its server
and started PKCE consent. Its shell command timed out; background retry
reached the client's own callback deadline while approval remained
pending. Latest instructions cover that handoff. **No completed Butter
read/delegation/result retrieval is claimed.**
- Cloud companion
https://github.com/paperclipai/paperclip-cloud/pull/672 passes
checks/review and deployed. Anonymous setup and device-protocol routing
verified. Core `2992ef271` deployed successfully and the actual Claude
web flow now reaches consent. Its extra JWT-bearer metadata is filtered
to implemented grants; unsupported token grants remain rejected. Final
`0637b9f1c` deployed successfully to Butter in
https://github.com/paperclipai/paperclip-cloud/actions/runs/37409195300;
live HTML and Markdown both contain the final guidance. The superseded
instruction-only build was canceled before deployment. This is a
core-only staging preview; private Cloud plugins are omitted. ChatGPT
web is signed out, so browser connector use is unverified.
- Screenshot gallery begins at Butter's dashboard and distinguishes real
setup/pending consent from local reuse and fixtures. It records the
timeout finding. New persistent access needs human confirmation before
the remaining actual-client acceptance work.
- Manual path: Connectors → Assistant Connection (MCP) → Copy invitation
→ paste into assistant → configure and start authorization → sign in and
approve → verify `paperclip_connection` → delegate → retrieve the saved
report later.
- Plan and instructions: `doc/plans/2026-10-05-assistant-invitations.md`
and `doc/public-mcp.md`.

## Risks

- Apply additive, replay-safe migration `0304_curvy_shadow_king.sql`
before using device authorization. The public setup link carries no
credential. Device codes and tokens stay private; neither sharing
instructions nor installing a plugin authorizes access.
- Apply additive migration `0305_chubby_vin_gonzales.sql` before
deploying the shared metadata admission gate. It retains at most 60
short-lived, hashed-source receipts per instance and rejects excess
attempts with 429.
- CIMD metadata fetching is a new external-input boundary. It requires
HTTPS, exact client ID and redirect validation, bounded responses and
guarded DNS/network access. Client names remain self-reported.
- Device support is per-instance. The central Cloud broker retains its
existing grant support. Host installation and tool reload capabilities
vary by client; instructions describe manual settings and restart
requirements.

- Consent names the registered client in its heading and displays its
identifying origin below. Known-origin icons are bundled; all other
origins show a neutral site icon without contacting client-selected
sites. Client names are self-reported; the callback origin is the
recipient check. The Cloud chooser also displays the original client and
receiving origin before tenant handoff.

- A user who accepts the preselected write permission can create tasks
and comments. Task creation and comments can start or wake agents and
use execution budget; the consent label uses the concise wording
explicitly requested by the maintainer. Scope requests, role checks, and
the final Connect action still apply.
- Migration `0303_supreme_garia.sql` adds one nullable UUID column with
`IF NOT EXISTS`. Requests without a company restriction keep the
direct-instance picker. The binding stays recorded if its company is
deleted; consent then fails closed.
- Deploy tenant support before the Cloud broker sends `company_id`.
Unknown or inaccessible organizations must never fall back to a
different company.
- The setting defaults off. Disabling access does not cancel work
already delegated. Existing tokens and unexpired subscriptions can
resume when enabled again; revocation remains separate.
- The catalog entry is visible for discovery while the feature is off.
Setup instructions, OAuth, and tool execution remain gated. No access is
granted by viewing the entry.
- Assistant sign-in starts in the external client so it owns PKCE and
callback state. Client command syntax can change and links to official
setup documentation are included.
- An authenticated instance and valid public URL are required. Hosting,
paid execution, and store publication remain separate rollout steps.

## Model Used

OpenAI GPT-6 in Codex, with tool use and code execution. The exact
serving model version and context window are not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused checks pass;
unrelated local full-suite timeouts are explicitly recorded above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
nightly/v2026.1006.0-nightly.0 beta/v2026.1006.0-beta.0 canary/v2026.1006.0-canary.19
2026-10-06 12:35:16 -05:00
Nicky LeachandPaperclip 0fe47882cf Allow configurable Runner listening ports (#15353)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Paperclip Runner supports authenticated provider ingress.
> - Each listening Runner currently binds port 43127.
> - Concurrent Runner processes on one computer need distinct listening
ports.
> - This pull request permits an explicit listening port and retains the
default.
> - Providers can route each run without changing the Runner protocol.

## Linked Issues or Issue Description

**Subsystem affected**

Paperclip Runner launch and durable transport.

**Problem or motivation**

Two listening Runner processes in one network namespace cannot bind the
same fixed port. The CLI accepts a port flag but rejects every value
except 43127.

**Proposed solution**

Accept `--listen-port` values from 1 through 65535. Default to 43127
when the flag is omitted. Preserve the wildcard bind address, exact run
path, PRP authentication, and secure frames. A warm attachment retains
its existing listening port.

**Alternatives considered**

Separate network namespaces or a shared Runner daemon need more changes.
Configurable launch ports preserve the existing process model.

**Roadmap alignment**

This extends existing Cloud / Sandbox agent support. The duplicate
search found no matching Runner listener-port change.

## What Changed

- Default an omitted listener port to 43127 and reject invalid values.
- Validate configurable ports in the durable transport.
- Reuse the selected port during warm attachment and reject port
changes.
- Cover default and explicit ports, invalid input, concurrent listeners,
and warm attachment.
- Update transport documentation. Daytona still uses its existing
default port.

## Verification

- Native `cargo test --locked --workspace` passed (two existing tests
ignored).
- Targeted listener and CLI tests passed, including executable launches
on two concurrent ports and warm listener retention.
- `pnpm -r typecheck` and `pnpm build` passed.
- Complete GitHub CI is green, including general and serialized test
suites, all browser shards, Runner Rust/Vitest lanes, builds, and
release checks.
- The additional full local `pnpm test:run` is still running; no final
local result is claimed.
- `git diff --check` passes.

## Risks

An explicit port can already be occupied. Runner fails its bind without
choosing a different port. Port allocation and ingress authorization
remain provider responsibilities. No schema or PRP wire format changes.
Existing explicit port 43127 callers continue to work.

## Model Used

OpenAI Codex, GPT-6 family, with reasoning, repository inspection, and
code execution. The session does not expose the exact model ID or
context window.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.1006.0-canary.18
2026-10-06 09:52:41 -07:00
DottaandPaperclip e34abee670 feat(mcp): connect assistants to a team with user OAuth (#14846)
## Thinking Path

> - Paperclip gives teams durable tasks, agent execution, budgets, and
approvals.
> - People also use assistants in Codex, Claude, and other MCP clients.
> - Those assistants need a scoped connection that preserves the
person’s permissions and attribution.
> - Delegating a task must not turn the assistant into the assigned
agent.
> - This PR adds opt-in user OAuth, ten first-party tools, browser
consent, and workflow packages.
> - Paid product evals verify the resulting tasks, documents,
attribution, retries, and access boundaries.
> - The team keeps working after the assistant conversation ends.

## Linked Issues or Issue Description

**Problem or motivation**

A person cannot connect an external assistant to an existing team
through browser consent and safely delegate durable work as themselves.

**Proposed solution**

Expose an opt-in `/mcp/paperclip` endpoint with individually described
first-party operations. Bind every connection to a person, client,
company, resource, and scopes. Reuse domain authorization and
scheduling. Package shared team-review, delegation, and follow-up
workflows for OpenAI/Codex and Claude.

**Alternatives considered**

Related PRs #9393 and #12549 cover earlier remote MCP and board-operator
approaches. This change uses user OAuth and a bounded public catalog. It
does not expose a generic executor, operator administration, static
shared board credentials, or external agent execution. Registry listing
work in #9851 is a separate distribution step.

**Roadmap alignment**

This maintainer-requested implementation extends the governed MCP
gateway, activity attribution, durable work products, and hosted
deployment direction in `ROADMAP.md`. It implements the first release of
the saved design plan; external agent participation and granted
third-party tools remain later releases.

## What Changed

- Add MCP 2.0 discovery and task status/comment/document Events on the
same authenticated endpoint. Persist subscriptions and delivery
receipts, verify HTTPS callbacks, sign Standard Webhooks, encrypt
callback material, recheck permissions/Cloud membership, and bound
retries/expiry. Older MCP clients keep their existing tools.
- Add discovery, dynamic client registration, S256 PKCE, resource
validation, rotating refresh tokens, revocation, and company consent.
Store credentials as hashes and recheck membership at execution.
- Add tools for connection identity, agents/projects, task
search/read/create, human comments, documents/deliverables, and
pending-approval links. Preserve current domain permissions and
scheduling.
- Add durable mutation receipts across reconnects. Matching retries
replay results; uncertain outcomes keep the same request ID and require
inspection.
- Add consent and connection-management pages, OAuth log redaction,
shared plugin workflows, and separate OpenAI/Codex and Claude package
outputs.
- Add eight paid Product E2E cases across three models, independent
durable-state grading, usage evidence, cleanup, and report integration.
Add task-document guidance and regenerate the runner capability
inventories.
- Add migrations 0301 and 0302, the dated implementation plan, result
notes, and direct-client setup instructions in `doc/public-mcp.md`.

## Verification

- Merge integration `e180b1948`: resolved conflicts with current master,
preserved both eval registries, regenerated capability catalogs, and
regenerated migrations as 0301/0302 while keeping the original
replay-safe SQL byte-identical. Local migration safety/snapshot tests
(26), MCP/OAuth tests (38), redaction/OpenAPI tests (71), and eval
catalog/grading tests (198) pass. Token and capability gates pass. Full
recursive typecheck passed. Fresh Greptile review is 5/5 with no
unresolved findings. CI is green on this exact head (55 successes, two
intentional skips, one neutral result): one unchanged Cursor sandbox
test timed out at 10 seconds, then passed locally in 856 ms. A single
retry of that failed shard and the aggregate workflow passed. Merge
remains blocked on the repository code-owner approval rule.

Earlier checks passed at `6aa0962d4fb715f2190bb7bb22efacab2e58495d`: 55
successes, two intentional skips and one neutral result. [The earlier CI
run](https://github.com/paperclipai/paperclip/actions/runs/36901592350)
includes all test shards, browser tests, typecheck, build and canary dry
run. Greptile was 5/5 on that commit with no unresolved review threads.
GitHub still requires code-owner review under the repository merge
rules; passing checks do not bypass that approval. Paid source
fingerprints remain separate below and in the dated result note.

- Paid Events qualification passes **3/3**: GPT-5.4 Mini, Claude Haiku
4.5 and Claude Sonnet 4.6. Each uses a real public HTTPS callback,
signature verification and report retrieval in a fresh conversation. A
final Mini regression passes after the quota/status fixes. All evidence
validates. Bounded tunnel startup retries occur before provider calls
and remain visible; failed earlier attempts retain their original
grades.
- The earlier complete seven-case matrix passes **21/21**, with a
separate **3/3** delegation regression. Two preceding matrices also
passed 21/21 each. A complete 24-cell matrix including Events has not
been run. [The dated
results](doc/plans/2026-10-01-public-mcp-paid-eval-results.md) retain
exact source fingerprints, failures, model IDs and partial costs.
- Node 24: repository-wide `pnpm -r typecheck` and `pnpm build` pass
after merging master. Server typecheck passes after the final
quota/status changes. Eval typecheck and all 892 eval-support tests
pass.
- All 33 real MCP/OAuth tests pass. The preceding combined MCP,
redaction, private-address and DNS-rebinding run passed 129 tests; two
later MCP regressions cover quota reuse and unchanged-status
suppression. All 28 adjacent issue-tree/stale-lock route tests pass. CI
then found a null checkout result in the existing concurrent-workspace
path; logging now uses optional status access. All 12 closed-workspace
tests and all 33 MCP tests pass after that correction. The exact-start
event calibration exposed a timestamp gap; scanning now includes the
subscription start, with all 33 MCP tests and server typecheck passing.
These two narrow corrections follow the paid regression.
- A real Core → Cloud → Core authority round trip passes OAuth, MCP 2.0
subscription/delivery, current membership loss, unsubscribe, legacy SDK
tools, refresh and revocation. Its callback transport is a fixture with
independent HMAC verification. The paid Events campaigns separately
prove public HTTPS delivery.
- Earlier component qualification passed UI 7,117 tests, CLI 502, shared
832, skills catalog 20, database 160 and OpenAPI 10. Token gates, module
boundaries, migration order and plugin regeneration passed. CI covers
general/serialized suites, eight browser shards, runner checks,
typecheck, build and canary dry run.
- **Local full-suite limitation:** the earlier monolithic run was not
clean. It encountered overlapping schema rebuilding, Mac database
shared-memory limits and isolated CLI/fixture failures. Targeted reruns
passed. The existing >32 MiB Git filename stress test still hit its
300-second Mac timeout. The additional serialized sweep stopped after 62
passing suites once CI passed. Original failures and partial logs
remain; this PR does not claim a wholly green local monolithic run.
- Local Codex CLI and Claude Code OAuth login and MCP SDK
interoperability were verified. Public-store installation, actual
ChatGPT Work Cloud Events UI, staging HTTPS client behavior and hosted
newcomer provisioning remain release gates.

Enablement is moving to **Settings → Experimental → Assistant
connections (MCP)** in the stacked follow-up
[#14933](https://github.com/paperclipai/paperclip/pull/14933). Merge
both for the intended setup experience. This foundation branch alone
still uses `PAPERCLIP_PUBLIC_MCP_ENABLED=true`. After deployment, set
`PAPERCLIP_PUBLIC_URL` to the authenticated instance's HTTPS origin, and
connect to `/mcp/paperclip`. Select a team and allow writes in browser
consent. Configure an available agent and budget, then delegate and
retrieve results later. For Events, rescan the deployed plugin catalog
in ChatGPT Work Cloud; the host supplies its webhook credentials when
the user asks to watch a task. See [the setup
runbook](doc/public-mcp.md).

## Risks

- Events are at-least-once and may arrive out of order. No replay cursor
is advertised. Clients must refresh finite subscriptions, read current
state and avoid comment feedback loops. Callback material uses the
instance secrets master key; hosted subscriptions require the updated
Cloud broker and are bounded to five minutes/the access proof expiry.
- ChatGPT Work Cloud/dot event UI, plugin rescan and a hosted staging
subscription remain deployment gates. Local signed-webhook and paid
model evidence does not claim those surfaces have been exercised.
- Disabled by default. Merging adds schema and opt-in code; it does not
deploy a public endpoint, publish a store listing, create a team, or
start paid agents.
- Migrations 0301 and 0302 are additive and idempotent. Their SQL is
unchanged from the earlier preview numbers, so hash-aware upgrade
reconciliation preserves prior staging applications. Normal instance
upgrades must apply it before enabling MCP.
- Task creation and comments can schedule paid agent work. Consent and
tool descriptions disclose that effect. Revocation blocks future calls
but does not undo delegated work.
- Public deployments need edge rate limits and credential-safe logging.
Internal dispatch is restricted to the closed catalog and carries a
request-local verified actor.
- Hosted onboarding requires the companion Cloud broker, encryption-key
configuration, and tenant rollout. Self-hosted direct connections can
use this PR alone.
- Store acceptance and agent-mode participation are not claimed.
Checked-in plugin endpoints are development defaults; rebuild packages
for a real deployment before installation.

## Model Used

OpenAI GPT-6 in Codex, with reasoning, tool use, and code execution. A
more specific serving version and context-window size were not exposed
by the session. Paid eval models: `gpt-5.4-mini-2026-03-17`,
`claude-haiku-4-5-20251001`, and `claude-sonnet-4-6`.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (targeted/component checks;
full local-run limitations are recorded above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 11:48:53 -05:00
Devin FoleyandPaperclip 202c2d307e fix(server): leave unclaimed warm Cloud databases idle (#15314)
## Thinking Path

> - Paperclip manages work performed by AI agents.
> - Managed deployments prepare empty applications before an owner
claims them.
> - Those applications start database pollers even though no company can
have work.
> - Health probes also query SQL, so idle databases cannot remain
suspended.
> - This pull request adds an explicit standby marker for empty,
unclaimed Cloud apps.
> - The existing signed, durable claim resumes normal processing without
restarting the app.

## Linked Issues or Issue Description

**What happened?**

An empty, unclaimed warm application runs recurring chat, email, plugin,
heartbeat, cleanup, and reconciliation queries. Its health route also
opens the database. This prevents idle database compute from suspending.

**Expected behavior**

An explicitly marked unclaimed application should keep its HTTP process
and sandbox provider plugins ready while leaving the database idle. A
successful signed claim should resume normal API behavior and background
processing. Claimed and self-hosted instances should keep their current
behavior.

**Steps to reproduce**

Start an empty Cloud-managed application and leave it unclaimed. Observe
database activity while repeatedly requesting `/api/health`. Before this
change, periodic queries continue without company data.

Related: #15153 reduces allocation during chat polling. This change
suppresses polling only for explicitly marked, empty, unclaimed Cloud
apps.

## What Changed

- Add `PAPERCLIP_CLOUD_WARM_STANDBY=1`. Check company emptiness once
after restoring the persisted Cloud runtime identity. Missing Cloud
configuration, existing data, or a persisted claim leaves normal
processing active.
- Gate recurring database pollers with an in-memory predicate. Keep
startup preparation and sandbox provider plugin loading intact.
- Serve unclaimed health probes without session or database reads and
report `warmStandby: true`. Serve standby pages/assets directly from the
UI router, bypassing session, bearer, tenant, and dynamic handlers.
Refuse API requests and all WebSocket upgrades before authentication can
query SQL or seed company data.
- Exit standby after the existing signed identity assertion commits.
Normal timers resume at their next tick; a restart restores the claim
even with stale provider variables.
- Document the marker, readiness semantics, rollout checks, and
rollback.

## Verification

- `pnpm -r typecheck` passed. A final server typecheck also passed after
adding tests.
- `pnpm build` passed.
- Focused standby, signed claim, restart, health, static/Vite routing,
hostname, HMR, and live-events suites: 72 passed after the review fixes.
Includes real HTTP upgrade admission before/after claim.
- `pnpm test:run` was attempted locally; both superseded runs were
stopped after encountering checkout/platform failures. A clean-checkout
rerun eliminated ancestor skill-directory lookup failures. The
company-skills/runtime-cache families encounter macOS
read-only-directory rename failures (`EACCES`); all three company-skills
failures reproduce on unmodified base `bf14f803d5`. The initial full run
also reported one native runner API test failure; an isolated comparison
on both revisions was blocked by local embedded PostgreSQL startup
failures. The full [Linux CI
run](https://github.com/paperclipai/paperclip/actions/runs/37423263993)
passed on final commit `5016c415ea`, including all server and workspace
test shards, browser suites, typecheck, build, and release canary. This
is not a claim that the full local suite passed.
- Isolated full server with local PostgreSQL: after startup and
connection expiry, 70 health probes, 70 page requests carrying valid
synthetic tenant credentials, and 70 rejected WebSocket upgrades over 70
seconds observed zero app database connections. The signed claim
completed in 62 ms and normal polling resumed (475 database transactions
over 12 seconds). Restart with stale provider variables restored the
durable claim. The latency is local-only, not a provider wake
measurement.
- Apex review: **5/5** on `5016c415ea`, both earlier threads resolved,
no open recommendations.
- No live-provider test or production deployment was performed. An
actual database suspension/resume canary remains required before
enabling the control-plane switch.

## Risks

- Standby health reports HTTP readiness rather than current database
connectivity. The signed claim still requires a durable database write;
claimed health checks retain the SQL probe and 503 failure behavior.
- Pollers resume at their usual intervals. A suspended database may add
claim latency. Validate the real provider before enabling the marker.
- Startup preparation and sandbox plugins remain loaded. New plugins or
background loops must respect the same standby contract.
- The marker is off by default. Remove it or set it to `0` and restart
to roll back. No schema migration or claimed-workspace inactivity policy
changes.

## Model Used

OpenAI Codex, based on GPT-6. The exact serving snapshot and configured
context-window size are not exposed in this session. Assistance included
source review, TypeScript changes, command execution, and PostgreSQL
tests.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass — targeted tests pass; full
local suite limitations are documented above, and full Linux CI is green
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.1006.0-canary.17
2026-10-06 09:11:25 -07:00
DottaandPaperclip f2715e02bb fix(native): surface model capacity errors and retry automatically (#15347)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Native runs preserve provider events and results, then decide task
state.
> - Codex can fail a turn because the selected model is temporarily at
capacity.
> - Paperclip displayed a generic native-session failure and left the
task in recovery.
> - A known capacity failure needs a clear message and a durable delayed
retry.
> - This pull request uses committed terminal evidence, the
status-effect ledger, and existing dispatch gates.
> - The task can continue automatically without an unbounded retry loop
or a model switch.

## Linked Issues or Issue Description

Related: #13993 adds launcher capacity deferrals before provider
startup. This change handles a committed native Codex `serverOverloaded`
terminal after provider work has started.

**What happened?**
A native Codex turn ended with `codexErrorInfo: serverOverloaded` and
“Selected model is at capacity. Please try a different model.” Paperclip
saved the result, showed a generic failure, and required recovery
instead of waiting and retrying.

**Expected behavior**
Show the capacity error directly. Schedule bounded retries after a short
delay. Preserve saved work and honor current execution gates.

**Steps to reproduce**
1. Run a native Codex task.
2. Commit a `turn.failed` event with `serverOverloaded`, then accept a
failed result for the same turn.
3. Inspect the task error and its recovery state. The regression test
reproduces this without a paid provider call.

**Paperclip version or commit**
Reproduced against master `16b7db35ffa0f9a95913c8cbdeea3d595435691f`.

## What Changed

- Classify capacity failures from committed runner events and the pinned
execution identity. Preserve accepted results and display the specific
capacity error.
- Atomically persist one scheduled successor with the status decision.
Retry after one minute, then two minutes. Share the existing failure
budget and stop after two automatic retries.
- Reuse dispatch gates for ownership, task holds, dependencies, budget,
and locks. Wait for predecessor execution, finalization, and cleanup
before claiming a retry.
- Preserve pending reviewer authority and suppress retries after
reassignment or a successor claim.
- Consume the failed run's resume receipt and delivered wake input.
Rebuild ordinary continuation from the failed run so explicitly resumed
tasks can retry without borrowing one-run authorization.
- Label scheduled retries “Model at capacity.” Add a Storybook example
and avoid duplicate punctuation in the existing retry card.
- Add regression coverage and document the runtime contract.

## Verification

- Targeted recovery and UI suites: 125 tests passed. Additional final
cleanup and lock regressions passed.
- Latest-head continuation and authorization regressions: 57 tests
passed, including all four capacity integration tests against an
isolated PostgreSQL database.
- Rechecked all four capacity integration tests with the full runner's
isolated `PAPERCLIP_HOME`, config, temporary directory, and serial fork
settings: passed.
- `pnpm -r typecheck`: passed. Final server typecheck passed.
- `pnpm build`: passed.
- `pnpm check:token-gates`: passed.
- Browser: checked the real Storybook card. It shows the capacity
message, automatic retry time, and existing Retry now action.
- `pnpm test:run`: attempted and restarted after an interruption. The
resumed run started before the final review correction and was stopped
after recorded workspace/native test failures and the five-minute Git
streaming timeout already documented on master. Exit 130; no complete
local full-suite pass is claimed. Current-head isolated capacity tests
and the complete CI suite pass.
- A combined local recovery-suite attempt also hit PostgreSQL
initialization failures at this macOS host's global shared-memory limit
(32 slots). The focused recovery suites passed separately; final-head
capacity tests also pass with the full runner's isolated environment
settings.
- Latest-head GitHub checks: all 56 checks green, including general and
serialized test shards, browser E2E, typecheck, build, Runner
verification, and canary dry run. No merge conflicts.
- Greptile: 5/5 on `813a470b8c18c05aeb7e31e63e573c3a3a2f5cac`, with zero
unresolved findings after fixing the resumed-task continuation issue.

## Risks

- Capacity retries can repeat a task turn after partial work. They start
a fresh provider session with task history and wait for predecessor
cleanup. They do not resume the failed turn.
- Retries retain the configured model unless an operator changes
configuration. Persistent overload consumes the existing failure budget
and then needs an explicit retry or model change.
- Only run-bound native Codex `serverOverloaded` failures qualify.
Usage-limit exhaustion, model/account incompatibility, unbound text, and
unknown provider failures retain their current recovery behavior.
- No schema migration is required.

## Model Used

- OpenAI GPT-6 through Codex. The exact model ID and context-window size
are not exposed to this session. Used reasoning, repository tools, code
execution, and browser verification. No subagents.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.1006.0-canary.16
2026-10-06 10:59:18 -05:00
DottaandPaperclip 16b7db35ff Shorten planning skills and measure task decomposition (#15296)
## Thinking Path

> - Paperclip manages work for AI agents.
> - Planning guidance helps agents choose owners and dependencies.
> - The runtime skill favors few tasks, but the catalog skill requires a
child-task breakdown.
> - Both add repeated process instructions that can distract from the
requested outcome.
> - This change keeps the ownership and dependency rules and removes the
required matrix and repeated checklist.
> - A bounded Product E2E comparison measures saved outcomes and task
handoffs before qualification.

## Linked Issues or Issue Description

Refs #11057. Related measurement work: #15218.

**What existing behavior does this improve?**
Planning and delegation through the runtime plan-to-tasks and bundled
task-planning skills.

**Current behavior**
The two skills contain about 1,900 words and conflicting guidance on
whether plans require child tasks.

**Proposed behavior**
Keep cohesive work with one owner. Split only for a real owner, parallel
output, dependency, independent review, or follow-up lifecycle. Preserve
existing authorization and planning mechanics.

## What Changed

- Shorten both skills to about 400 words combined. Preserve their keys
and installed-version behavior.
- Remove the duplicate operational-skill pointer and regenerate affected
source metadata.
- Add twelve explicit Product E2E cells: four scenarios with current,
short and disabled planning skills.
- Use the current task composer and actual create-response ID; calibrate
public skill APIs and browser creation without providers.
- Eliminate an observed collision in chat-test company prefixes with a
per-suite sequence.
- Grade saved documents, exact author/run attribution, child count,
prerequisite execution order, review boundaries and completion handoffs.
- Retain current skill bytes and report source, selections, run
accounting and failures.

## Verification

- `pnpm test:e2e:runner:typecheck`: pass.
- `pnpm test:e2e:runner:unit`: 1,287 Vitest tests and 128 Node checks
pass.
- `pnpm test:e2e:runner -- --list --suite plan-task-guidance`: twelve
local Codex cells.
- Archived current skills match master
`72ff3a9f27e581a27acb49771e8658bbb0bbaa47` exactly.
- Corrected fixture: three real public-API/database calibrations pass
with zero provider runs; all 35 evaluator checks and Product E2E
typecheck pass.
- Setup campaign
[37399550253](https://github.com/paperclipai/paperclip/actions/runs/37399550253)
was canceled after source review found unsupported bundled edits and
automatic core reinstallation. Its paid-cell step was skipped: zero
provider runs, no behavioral grade.
- The next setup
[37401094799](https://github.com/paperclipai/paperclip/actions/runs/37401094799)
failed before task creation on the old title-field selector: zero actual
runs, original FAIL retained, cleanup passed. A real browser/API
calibration of the new helper passes with paused non-provider agents and
zero runs.
- Full local typecheck/build pass. Full local tests retain one unchanged
five-minute Git streaming timeout (also fails isolated), 9,591 passes
and 5,796 skips. CI's chat failure was a proven random fixture-prefix
collision; five affected cases pass after the test-only repair.
- Paid behavior comparison and new-head CI/review remain pending. This
PR remains a draft.

## Risks

- The shorter text may change delegation decisions. Live outcomes are
not yet qualified.
- The initial comparison uses one profile and one attempt per cell. It
cannot establish cross-model reliability or cost trends.
- Disabled means unassigned company-owned copies; the company library
remains discoverable. This does not qualify global removal, automatic
accepted-plan wiring changes, or installed-copy migration.
- Skill availability does not prove a model read or cognitively used it.
- No provider/tool protocol, permission, timeout or runtime lifecycle
behavior changes in production.

## Model Used

OpenAI Codex (GPT-6), with repository inspection, code editing and tool
use. The exact backend model ID and context-window size are not exposed
in this session. The declared eval model is native Codex `gpt-5.6-sol`.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 09:36:01 -05:00
Nicky LeachandPaperclip 44e4979d23 Capture project and repository update lifecycle events (#15306)
## Thinking Path

> - Existing lifecycle capture records agent transitions and project
creation.
> - Project and repository edits need matching project update records.
> - The record must commit with the mutation so failed writes cannot
lose a hook.
> - Creation with repositories is one creation, and repository
replacement is one update.
> - Archive-only changes preserve project state without creating a hook.
> - Plugin consumption and provider behavior are separate work.

## Linked Issues or Issue Description

**Problem or motivation**

Project edits and repository/workspace changes lack durable lifecycle
records. A future resource plugin needs those changes captured alongside
the existing project creation hook. Archive-only changes must produce no
hook.

**Proposed solution**

Allow project `update` records in the existing lifecycle journal. Record
project and workspace mutations in their database transaction while
holding the project row lock. Suppress intermediate workspace hooks
during project creation and aggregate repository replacement.

**Alternatives considered**

Route-only hooks miss shared service callers. Recording after commit can
lose an event. Emitting a hook for each child mutation exposes
intermediate repository state.

**Roadmap alignment**

This completes project lifecycle capture begun in #15280. Plugin
delivery, VM/volume provisioning, and backfill remain separate. Searches
found no duplicate project lifecycle work; related #13306 concerns
decision events on the in-process plugin bus.

## What Changed

- Record project edits and workspace additions, updates, and removals as
project `update` events.
- Commit each event atomically with its mutation under the project row
lock.
- Keep project creation with repositories to one creation event and
repository replacement to one aggregate update.
- Ignore archive-only changes and retain workspace records.
- Extend the journal action constraint and document project update
capture.

## Verification

- `pnpm -r typecheck` passed on the narrowed scope.
- 57 tests passed across six lifecycle, project, repository, and
chat-project suites.
- Seven managed-sandbox workspace route tests and the CLI
lagging-worktree migration regression passed (65 targeted tests total).
- Full GitHub CI passed on `8ba6f97f22`; all required gates are green.
- Greptile scored the final project-only commit 5/5 with zero unresolved
review threads.
- The branch is current with `master` and has no merge conflicts.
- `git diff --check` and a local secret/PII scan passed.

## Risks

- Apply migration `0300_chunky_chamber.sql` before running the new
server. It permits project update actions and tolerates older JavaScript
worktree backups that omitted the prior CHECK constraint.
- Event-write failure intentionally rolls back the project or repository
mutation.
- Records contain identity and action; future consumers must load
current authorized project/workspace data.
- Plugin consumption, provider calls, volume cleanup, and backfill are
outside this PR.

## Model Used

OpenAI Codex, based on GPT-6, with reasoning, repository inspection,
code execution, and tool use. The exact deployment model ID and context
window are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.1006.0-canary.15
2026-10-06 07:33:18 -07:00
1477d1ecea test: remove the no-op sequential describe modifier (#15286)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Paperclip runs the server and runner test suites with Vitest.
> - Vitest 5 removes the deprecated `describe.sequential` property, so
the pending Vitest 5 upgrade fails the type-check and test jobs.
> - `describe.sequential` only changes behaviour inside a
`describe.concurrent` suite, or when `sequence.concurrent` is on.
> - This repository has neither, so the modifier changed nothing at run
time.
> - The benefit is that the Vitest 5 upgrade can land, and the test
files lose a modifier that did no work.

## Linked Issues or Issue Description

Refs: #12969

## What Changed

- Replace every `describe.sequential` use with a plain `describe` call.
- Drop the `{ concurrent: false }` suite option from the two runner test
files.
- Add a comment to `server/vitest.config.ts` that records why these
suites must run one test at a time.
- Leave the package manifests and the lockfile unchanged.

## Why the modifier did nothing

The Vitest documentation states that `describe.sequential` is useful to
run tests in sequence inside a `describe.concurrent` suite, or with the
`--sequence.concurrent` option. `sequence.concurrent` defaults to
`false`.

This repository satisfies neither condition:

- No test file uses `describe.concurrent`, `it.concurrent`, or
`test.concurrent`.
- `server/vitest.config.ts` sets `sequence.concurrent: false`, with
`maxWorkers: 1`, `maxConcurrency: 1`, and `isolate: true`.
- `packages/paperclip-runner/vitest.config.ts` sets no `sequence` block,
so the `false` default applies.

`packages/db` and `cli` already run the same embedded-Postgres suites
with a plain `describe`, and those jobs are green. The server package
was the only outlier.

The modifier did carry one real piece of knowledge: these suites need
their tests to run one at a time. The new comment in
`server/vitest.config.ts` records that reason next to the setting that
enforces it.

## Verification

- `git grep` for `describe.sequential` returns nothing outside
`node_modules`.
- The author ran the changed server test files under the installed
Vitest 4, and the results match the results without this change.
- Two very large embedded-Postgres test files exceeded the author's
local memory limit, so the CI test jobs cover those two.
- The two changed runner test files have pre-existing local failures
caused by a missing Rust toolchain and a missing global `pnpm` binary.
The failures are identical with and without this change.
- The author type-checked the changed files and found no new error.
- CI must pass the typecheck, build, server test, and runner verify
jobs.

## Risks

- Low risk. Suite execution stays serial, because the Vitest config
enforces it.
- The change adds no dependency and changes no package manifest or
lockfile.
- A future change that turns `sequence.concurrent` on would break these
suites. The new config comment warns against it.

## Model Used

- Claude Sonnet 5 — code edits and local verification.
- OpenAI Codex, GPT-5 — the earlier revision of this branch.

## Test plan

- [x] Every CI check reaches a terminal green state. A pending or queued
check is not a pass.
- [x] The `Typecheck + Release Registry` job passes. This change must
not introduce a type error.
- [x] The `Build` job passes.
- [x] The server test jobs and the runner verify jobs pass.
- [x] Greptile re-reviews this commit set and posts a passing verdict.
The dependabot waiver does not apply to this pull request.
- [x] `mergeable` reads `MERGEABLE` as a terminal value.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have linked the related public issue with `Refs: #12969`
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-Authored-By: Priya Raman <priya.raman@paperclip.ing>

---------

Co-authored-by: Priya Raman <priya.raman@paperclip.ing>
Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Co-authored-by: nickyleach <331803+nickyleach@users.noreply.github.com>
2026-10-06 07:19:15 -07:00
Devin FoleyandPaperclip 90182b4f8b Report bounded Cloud portfolio failure diagnostics (#15340)
## Thinking Path

> - Paperclip manages work for AI agents and their companies.
> - Cloud instances proxy a trusted user's portfolio through the control
plane.
> - A failed request returns a generic error to the client.
> - That replacement error loses the failure phase and network code in
Sentry.
> - Operators need bounded evidence without upstream messages or
credentials.
> - This PR adds safe diagnostics while preserving the existing request
behavior.

## Linked Issues or Issue Description

Refs #10850, which added the portfolio proxy. No open PR for this
diagnostic gap was found.

**What happened?**

A rejected portfolio fetch becomes a generic 502 in Sentry. The event
cannot distinguish a connection reset, deadline, HTTP response failure,
or body failure. The route's previous warning also included the original
error and a stack identifier.

**Steps to reproduce**

Make the portfolio proxy's fetch reject with a TypeError whose cause has
`code: ECONNRESET`. The client correctly receives the generic 502, but
the captured replacement error loses that code.

**Expected behavior**

Keep the existing client response. Attach only bounded server-side
diagnostic fields to the failure event. Do not retry the request or
expose the original error.

## What Changed

- Add a typed portfolio error with a private frozen diagnostic record:
phase, upstream HTTP status, elapsed milliseconds, and an allowlisted
network code.
- Read at most four error/cause objects through own data properties.
Unknown codes, messages, getters, and out-of-range values do not enter
the record.
- Replace the route's raw-error warnings with safe fields. Send a plain
error plus event-local context through the existing optional Sentry
gate. Keep the route callsite and default fingerprint policy.
- Preserve authentication, trusted headers, cookies, exact HTTP error
bodies, cache behavior, the ten-second deadline, and one fetch per
request. Public responses receive no diagnostic fields.
- Document the fields and test HTTP behavior, privacy, and event
isolation with the real Sentry SDK.

## Verification

- `PAPERCLIP_REQUIRE_SENTRY_TEST_SDK=1 pnpm exec vitest run
server/src/__tests__/cloud-portfolio-error.test.ts
server/src/__tests__/cloud-routes.test.ts
server/src/__tests__/sentry.test.ts
server/src/__tests__/run-failure-sentry-real-sdk.test.ts`: 65 passed,
with the audited optional SDK installed.
- Full local `pnpm -r typecheck` and `pnpm build` passed on Node 24.21.0
and pnpm 9.15.4.
- Independent review found no blockers and independently passed all 65
tests, including the real SDK checks, on this exact commit.
- Full Linux CI passed on this exact commit and provides aggregate suite
coverage (54 successful checks, 2 intentional skips). A duplicate full
local aggregate was not run.
- The first SDK contract job failed before tests when npm could not
resolve an OpenTelemetry transitive package. A subsequent empty-cache
install first encountered a missing tarball, then succeeded after the
registry artifact became available. All 6 real-SDK tests passed against
that fresh install; the single unchanged-head CI retry passed. The SDK
pin, workflow, and dependency files are unchanged.
- The first browser shard 8 run timed out waiting for the inbox retry
reply after 45 seconds. The unchanged isolated case passed (1/1), and
the test, UI handler, fixture, and recovery files match the base commit.
The failed log contains no wakeup POST before the test's immediate
navigation; a navigation/request timing race is suspected but unproven
without a trace. The single unchanged-head shard retry passed (20
passed, 1 skipped); no timeout or source change was made.
- Greptile reviewed this exact commit at 5/5 with no unresolved review
threads.
- No live portfolio request was replayed. Route tests use controlled
local upstream responses.
- The added diff passed the secret and PII scan and `git diff --check`.

## Risks

This is a diagnostic change. It does not identify or repair the origin
of a connection reset. Unknown transport failures remain `unknown`.
Elapsed values outside 0–60,000 ms become null. Default Sentry
fingerprinting remains enabled; exact historical group membership is not
guaranteed. No retry, migration, deployment, or configuration change is
included.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository editing, code
execution, and independent agent review. The exact deployment model ID
and context window are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 06:58:36 -07:00
Devin FoleyandPaperclip 3290d97417 Scope the release smoke opening question to its card (#15339)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The release lane checks onboarding before it promotes a nightly
candidate.
> - The first task shows an unanswered-question summary and an open
question card.
> - Both display the same prompt, so the smoke test's page-wide text
locator fails strict matching.
> - This pull request scopes that assertion to the open interaction
card.
> - The smoke still requires the actual prompt and verifies that
onboarding does not start an agent run.

## Linked Issues or Issue Description

**What happened?**

[Scheduled release run
37441718130](https://github.com/paperclipai/paperclip/actions/runs/37441718130)
failed `smoke_nightly / smoke` on both attempts. The page-wide
`getByText("What would you like to do?")` matched two elements: the
timeline summary and the open question card. This skipped
`publish_nightly`.

**Expected behavior**

The test confirms that the seeded task's open question card displays the
exact prompt. A summary row alone must not satisfy that check.

**Steps to reproduce**

Run the Docker onboarding smoke against published candidate
`2026.1006.0-canary.11`, then run the release-smoke Playwright spec. The
old assertion fails after the greeting appears.

**Paperclip version or commit**

Release workflow commit `f858207161ba29c01c82f4674aef83d91b74480f`;
published candidate `2026.1006.0-canary.11`. The assertion is unchanged
on the current master base `9b3fe260bac576d622ffdfefb923db73cc5273f7`.

**Deployment mode**

Docker smoke harness, authenticated/private, with its existing mock
provider.

Related: #13166 added the first-task chat assertion. I also reviewed
open #12316, which fixes a separate bootstrap race and does not change
this selector. Searches found no duplicate selector fix or public issue.

## What Changed

- Scope the exact opening-prompt text to `task-chat-interaction`.
- Explain why the timeline summary cannot satisfy the assertion.
- Preserve the rest of the smoke flow, including the 15-second check for
no agent runs.

## Verification

- Local Docker harness ran the exact failed published candidate,
`2026.1006.0-canary.11`, with the existing mock provider.
- The old spec reproduced the same two-element strict-mode error in
Chrome.
- The fixed spec passed the full authenticated onboarding flow and the
15-second no-run check: 1/1, 32.6 seconds. The executed spec copy was
byte-identical to the changed repository file; only artifact output
paths and the local port were overridden.
- Independent Chrome checks: the scoped locator passes with both summary
and card present. Summary-only, wrong prompt, hidden prompt, and
duplicate active prompts each fail as intended.
- `pnpm typecheck` and `pnpm build` passed on Node 24.21.0. Canonical
Linux CI passed all 55 checks (53 success, 2 intentional Storybook
skips), including the full aggregate tests, build, typecheck, all eight
E2E shards, and post-ready security scan.
- Greptile scored the exact head
`5a30e667396600a052a4041c819652c19b8c68ff` at 5/5 with no actionable
issues. Independent review found no issues; there are no review threads.
- `git diff --check` passed. The diff and PR text were scanned for
secrets and private identifiers.

## Risks

Low risk: one assertion in a release smoke test changes. A missing or
hidden prompt still fails, and multiple matching prompts in interaction
cards still fail strict mode. This does not change the product, release
selection, or publishing. The scheduled release lane still needs its
next normal successful run to prove nightly recovery.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, shell tools, and an
independent Codex review. The exact serving model identifier and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 06:42:21 -07:00
DottaandPaperclip 9b3fe260ba fix(tasks): surface Codex ChatGPT model rejection (#15299)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for work.
> - Native Runner runs report provider errors and can also save a structured failure result.
> - Codex rejects a model that the user's ChatGPT account cannot use.
> - Master now diagnoses this rejection, but the task's compact run data omits its message. The thread can still call it a generic run failure.
> - Users need the account restriction and a clear step to repair the model selection.
> - This pull request shows the existing diagnosis as “Model unavailable” on the task.

## Linked Issues or Issue Description

**What happened?**

Codex returns HTTP 400 with `invalid_request_error` and the message `The 'gpt-6.1-sol' model is not supported when using Codex with a ChatGPT account.` Master now stores an actionable diagnosis for this rejection. The task thread still labels it “Run failed” and does not receive its error text in compact run data.

**Expected behavior**

Show the account restriction on the task and in the run error. Tell the user to choose a supported model or clear the task's model override before retrying.

**Steps to reproduce**

1. Use a Native Runner Codex agent signed in with a ChatGPT account.
2. Select a model that produces the rejection above and start a task.
3. Let the runner save its generic failed result. Inspect the task's failure marker and recovery notice.

**Additional context**

Refs: #15304. That merged PR diagnoses the provider failure and preserves worker and review recovery rules. This PR adds its task-facing message and guidance without changing that diagnosis or those rules.

Refs: #13134. That PR improves model discovery for ChatGPT accounts. This PR exposes the rejection when a configured model still fails at execution time.

## What Changed

- Return a bounded model rejection message in compact issue-run data.
- Show “Model unavailable” and model-change guidance in the task thread and recovery notice.
- Add database and UI regression tests, including both the original rejection text and master's fixed diagnosis. Document the new failure message.

## Verification

- Focused merged-branch validation: 279 tests passed across activity service, native provider failure observation and PostgreSQL integration, TaskChatThread, and ExecutionBlockerNotice.
- `pnpm -r typecheck`, `pnpm build`, and `pnpm check:token-gates` passed on the merged source tree.
- Greptile reviewed conflict-resolution commit `1db39c4e636a92e2e76ba224b71c6f1e2f556c81` at 5/5 with no findings or inline comments.
- All 55 check runs completed without failure on that commit. The two optional Storybook jobs were skipped. The legacy Snyk status passed. The branch has no merge conflicts.
- UI regression assertion: the task's failure marker says “Model unavailable” and retains the account restriction. It no longer says that this failure happened after a final response.

## Risks

- Provider recognition and recovery are owned by the existing master implementation. This PR exposes only bounded error text for failed runs with `native_provider_model_rejected`.
- Historical runs with the generic `adapter_failed` code are not reclassified. No stored run is rewritten.
- Existing retry and reconciliation gates remain in place. No schema migration or model configuration change is required.

## Model Used

OpenAI Codex, based on GPT-6, with reasoning, code execution, and browser inspection. The exact deployment model ID and context window are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context to this change
- [x] I have specified the model used (with version and capability details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before requesting merge

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-06 08:05:25 -05:00
DottaandPaperclip 63f3aa2dbf fix(runner): continue restart-interrupted Codex turns (#15297)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Native runs retain their provider conversation across server
restarts.
> - A dead runner can restore a Codex conversation after its active turn
is lost.
> - The old recovery path synthesized a failed task result from that
interruption.
> - The task then required operator action even though its conversation
and workspace were available.
> - This pull request preserves the interruption cause and uses the
admitted restart attempt for one continuation in the same conversation.
> - The agent can reconcile unfinished actions and complete the current
request without resending the original task.

## Linked Issues or Issue Description

Related: #12845 added native restart recovery. #15042 covers admission
during shutdown. #14796 covers legacy shutdown recovery. This change
covers a lost native Codex turn after successful conversation
restoration.

**What happened?**

After a server restart killed the local runner, Paperclip restored the
saved Codex thread. Runnerd found that the old turn was no longer
active. It synthesized a failed terminal and a needs-review result from
the last progress message. The transport also discarded the terminal
error when it reconstructed thread history. The task failed instead of
continuing.

**Expected behavior**

After proving that the old process stopped and admitting a bounded
recovery attempt, resume the current request in the same conversation.
Preserve the workspace. Inspect unfinished actions before proceeding.
Keep real provider failures, accepted results, intentional stops,
unknown unreconciled effects, and exhausted attempts subject to their
existing rules.

**Steps to reproduce**

1. Start a local native Codex run and leave its turn active.
2. Kill the isolated runner and provider processes, as can happen during
a server restart.
3. Restore the same provider thread with no active turn.
4. Observe the synthetic task failure. The new real-process regression
reproduces this boundary with a scripted provider.

## What Changed

- Record an explicit recoverable process-loss cause without inventing a
task result.
- Preserve terminal errors and prior turns in reconstructed provider
history. Recover the authoritative saved result when adopting an
accepted continuation.
- Send one continuation in the same conversation for an admitted
dead-runner recovery. Require reconciliation of unfinished commands and
external actions.
- Persist the interrupted terminal before submission and retain the
existing recovery marker across another controller loss.
- Keep provider attempt limits, terminal failures, and intentional
cancellation behavior.
- Add red/green regressions, real process-kill coverage, restart
checkpoint coverage, and retry-budget coverage. Document the behavior
and run-log evidence.

## Verification

- Red: the new native runtime regression rejected with
`NativeProviderTerminalFailure` on the original code; the Rust restore
regression found a missing recovery cause.
- Red/green: if restoring the conversation fails and replacement is
allowed, the replacement receives the full task and fresh-session
handoff. Both prepared and legacy execution inputs are covered.
- Red: a second controller crash after the provider accepted the
continuation caused an extra `turn/start`. The regression now proves
there are exactly two submissions total: the original and its
continuation.
- Green: focused runtime, backend, driver recovery, and real-process
restart suites (207 tests). After the final history/result changes,
driver recovery and real-process restart suites passed again (36 tests).
- Green: complete Codex transport suite (186 tests), server restart
classification/database integration suites (34 tests), and Rust Codex
provider suite (92 passed, 2 ignored).
- Full `pnpm -r typecheck` and `pnpm build` passed on `60141e649`. The
subsequent replacement-prompt guard passed the Runner TypeScript check
and the complete runtime plus process-restart suites (148 tests).
- Local full-suite attempt: `pnpm test:run` reported two failures in the
untouched chat integration suite. Both passed individually, and the
complete chat suite passed on rerun (1,063 tests). After all remote test
shards passed, the duplicate serial local run was stopped with SIGINT;
it is not claimed as a full local-suite pass.
- Latest-head CI (`b1297dcd4`): 55 successful checks and 4 intentionally
skipped checks, including all test shards, typecheck, build, native
Runner verification, end-to-end tests, and canary dry run. [CI
run](https://github.com/paperclipai/paperclip/actions/runs/37405346303).
- Greptile: 5/5 on the latest head, with no unresolved review threads.
The PR is mergeable.
- The process tests use the real runner binary and a scripted Codex
provider. They do not call a live model service.

## Risks

- This changes local Codex recovery after process loss. A continuation
can execute more work in the retained conversation. Its prompt requires
state inspection before repeating an uncertain action; the system does
not replay tool calls.
- Recovery shares the existing three-attempt budget and one-shot
continuation marker. Real failures and older unmarked failed checkpoints
are not reopened.
- No database migration or API change is required.

## Model Used

- OpenAI Codex, based on GPT-6. The exact model ID and context-window
size are not exposed in this session.
- Capabilities: reasoning, source inspection, tool use, code editing,
code execution, and test analysis.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.1006.0-canary.13
2026-10-06 07:49:47 -05:00
DottaandPaperclip a1ab55a56d fix(agents): grant configuration access by default (#15283)
## Thinking Path

> - Paperclip manages agent work and permissions for a company.
> - Agent setup can require one agent to configure another agent.
> - New standard agents could create agents, but they had no direct
configuration grant.
> - Existing agents must keep their current permissions after an
upgrade.
> - The requested new-agent defaults include 13 more direct permissions
for suggestions, skills, tools, audit, inbox, and task assignment.
> - This pull request adds the 14-grant set at creation and approval
activation, retains it through invitations, and keeps existing agents
unchanged.
> - The change keeps low trust and built-in agents on their narrower
permissions.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

The agent creation and approval flows, and the agent configuration
authorization path.

**Current behavior**

A new standard agent can create an agent. It cannot make protected
changes to a peer agent unless an operator adds an `agents:configure`
grant.

**Proposed behavior**

Add direct grants for `agents:configure`, `agents:suggest-changes`,
`skills:create`, `skills:suggest-changes`, `tools:manage_connections`,
`tools:manage_profiles`, `tools:view_audit`, `audit:view_agent_actions`,
`tools:use`, `tools:manage_runtime`, `inbox:manage`, `tasks:assign`,
`tasks:assign_scope`, and `tasks:manage_active_checkouts` to new
standard agents. Scope `tasks:assign_scope` to the agent’s reporting
subtree. Keep existing agents and their grants unchanged.

**Reason and benefit**

New standard agents can complete agent setup and the requested tool,
skill, inbox, and task workflows under existing route, scope, and
approval checks.

**Breaking changes**

Existing agents keep their current permissions. There is no permission
migration.

Related PR: #12212 adds scoped grant routes. This PR changes the default
grant.

## What Changed

- Add the 14 requested direct grants in agent creation and approval
activation. Keep them when invitation approval replaces grants,
preserving explicit scopes. Remove them when an agent is deleted.
- Apply defaults only to new agents. Remove the existing-agent
permission migration, its snapshot, and its journal entry. Preserve
existing scoped grants.
- Add permission, scope, invitation, and existing-agent regression
tests. Document all default and excluded permissions.
- Prevent agent keys from creating, changing, or restoring host-executed
process or local adapter command settings, and from restoring workspace
commands through rollback.

## Verification

- Run `node_modules/.bin/vitest run
server/src/__tests__/agent-default-configure-grants.test.ts
server/src/__tests__/invite-join-grants.test.ts
packages/db/src/migration-snapshot-drift.test.ts`. All 15 tests pass.
These cover new-agent defaults, unchanged existing grants, pending
approval, invitation grants, and migration history.
- Run `node_modules/.bin/tsc -p server/tsconfig.json --noEmit`. It
passes.
- Run `packages/db/node_modules/.bin/tsx
packages/db/src/check-migration-numbering.ts` and
`packages/db/node_modules/.bin/tsx
packages/db/src/check-migration-safety.ts`. Both pass.
- Confirm that this PR has no files under `packages/db/src/migrations/`
in its final diff.
- GitHub CI at `e8173b2c41` passes build, typecheck, server suites,
browser shards, and canary verification. One unchanged OpenCode
transport test reached its five-second timeout in the first Runner shard
run. That test passes locally in 2.55 seconds. The shard passed on one
retry. All 54 checks pass, with two expected skips.
- Greptile gives this exact head 5/5. Security review passes. The PR has
no unresolved review threads or merge conflicts.

## Risks

- The 14 default permissions apply only to new standard agents. Existing
permissions stay unchanged. Low trust and managed built-in agents are
excluded.
- New grants include connection, runtime, and active-checkout
management. Existing company, responsible-user, scope, and approval
checks remain in force. The default inbox grant carries a
responsible-user-only scope so it cannot override another user's inbox
settings.
- Protected changes still need the responsible user's authority when
that check applies. Company boundaries and approval gates still apply.
- Agent-authenticated requests cannot configure process adapters or host
command settings on local adapters. Known provider credential references
remain allowed, as do narrow plain authentication overrides on new peers
when the caller uses an AI connection pool. Arbitrary environment
settings remain blocked. Board operators retain the host configuration
paths.

> I checked `ROADMAP.md`. This is a narrow fix to existing agent
configuration behavior.

## Model Used

- OpenAI Codex, GPT-6 series. This runtime did not expose its exact
hosted model ID or context window. This revision uses OpenAI GPT-6
through Codex with tool use, code execution, and tests. The runtime does
not expose the exact hosted model ID or context window.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.1006.0-canary.12
2026-10-06 05:28:48 -05:00
Devin FoleyandPaperclip f858207161 Guard routine UUID lookups without narrowing valid inputs (#15313)
## Thinking Path

> - Paperclip manages work for AI agents and their companies.
> - Routines expose details, triggers, and management actions through
resource IDs.
> - These resource IDs are UUIDs. The UI and CLI do not resolve short
prefixes.
> - A malformed ID reaches a PostgreSQL UUID comparison and returns a
server error.
> - A lookup guard can return the existing not-found response before
that query.
> - This PR preserves valid PostgreSQL UUID input forms and the existing
access checks.

## Linked Issues or Issue Description

Refs #11471. This is a credited continuation of the lookup-guard
approach from @mv2woods. That older PR remains open and unchanged. Its
review requested the required PR description sections. This continuation
uses current master and adds UUID compatibility and authorization
coverage.

The earlier `isUuidLike` guard restricts UUID versions to 1–5 and trims
input. The database currently accepts other UUID values and input forms.
This guard preserves the existing [PostgreSQL UUID input
contract](https://www.postgresql.org/docs/current/datatype-uuid.html),
including uppercase, paired braces, omitted hyphens, and hyphens after
groups of four digits. It rejects whitespace and short prefixes without
rewriting the value sent to the database.

## What Changed

- Guard the shared routine and private-trigger UUID lookups. Malformed
resource IDs return null, so existing routes return their normal 404
response.
- Keep company access, assignee permissions, body validation order, and
database error propagation unchanged.
- Add real PostgreSQL tests for stored UUIDv4, UUIDv7, nil, and max
values in six input forms. Add no-query checks for malformed input and
HTTP coverage across all 15 root-resource route handlers.
- Document complete resource IDs and distinguish them from opaque public
webhook IDs.

## Verification

- `pnpm exec vitest run server/src/__tests__/routines-service.test.ts
server/src/__tests__/routines-e2e.test.ts`: 91 passed, including the new
PostgreSQL and HTTP regressions.
- Independent review found no blockers and passed all 9 focused
PostgreSQL and HTTP regression cases on this commit.
- Full local `pnpm -r typecheck` and `pnpm build` passed on Node 24.21.0
with pnpm 9.15.4.
- Full Linux CI passed on `0a2990eca2f904335821027a322e4db868151eb3`: 53
successful checks and 2 inapplicable Storybook skips. No retries were
needed. This provides aggregate suite coverage; a duplicate full local
aggregate was not run.
- [Greptile reviewed this exact
commit](https://github.com/paperclipai/paperclip/pull/15313#issuecomment-6010222312)
at 5/5 with no findings or unresolved threads. The PR is ready for
review and has no merge conflicts.
- `git diff --check` and the added-diff secret and PII scan passed.

## Risks

Malformed routine and private-trigger IDs now return 404 instead of a
database error. Full UUIDs retain the existing company and assignee
checks. This change does not add prefix lookup, alter public webhook
IDs, or validate unrelated nested revision/thread IDs or query
parameters. No migration is required.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository editing, code
execution, and independent agent review. The exact deployment model ID
and context window are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.1006.0-canary.11
2026-10-05 22:54:25 -07:00
Barış ÖZDEMİR bf14f803d5 fix(ssh): transport project repositories as their own git checkouts (#14782)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - A project can attach more than one repository. The task workspace
keeps the selected repository at its root and puts the other project
repositories under `.paperclip-repositories/<name>-<key>`, each with its
own `.git`
> - Agents can run on an SSH execution environment. Paperclip copies the
task workspace to the remote host before the run and restores it after
the run
> - The SSH copy excludes `.git` at every depth, but the restore
baseline excludes it only at the workspace root
> - So the other project repositories reach the remote host without Git,
and the restore then deletes their `.git` directories on the Paperclip
host
> - The next run of the same task fails during workspace setup, and the
agent cannot commit to those repositories on the remote host
> - This pull request transports each project repository as a Git
workspace of its own, the same way the sandbox path already handles them
> - The benefit is that multi-repository projects work on SSH
environments across consecutive runs

## Linked Issues or Issue Description

Refs #11632 (SSH workspace transfer exclude list). Related SSH workspace
PRs: #14233, #14428, #14472. I found no issue or PR for this bug.

**What happened?**

A project has two repositories and its agent runs on an SSH environment.
After the first run, the second repository under
`.paperclip-repositories/` has no `.git` directory on the Paperclip
host. The next run of the same task fails during setup with `Managed
workspace path "…/.paperclip-repositories/<repo>" already exists but is
not a git checkout.` On the remote host, `git` inside that repository
resolves to the parent repository.

**Expected behavior**

Each project repository reaches the remote host as a Git checkout with
its local changes. Remote commits and edits come back after the run. The
next run of the same task starts normally.

**Steps to reproduce**

1. Create a project with two repositories.
2. Configure an SSH execution environment and make it the agent's
default environment.
3. Assign a task to the agent and let it run once.
4. Look at `.paperclip-repositories/<repo>` in the task workspace:
`.git` is gone.
5. Wake the agent on the same task again: the run fails with
`setup_failed`.

**Paperclip version or commit**

Reproduced on `v2026.916.1` and on `master` (`5edf55d73`).

**Deployment mode**

Self-hosted (Docker), authenticated, with an SSH execution environment.

## What Changed

- `ssh.ts`: `prepareWorkspaceForSshExecution` lists the project
repositories under `.paperclip-repositories/`. It applies the discovery
rules of `readGitWorkspaceSnapshot`: each entry must be a directory with
a valid name and must be a Git repository root, else the prepare step
fails before any transfer.
- `ssh.ts`: the anchor copy leaves `.paperclip-repositories/` out. Each
project repository then gets the same import, sync, and deleted-path
steps as the anchor. The remote anchor repository ignores
`/.paperclip-repositories/`, as the local checkout does.
- `ssh.ts`: `prepareWorkspaceForSshExecution` returns the transported
repositories (the field is present only when there are repositories).
`restoreWorkspaceFromSshExecution` accepts them with their baselines. It
validates each path and baseline first, then restores the repositories
before the anchor and stops at the first failure, as the sandbox restore
does.
- `remote-managed-runtime.ts`: the anchor baseline excludes
`.paperclip-repositories/`, and each project repository gets its own
baseline for the restore merge.
- `ssh-fixture.test.ts`: regression tests for two consecutive managed
runs and for the direct restore path, on a workspace with a project
repository (commits, dirty edits, and a deleted file). Two tests for the
new validation.
-
`docs/guides/board-operator/execution-workspaces-and-runtime-services.md`:
one line about project repositories in the SSH round trip.

## Verification

- The new regression test fails on `master` (`expected 'backend
initial\n?? ../\n' to contain 'frontend initial'`) and passes with this
change.
- `PAPERCLIP_ENABLE_DARWIN_SSH_ENV_LAB=1 npx vitest run
packages/adapter-utils/src/ssh-fixture.test.ts
packages/adapter-utils/src/remote-managed-runtime.test.ts`: 32 passed,
with the sshd fixture running.
- `tsc --noEmit` passes for `packages/adapter-utils` and `server`, and
`pnpm -r typecheck` passes for the other workspaces. The Rust step of
`@paperclipai/paperclip-runner` did not run locally because `cargo` is
not installed.
- `node ./scripts/check-no-git-push.mjs` and `pnpm
check:module-boundaries` pass.
- `pnpm test:run` did not complete locally. Before it stopped, 5 tests
failed: 2 in `server/src/__tests__/workspace-runtime.test.ts` and 3 in
`server/src/__tests__/company-skills-service.test.ts`. The same 5 tests
also fail on the base commit `5edf55d73` without this change. CI runs
the full suite.
- `pnpm build` passes for all workspaces except
`@paperclipai/paperclip-runner` and `server`, because their build
compiles the Rust runner binary and `cargo` is not installed. `tsc
--noEmit` passes for `server`.
- Manual test on a self-hosted `v2026.916.1` instance with the same
change applied: a project with two repositories and an SSH environment.
Two runs on the same task passed. After each run, the second repository
keeps its `.git` on the host. On the remote host it is a Git checkout,
and the remote anchor ignores it.

## Risks

- Low risk. Workspaces without `.paperclip-repositories/` take the same
path as before, and the return value is unchanged for them.
- A workspace with an invalid entry under `.paperclip-repositories/` now
fails the SSH prepare step. The sandbox path already rejects such
entries.
- If one repository fails to restore, the restore stops, as in the
sandbox path. The remote run directory keeps the agent's work.
- Each project repository adds one bundle import and one restore per
run. The time grows with the number and size of the repositories.
- Out of scope: other nested `.git` directories (for example a vendored
checkout inside a repository) keep the existing SSH behavior.

## Model Used

- Provider and model: Anthropic Claude Opus 5.5 (`claude-opus-5-5`), in
Claude Code.
- Capabilities: extended thinking, tool use, and code execution. The
context window size was not recorded.
- Use: the model investigated the bug, wrote the change and the tests,
and ran the checks. A separate Claude Code agent reviewed the diff. The
author reviewed the change. The manual test ran on the author's
self-hosted instance.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
canary/v2026.1006.0-canary.10
2026-10-05 22:21:00 -07:00