Commit Graph
4873 Commits
Author SHA1 Message Date
DottaandPaperclip fcff93abdd Merge provider-routing core from master before UI qualification
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-06 13:30:37 -05:00
DottaandPaperclip 0d43238bae fix(setup): keep sign-in environment selection and cleanup evidence reachable
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-06 13:29:40 -05:00
DottaandPaperclip 9f7057e122 feat(connections): configure custom model providers across agent harnesses (#14970)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents use a harness, a model, and a credential to run tasks.
> - Connections already store credentials and control who can use them.
> - Custom providers also need an endpoint and a supported API format.
> - A per-agent endpoint would duplicate credentials and access rules.
> - This pull request stores routing on the connection and projects it
into the harness.
> - Isolated credentials and protocol checks keep the selected
connection authoritative.

## Linked Issues or Issue Description

Refs #37, #13083, #14104, #14565, #12692.

#14016 is a reference only. This PR has its own schema, vault
persistence, routing validation, runtime projection, and tests. None of
the five commits in #14016 is an ancestor of this branch. We do not
depend on or plan to merge it. #14967 addresses task-pinned account
pools. #14422 addresses another provider integration.

This is the first of two linked PRs. Merge this connection change before
#15341, which refines agent setup and adds the qualification harness.
The split keeps each review below 100 changed files. Provider catalog
entries have usable setup forms in this PR. Local browser subscription
sign-in is included.

## What Changed

- Store non-secret routing metadata on AI connections. Vault provider
API keys, including Bedrock bearer API keys. Reject general AWS access
keys.
- Enforce company, owner, human audience, agent access, connection
status, and protocol checks before resolving credentials. Keep reconnect
destinations immutable and retain connection identity during key
rotation.
- Project OpenRouter and compatible custom endpoints into Codex, Claude,
OpenCode, and local Hermes. Carry these settings through both legacy and
native runner transports. Clear conflicting host credentials and redact
keys from diagnostics.
- Preserve older OpenRouter accounts and native personal defaults. Add
Google API-key accounts and migration `0306` for the two
provider-default constraints.
- Run local Claude and Codex subscription sign-in behind the existing
browser sign-in card. Use private attempt homes and owner-bound
completion instead of a copied terminal command.
- Seed isolated Gemini authentication and preserve OpenCode workspace
permissions. Keep the selected connection authoritative. The independent
Gemini and Grok workflow fixes are in #15341.
- Keep native OpenCode custom gateway keys in a runner-owned
selected-model proxy; the harness config contains only a session-scoped
capability. Honor runtime outgoing proxy and certificate settings.
Preserve streamed responses and revoke the proxy on close or startup
failure.
- Allow ordinary members to connect native personal accounts before an
agent exists.
- Repair routed accounts from task cards using the saved provider
destination, protocol, model aliases, and connection identity.
- Add provider catalog definitions, model discovery, pinned logos, and
complete native and routed setup forms. Allow a personal routed
connection before a new agent exists. Keep endpoint authentication keys
out of Hermes terminal children.
- Recover cancelled or restarted browser sign-in with a clear restart
action. Support no-auth endpoints without a vault credential. Add
isolation and recovery regressions and runtime documentation.

## Verification

- Updated with `origin/master` at `22a3ea341`. Migration `0306` follows
the new master migration and passes migration and snapshot checks.
- The integrated connection regressions passed 152 tests and 50 native
OpenCode driver tests, including key-free child-shell configuration
reads, authenticated/no-auth forwarding, streaming, model/path
restrictions, cancellation, outgoing proxy routing, and NO_PROXY bypass.
Provider setup has 14 passing tests. The pinned real OpenCode 1.18.34
executable also completed a turn through the proxy against a local
synthetic provider; the reusable key was absent from its config. A
second real-executable smoke passed with an HTTPS CONNECT proxy and
runtime-specific synthetic certificate trust. Certificate-file and
certificate-directory regressions pass.
- Task-card repair passed 48 tests, including OpenRouter, Bedrock, and
custom gateway reconnect cases. UI typecheck and token gates passed.
- The prior core regression set passed 133 tests across new-agent setup,
provider forms, browser sign-in, routing projection, and connection
authorization. Token gates and UI typecheck passed.
- Full workspace typecheck and production build passed again after the
latest integration and credential-proxy fix. The merged deterministic
runner E2E suite passed 1,400 Vitest tests and 128 Node tests.
- Full workspace typecheck passed on the prior linked combined
implementation. Production build, Storybook build, 1,316 browser-harness
Vitest tests, and 128 Node tests passed. Head `9d964c8d9` includes the
latest master integration and regenerated migration. This exact head
passed 54 remote checks with four expected skips and Greptile 5/5; no
review threads remain open. An unchanged server fixture had a random
six-character issue-prefix collision on its first attempt. All 245 tests
passed locally and the single CI retry passed.
- A provider-free terminal check used the cited supported Hermes source
and dummy keys. Gateway and OpenRouter terminal children could not read
the selected key.
- The broad local Vitest attempt passed 15,442 tests but was not green.
It had an embedded-Postgres startup failure, an HTTP logger timeout, an
origin socket error, and a browser cancellation wait timeout. The
cancellation wait was corrected. The relevant connection tests and the
full origin test file passed separately. Latest-head CI must pass before
merge.
- Prior credential-backed acceptance exercised task creation, tool use,
artifact delivery, completion, and context-dependent follow-up. Claude
legacy and native runners passed Bedrock with `us-east-1` and
`us.anthropic.claude-sonnet-4-6`.
- Historical local qualification retained 43 passing API/gateway cells
out of 46. Those attempts span earlier builds. They do not qualify this
exact commit or staging. All subscription combinations and staging
remain unqualified.
- Verify native subscription and API-key setup. Connect a regular
provider catalog row. Verify an incompatible harness and a changed
reconnect URL are rejected. Use #15341 for the complete browser
campaign.

## Risks

- Migration `0306` changes two check constraints. It preserves rows and
is safe to reapply. It takes normal constraint-change locks.
- Credential projection touches several harnesses. CLI upgrades can
change provider configuration and session behavior.
- The native OpenCode proxy adds a loopback hop, pins requests to the
selected model, limits request bodies to 16 MiB, rejects redirects, and
expires at session close. It prevents reusable keys in the child
configuration; it is not an OS isolation boundary against a process
debugger running as the same user.
- Custom endpoints must be reachable from the agent environment. Saving
a connection does not prove connectivity. Bedrock keys require rotation
before expiry.
- Gemini CLI 0.58.0 has an upstream ACP new-file error conversion
defect. Provider overloads and an unresolved follow-up timeout also
affect live Gemini qualification. We have not patched the installed CLI
or marked those cases as passing.
- OpenClaw Gateway, Hermes Gateway, Claude Managed, AWS AgentCore,
Process, HTTP, and legacy ACPX local are excluded. Vertex, ambient AWS
identity, arbitrary auth headers, and custom routing for other harnesses
are excluded.
- These PRs do not establish production or staging qualification for
every provider and login method.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, code
execution, and browser testing. The exact deployment model ID and
context window size were not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 13:29:03 -05:00
Devin Foleyandgithub-actions[bot] 7be51f30e1 docs(release): canonicalize stable notes for v2026.1005.0 (#15357)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The release process keeps stable notes under a beta-keyed path
during the soak and renames them to the versioned path after the stable
ships
> - Stable v2026.1005.0 is published. The `canonicalize_stable_notes`
job pushed the rename branch but it does not open a pull request
> - Until the rename merges, `releases/v2026.1005.0.md` does not exist
on master and the announcement links do not resolve
> - This pull request merges the workflow's rename commit. It is a pure
rename with no content changes
> - The benefit is that the repository returns to the canonical
release-notes layout and the notes link in the announcements resolves

## Linked Issues or Issue Description

**Issue type**

Missing content

**Where is the issue?**

`releases/` — the stable notes for v2026.1005.0 still live at the
beta-keyed path `releases/beta/v2026.1002.0-beta.0.md`.

**What's wrong?**

The `canonicalize_stable_notes` job in the stable release run pushed
branch `release-notes/v2026.1005.0-canonicalize` with the rename, but it
does not open a pull request. The versioned path
`releases/v2026.1005.0.md` does not exist on master until this merges.

**Suggested fix**

Merge the workflow's rename commit. The notes were already corrected
before the stable dispatch in #15251, so no content change is needed
here.

## What Changed

- Renamed `releases/beta/v2026.1002.0-beta.0.md` to
`releases/v2026.1005.0.md` (workflow commit `ecc10236`, authored by
`github-actions[bot]`)
- No content changed. This is a pure rename. The file already carries
the correct header, commit count, and Contributors section from #14928
and #15251

## Verification

- The compare view for this branch against master shows one commit and
one file with status `renamed`, zero line changes
- The GitHub Release body for v2026.1005.0 is identical to this file
(one trailing newline differs)
- `git rev-list --count v2026.1001.0..v2026.1005.0` returns 179, which
matches the Contributors section
- After merge,
https://github.com/paperclipai/paperclip/blob/master/releases/v2026.1005.0.md
returns 200

## Risks

- Low risk: a documentation-only rename with no content changes.

## Model Used

- Claude (Anthropic), model ID `claude-fable-5-1` (Claude Fable 5.1),
extended thinking enabled, tool use via Claude Code

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
canary/v2026.1006.0-canary.20
2026-10-06 11:10:54 -07:00
DottaandPaperclip f485863c21 Merge verified provider-routing task-card repair
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-06 13:05:36 -05:00
DottaandPaperclip 9d964c8d91 fix(connections): preserve provider routing during task-card repair
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-06 13:03:17 -05:00
DottaandPaperclip c582338692 merge: preserve provider setup alongside master assistant connections
Combine the provider and public MCP qualification matrices and retain both connection catalog regression groups.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-06 12:49:09 -05:00
DottaandPaperclip a97c39874f merge: align provider migration with latest master schema
Regenerate the provider-default constraints as migration 0306 after master public MCP invitation migrations; preserve replay safety and generated metadata.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-06 12:45:51 -05:00
DottaandPaperclip a96b0722e1 merge: include proxy and certificate trust regression coverage
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-06 12:43:47 -05:00
DottaandPaperclip bb11e5779a test(opencode): cover selected certificate trust and proxy bypass
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-06 12:43:42 -05:00
DottaandPaperclip 0adf8e5fd8 merge: include custom certificate trust correction
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-06 12:42:29 -05:00
DottaandPaperclip e76affdd11 fix(opencode): copy default trust before extending runtime certificates
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-06 12:42:24 -05:00
DottaandPaperclip a150e1bcfe merge: include runtime proxy preservation fix
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-06 12:40:34 -05:00
DottaandPaperclip 5b6945401c fix(opencode): retain runtime proxy and certificate settings
Use isolated Node HTTP and HTTPS agents for authenticated gateway requests, preserving proxy fallbacks, bypasses, and selected certificate trust. Keep harness loopback traffic outside outbound proxies.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-06 12:40:11 -05:00
DottaandPaperclip 22a3ea3414 Invite assistants from Connections with scoped browser and device consent (#14933)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Public MCP lets people use their organization from an external
assistant.
> - Operators need a visible control for this experimental access.
> - Hosted users should select an organization once and then approve its
permissions.
> - This pull request adds the setting, invitation-first setup, and
browser or device consent.
> - Connections provides a copyable invitation with public instructions
that grant no access.
> - Users reach browser consent from their assistant and return to
inspect or revoke access.

## Linked Issues or Issue Description

Builds on merged foundation #14846. This PR now targets master. Related
settings convention: #13905.

**Current behavior**

The preview uses an environment variable to enable MCP. Hosted consent
repeats organization selection. Assistant access has no entry in
Connections, so users must already know the endpoint and how to reach
consent.

**Proposed behavior**

An administrator enables Settings → Experimental → Assistant connections
(MCP). A hosted connection shows the selected organization and its icon,
then asks for permissions. Requested write access starts checked when
the user’s role permits it; the user can opt out before connecting.
Direct instance connections show an organization picker with the first
available organization selected. The selection stays fixed across
refetches and still requires an explicit Connect action. Connections
includes Assistant Connection (MCP). Its setup page explains the
canonical endpoint, client configuration, browser authentication, and
connected access. It connects as the current person and does not select
or impersonate an agent.

**Reason and benefit**

Operators manage access with the other experiments. Users select one
organization, and both the UI and server enforce that choice.

**Breaking changes**

The old enable variable has no effect. Preview operators must enable the
setting once. Apply the additive consent-request migration before
deploying the tenant, then deploy the compatible Cloud broker. Existing
direct requests and grants keep their behavior.

## What Changed

- Simplify OAuth and device consent: show the Paperclip logo beside
“Connect {client} to Paperclip”, fall back to “your assistant”, and show
the identifying origin plus its favicon below, with the callback URL
also visible when different. Remove the hosted-organization creation
action. Default to the first available organization without silently
changing it on refetch; preserve company restrictions and write
opt-outs. Keep the button row contained on narrow screens.
- Make Copy invitation the primary action, using the shared animated
AgentSetupPrompt and a collapsed manual setup section with icon-labeled
line tabs. Remove redundant link actions, copy-status text, the extra
first-prompt well and revocation explanation from the setup page. Serve
shared version-aware HTML and Markdown instructions without private
organization data.
- Support guarded Client ID Metadata Documents alongside dynamic
registration, and include authorization response issuer identification.
- Add RFC 8628 device authorization with separately hashed codes,
expiry, shared request quotas, persistent polling backoff and atomic
redemption. Reuse human consent, role checks, scoped grants, audit and
revocation.
- Add CLI device login and a local stdio bridge. Store credentials
separately with private permissions and serialize rotating refreshes.
- Add device consent stories and five cold-start paid Product E2E cases
with independent grant, configuration and durable-work assertions.

- Add `enablePublicMcp` to the settings validator, normalizer, feature
catalog, and toggle UI. Check it live for OAuth, tools, subscriptions,
and event delivery. Keep connection management and revocation available
while disabled.
- Default the MCP origin to the existing auth public URL, with strict
validation and an explicit override.
- Persist the optional OAuth `company_id` restriction. Describe only
that company and reject approval for any other company, even if the
person belongs to both. Keep active-membership and role checks.
- Show the Paperclip icon and a large organization icon during consent.
Return the saved company logo through the company-scoped request
response and reuse the standard fallback icon. Use the requested concise
permission labels: “Read all of your Paperclip data” and “Allow write
access and creating tasks as me”. Use concise permission copy, retain a
compact client and callback-origin disclosure, and remove the footer
link.
- Default requested write access on for eligible roles. Preserve opt-out
across organization changes and refetch, reset defaults for a new
request, and submit read-only access when the request or role does not
allow writes. Align the shared checkbox with its label.
- Use organization wording in consent, management, settings, and
walkthroughs. Keep the organization fixed for hosted requests and retain
direct-instance choice.
- Add an Assistant Connection (MCP) card to the Connectors catalog, a
setup page in the app shell, and a return link from Experimental
settings. Include Codex, Claude Code, OpenCode, and generic remote MCP
instructions.
- Read the live gate and canonical server URL through authenticated
setup metadata. Show only the current person’s grants for the selected
organization, refresh after consent, and support revocation. Surface
catalog status failures with an explicit retry action; do not present
them as an empty connection list. Opening setup grants no authority.
- Start the eight guided chapters in Connections. Keep presenter notes
and chapter controls around real product pages in the app shell. Explain
the terminal, consent, delegation, retrieval, and revocation handoffs.
Mark conversation examples as illustrative. Cover first use, client
setup, connected, loading, and error states. Keep the existing consent
and management stories.
- Keep the paid-eval setup and browser helper aligned with the setting
and consent button.

## Verification

- Warm-standby integration fix `ec64ea05e`: public MCP ingress now
follows the Cloud claim guard; MCP and discovery paths return 503
instead of SPA HTML while unclaimed. Event polling checks the in-memory
claim before reading the persisted experimental setting. All 97 focused
OAuth/Cloud tests and server typecheck pass, including new request and
timer regressions for idle-before-claim and resume-after-claim behavior.
Fresh review is 5/5 with no unresolved threads, and all security scans
pass on this final head. All browser shards, typecheck, build, canary
installation and other test groups passed on the first attempt. The
unchanged Cursor sandbox default-command test timed out at 10 seconds;
the exact test passed locally without edits in 587 ms. The single
failed-job retry passed, with the original failure retained in workflow
37500711895. All 54 final-head checks pass on
`ec64ea05e9a03e2179d4e2f84c2de03761f7ce26` (two optional Storybook jobs
are intentionally skipped).

- Final master integration `8457828fc`: merged foundation #14846 and
current master, preserving the invitation changes and all 33 files from
the two newer upstream changes. No migration renumbering was required.
All 95 focused OAuth/Cloud integration tests, full recursive typecheck
and token gates pass. All CI gates passed on that integration head;
review identified the warm-standby issue fixed above.

- Security-review fix `8c1d0b696`: commit shared global/per-source
admission before outbound CIMD work, preserve failed-attempt receipts,
and validate resource/scope before fetching. Added migration
`0305_chubby_vin_gonzales.sql` and six concurrent/adversarial regression
cases. All 69 OAuth/metadata tests, 26 migration checks, full recursive
typecheck and production build pass. The security scanner passed that
commit. Follow-up `87f9658e7` limits only actual cache-miss fetches; 18
authorization requests sharing one proxy across two service instances
use just two fetches. All 70 OAuth/metadata tests and server typecheck
pass after that refinement. Final follow-up `8ebeae84c` reports
admission-storage failures as retryable HTTP 503 instead of invalid
client metadata. Its regression proves no outbound request before
admission and successful retry after storage recovers. All 71
OAuth/metadata tests and server typecheck pass. Final-head security
scanning passes; Greptile is 5/5 with no unresolved findings. CI passed
all browser shards, typecheck, build, token gates and canary
installation. One unchanged adapter-utils bridge test raced a
response-file write (expected a JSON error, received the safe
file-changed error). The exact test passed locally without edits. The
single failed-job retry passed; the original failure is retained in
workflow 37490609192. All 54 checks now pass on final head
`8ebeae84ca77c0cf7ac12c2006f0f8743fe50e0b`, with security scan and fresh
Greptile 5/5 and no unresolved threads. Foundation #14846 subsequently
merged as `e34abee670069cca84afb2efb86041bce7dccbec`; the final
integration above now targets master.

- Integration with current master: preserved the new Connections source
filters and pagination, kept all eval suites, and regenerated the
consent/device snapshots as migrations 0303/0304. All four MCP migration
SQL hashes are unchanged from the staging versions. Full recursive
typecheck and production build, 132 focused UI tests (including catalog
filtering), 89 server authorization/settings tests, 26 migration tests,
120 eval calibration tests and token gates pass. Review follow-up
`4820ce74c` also keeps active assistant grants in Installed, with
pending/error recovery and revocation/company-isolation coverage. All 76
setup/catalog tests, UI typecheck and token gates pass after that fix.
The unchanged signoff browser test timed out waiting for a heartbeat in
CI at `4820ce74c`; the exact test passed locally without code changes,
and the preceding CI head passed that shard. That same unchanged test
failed at the reviewer stage in the next CI run. All five signoff tests
passed three times locally (15/15), without test changes. All eight
browser shards pass at final head `8ebeae84c`; no browser-test edits or
failed-browser-job retries were needed.

- Setup-page refinement at `9ab009178`: all 17 focused setup/consent
tests pass, along with UI typecheck, production build, Storybook build
and token gates. Browser exercised the shared prompt preview and client
tab switching, and the updated InvitationCopied Storybook interaction
checks its clipboard fixture. All final-head CI checks pass at
`9ab009178`, with no unresolved review findings. Deployed successfully
to Butter in
https://github.com/paperclipai/paperclip-cloud/actions/runs/37475189524.
Verified the actual page, tab switching and line styling, removed
actions/copy, and successful native copy/paste of the complete Butter
invitation into a local-only test field. The existing Claude grant was
left intact.
- Consent follow-up at `dc8e9fd11`: all 10 consent tests and token gates
pass. UI typecheck and production build passed again at `4e4d5e4d9`;
Storybook build and eval-helper typecheck passed for `28101cf91`.
Follow-ups let the primary button wrap on narrow screens, preserve a
distinct callback URL, and use only bundled icons to avoid pre-consent
requests to client-selected sites. Browser-verified the real consent
component in desktop and 320px mobile stories, including default
selection, write access and preserved opt-out. Updated E2E
heading/default-selection helpers. All CI checks passed at `dc8e9fd11`,
with review 5/5 and no unresolved threads. The Butter preview
publication needed a retry because npm initially accepted the DB package
before making it visible; the retry succeeded and `dc8e9fd11` deployed.
Verified a fresh, unapproved native Codex CIMD request on Butter:
default organization/write selection, known-client heading and icon,
distinct callback origin, and removed creation action. No grant was
approved for this UI check. Prior paid runs below retain their exact
source provenance; this UI-only follow-up did not rerun paid
qualification.
- Source-pinned paid matrix at
`2992ef2710f47230e7f484c709c6ba02524f884c`: **15/15 passed**, five cases
each on GPT-5.4 Mini, Claude Haiku and Sonnet. Campaign
`local-2026-10-06T02-41-14-462Z`. Covers cold start, existing config,
unavailable host, denied consent and reconnect/later retrieval, with
independent configuration/grant/task/run/document assertions. Original
failures, transcripts, source fingerprints and billing remain retained.
- Final instruction follow-up `cda8178af`: **3/3 cold starts passed** on
Mini, Haiku and Sonnet. Campaign `local-2026-10-06T02-58-30-041Z`.
Latest `0637b9f1c` shares that same guidance across HTML, Markdown and
manual UI after review; generated Markdown is verified byte-identical to
the paid-evaluated version. Shared build, server/UI typechecks, token
gates and 63 auth/metadata tests passed again. Every CI gate passed at
prior HEAD `0637b9f1c`, with review 5/5 and no unresolved threads.
- Other focused checks: 11 CLI credential/refresh-lock tests, 120 eval
calibration tests, server/UI/eval typechecks, token gates and Storybook
build passed. Full recursive typecheck and production build passed
during implementation; CI also passed them at `2992ef271`.
- Local full-suite limitations: a large-file Git streaming test times
out on this Mac, and broader CLI/route runs hit DB hook timeouts. Fresh
MCP reruns passed, and the corresponding CI groups passed. No claim that
the local full suite is green.
- Actual clients: Codex 0.153.4 and Claude Code 2.1.245 reach CIMD
consent; device CLI reaches verification/consent. New grants await human
approval. Existing local OpenCode retrieved a saved result in a fresh
conversation through its previously approved grant.
- Fresh OpenCode 1.18.17 on Butter: started with no MCP config, received
the exact copied invitation, read public setup, configured its server
and started PKCE consent. Its shell command timed out; background retry
reached the client's own callback deadline while approval remained
pending. Latest instructions cover that handoff. **No completed Butter
read/delegation/result retrieval is claimed.**
- Cloud companion
https://github.com/paperclipai/paperclip-cloud/pull/672 passes
checks/review and deployed. Anonymous setup and device-protocol routing
verified. Core `2992ef271` deployed successfully and the actual Claude
web flow now reaches consent. Its extra JWT-bearer metadata is filtered
to implemented grants; unsupported token grants remain rejected. Final
`0637b9f1c` deployed successfully to Butter in
https://github.com/paperclipai/paperclip-cloud/actions/runs/37409195300;
live HTML and Markdown both contain the final guidance. The superseded
instruction-only build was canceled before deployment. This is a
core-only staging preview; private Cloud plugins are omitted. ChatGPT
web is signed out, so browser connector use is unverified.
- Screenshot gallery begins at Butter's dashboard and distinguishes real
setup/pending consent from local reuse and fixtures. It records the
timeout finding. New persistent access needs human confirmation before
the remaining actual-client acceptance work.
- Manual path: Connectors → Assistant Connection (MCP) → Copy invitation
→ paste into assistant → configure and start authorization → sign in and
approve → verify `paperclip_connection` → delegate → retrieve the saved
report later.
- Plan and instructions: `doc/plans/2026-10-05-assistant-invitations.md`
and `doc/public-mcp.md`.

## Risks

- Apply additive, replay-safe migration `0304_curvy_shadow_king.sql`
before using device authorization. The public setup link carries no
credential. Device codes and tokens stay private; neither sharing
instructions nor installing a plugin authorizes access.
- Apply additive migration `0305_chubby_vin_gonzales.sql` before
deploying the shared metadata admission gate. It retains at most 60
short-lived, hashed-source receipts per instance and rejects excess
attempts with 429.
- CIMD metadata fetching is a new external-input boundary. It requires
HTTPS, exact client ID and redirect validation, bounded responses and
guarded DNS/network access. Client names remain self-reported.
- Device support is per-instance. The central Cloud broker retains its
existing grant support. Host installation and tool reload capabilities
vary by client; instructions describe manual settings and restart
requirements.

- Consent names the registered client in its heading and displays its
identifying origin below. Known-origin icons are bundled; all other
origins show a neutral site icon without contacting client-selected
sites. Client names are self-reported; the callback origin is the
recipient check. The Cloud chooser also displays the original client and
receiving origin before tenant handoff.

- A user who accepts the preselected write permission can create tasks
and comments. Task creation and comments can start or wake agents and
use execution budget; the consent label uses the concise wording
explicitly requested by the maintainer. Scope requests, role checks, and
the final Connect action still apply.
- Migration `0303_supreme_garia.sql` adds one nullable UUID column with
`IF NOT EXISTS`. Requests without a company restriction keep the
direct-instance picker. The binding stays recorded if its company is
deleted; consent then fails closed.
- Deploy tenant support before the Cloud broker sends `company_id`.
Unknown or inaccessible organizations must never fall back to a
different company.
- The setting defaults off. Disabling access does not cancel work
already delegated. Existing tokens and unexpired subscriptions can
resume when enabled again; revocation remains separate.
- The catalog entry is visible for discovery while the feature is off.
Setup instructions, OAuth, and tool execution remain gated. No access is
granted by viewing the entry.
- Assistant sign-in starts in the external client so it owns PKCE and
callback state. Client command syntax can change and links to official
setup documentation are included.
- An authenticated instance and valid public URL are required. Hosting,
paid execution, and store publication remain separate rollout steps.

## Model Used

OpenAI GPT-6 in Codex, with tool use and code execution. The exact
serving model version and context window are not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused checks pass;
unrelated local full-suite timeouts are explicitly recorded above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
nightly/v2026.1006.0-nightly.0 beta/v2026.1006.0-beta.0 canary/v2026.1006.0-canary.19
2026-10-06 12:35:16 -05:00
DottaandPaperclip b7babda93a merge: include reviewed provider credential and setup fixes
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-06 12:24:07 -05:00
DottaandPaperclip 7fed41a66d fix(connections): protect native gateway keys and allow personal provider setup
Keep custom OpenCode authentication in a session-scoped selected-model proxy rather than an agent-readable config. Permit ordinary members to create native personal accounts before an agent exists.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-06 12:22:50 -05:00
DottaandPaperclip e2891ef538 merge: integrate current master with provider setup qualification
Preserve provider-connection and public MCP suites together, including their independent fixture documentation and matrix expectations.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-06 12:17:53 -05:00
DottaandPaperclip 07ebd89ac2 merge: refresh provider connections on current master
Regenerate provider constraints as migration 0303 after the new master migrations, retaining replay safety and synchronized schema metadata.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-06 12:08:14 -05:00
Nicky LeachandPaperclip 0fe47882cf Allow configurable Runner listening ports (#15353)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Paperclip Runner supports authenticated provider ingress.
> - Each listening Runner currently binds port 43127.
> - Concurrent Runner processes on one computer need distinct listening
ports.
> - This pull request permits an explicit listening port and retains the
default.
> - Providers can route each run without changing the Runner protocol.

## Linked Issues or Issue Description

**Subsystem affected**

Paperclip Runner launch and durable transport.

**Problem or motivation**

Two listening Runner processes in one network namespace cannot bind the
same fixed port. The CLI accepts a port flag but rejects every value
except 43127.

**Proposed solution**

Accept `--listen-port` values from 1 through 65535. Default to 43127
when the flag is omitted. Preserve the wildcard bind address, exact run
path, PRP authentication, and secure frames. A warm attachment retains
its existing listening port.

**Alternatives considered**

Separate network namespaces or a shared Runner daemon need more changes.
Configurable launch ports preserve the existing process model.

**Roadmap alignment**

This extends existing Cloud / Sandbox agent support. The duplicate
search found no matching Runner listener-port change.

## What Changed

- Default an omitted listener port to 43127 and reject invalid values.
- Validate configurable ports in the durable transport.
- Reuse the selected port during warm attachment and reject port
changes.
- Cover default and explicit ports, invalid input, concurrent listeners,
and warm attachment.
- Update transport documentation. Daytona still uses its existing
default port.

## Verification

- Native `cargo test --locked --workspace` passed (two existing tests
ignored).
- Targeted listener and CLI tests passed, including executable launches
on two concurrent ports and warm listener retention.
- `pnpm -r typecheck` and `pnpm build` passed.
- Complete GitHub CI is green, including general and serialized test
suites, all browser shards, Runner Rust/Vitest lanes, builds, and
release checks.
- The additional full local `pnpm test:run` is still running; no final
local result is claimed.
- `git diff --check` passes.

## Risks

An explicit port can already be occupied. Runner fails its bind without
choosing a different port. Port allocation and ingress authorization
remain provider responsibilities. No schema or PRP wire format changes.
Existing explicit port 43127 callers continue to work.

## Model Used

OpenAI Codex, GPT-6 family, with reasoning, repository inspection, and
code execution. The session does not expose the exact model ID or
context window.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.1006.0-canary.18
2026-10-06 09:52:41 -07:00
DottaandPaperclip e34abee670 feat(mcp): connect assistants to a team with user OAuth (#14846)
## Thinking Path

> - Paperclip gives teams durable tasks, agent execution, budgets, and
approvals.
> - People also use assistants in Codex, Claude, and other MCP clients.
> - Those assistants need a scoped connection that preserves the
person’s permissions and attribution.
> - Delegating a task must not turn the assistant into the assigned
agent.
> - This PR adds opt-in user OAuth, ten first-party tools, browser
consent, and workflow packages.
> - Paid product evals verify the resulting tasks, documents,
attribution, retries, and access boundaries.
> - The team keeps working after the assistant conversation ends.

## Linked Issues or Issue Description

**Problem or motivation**

A person cannot connect an external assistant to an existing team
through browser consent and safely delegate durable work as themselves.

**Proposed solution**

Expose an opt-in `/mcp/paperclip` endpoint with individually described
first-party operations. Bind every connection to a person, client,
company, resource, and scopes. Reuse domain authorization and
scheduling. Package shared team-review, delegation, and follow-up
workflows for OpenAI/Codex and Claude.

**Alternatives considered**

Related PRs #9393 and #12549 cover earlier remote MCP and board-operator
approaches. This change uses user OAuth and a bounded public catalog. It
does not expose a generic executor, operator administration, static
shared board credentials, or external agent execution. Registry listing
work in #9851 is a separate distribution step.

**Roadmap alignment**

This maintainer-requested implementation extends the governed MCP
gateway, activity attribution, durable work products, and hosted
deployment direction in `ROADMAP.md`. It implements the first release of
the saved design plan; external agent participation and granted
third-party tools remain later releases.

## What Changed

- Add MCP 2.0 discovery and task status/comment/document Events on the
same authenticated endpoint. Persist subscriptions and delivery
receipts, verify HTTPS callbacks, sign Standard Webhooks, encrypt
callback material, recheck permissions/Cloud membership, and bound
retries/expiry. Older MCP clients keep their existing tools.
- Add discovery, dynamic client registration, S256 PKCE, resource
validation, rotating refresh tokens, revocation, and company consent.
Store credentials as hashes and recheck membership at execution.
- Add tools for connection identity, agents/projects, task
search/read/create, human comments, documents/deliverables, and
pending-approval links. Preserve current domain permissions and
scheduling.
- Add durable mutation receipts across reconnects. Matching retries
replay results; uncertain outcomes keep the same request ID and require
inspection.
- Add consent and connection-management pages, OAuth log redaction,
shared plugin workflows, and separate OpenAI/Codex and Claude package
outputs.
- Add eight paid Product E2E cases across three models, independent
durable-state grading, usage evidence, cleanup, and report integration.
Add task-document guidance and regenerate the runner capability
inventories.
- Add migrations 0301 and 0302, the dated implementation plan, result
notes, and direct-client setup instructions in `doc/public-mcp.md`.

## Verification

- Merge integration `e180b1948`: resolved conflicts with current master,
preserved both eval registries, regenerated capability catalogs, and
regenerated migrations as 0301/0302 while keeping the original
replay-safe SQL byte-identical. Local migration safety/snapshot tests
(26), MCP/OAuth tests (38), redaction/OpenAPI tests (71), and eval
catalog/grading tests (198) pass. Token and capability gates pass. Full
recursive typecheck passed. Fresh Greptile review is 5/5 with no
unresolved findings. CI is green on this exact head (55 successes, two
intentional skips, one neutral result): one unchanged Cursor sandbox
test timed out at 10 seconds, then passed locally in 856 ms. A single
retry of that failed shard and the aggregate workflow passed. Merge
remains blocked on the repository code-owner approval rule.

Earlier checks passed at `6aa0962d4fb715f2190bb7bb22efacab2e58495d`: 55
successes, two intentional skips and one neutral result. [The earlier CI
run](https://github.com/paperclipai/paperclip/actions/runs/36901592350)
includes all test shards, browser tests, typecheck, build and canary dry
run. Greptile was 5/5 on that commit with no unresolved review threads.
GitHub still requires code-owner review under the repository merge
rules; passing checks do not bypass that approval. Paid source
fingerprints remain separate below and in the dated result note.

- Paid Events qualification passes **3/3**: GPT-5.4 Mini, Claude Haiku
4.5 and Claude Sonnet 4.6. Each uses a real public HTTPS callback,
signature verification and report retrieval in a fresh conversation. A
final Mini regression passes after the quota/status fixes. All evidence
validates. Bounded tunnel startup retries occur before provider calls
and remain visible; failed earlier attempts retain their original
grades.
- The earlier complete seven-case matrix passes **21/21**, with a
separate **3/3** delegation regression. Two preceding matrices also
passed 21/21 each. A complete 24-cell matrix including Events has not
been run. [The dated
results](doc/plans/2026-10-01-public-mcp-paid-eval-results.md) retain
exact source fingerprints, failures, model IDs and partial costs.
- Node 24: repository-wide `pnpm -r typecheck` and `pnpm build` pass
after merging master. Server typecheck passes after the final
quota/status changes. Eval typecheck and all 892 eval-support tests
pass.
- All 33 real MCP/OAuth tests pass. The preceding combined MCP,
redaction, private-address and DNS-rebinding run passed 129 tests; two
later MCP regressions cover quota reuse and unchanged-status
suppression. All 28 adjacent issue-tree/stale-lock route tests pass. CI
then found a null checkout result in the existing concurrent-workspace
path; logging now uses optional status access. All 12 closed-workspace
tests and all 33 MCP tests pass after that correction. The exact-start
event calibration exposed a timestamp gap; scanning now includes the
subscription start, with all 33 MCP tests and server typecheck passing.
These two narrow corrections follow the paid regression.
- A real Core → Cloud → Core authority round trip passes OAuth, MCP 2.0
subscription/delivery, current membership loss, unsubscribe, legacy SDK
tools, refresh and revocation. Its callback transport is a fixture with
independent HMAC verification. The paid Events campaigns separately
prove public HTTPS delivery.
- Earlier component qualification passed UI 7,117 tests, CLI 502, shared
832, skills catalog 20, database 160 and OpenAPI 10. Token gates, module
boundaries, migration order and plugin regeneration passed. CI covers
general/serialized suites, eight browser shards, runner checks,
typecheck, build and canary dry run.
- **Local full-suite limitation:** the earlier monolithic run was not
clean. It encountered overlapping schema rebuilding, Mac database
shared-memory limits and isolated CLI/fixture failures. Targeted reruns
passed. The existing >32 MiB Git filename stress test still hit its
300-second Mac timeout. The additional serialized sweep stopped after 62
passing suites once CI passed. Original failures and partial logs
remain; this PR does not claim a wholly green local monolithic run.
- Local Codex CLI and Claude Code OAuth login and MCP SDK
interoperability were verified. Public-store installation, actual
ChatGPT Work Cloud Events UI, staging HTTPS client behavior and hosted
newcomer provisioning remain release gates.

Enablement is moving to **Settings → Experimental → Assistant
connections (MCP)** in the stacked follow-up
[#14933](https://github.com/paperclipai/paperclip/pull/14933). Merge
both for the intended setup experience. This foundation branch alone
still uses `PAPERCLIP_PUBLIC_MCP_ENABLED=true`. After deployment, set
`PAPERCLIP_PUBLIC_URL` to the authenticated instance's HTTPS origin, and
connect to `/mcp/paperclip`. Select a team and allow writes in browser
consent. Configure an available agent and budget, then delegate and
retrieve results later. For Events, rescan the deployed plugin catalog
in ChatGPT Work Cloud; the host supplies its webhook credentials when
the user asks to watch a task. See [the setup
runbook](doc/public-mcp.md).

## Risks

- Events are at-least-once and may arrive out of order. No replay cursor
is advertised. Clients must refresh finite subscriptions, read current
state and avoid comment feedback loops. Callback material uses the
instance secrets master key; hosted subscriptions require the updated
Cloud broker and are bounded to five minutes/the access proof expiry.
- ChatGPT Work Cloud/dot event UI, plugin rescan and a hosted staging
subscription remain deployment gates. Local signed-webhook and paid
model evidence does not claim those surfaces have been exercised.
- Disabled by default. Merging adds schema and opt-in code; it does not
deploy a public endpoint, publish a store listing, create a team, or
start paid agents.
- Migrations 0301 and 0302 are additive and idempotent. Their SQL is
unchanged from the earlier preview numbers, so hash-aware upgrade
reconciliation preserves prior staging applications. Normal instance
upgrades must apply it before enabling MCP.
- Task creation and comments can schedule paid agent work. Consent and
tool descriptions disclose that effect. Revocation blocks future calls
but does not undo delegated work.
- Public deployments need edge rate limits and credential-safe logging.
Internal dispatch is restricted to the closed catalog and carries a
request-local verified actor.
- Hosted onboarding requires the companion Cloud broker, encryption-key
configuration, and tenant rollout. Self-hosted direct connections can
use this PR alone.
- Store acceptance and agent-mode participation are not claimed.
Checked-in plugin endpoints are development defaults; rebuild packages
for a real deployment before installation.

## Model Used

OpenAI GPT-6 in Codex, with reasoning, tool use, and code execution. A
more specific serving version and context-window size were not exposed
by the session. Paid eval models: `gpt-5.4-mini-2026-03-17`,
`claude-haiku-4-5-20251001`, and `claude-sonnet-4-6`.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (targeted/component checks;
full local-run limitations are recorded above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 11:48:53 -05:00
DottaandPaperclip 3a4939ba81 fix: stop provider qualification campaigns across their entire lifecycle
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-06 11:12:17 -05:00
DottaandPaperclip 31da0558d1 fix: preserve Grok session metadata when history exceeds retention limits
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-06 11:12:17 -05:00
DottaandPaperclip 797e578452 fix: restore provider workflow fixes and persist ACP resource errors
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-06 11:12:17 -05:00
Devin FoleyandPaperclip 202c2d307e fix(server): leave unclaimed warm Cloud databases idle (#15314)
## Thinking Path

> - Paperclip manages work performed by AI agents.
> - Managed deployments prepare empty applications before an owner
claims them.
> - Those applications start database pollers even though no company can
have work.
> - Health probes also query SQL, so idle databases cannot remain
suspended.
> - This pull request adds an explicit standby marker for empty,
unclaimed Cloud apps.
> - The existing signed, durable claim resumes normal processing without
restarting the app.

## Linked Issues or Issue Description

**What happened?**

An empty, unclaimed warm application runs recurring chat, email, plugin,
heartbeat, cleanup, and reconciliation queries. Its health route also
opens the database. This prevents idle database compute from suspending.

**Expected behavior**

An explicitly marked unclaimed application should keep its HTTP process
and sandbox provider plugins ready while leaving the database idle. A
successful signed claim should resume normal API behavior and background
processing. Claimed and self-hosted instances should keep their current
behavior.

**Steps to reproduce**

Start an empty Cloud-managed application and leave it unclaimed. Observe
database activity while repeatedly requesting `/api/health`. Before this
change, periodic queries continue without company data.

Related: #15153 reduces allocation during chat polling. This change
suppresses polling only for explicitly marked, empty, unclaimed Cloud
apps.

## What Changed

- Add `PAPERCLIP_CLOUD_WARM_STANDBY=1`. Check company emptiness once
after restoring the persisted Cloud runtime identity. Missing Cloud
configuration, existing data, or a persisted claim leaves normal
processing active.
- Gate recurring database pollers with an in-memory predicate. Keep
startup preparation and sandbox provider plugin loading intact.
- Serve unclaimed health probes without session or database reads and
report `warmStandby: true`. Serve standby pages/assets directly from the
UI router, bypassing session, bearer, tenant, and dynamic handlers.
Refuse API requests and all WebSocket upgrades before authentication can
query SQL or seed company data.
- Exit standby after the existing signed identity assertion commits.
Normal timers resume at their next tick; a restart restores the claim
even with stale provider variables.
- Document the marker, readiness semantics, rollout checks, and
rollback.

## Verification

- `pnpm -r typecheck` passed. A final server typecheck also passed after
adding tests.
- `pnpm build` passed.
- Focused standby, signed claim, restart, health, static/Vite routing,
hostname, HMR, and live-events suites: 72 passed after the review fixes.
Includes real HTTP upgrade admission before/after claim.
- `pnpm test:run` was attempted locally; both superseded runs were
stopped after encountering checkout/platform failures. A clean-checkout
rerun eliminated ancestor skill-directory lookup failures. The
company-skills/runtime-cache families encounter macOS
read-only-directory rename failures (`EACCES`); all three company-skills
failures reproduce on unmodified base `bf14f803d5`. The initial full run
also reported one native runner API test failure; an isolated comparison
on both revisions was blocked by local embedded PostgreSQL startup
failures. The full [Linux CI
run](https://github.com/paperclipai/paperclip/actions/runs/37423263993)
passed on final commit `5016c415ea`, including all server and workspace
test shards, browser suites, typecheck, build, and release canary. This
is not a claim that the full local suite passed.
- Isolated full server with local PostgreSQL: after startup and
connection expiry, 70 health probes, 70 page requests carrying valid
synthetic tenant credentials, and 70 rejected WebSocket upgrades over 70
seconds observed zero app database connections. The signed claim
completed in 62 ms and normal polling resumed (475 database transactions
over 12 seconds). Restart with stale provider variables restored the
durable claim. The latency is local-only, not a provider wake
measurement.
- Apex review: **5/5** on `5016c415ea`, both earlier threads resolved,
no open recommendations.
- No live-provider test or production deployment was performed. An
actual database suspension/resume canary remains required before
enabling the control-plane switch.

## Risks

- Standby health reports HTTP readiness rather than current database
connectivity. The signed claim still requires a durable database write;
claimed health checks retain the SQL probe and 503 failure behavior.
- Pollers resume at their usual intervals. A suspended database may add
claim latency. Validate the real provider before enabling the marker.
- Startup preparation and sandbox plugins remain loaded. New plugins or
background loops must respect the same standby contract.
- The marker is off by default. Remove it or set it to `0` and restart
to roll back. No schema migration or claimed-workspace inactivity policy
changes.

## Model Used

OpenAI Codex, based on GPT-6. The exact serving snapshot and configured
context-window size are not exposed in this session. Assistance included
source review, TypeScript changes, command execution, and PostgreSQL
tests.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass — targeted tests pass; full
local suite limitations are documented above, and full Linux CI is green
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.1006.0-canary.17
2026-10-06 09:11:25 -07:00
DottaandPaperclip f2715e02bb fix(native): surface model capacity errors and retry automatically (#15347)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Native runs preserve provider events and results, then decide task
state.
> - Codex can fail a turn because the selected model is temporarily at
capacity.
> - Paperclip displayed a generic native-session failure and left the
task in recovery.
> - A known capacity failure needs a clear message and a durable delayed
retry.
> - This pull request uses committed terminal evidence, the
status-effect ledger, and existing dispatch gates.
> - The task can continue automatically without an unbounded retry loop
or a model switch.

## Linked Issues or Issue Description

Related: #13993 adds launcher capacity deferrals before provider
startup. This change handles a committed native Codex `serverOverloaded`
terminal after provider work has started.

**What happened?**
A native Codex turn ended with `codexErrorInfo: serverOverloaded` and
“Selected model is at capacity. Please try a different model.” Paperclip
saved the result, showed a generic failure, and required recovery
instead of waiting and retrying.

**Expected behavior**
Show the capacity error directly. Schedule bounded retries after a short
delay. Preserve saved work and honor current execution gates.

**Steps to reproduce**
1. Run a native Codex task.
2. Commit a `turn.failed` event with `serverOverloaded`, then accept a
failed result for the same turn.
3. Inspect the task error and its recovery state. The regression test
reproduces this without a paid provider call.

**Paperclip version or commit**
Reproduced against master `16b7db35ffa0f9a95913c8cbdeea3d595435691f`.

## What Changed

- Classify capacity failures from committed runner events and the pinned
execution identity. Preserve accepted results and display the specific
capacity error.
- Atomically persist one scheduled successor with the status decision.
Retry after one minute, then two minutes. Share the existing failure
budget and stop after two automatic retries.
- Reuse dispatch gates for ownership, task holds, dependencies, budget,
and locks. Wait for predecessor execution, finalization, and cleanup
before claiming a retry.
- Preserve pending reviewer authority and suppress retries after
reassignment or a successor claim.
- Consume the failed run's resume receipt and delivered wake input.
Rebuild ordinary continuation from the failed run so explicitly resumed
tasks can retry without borrowing one-run authorization.
- Label scheduled retries “Model at capacity.” Add a Storybook example
and avoid duplicate punctuation in the existing retry card.
- Add regression coverage and document the runtime contract.

## Verification

- Targeted recovery and UI suites: 125 tests passed. Additional final
cleanup and lock regressions passed.
- Latest-head continuation and authorization regressions: 57 tests
passed, including all four capacity integration tests against an
isolated PostgreSQL database.
- Rechecked all four capacity integration tests with the full runner's
isolated `PAPERCLIP_HOME`, config, temporary directory, and serial fork
settings: passed.
- `pnpm -r typecheck`: passed. Final server typecheck passed.
- `pnpm build`: passed.
- `pnpm check:token-gates`: passed.
- Browser: checked the real Storybook card. It shows the capacity
message, automatic retry time, and existing Retry now action.
- `pnpm test:run`: attempted and restarted after an interruption. The
resumed run started before the final review correction and was stopped
after recorded workspace/native test failures and the five-minute Git
streaming timeout already documented on master. Exit 130; no complete
local full-suite pass is claimed. Current-head isolated capacity tests
and the complete CI suite pass.
- A combined local recovery-suite attempt also hit PostgreSQL
initialization failures at this macOS host's global shared-memory limit
(32 slots). The focused recovery suites passed separately; final-head
capacity tests also pass with the full runner's isolated environment
settings.
- Latest-head GitHub checks: all 56 checks green, including general and
serialized test shards, browser E2E, typecheck, build, Runner
verification, and canary dry run. No merge conflicts.
- Greptile: 5/5 on `813a470b8c18c05aeb7e31e63e573c3a3a2f5cac`, with zero
unresolved findings after fixing the resumed-task continuation issue.

## Risks

- Capacity retries can repeat a task turn after partial work. They start
a fresh provider session with task history and wait for predecessor
cleanup. They do not resume the failed turn.
- Retries retain the configured model unless an operator changes
configuration. Persistent overload consumes the existing failure budget
and then needs an explicit retry or model change.
- Only run-bound native Codex `serverOverloaded` failures qualify.
Usage-limit exhaustion, model/account incompatibility, unbound text, and
unknown provider failures retain their current recovery behavior.
- No schema migration is required.

## Model Used

- OpenAI GPT-6 through Codex. The exact model ID and context-window size
are not exposed to this session. Used reasoning, repository tools, code
execution, and browser verification. No subagents.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.1006.0-canary.16
2026-10-06 10:59:18 -05:00
DottaandPaperclip a259543902 docs(connections): explain encrypted history and qualification environment checks
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-06 10:45:29 -05:00
DottaandPaperclip ec1e5a330b merge: apply core provider review fixes to setup and qualification
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-06 10:40:37 -05:00
DottaandPaperclip fe1da1ca25 fix(connections): recover browser logins and isolate Hermes gateway keys
Allow a personal gateway to be saved before a new agent exists. Update the new-agent connection selector regression.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-06 10:39:50 -05:00
DottaandPaperclip 925d203823 test(connections): cover bounded history and fix Storybook selector types
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-06 10:39:05 -05:00
DottaandPaperclip e1e83b51dd merge: synchronize provider UI and evaluations with current master
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-06 10:21:50 -05:00
DottaandPaperclip 28c4615258 fix(evals): verify execution environment and cancel on hangup
Block unavailable requested environments, verify saved and actual execution targets, handle SIGHUP during paid campaigns, and correct the gateway Storybook submit selector.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-06 09:57:55 -05:00
DottaandPaperclip 586950c9c9 Merge origin/master and renumber provider-default migration to 0301
Preserve the new master schema snapshot and use a later timestamp for the idempotent Google constraint migration.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-06 09:51:52 -05:00
DottaandPaperclip 3bb36afcd9 fix(connections): wire provider catalog setup and no-auth execution
Keep the credential form and compatible personal-account picker with their catalog definitions. Add no-auth runtime preparation and abort-aware Claude code input. Move independent Grok and Gemini live-workflow fixes to the qualification follow-up.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-06 09:51:13 -05:00
DottaandPaperclip 20740a710d test(connections): use valid protocol contracts in no-auth cases
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-06 09:48:28 -05:00
DottaandPaperclip c00dbd8ffe fix(connections): secure Grok history and repair no-auth and login cancellation
Store Grok history encrypted on the authorized task checkpoint. Restore regular files into a temporary home without a shared host symlink. Permit no-auth routes without a vault credential and keep their session identity stable. Abort pending Claude browser-code input so cancellation and timeout can dispose the login process.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-06 09:47:34 -05:00
DottaandPaperclip 16b7db35ff Shorten planning skills and measure task decomposition (#15296)
## Thinking Path

> - Paperclip manages work for AI agents.
> - Planning guidance helps agents choose owners and dependencies.
> - The runtime skill favors few tasks, but the catalog skill requires a
child-task breakdown.
> - Both add repeated process instructions that can distract from the
requested outcome.
> - This change keeps the ownership and dependency rules and removes the
required matrix and repeated checklist.
> - A bounded Product E2E comparison measures saved outcomes and task
handoffs before qualification.

## Linked Issues or Issue Description

Refs #11057. Related measurement work: #15218.

**What existing behavior does this improve?**
Planning and delegation through the runtime plan-to-tasks and bundled
task-planning skills.

**Current behavior**
The two skills contain about 1,900 words and conflicting guidance on
whether plans require child tasks.

**Proposed behavior**
Keep cohesive work with one owner. Split only for a real owner, parallel
output, dependency, independent review, or follow-up lifecycle. Preserve
existing authorization and planning mechanics.

## What Changed

- Shorten both skills to about 400 words combined. Preserve their keys
and installed-version behavior.
- Remove the duplicate operational-skill pointer and regenerate affected
source metadata.
- Add twelve explicit Product E2E cells: four scenarios with current,
short and disabled planning skills.
- Use the current task composer and actual create-response ID; calibrate
public skill APIs and browser creation without providers.
- Eliminate an observed collision in chat-test company prefixes with a
per-suite sequence.
- Grade saved documents, exact author/run attribution, child count,
prerequisite execution order, review boundaries and completion handoffs.
- Retain current skill bytes and report source, selections, run
accounting and failures.

## Verification

- `pnpm test:e2e:runner:typecheck`: pass.
- `pnpm test:e2e:runner:unit`: 1,287 Vitest tests and 128 Node checks
pass.
- `pnpm test:e2e:runner -- --list --suite plan-task-guidance`: twelve
local Codex cells.
- Archived current skills match master
`72ff3a9f27e581a27acb49771e8658bbb0bbaa47` exactly.
- Corrected fixture: three real public-API/database calibrations pass
with zero provider runs; all 35 evaluator checks and Product E2E
typecheck pass.
- Setup campaign
[37399550253](https://github.com/paperclipai/paperclip/actions/runs/37399550253)
was canceled after source review found unsupported bundled edits and
automatic core reinstallation. Its paid-cell step was skipped: zero
provider runs, no behavioral grade.
- The next setup
[37401094799](https://github.com/paperclipai/paperclip/actions/runs/37401094799)
failed before task creation on the old title-field selector: zero actual
runs, original FAIL retained, cleanup passed. A real browser/API
calibration of the new helper passes with paused non-provider agents and
zero runs.
- Full local typecheck/build pass. Full local tests retain one unchanged
five-minute Git streaming timeout (also fails isolated), 9,591 passes
and 5,796 skips. CI's chat failure was a proven random fixture-prefix
collision; five affected cases pass after the test-only repair.
- Paid behavior comparison and new-head CI/review remain pending. This
PR remains a draft.

## Risks

- The shorter text may change delegation decisions. Live outcomes are
not yet qualified.
- The initial comparison uses one profile and one attempt per cell. It
cannot establish cross-model reliability or cost trends.
- Disabled means unassigned company-owned copies; the company library
remains discoverable. This does not qualify global removal, automatic
accepted-plan wiring changes, or installed-copy migration.
- Skill availability does not prove a model read or cognitively used it.
- No provider/tool protocol, permission, timeout or runtime lifecycle
behavior changes in production.

## Model Used

OpenAI Codex (GPT-6), with repository inspection, code editing and tool
use. The exact backend model ID and context-window size are not exposed
in this session. The declared eval model is native Codex `gpt-5.6-sol`.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 09:36:01 -05:00
Nicky LeachandPaperclip 44e4979d23 Capture project and repository update lifecycle events (#15306)
## Thinking Path

> - Existing lifecycle capture records agent transitions and project
creation.
> - Project and repository edits need matching project update records.
> - The record must commit with the mutation so failed writes cannot
lose a hook.
> - Creation with repositories is one creation, and repository
replacement is one update.
> - Archive-only changes preserve project state without creating a hook.
> - Plugin consumption and provider behavior are separate work.

## Linked Issues or Issue Description

**Problem or motivation**

Project edits and repository/workspace changes lack durable lifecycle
records. A future resource plugin needs those changes captured alongside
the existing project creation hook. Archive-only changes must produce no
hook.

**Proposed solution**

Allow project `update` records in the existing lifecycle journal. Record
project and workspace mutations in their database transaction while
holding the project row lock. Suppress intermediate workspace hooks
during project creation and aggregate repository replacement.

**Alternatives considered**

Route-only hooks miss shared service callers. Recording after commit can
lose an event. Emitting a hook for each child mutation exposes
intermediate repository state.

**Roadmap alignment**

This completes project lifecycle capture begun in #15280. Plugin
delivery, VM/volume provisioning, and backfill remain separate. Searches
found no duplicate project lifecycle work; related #13306 concerns
decision events on the in-process plugin bus.

## What Changed

- Record project edits and workspace additions, updates, and removals as
project `update` events.
- Commit each event atomically with its mutation under the project row
lock.
- Keep project creation with repositories to one creation event and
repository replacement to one aggregate update.
- Ignore archive-only changes and retain workspace records.
- Extend the journal action constraint and document project update
capture.

## Verification

- `pnpm -r typecheck` passed on the narrowed scope.
- 57 tests passed across six lifecycle, project, repository, and
chat-project suites.
- Seven managed-sandbox workspace route tests and the CLI
lagging-worktree migration regression passed (65 targeted tests total).
- Full GitHub CI passed on `8ba6f97f22`; all required gates are green.
- Greptile scored the final project-only commit 5/5 with zero unresolved
review threads.
- The branch is current with `master` and has no merge conflicts.
- `git diff --check` and a local secret/PII scan passed.

## Risks

- Apply migration `0300_chunky_chamber.sql` before running the new
server. It permits project update actions and tolerates older JavaScript
worktree backups that omitted the prior CHECK constraint.
- Event-write failure intentionally rolls back the project or repository
mutation.
- Records contain identity and action; future consumers must load
current authorized project/workspace data.
- Plugin consumption, provider calls, volume cleanup, and backfill are
outside this PR.

## Model Used

OpenAI Codex, based on GPT-6, with reasoning, repository inspection,
code execution, and tool use. The exact deployment model ID and context
window are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.1006.0-canary.15
2026-10-06 07:33:18 -07:00
DottaandPaperclip 88e35f27fe feat(connections): add advanced setup and repeatable live qualification
Reuse connector rows, access controls, and agent setup components. Add persistent subscription, API key, and custom gateway choices with matching dropdowns and provider logos. Include grouped review stories and an opt-in browser campaign for local or staging targets with private credential and evidence handling.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-06 09:23:24 -05:00
DottaandPaperclip fa9b3d22a1 refactor(connections): keep runtime and browser-login changes independently usable
Keep the provider runtime PR below the review file limit. Preserve browser sign-in compatibility and provider branding here; ship advanced setup, stories, and the qualification harness in the linked UI PR.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-06 09:21:04 -05:00
1477d1ecea test: remove the no-op sequential describe modifier (#15286)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Paperclip runs the server and runner test suites with Vitest.
> - Vitest 5 removes the deprecated `describe.sequential` property, so
the pending Vitest 5 upgrade fails the type-check and test jobs.
> - `describe.sequential` only changes behaviour inside a
`describe.concurrent` suite, or when `sequence.concurrent` is on.
> - This repository has neither, so the modifier changed nothing at run
time.
> - The benefit is that the Vitest 5 upgrade can land, and the test
files lose a modifier that did no work.

## Linked Issues or Issue Description

Refs: #12969

## What Changed

- Replace every `describe.sequential` use with a plain `describe` call.
- Drop the `{ concurrent: false }` suite option from the two runner test
files.
- Add a comment to `server/vitest.config.ts` that records why these
suites must run one test at a time.
- Leave the package manifests and the lockfile unchanged.

## Why the modifier did nothing

The Vitest documentation states that `describe.sequential` is useful to
run tests in sequence inside a `describe.concurrent` suite, or with the
`--sequence.concurrent` option. `sequence.concurrent` defaults to
`false`.

This repository satisfies neither condition:

- No test file uses `describe.concurrent`, `it.concurrent`, or
`test.concurrent`.
- `server/vitest.config.ts` sets `sequence.concurrent: false`, with
`maxWorkers: 1`, `maxConcurrency: 1`, and `isolate: true`.
- `packages/paperclip-runner/vitest.config.ts` sets no `sequence` block,
so the `false` default applies.

`packages/db` and `cli` already run the same embedded-Postgres suites
with a plain `describe`, and those jobs are green. The server package
was the only outlier.

The modifier did carry one real piece of knowledge: these suites need
their tests to run one at a time. The new comment in
`server/vitest.config.ts` records that reason next to the setting that
enforces it.

## Verification

- `git grep` for `describe.sequential` returns nothing outside
`node_modules`.
- The author ran the changed server test files under the installed
Vitest 4, and the results match the results without this change.
- Two very large embedded-Postgres test files exceeded the author's
local memory limit, so the CI test jobs cover those two.
- The two changed runner test files have pre-existing local failures
caused by a missing Rust toolchain and a missing global `pnpm` binary.
The failures are identical with and without this change.
- The author type-checked the changed files and found no new error.
- CI must pass the typecheck, build, server test, and runner verify
jobs.

## Risks

- Low risk. Suite execution stays serial, because the Vitest config
enforces it.
- The change adds no dependency and changes no package manifest or
lockfile.
- A future change that turns `sequence.concurrent` on would break these
suites. The new config comment warns against it.

## Model Used

- Claude Sonnet 5 — code edits and local verification.
- OpenAI Codex, GPT-5 — the earlier revision of this branch.

## Test plan

- [x] Every CI check reaches a terminal green state. A pending or queued
check is not a pass.
- [x] The `Typecheck + Release Registry` job passes. This change must
not introduce a type error.
- [x] The `Build` job passes.
- [x] The server test jobs and the runner verify jobs pass.
- [x] Greptile re-reviews this commit set and posts a passing verdict.
The dependabot waiver does not apply to this pull request.
- [x] `mergeable` reads `MERGEABLE` as a terminal value.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have linked the related public issue with `Refs: #12969`
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-Authored-By: Priya Raman <priya.raman@paperclip.ing>

---------

Co-authored-by: Priya Raman <priya.raman@paperclip.ing>
Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Co-authored-by: nickyleach <331803+nickyleach@users.noreply.github.com>
2026-10-06 07:19:15 -07:00
Devin FoleyandPaperclip 90182b4f8b Report bounded Cloud portfolio failure diagnostics (#15340)
## Thinking Path

> - Paperclip manages work for AI agents and their companies.
> - Cloud instances proxy a trusted user's portfolio through the control
plane.
> - A failed request returns a generic error to the client.
> - That replacement error loses the failure phase and network code in
Sentry.
> - Operators need bounded evidence without upstream messages or
credentials.
> - This PR adds safe diagnostics while preserving the existing request
behavior.

## Linked Issues or Issue Description

Refs #10850, which added the portfolio proxy. No open PR for this
diagnostic gap was found.

**What happened?**

A rejected portfolio fetch becomes a generic 502 in Sentry. The event
cannot distinguish a connection reset, deadline, HTTP response failure,
or body failure. The route's previous warning also included the original
error and a stack identifier.

**Steps to reproduce**

Make the portfolio proxy's fetch reject with a TypeError whose cause has
`code: ECONNRESET`. The client correctly receives the generic 502, but
the captured replacement error loses that code.

**Expected behavior**

Keep the existing client response. Attach only bounded server-side
diagnostic fields to the failure event. Do not retry the request or
expose the original error.

## What Changed

- Add a typed portfolio error with a private frozen diagnostic record:
phase, upstream HTTP status, elapsed milliseconds, and an allowlisted
network code.
- Read at most four error/cause objects through own data properties.
Unknown codes, messages, getters, and out-of-range values do not enter
the record.
- Replace the route's raw-error warnings with safe fields. Send a plain
error plus event-local context through the existing optional Sentry
gate. Keep the route callsite and default fingerprint policy.
- Preserve authentication, trusted headers, cookies, exact HTTP error
bodies, cache behavior, the ten-second deadline, and one fetch per
request. Public responses receive no diagnostic fields.
- Document the fields and test HTTP behavior, privacy, and event
isolation with the real Sentry SDK.

## Verification

- `PAPERCLIP_REQUIRE_SENTRY_TEST_SDK=1 pnpm exec vitest run
server/src/__tests__/cloud-portfolio-error.test.ts
server/src/__tests__/cloud-routes.test.ts
server/src/__tests__/sentry.test.ts
server/src/__tests__/run-failure-sentry-real-sdk.test.ts`: 65 passed,
with the audited optional SDK installed.
- Full local `pnpm -r typecheck` and `pnpm build` passed on Node 24.21.0
and pnpm 9.15.4.
- Independent review found no blockers and independently passed all 65
tests, including the real SDK checks, on this exact commit.
- Full Linux CI passed on this exact commit and provides aggregate suite
coverage (54 successful checks, 2 intentional skips). A duplicate full
local aggregate was not run.
- The first SDK contract job failed before tests when npm could not
resolve an OpenTelemetry transitive package. A subsequent empty-cache
install first encountered a missing tarball, then succeeded after the
registry artifact became available. All 6 real-SDK tests passed against
that fresh install; the single unchanged-head CI retry passed. The SDK
pin, workflow, and dependency files are unchanged.
- The first browser shard 8 run timed out waiting for the inbox retry
reply after 45 seconds. The unchanged isolated case passed (1/1), and
the test, UI handler, fixture, and recovery files match the base commit.
The failed log contains no wakeup POST before the test's immediate
navigation; a navigation/request timing race is suspected but unproven
without a trace. The single unchanged-head shard retry passed (20
passed, 1 skipped); no timeout or source change was made.
- Greptile reviewed this exact commit at 5/5 with no unresolved review
threads.
- No live portfolio request was replayed. Route tests use controlled
local upstream responses.
- The added diff passed the secret and PII scan and `git diff --check`.

## Risks

This is a diagnostic change. It does not identify or repair the origin
of a connection reset. Unknown transport failures remain `unknown`.
Elapsed values outside 0–60,000 ms become null. Default Sentry
fingerprinting remains enabled; exact historical group membership is not
guaranteed. No retry, migration, deployment, or configuration change is
included.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository editing, code
execution, and independent agent review. The exact deployment model ID
and context window are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 06:58:36 -07:00
DottaandPaperclip ab56046a8e refactor(connections): split UI and qualification into follow-up PR
Keep the provider and runtime implementation in this PR. Preserve the completed interface, review stories and live browser qualification on a linked follow-up branch for a bounded review.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-06 08:44:40 -05:00
Dotta 7cbd61349a Merge remote-tracking branch 'origin/master' into codex/provider-routing-connections
* origin/master:
  fix(tasks): surface Codex ChatGPT model rejection (#15299)
  fix(runner): continue restart-interrupted Codex turns (#15297)
  fix(agents): grant configuration access by default (#15283)
  Guard routine UUID lookups without narrowing valid inputs (#15313)
  fix(ssh): transport project repositories as their own git checkouts (#14782)
  feat(exe-dev): copy a source VM with exe.dev cp (#14975)
  Diagnose native model rejections and preserve repairable reviews (#15304)
  Verify native semantic input against its raw wire digest (#15301)
  Retry pooled execution projection reads after disconnects (#15300)
  Fix browser polling for unsaved agent chats (#15298)
  test(ui): keep initial reasoning with the saved reply (#15107)
  docs: add a product feature map (#15288)
2026-10-06 08:42:29 -05:00
Devin FoleyandPaperclip 3290d97417 Scope the release smoke opening question to its card (#15339)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The release lane checks onboarding before it promotes a nightly
candidate.
> - The first task shows an unanswered-question summary and an open
question card.
> - Both display the same prompt, so the smoke test's page-wide text
locator fails strict matching.
> - This pull request scopes that assertion to the open interaction
card.
> - The smoke still requires the actual prompt and verifies that
onboarding does not start an agent run.

## Linked Issues or Issue Description

**What happened?**

[Scheduled release run
37441718130](https://github.com/paperclipai/paperclip/actions/runs/37441718130)
failed `smoke_nightly / smoke` on both attempts. The page-wide
`getByText("What would you like to do?")` matched two elements: the
timeline summary and the open question card. This skipped
`publish_nightly`.

**Expected behavior**

The test confirms that the seeded task's open question card displays the
exact prompt. A summary row alone must not satisfy that check.

**Steps to reproduce**

Run the Docker onboarding smoke against published candidate
`2026.1006.0-canary.11`, then run the release-smoke Playwright spec. The
old assertion fails after the greeting appears.

**Paperclip version or commit**

Release workflow commit `f858207161ba29c01c82f4674aef83d91b74480f`;
published candidate `2026.1006.0-canary.11`. The assertion is unchanged
on the current master base `9b3fe260bac576d622ffdfefb923db73cc5273f7`.

**Deployment mode**

Docker smoke harness, authenticated/private, with its existing mock
provider.

Related: #13166 added the first-task chat assertion. I also reviewed
open #12316, which fixes a separate bootstrap race and does not change
this selector. Searches found no duplicate selector fix or public issue.

## What Changed

- Scope the exact opening-prompt text to `task-chat-interaction`.
- Explain why the timeline summary cannot satisfy the assertion.
- Preserve the rest of the smoke flow, including the 15-second check for
no agent runs.

## Verification

- Local Docker harness ran the exact failed published candidate,
`2026.1006.0-canary.11`, with the existing mock provider.
- The old spec reproduced the same two-element strict-mode error in
Chrome.
- The fixed spec passed the full authenticated onboarding flow and the
15-second no-run check: 1/1, 32.6 seconds. The executed spec copy was
byte-identical to the changed repository file; only artifact output
paths and the local port were overridden.
- Independent Chrome checks: the scoped locator passes with both summary
and card present. Summary-only, wrong prompt, hidden prompt, and
duplicate active prompts each fail as intended.
- `pnpm typecheck` and `pnpm build` passed on Node 24.21.0. Canonical
Linux CI passed all 55 checks (53 success, 2 intentional Storybook
skips), including the full aggregate tests, build, typecheck, all eight
E2E shards, and post-ready security scan.
- Greptile scored the exact head
`5a30e667396600a052a4041c819652c19b8c68ff` at 5/5 with no actionable
issues. Independent review found no issues; there are no review threads.
- `git diff --check` passed. The diff and PR text were scanned for
secrets and private identifiers.

## Risks

Low risk: one assertion in a release smoke test changes. A missing or
hidden prompt still fails, and multiple matching prompts in interaction
cards still fail strict mode. This does not change the product, release
selection, or publishing. The scheduled release lane still needs its
next normal successful run to prove nightly recovery.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, shell tools, and an
independent Codex review. The exact serving model identifier and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 06:42:21 -07:00
DottaandPaperclip 4b1caf043e test(connections): add repeatable live provider qualification
Exercise Apps and new-agent setup with real credentials, isolated targets, browser login handoffs, independently verified artifacts and follow-ups. Preserve safe failure diagnostics, private evidence, cleanup, provenance and incomplete coverage.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-06 08:42:03 -05:00
DottaandPaperclip 01e756bc61 fix(connections): preserve provider workspaces and follow-up context
Keep selected model and skill paths current across disposable connection homes. Preserve private Grok history and assigned OpenCode and Hermes workspaces. Fix artifact helper paths and cover the boundaries with regression checks.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-06 08:41:41 -05:00
DottaandPaperclip 9b3fe260ba fix(tasks): surface Codex ChatGPT model rejection (#15299)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for work.
> - Native Runner runs report provider errors and can also save a structured failure result.
> - Codex rejects a model that the user's ChatGPT account cannot use.
> - Master now diagnoses this rejection, but the task's compact run data omits its message. The thread can still call it a generic run failure.
> - Users need the account restriction and a clear step to repair the model selection.
> - This pull request shows the existing diagnosis as “Model unavailable” on the task.

## Linked Issues or Issue Description

**What happened?**

Codex returns HTTP 400 with `invalid_request_error` and the message `The 'gpt-6.1-sol' model is not supported when using Codex with a ChatGPT account.` Master now stores an actionable diagnosis for this rejection. The task thread still labels it “Run failed” and does not receive its error text in compact run data.

**Expected behavior**

Show the account restriction on the task and in the run error. Tell the user to choose a supported model or clear the task's model override before retrying.

**Steps to reproduce**

1. Use a Native Runner Codex agent signed in with a ChatGPT account.
2. Select a model that produces the rejection above and start a task.
3. Let the runner save its generic failed result. Inspect the task's failure marker and recovery notice.

**Additional context**

Refs: #15304. That merged PR diagnoses the provider failure and preserves worker and review recovery rules. This PR adds its task-facing message and guidance without changing that diagnosis or those rules.

Refs: #13134. That PR improves model discovery for ChatGPT accounts. This PR exposes the rejection when a configured model still fails at execution time.

## What Changed

- Return a bounded model rejection message in compact issue-run data.
- Show “Model unavailable” and model-change guidance in the task thread and recovery notice.
- Add database and UI regression tests, including both the original rejection text and master's fixed diagnosis. Document the new failure message.

## Verification

- Focused merged-branch validation: 279 tests passed across activity service, native provider failure observation and PostgreSQL integration, TaskChatThread, and ExecutionBlockerNotice.
- `pnpm -r typecheck`, `pnpm build`, and `pnpm check:token-gates` passed on the merged source tree.
- Greptile reviewed conflict-resolution commit `1db39c4e636a92e2e76ba224b71c6f1e2f556c81` at 5/5 with no findings or inline comments.
- All 55 check runs completed without failure on that commit. The two optional Storybook jobs were skipped. The legacy Snyk status passed. The branch has no merge conflicts.
- UI regression assertion: the task's failure marker says “Model unavailable” and retains the account restriction. It no longer says that this failure happened after a final response.

## Risks

- Provider recognition and recovery are owned by the existing master implementation. This PR exposes only bounded error text for failed runs with `native_provider_model_rejected`.
- Historical runs with the generic `adapter_failed` code are not reclassified. No stored run is rewritten.
- Existing retry and reconciliation gates remain in place. No schema migration or model configuration change is required.

## Model Used

OpenAI Codex, based on GPT-6, with reasoning, code execution, and browser inspection. The exact deployment model ID and context window are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context to this change
- [x] I have specified the model used (with version and capability details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before requesting merge

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-06 08:05:25 -05:00
DottaandPaperclip 63f3aa2dbf fix(runner): continue restart-interrupted Codex turns (#15297)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Native runs retain their provider conversation across server
restarts.
> - A dead runner can restore a Codex conversation after its active turn
is lost.
> - The old recovery path synthesized a failed task result from that
interruption.
> - The task then required operator action even though its conversation
and workspace were available.
> - This pull request preserves the interruption cause and uses the
admitted restart attempt for one continuation in the same conversation.
> - The agent can reconcile unfinished actions and complete the current
request without resending the original task.

## Linked Issues or Issue Description

Related: #12845 added native restart recovery. #15042 covers admission
during shutdown. #14796 covers legacy shutdown recovery. This change
covers a lost native Codex turn after successful conversation
restoration.

**What happened?**

After a server restart killed the local runner, Paperclip restored the
saved Codex thread. Runnerd found that the old turn was no longer
active. It synthesized a failed terminal and a needs-review result from
the last progress message. The transport also discarded the terminal
error when it reconstructed thread history. The task failed instead of
continuing.

**Expected behavior**

After proving that the old process stopped and admitting a bounded
recovery attempt, resume the current request in the same conversation.
Preserve the workspace. Inspect unfinished actions before proceeding.
Keep real provider failures, accepted results, intentional stops,
unknown unreconciled effects, and exhausted attempts subject to their
existing rules.

**Steps to reproduce**

1. Start a local native Codex run and leave its turn active.
2. Kill the isolated runner and provider processes, as can happen during
a server restart.
3. Restore the same provider thread with no active turn.
4. Observe the synthetic task failure. The new real-process regression
reproduces this boundary with a scripted provider.

## What Changed

- Record an explicit recoverable process-loss cause without inventing a
task result.
- Preserve terminal errors and prior turns in reconstructed provider
history. Recover the authoritative saved result when adopting an
accepted continuation.
- Send one continuation in the same conversation for an admitted
dead-runner recovery. Require reconciliation of unfinished commands and
external actions.
- Persist the interrupted terminal before submission and retain the
existing recovery marker across another controller loss.
- Keep provider attempt limits, terminal failures, and intentional
cancellation behavior.
- Add red/green regressions, real process-kill coverage, restart
checkpoint coverage, and retry-budget coverage. Document the behavior
and run-log evidence.

## Verification

- Red: the new native runtime regression rejected with
`NativeProviderTerminalFailure` on the original code; the Rust restore
regression found a missing recovery cause.
- Red/green: if restoring the conversation fails and replacement is
allowed, the replacement receives the full task and fresh-session
handoff. Both prepared and legacy execution inputs are covered.
- Red: a second controller crash after the provider accepted the
continuation caused an extra `turn/start`. The regression now proves
there are exactly two submissions total: the original and its
continuation.
- Green: focused runtime, backend, driver recovery, and real-process
restart suites (207 tests). After the final history/result changes,
driver recovery and real-process restart suites passed again (36 tests).
- Green: complete Codex transport suite (186 tests), server restart
classification/database integration suites (34 tests), and Rust Codex
provider suite (92 passed, 2 ignored).
- Full `pnpm -r typecheck` and `pnpm build` passed on `60141e649`. The
subsequent replacement-prompt guard passed the Runner TypeScript check
and the complete runtime plus process-restart suites (148 tests).
- Local full-suite attempt: `pnpm test:run` reported two failures in the
untouched chat integration suite. Both passed individually, and the
complete chat suite passed on rerun (1,063 tests). After all remote test
shards passed, the duplicate serial local run was stopped with SIGINT;
it is not claimed as a full local-suite pass.
- Latest-head CI (`b1297dcd4`): 55 successful checks and 4 intentionally
skipped checks, including all test shards, typecheck, build, native
Runner verification, end-to-end tests, and canary dry run. [CI
run](https://github.com/paperclipai/paperclip/actions/runs/37405346303).
- Greptile: 5/5 on the latest head, with no unresolved review threads.
The PR is mergeable.
- The process tests use the real runner binary and a scripted Codex
provider. They do not call a live model service.

## Risks

- This changes local Codex recovery after process loss. A continuation
can execute more work in the retained conversation. Its prompt requires
state inspection before repeating an uncertain action; the system does
not replay tool calls.
- Recovery shares the existing three-attempt budget and one-shot
continuation marker. Real failures and older unmarked failed checkpoints
are not reopened.
- No database migration or API change is required.

## Model Used

- OpenAI Codex, based on GPT-6. The exact model ID and context-window
size are not exposed in this session.
- Capabilities: reasoning, source inspection, tool use, code editing,
code execution, and test analysis.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.1006.0-canary.13
2026-10-06 07:49:47 -05:00
DottaandPaperclip a1ab55a56d fix(agents): grant configuration access by default (#15283)
## Thinking Path

> - Paperclip manages agent work and permissions for a company.
> - Agent setup can require one agent to configure another agent.
> - New standard agents could create agents, but they had no direct
configuration grant.
> - Existing agents must keep their current permissions after an
upgrade.
> - The requested new-agent defaults include 13 more direct permissions
for suggestions, skills, tools, audit, inbox, and task assignment.
> - This pull request adds the 14-grant set at creation and approval
activation, retains it through invitations, and keeps existing agents
unchanged.
> - The change keeps low trust and built-in agents on their narrower
permissions.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

The agent creation and approval flows, and the agent configuration
authorization path.

**Current behavior**

A new standard agent can create an agent. It cannot make protected
changes to a peer agent unless an operator adds an `agents:configure`
grant.

**Proposed behavior**

Add direct grants for `agents:configure`, `agents:suggest-changes`,
`skills:create`, `skills:suggest-changes`, `tools:manage_connections`,
`tools:manage_profiles`, `tools:view_audit`, `audit:view_agent_actions`,
`tools:use`, `tools:manage_runtime`, `inbox:manage`, `tasks:assign`,
`tasks:assign_scope`, and `tasks:manage_active_checkouts` to new
standard agents. Scope `tasks:assign_scope` to the agent’s reporting
subtree. Keep existing agents and their grants unchanged.

**Reason and benefit**

New standard agents can complete agent setup and the requested tool,
skill, inbox, and task workflows under existing route, scope, and
approval checks.

**Breaking changes**

Existing agents keep their current permissions. There is no permission
migration.

Related PR: #12212 adds scoped grant routes. This PR changes the default
grant.

## What Changed

- Add the 14 requested direct grants in agent creation and approval
activation. Keep them when invitation approval replaces grants,
preserving explicit scopes. Remove them when an agent is deleted.
- Apply defaults only to new agents. Remove the existing-agent
permission migration, its snapshot, and its journal entry. Preserve
existing scoped grants.
- Add permission, scope, invitation, and existing-agent regression
tests. Document all default and excluded permissions.
- Prevent agent keys from creating, changing, or restoring host-executed
process or local adapter command settings, and from restoring workspace
commands through rollback.

## Verification

- Run `node_modules/.bin/vitest run
server/src/__tests__/agent-default-configure-grants.test.ts
server/src/__tests__/invite-join-grants.test.ts
packages/db/src/migration-snapshot-drift.test.ts`. All 15 tests pass.
These cover new-agent defaults, unchanged existing grants, pending
approval, invitation grants, and migration history.
- Run `node_modules/.bin/tsc -p server/tsconfig.json --noEmit`. It
passes.
- Run `packages/db/node_modules/.bin/tsx
packages/db/src/check-migration-numbering.ts` and
`packages/db/node_modules/.bin/tsx
packages/db/src/check-migration-safety.ts`. Both pass.
- Confirm that this PR has no files under `packages/db/src/migrations/`
in its final diff.
- GitHub CI at `e8173b2c41` passes build, typecheck, server suites,
browser shards, and canary verification. One unchanged OpenCode
transport test reached its five-second timeout in the first Runner shard
run. That test passes locally in 2.55 seconds. The shard passed on one
retry. All 54 checks pass, with two expected skips.
- Greptile gives this exact head 5/5. Security review passes. The PR has
no unresolved review threads or merge conflicts.

## Risks

- The 14 default permissions apply only to new standard agents. Existing
permissions stay unchanged. Low trust and managed built-in agents are
excluded.
- New grants include connection, runtime, and active-checkout
management. Existing company, responsible-user, scope, and approval
checks remain in force. The default inbox grant carries a
responsible-user-only scope so it cannot override another user's inbox
settings.
- Protected changes still need the responsible user's authority when
that check applies. Company boundaries and approval gates still apply.
- Agent-authenticated requests cannot configure process adapters or host
command settings on local adapters. Known provider credential references
remain allowed, as do narrow plain authentication overrides on new peers
when the caller uses an AI connection pool. Arbitrary environment
settings remain blocked. Board operators retain the host configuration
paths.

> I checked `ROADMAP.md`. This is a narrow fix to existing agent
configuration behavior.

## Model Used

- OpenAI Codex, GPT-6 series. This runtime did not expose its exact
hosted model ID or context window. This revision uses OpenAI GPT-6
through Codex with tool use, code execution, and tests. The runtime does
not expose the exact hosted model ID or context window.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.1006.0-canary.12
2026-10-06 05:28:48 -05:00