Commit Graph
403 Commits
Author SHA1 Message Date
DottaandPaperclip 3d640b4ce3 Bind consolidated Pi cold-start qualification source
Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-02 10:39:20 -05:00
DottaandPaperclip 353ea98b81 Measure retained bba Intel startup phases
Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-02 09:09:30 -05:00
DottaandPaperclip 903242be1b Exclude retained Intel checks from EC2 fallback
Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-02 08:27:49 -05:00
DottaandPaperclip 2617947406 Verify canonical Pi13 startup on exact retained Intel build
Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-02 08:14:20 -05:00
DottaandPaperclip 0f48b6d6c5 Bound normal Intel qualification process inspection by lifecycle phase
Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-02 07:54:05 -05:00
DottaandPaperclip f07200505b Bind fresh Pi13 verification to final Darwin launch repair
Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-02 07:00:00 -05:00
DottaandPaperclip 4ba81672ae ci: evaluate isolated Pi admission overlap
Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-02 05:31:16 -05:00
DottaandPaperclip fed3c8b9ba ci: time retained Pi13 Intel startup phases
Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-02 04:23:49 -05:00
DottaandPaperclip 94848f0415 ci: verify frozen Pi 13 with fresh native Intel build
Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-02 03:28:42 -05:00
DottaandPaperclip 84b3a9ea6b ci: retain exact staged daemon holder evidence
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-02 02:42:46 -05:00
DottaandPaperclip 2ebf3795eb ci(runner): verify normal Pi with audited native daemon
Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-02 00:40:43 -05:00
DottaandPaperclip 4286d26d05 ci(runner): compare retained and streaming snapshots
Add a bounded native Intel ABBA diagnostic using the exact retained baseline and separately compiled unpublished streaming candidate. Keep full output integrity, process ownership, and diagnostic-only attribution.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-02 00:29:34 -05:00
DottaandPaperclip 3f23a76698 ci(runner): add exact snapshot pool comparison
Measure four retained B9 snapshots with process-start pool4/16 ABBA, full sealed-output integrity and bounded owned cleanup. Keep diagnostic evidence distinct from native startup qualification.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-01 23:47:38 -05:00
DottaandPaperclip 2c5a21762e ci(runner): verify normal qualified Pi installation source
Pin normal Intel verification to c9 and require the default Pi-only qualified provider pack. Preserve the historical baseline diagnostic and all startup deadlines.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-01 23:32:33 -05:00
DottaandPaperclip 07af1ef5b0 ci: diagnose exact Pi12 Intel startup and retain final integrity
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-01 23:17:26 -05:00
DottaandPaperclip 9541eb1878 ci(runner): build fresh Pi12 Intel daemon from frozen source
Pin the final source and profile, restore fresh Rust compilation with SDK provenance, and keep manual immutable-source checks independent.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-01 22:20:22 -05:00
DottaandPaperclip 2b3924f91d ci: diagnose retained W native Intel startup
Bind the existing bounded diagnostic to the exact failed W pack and add sparse native-manifest and layout timings without changing startup assertions, deadlines, or cleanup.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-01 21:19:07 -05:00
DottaandPaperclip 31e7bc15b0 ci: bound complete Runner evidence and verify W source
Admit complete source and Runner artifacts before upload, stage the exact Runner evidence closure without following directory symlinks, and audit immutable source and lock inputs after full source checks. Bind native Intel preparation to the reviewed W revision while preserving verification commands and deadlines.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-01 20:19:06 -05:00
DottaandPaperclip 1e57cf95c1 ci: qualify Pi single-pass startup on native Intel
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-01 20:01:39 -05:00
DottaandPaperclip 05f5c9a4e0 fix(ci): bound diagnostic process inspection within phase deadlines
Reduce diagnostic polling pressure while retaining fatal inspection failures and exact process cleanup. Keep shared inspection defaults and original startup assertions unchanged.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-01 19:19:52 -05:00
DottaandPaperclip 202cd9a222 ci: diagnose retained Pi startup phases on native Intel
Reuse the exact failed qualification pack and daemon, retaining an additive sidecar timing delta and unchanged no-key startup assertions. Bound complete diagnostic evidence to 256 MiB for seven days.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-01 18:54:51 -05:00
DottaandPaperclip 9b7a91e3fc ci(runner): reuse pinned Intel daemon for cold-start proof
Build current TypeScript and provider pack while admitting the exact retained native daemon through source-input equality and original compiler provenance. Keep original startup assertions and evidence bounds.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-01 18:12:23 -05:00
DottaandPaperclip d0e835100d ci(runner): add bounded hosted E2E support recheck
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-01 17:45:19 -05:00
DottaandPaperclip b58907578e ci(runner): verify closed pack hardlinks and asset presence
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-01 17:10:34 -05:00
DottaandPaperclip 39ca241d26 ci(runner): preserve rustup multicall invocation name
Keep the rustup PATH alias instead of invoking its rustup-init target directly. Audit resolved target bytes and recheck them before invocation; add a regression reproducing symlink-based CLI dispatch.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-01 16:32:42 -05:00
DottaandPaperclip 2f90cb31ae ci(runner): retain Intel pack and stop on uncertain cleanup
Archive and verify the complete tested pack before startup checks. Make cleanup uncertainty sticky across test cases, retain scratch after interrupted launches, and enforce the reviewed 16 GiB seven-day artifact storage bound without truncation.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-01 16:21:11 -05:00
DottaandPaperclip d472a3c95e ci(runner): verify frozen Pi on native Intel macOS
Add an exclusive manual no-provider lane that rebuilds the reviewed Pi 1.0 source and lock closure, checks native Intel execution, and retains unchanged startup/SDK test evidence with owned cleanup.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-01 16:12:38 -05:00
DottaandPaperclip 3b4b270650 fix(adapters): preserve ACP terminal failure diagnostics (#14573)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The shared ACP adapter engine records agent failures for operators.
> - ACP providers can report a failure category, title, and detailed
cause.
> - Our patch kept only the category in the saved error, so an operator
could not diagnose a failure when tracing was off.
> - This pull request preserves redacted provider diagnostics in the run
error, transcript, and structured run result.
> - Operators can now inspect the provider message and any supplied
request ID or stack trace after the run ends.

## Linked Issues or Issue Description

Refs #13889 (the diagnostic gap; this PR does not update the bundled
Claude version).
Refs #14484 (related model-refusal classification; this PR retains
diagnostics for all terminal failure categories).

**What happened?**
An ACP turn failed with only `ACP agent reported a terminal service
failure.` The provider's title and details were available in memory but
absent from the saved error and transcript.

**Expected behavior**
The run retains useful provider diagnostics even when raw tracing is
disabled. Credentials remain redacted. A size limit must report
truncation instead of silently removing the cause.

**Steps to reproduce**
1. Run an ACP agent that returns an error-severity typed session
failure.
2. Include an HTTP error, request ID, and stack text in its title and
details.
3. Inspect the failed run with tracing disabled. Before this change,
only the category survives.

## What Changed

- Both pinned ACPX patches pass complete error text to the in-memory
callback, so redaction happens before truncation.
- The shared engine retains the sanitized category, title, and details
in `resultJson.terminalSessionFailure` and includes the text in the run
error and error transcript.
- Diagnostics redact configured environment values even under arbitrary
names, unknown launch-environment values, connection URL passwords, run
credentials, and common credential syntax. Known boolean settings remain
readable, while credential values are redacted even when embedded in
other text. Diagnostics remove control characters and invalid Unicode.
- Title and detail limits keep escaped transcript JSON below the
server's chunk limit. Truncated fields include an omission count. The
safe run-result projection preserves a byte-bounded diagnostic preview
when the result exceeds its byte budget, with an explicit pointer to the
full adapter-bounded run error and transcript.
- The existing UI and CLI display the error. Diagnostics do not become
assistant output. Issue continuation summaries and session-compaction
prompts receive only the generic category, preventing provider text from
becoming handoff instructions. Existing quota classification, warnings,
timeout precedence, and control-channel failure precedence remain in
place.
- Regression tests cover real ACP child processes with both pinned
versions in one-shot and persistent modes, credential redaction, request
IDs after the old 4 KiB cutoff, transcript parsing, storage bounds, and
database retrieval of oversized multibyte diagnostics.

## Verification

- Full CI on `20ad4f5f1f66c46d2c260e6ad0339cbea607b4cf`: **54 passed, 2
intentionally skipped, no pending or failing checks**. Includes
typechecking, build, all Vitest shards, Runner checks, browser E2E, and
the canary packaging/public-install dry run.
- Greptile: **5/5** on this commit. Superagent security scan passes. All
review threads are resolved.
- Local verification passed: shared ACP engine suite (395 tests); real
Claude ACP child-process and diagnostic regressions across both pinned
runtimes and both execution modes; run retrieval and model-handoff
regressions (59 tests); ACPX patch packaging (16 tests); full typecheck
and build. Affected package typechecks and focused tests were rerun
after review fixes.
- The broad local `pnpm test:run` was stopped after review edits made
its cached imports stale. Fresh targeted runs pass, including both
affected server suites. Cold-build import failures were also rerun after
dependency builds: chat integration (1,063 tests) and tool access (351
tests) pass. The final commit's complete CI matrix is green.

## Risks

- Provider diagnostic text is untrusted. This change retains more of it
in company-scoped run records. Redaction and size bounds apply before
persistence.
- Diagnostics are limited to fields the provider supplies. Old runs
cannot recover discarded error text.
- No schema migration, recovery-policy change, or new Telemetry or
OpenTelemetry export.

## Model Used

- OpenAI GPT-6 through Codex, with reasoning, repository inspection,
code editing, and test execution. The exact serving model ID and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-29 10:18:30 -05:00
DottaandPaperclip 24beb00575 feat(runner): add rich ACP transport and durable interaction foundation (#14430)
Add shared rich ACP transport, durable questions and permissions, verified provider packaging, and bounded activity and plan presentation. Keep Cursor, Copilot, and Pi pending their separate provider qualification.

Persist interaction settlement before publication, fence failed writes until fresh recovery, and preserve owned-process cleanup. Incorporate reviewed mainline integration with extended harness coverage.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-29 08:56:21 -05:00
0be2afcca6 feat(ui): improve task composer controls and pending input (#14322)
## Thinking Path

> - Paperclip lets operators assign tasks to AI agents and review their
work.
> - The task composer controls the next message and its assigned agent.
> - Operators needed a way to choose that agent's model and effort
without leaving the composer.
> - The old mode selector, upload button, and input cards made the
mobile composer crowded and hid normal messaging during a pending
decision.
> - Harnesses publish different model and effort capabilities, so the
picker must follow the selected agent.
> - This pull request adds one responsive composer flow, keeps pending
cards visible above it, and protects Codex ACP authentication in the
local test path.
> - Operators can choose run settings, send a message, and answer a
pending card as separate actions.

## Linked Issues or Issue Description

**Subsystem affected**

Task composer UI, issue thread interactions, Codex ACP credential
handling, and Storybook.

**Problem or motivation**

The composer did not expose model or effort for the selected agent.
Mobile actions wrapped poorly. Pending questions and confirmations
replaced the composer. A local Codex ACP test could also reuse host
authentication after the managed key was removed.

**Proposed solution**

Put assignee search, model search, exact model IDs, effort, and fast
mode in one picker. Use a mobile dialog. Replace the direct-upload plus
action and separate mode selector with an Add menu and removable Plan or
Ask chips. Place pending interaction cards above the usable composer.
Keep these cards pending after an ordinary message unless their creator
asks for comment superseding. Replace managed ACP auth files atomically
and isolate the test key from host credentials.

**Roadmap alignment**

ROADMAP.md does not list an overlapping composer milestone. This change
improves the existing task and review flows.

## What Changed

- Added the combined assignee, model, and effort picker to both task
composers. Search matches agent name, role, and harness. The server uses
a curated Codex list by default and honors instance-declared models.
Manual IDs remain available.
- Added an effort slider for known model capabilities, a conditional
Codex fast control, and reset. The picker opens in a modal on mobile.
- Added the Add menu for files, supported goals, Plan mode, and Ask
mode. Plan and Ask are exclusive removable chips. Keyboard mode cycling
remains available.
- Adjusted mobile spacing, avatars, wrapping, and Send placement.
Removed the composer divider.
- Moved pending question, confirmation, review, and related cards above
the composer. Ordinary comments now leave question and confirmation
cards pending by default. The onboarding prompt retains explicit comment
superseding.
- Updated the Storybook composer group with responsive states and the
production picker. Added UI, service, route, and browser regression
coverage.
- Isolated Codex ACP API-key authentication, skipped subscription auth
merge and shared-home copy-back for remote API-key runs, and replaced
the managed auth file atomically.

## Verification

- `pnpm -r typecheck` — passed on the final local head.
- `pnpm check:token-gates` — passed on the final local head.
- `pnpm exec vitest run server/src/__tests__/adapter-models.test.ts
ui/src/components/task-chat/ComposerRunSettingsPicker.test.tsx` — 31
tests passed, including role and harness search, declared Codex models,
and filtering general OpenAI models.
- `pnpm exec vitest run
server/src/__tests__/issue-thread-interactions-service.test.ts` — 74
tests passed.
- `pnpm exec vitest run
packages/adapters/codex-local/src/server/acp.test.ts` — 42 tests passed,
including remote API-key copy-back isolation.
- `pnpm test:run` — attempted locally; the embedded PostgreSQL test
database could not initialize on macOS. The isolated
`heartbeat-run-event-sequencing` suite reproduced that environment
failure. GitHub CI runs the full test matrix for this head.
- `pnpm build` — passed on the final head. `pnpm build-storybook` passed
after the last UI change; only server code, tests, and docs changed
afterward.
- Live local test drive — Codex ACP ran a task with a managed API key.
The test agent was restored to its default ACP configuration afterward.
- Review the interactive stories under the top-level Composer group with
`pnpm storybook`. Check a narrow desktop width and mobile Plan, Ask,
picker, and pending-question states.

## Risks

- A pending card stays open when an ordinary comment changes the
discussion. Its creator can set `supersedeOnUserComment: true` when a
new comment should replace it.
- Model and effort overrides persist on the task until reset or changed.
An unlisted manual model ID may fail when the provider runs it.
- Some harness catalogs do not report effort support. The picker hides
effort for those models.
- No database migration is required.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

OpenAI GPT-6 via Codex. This runtime does not expose the exact model ID
or context window to the task. The model used code execution and browser
tools.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: OpenAI Codex <codex@openai.com>
2026-09-28 23:02:39 +00:00
DottaandPaperclip 18e8c121d9 fix(runner): include Grok support in public installs with sandbox prerequisites (#14024)
## Thinking Path

> - Paperclip manages agents through a shared native runner.
> - Built-in harness support should ship with Paperclip's public
distribution.
> - Grok already speaks ACP; it does not require a new public bridge
package.
> - Sandbox provisioning owns the native executable and its pinned
version.
> - The runner must verify that prerequisite without downloading it
during npm installation.
> - This change separates built-in launcher identity from external
runtime identity.
> - Clean npm installation and live staging checks verify the
distribution boundary.

## Linked Issues or Issue Description

Refs #13882, #13973, #13977, #13979.

This follow-up now targets master after #13882 was squash-merged. It
replaces the private `@paperclipai/grok-acp` workspace package with
runner-owned assets. Current master is included so the branch also
contains the merged scheduler, complete-event capture, and durable
cleanup fixes.

## What Changed

- Ship Grok launcher and qualification metadata inside the runner's
compiled output and the public server's vendored runner tree.
- Remove the separate Grok npm package and all package-manager install
hooks for this runtime.
- Require the checksum-verified Grok Build 1.0.13 binary at
`/opt/paperclip/providers/grok/1.0.13/grok` in the selected execution
environment. Provision it explicitly in the Daytona image and CI setup.
- Keep native binaries outside the provider pack. Bind the built-in
launcher into the pack manifest.
- Preserve executable leases, descriptor-backed startup, credential
fences, permissions, and exact ACP model admission.
- Use `builtin:grok-acp` and `native:grok` as profile identities.
Historical package-profile sessions fail closed on resume rather than
being silently reinterpreted.
- Resolve built-in assets from the authenticated sidecar location,
including public server npm layouts. Keep the controller path out of
provider environments.
- Add clean npm tarball installation verification to the existing
trusted canary CI job and the admitted manual EC2 verification path. It
stages a unified release version and runs npm lifecycle scripts, then
verifies missing-prerequisite rejection and admission after separate
provisioning without credentials or inference.
- Include the controller-owned provider pack in stamped Cloud images.
Unstamped local images omit the pack and remain usable; remote ACPX
requires full source provenance.
- Correct CLI approval-page metadata for an already authenticated Cloud
board user; approval authorization remains unchanged.
- Honor explicit native-runner enablement in the Cloud agent picker and
direct setup page, keeping the flag disabled by default.
- Allow selecting the execution environment before connecting
credentials. Include Grok in the existing authenticated hello-probe
flow, targeting its pinned native prerequisite for runner setup.
- Recover an existing subscription sign-in conflict through an explicit
cancel-and-retry action, serialized after cancellation succeeds.
- Preserve the selected ACPX harness before normalizing config fields,
so new Grok agents use the Grok default model.
- Keep the credential-free Cloud provider pack root-owned and readable
after runtime UID remapping; verify manifest and referenced asset access
under an unrelated unprivileged UID during image builds.
- Archive prior failover backups alongside explicitly replaced harness
state, preserving evidence while preventing stale backups from blocking
a fresh replacement.
- Update Daytona image content inputs and contract tests for the
built-in assets and explicit provisioner.
- Document and regression-test the shared `approve-all` default for Grok
setup, saved configuration, and native execution. Explicitly saved
restrictions remain unchanged.

## Verification

Current merge-repair head `df09eb3e1a619430ad8419a0ee9aedd486689b05`
incorporates master `f1a394bd30cb56fb9e479f98b9f50176fe921858` after the
base PR was squash-merged. All 12 conflicts came from incoming files
identical to the tested pre-squash base. The final tree exactly matches
a three-way merge using that original base, preserving built-in Grok
distribution and removal of the obsolete private package. All 252
focused runner/UI tests, six npm-isolation tests, and token gates pass.
Fresh exact-head Greptile review is 5/5 with no outstanding findings;
security scans and EC2 native compilation pass. All current-head CI is
green: 56 successful checks/statuses and four intentional skips ([run
36468768035](https://github.com/paperclipai/paperclip/actions/runs/36468768035)).
The repository owner explicitly authorized bypassing code-owner approval
after all checks passed; no CI checks or repository protection settings
are bypassed or changed. The only remaining PR was removed from the
completed stack metadata to permit native auto-merge.

Earlier integration head `78cb306ecc41b5c96577c26c1d89153b0ef865a1`
includes master `3447609d2247e75e55d91493dda91a608364f672` (2026-09-28).
Two master advances during verification overlapped the eval catalog; the
final merge preserves Grok qualification, completion updates, and
bounded API-response reading in all 348 cells. All 77 focused
catalog/eval/workflow tests pass. Both native stack layers (#14397) are
mergeable, and both exact-head Greptile reviews are 5/5 with successful
security scans and no unresolved review threads. All current-head CI is
green: 56 successful checks/statuses and four intentional skips ([CI
attempts](https://github.com/paperclipai/paperclip/actions/runs/36447124691)).
The initial attempt lost two EC2 runners to shutdown signals and stalled
a third shard during dependency preparation; all three passed the
same-commit failed-job-only retry. Trunk code-owner requirements remain
enforced. The review summary’s non-blocking saved-asset offset
classification note concerns code already merged in #14301; those
runtime files are identical to master and outside this stack’s diff.
Historical live evidence below retains its original source revisions.
[Final public npm
verification](https://github.com/paperclipai/paperclip/actions/runs/36445542764)
passed on `76ea70cd4d13786a042af9df82f0fd7a8c85ae30`: 17 public
packages, an executed offline lifecycle sentinel, unchanged consumer
lock, built-in launcher, missing-prerequisite rejection, and verified
separately provisioned binary/command lease. Provisioning and cleanup
require no host privilege elevation; only the positive probe mounts the
temporary native binary read-only. The verifier is unchanged by the
final master merge. All six isolation tests and an offline npm smoke
test pass. The prior head had 56 green CI checks and a 5/5 review after
two unchanged tests timed out and passed a failed-job-only retry ([CI
attempts](https://github.com/paperclipai/paperclip/actions/runs/36444597313)).
All 56 recovery-display/lineage tests pass; re-review cleared the
already-covered missed-retry concern. Earlier EC2 failures remain
retained: [npm lockfile
rejection](https://github.com/paperclipai/paperclip/actions/runs/36436311203),
[missing compiler in the slim
image](https://github.com/paperclipai/paperclip/actions/runs/36440210984),
and the aggregate 15-minute test timeouts in those broad runs. Both
broad attempts passed typecheck, token gates, Product E2E type/unit
checks and build. The focused EC2 lane preserves the existing
trusted-actor and immutable-source gates.


Earlier documentation/test checkpoint
`ff244c4fd78a7ede5a3e00efe09f475f133ef33e` leaves runtime behavior
unchanged. 154 focused tests pass across configuration building, native
provider resolution, permission policy, credentials, UI configuration,
and new-agent setup (including both Grok auth modes); token gates pass.
All fresh CI is green for this head: 56 successful checks/statuses and
two intentional skips ([run
36367065119](https://github.com/paperclipai/paperclip/actions/runs/36367065119)).
Greptile is 5/5 with no new findings. Grok already inherits the shared
`approve-all` default, so unattended setup requires no manual permission
change.

Runtime head `bb5a9307991f1ac567b781970ef11b39d518e19b` fixes a final
staging continuation failure before provider startup: explicit
replacement archived the old harness but left its failover backups
active, which caused `runner_harness_state_mismatch`. The regression
fails before the fix and passes after it; all eight adjacent
recovery-safety cases also pass. Old backups remain inspectable inside
the continuity archive. All fresh CI is green at this head ([run
36360839248](https://github.com/paperclipai/paperclip/actions/runs/36360839248)),
with a 5/5 review. One unrelated Cursor test timed out in the initial
server shard; the same-commit failed-job rerun passed, and both attempts
are retained. Staging deployment is confirmed healthy on this revision.
The controller image is
`ghcr.io/paperclipai/paperclip@sha256:6ad91c487910ccd2596ff7aed0a3a3ea5233d12b51b83cd6e1402237749b9673`.
The final browser-created staging task passed on this exact revision
with API authentication: context read → structured human question →
controller restart → answer submission → same native provider session
resumed → document saved → task Done. The two turns took approximately
119s and 77s. The actual write receipt was applied, and the saved
document has exactly one revision containing the selected answer and
requested marker. Usage and cost were not reported. [Controller image
build](https://github.com/paperclipai/paperclip/actions/runs/36360889243).

- Previous integration head `a44f7dbb6b6f77cd9ed893756ca453307f281e5f`:
all CI green (53 successful checks/statuses, two intentional skips),
including repository typecheck/build/tests, native Runner tests, browser
shards, and canary installation checks. [CI run
36358672529](https://github.com/paperclipai/paperclip/actions/runs/36358672529).
Greptile is 5/5 with no unresolved findings.
- Focused checks cover Grok credentials, executable admission, launcher
assets, provider-pack paths/permissions, workflow contracts, setup
defaults, CLI authorization, and subscription conflict recovery. All 39
protocol definitions validate. Final integration checks pass 124
catalog/evidence/cache tests and nine project-form tests; token gates
pass. Some local dependency checks could not load the stale installed
dependency tree; the corresponding fresh EC2 checks pass.
- Clean public npm installation passed on EC2 at
`8b172ebcf8e02e30662d830c00f3961e3bd459ec` ([run
36164964900](https://github.com/paperclipai/paperclip/actions/runs/36164964900)):
17 unified-version packages, lifecycle scripts enabled, built-in
launcher present, no separate Grok package or npm-downloaded binary,
missing prerequisite rejected, separately provisioned native executable
and command lease verified. No credentials or inference were used.
Subsequent changes preserve this npm asset layout.
- The immutable Daytona prerequisite image is
`ghcr.io/paperclipai/paperclip-daytona-runner@sha256:98957d5be0ac774d086b6402b5849e8e6356fec70fb8c09fca6eb4ed6de918e0`,
built from `5a2db471f3ddabe77f9f80e76ed27f996cb97fba`. The previous
Cloud controller image was
`ghcr.io/paperclipai/paperclip@sha256:fd914e1ab1e45f741e8e078ff452d16f082d7ac05f9b4b3506d3a3c64150d204`,
built from `a44f7dbb6b6f77cd9ed893756ca453307f281e5f`; it is superseded
by the latest image above. Its EC2 build verified provider-pack access
under an unrelated unprivileged UID.
- Browser staging at `40f898bc4cba73c1dff4e6344a3983ba0fb247ef` passed
full Grok onboarding with the correct `grok-4.7` model, saved credential
delivery, and pinned Daytona execution. A browser-created task read
context and asked the structured human question. After a controller
restart, answering the persisted question resumed the same native
provider session, saved the requested document, and completed the task.
Actual tool outcomes and durable state agree: one question and one
document revision. The two successful turns took 42.7s and 63.1s; usage
and cost were not reported.
- Restricted policy returned the expected `approval_required` outcome.
Functional staging tests explicitly selected `approve-all`; controller
authorization and governed approvals remain enforced. Temporary board
CLI access was revoked and verified rejected (HTTP 401), and the
disposable onboarding agent was paused.

Failures remain retained: the pre-fix continuation failure (its task
remains blocked; the passing final task is fresh), the original Cloud
provider-pack permission failure, the expected restricted-policy denial,
the superseded npm staging failure, and an earlier monolithic CI
infrastructure timeout. Browser CI exposed a project alias/form race;
the final stack uses master's stronger draft-preservation fix and all
browser shards pass. Historical full subscription/API protocol and
Product rosters retain their original source revisions and do not
qualify this packaging revision. No local Docker or Rust build was used.

## Risks

The branch includes master’s draft-preservation fix for project URL
aliases. It keeps the same project’s edit form mounted and clears prior
data when the project or company changes.

Custom sandboxes and local execution hosts must provision the pinned
binary before Grok starts. Missing, changed, unsupported-platform, and
symlinked executables fail admission. The new builtin profile cannot
resume sessions created with the former private-package profile.
Existing Claude/Codex npm bridge profiles retain their package pins.
Grok restricted modes preserve the selected policy but cannot
automatically admit Paperclip calls: ACP permission metadata does not
independently bind tool authority, so those calls stop with
`approval_required`. New Grok configurations default to `approve-all`,
including API configurations that omit the mode. Existing explicitly
restricted configurations remain restricted; controller authorization
and governed approvals remain enforced.

## Model Used

OpenAI GPT-6 through Codex, with tool use and code execution. The exact
serving model identifier and context-window size are not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 14:54:44 -05:00
DottaandPaperclip 18dac1e1ef feat(connections): add experimental memory providers and remote MCP access (#13942)
## Thinking Path

> - Paperclip manages AI agents and their work.
> - Connections give agents governed access to external tools.
> - Agents need durable memory across tasks and execution environments.
> - Mem0, Zep, Supermemory, Cognee, and Honcho provide hosted memory
tools.
> - This pull request adds their setup flows behind an experimental
toggle.
> - It also delivers assigned MCP tools through the native remote Codex
runner.
> - Operators can connect a provider once and use the same governed
tools locally or in Daytona.

## Linked Issues or Issue Description

**Problem or motivation**

The Apps catalog lacks a complete set of memory providers. Remote native
Codex agents also need access to assigned managed MCP tools without
receiving provider credentials.

**Proposed solution**

Add five memory connectors behind the disabled-by-default Experimental
memory connectors setting. Use the existing connection setup and
permissions UI. Default their tools to Allowed. Preserve company
boundaries, operator permission changes, provider scopes, and audit
attribution.

**Alternatives considered**

Direct provider credentials in each sandbox would duplicate setup and
bypass the managed gateway. The remote runner instead uses its existing
protocol channel to call the gateway on the server.

**Roadmap alignment**

This maintainer-requested experiment supports the Memory / Knowledge and
Connected Apps roadmap areas. It adds provider connections without
introducing a separate memory UI. Related connector authoring
documentation is tracked in #13692; no duplicate memory-provider
implementation was found.

## What Changed

- Add provider definitions, official branding, and the experimental
setting for all five providers.
- Use OAuth for Zep and Supermemory, API credentials for Mem0 and
Honcho, and a bundled Cloud API bridge for Cognee with no runtime
downloads or subprocesses.
- Default memory tools to Allowed and classify destructive actions
explicitly.
- Fix personal remote credential resolution and propagate provider tool
errors.
- Relay assigned managed MCP tools to remote native Codex through the
runner protocol. Recheck current authority for each call and rotate
stale tool contracts.
- Add 21 Storybook states and complete OAuth walkthrough fixtures.
- Document provider research, sanitized tool inventory, and live local
and Daytona proof.

## Verification

- Passed workspace typecheck: `pnpm -r typecheck`.
- Passed production build: `pnpm build`.
- Full local `pnpm test:run`: 13,273 passed, with failures from
process/readiness timeouts under parallel load. Reran all 16 affected
suites with one worker: 456 passed, leaving two macOS `/var` versus
`/private/var` path assertions. Both passed with `TMPDIR=/private/tmp`.
No test failures remain unverified. Latest-head remote CI passes all 54
checks (two optional Storybook jobs skipped). Greptile is 5/5 with no
unresolved findings.
- Latest Cognee gateway regression: 71 passed, including public
deployment without a runtime host and immediate recovery after a
provider error. Bundled bridge tests: 25 passed.
- Browser setup and real agent tasks exercised all five providers. Mem0,
Cognee, Zep, and Honcho have successful store/retrieve proof.
- Supermemory now has scoped read/write consent. Local storage and real
Daytona write, document read, and semantic recall passed; indexing
completion was verified before claiming success.
- All five providers were exercised through a real Daytona sandbox and
its native runner MCP relay. Provider credentials remained on the
server. The final bundled Cognee bridge also passed a fresh Daytona
store/recall run. All disposable sandboxes were removed and verified
absent after testing.
- Storybook is rebuilt and contains the experimental toggle, catalog,
setup, permissions, and error states. The Zep and Supermemory access
steps advance correctly.

## Risks

- Provider OAuth scopes and plan limits remain independent of Paperclip
tool permissions. An Allowed tool can still be rejected by the provider.
- Providers can queue memory indexing; save acceptance does not prove
that semantic recall is ready.
- Remote tool contracts must stay synchronized with current connection
authority. Regression tests cover revocation and stale contracts.
- This change has no database migration. Existing connections remain
usable when the experimental catalog toggle is disabled.

## Model Used

OpenAI Codex, GPT-6 Astra (`gpt-6-astra`), with reasoning, repository
edits, code execution, and browser automation. The exact context window
size is not exposed in this session. Live acceptance agents used
`gpt-5.6-sol`.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-24 15:54:38 -05:00
Devin FoleyandPaperclip 8ee8f1fd6e ci: retire recurring public cloud image builds (#13827)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Core publishes standard images and source verification for
downstream services.
> - Managed services can now compose private images from the signed
standard image.
> - Core still builds a second public cloud image on every master push
and release.
> - That duplicate producer consumes build capacity and retains an
obsolete readiness contract.
> - This pull request retires recurring cloud publication while
preserving the standard producer and rollback artifacts.

## Linked Issues or Issue Description

Refs #13797 and #13789. Related: #12856 changes image dependency
packaging; it does not retire this producer.

**What existing behavior does this improve?**

Core's recurring Docker publication and Cloud readiness workflow.

**Current behavior**

Master pushes call the legacy cloud publisher from Cloud readiness.
Release tags and manual Docker runs call it too. Canary promotion also
requires the legacy image.

**Proposed behavior**

Publish standard Core images and retain `Cloud source verified v1`. Let
downstream services build their managed image. Keep explicit commit
previews and existing images available.

## What Changed

- Remove `docker-cloud.yml`, its master and release callers, and its
unused cache selector.
- Remove the legacy image/migrator wait and `Cloud deployable v1` job.
Keep the full source verification workflow and exact source-proof name.
- Make canary promotion inspect and promote the standard image only.
- Preserve signed standard-image publication, direct migrator
publication, and explicit `release.yml` previews. The preview path still
uses the Dockerfile `cloud` target.
- Update workflow, preview, build-stamp, and packaging tests. Exercise
the promotion shell with mocked registry commands, including
missing-image and missing-tag cases.
- Document frozen legacy aliases, consumer requirements, preview
compatibility, and rollback retention.

## Verification

- All 377 workflow tests pass: `node --test
.github/scripts/tests/*.test.mjs`.
- All 129 release-registry tests pass: `pnpm test:release-registry`.
- Focused source-proof, standard-image, preview, and workflow tests
pass: 256 tests.
- Focused image packaging/build-stamp tests pass: 16 tests.
- Actionlint passes on all three changed workflow files. `git diff
--check` passes.
- Full local `pnpm build` and `pnpm -r typecheck` pass.
- The policy follow-up updates an old assertion that required the
removed readiness job. All 37 source-proof/release-workflow tests pass
locally.
- Full local `pnpm test:run` did not complete successfully while the Mac
ran out of disk space. No full-suite pass is claimed. Removed 1.2 GiB of
generated Cargo output from this isolated worktree with `cargo clean`.
GitHub CI passed on the final head: 52 successful checks and 2 optional
skips.
- Fresh Greptile review for `4f5fe1951f0bd7f7739cf6655d395ff78f1ed944`:
**5/5**, successful current-head check, zero review threads.
- September 23 refresh: the unchanged PR head merges cleanly with
current master `db8f8fe5b73a2697684a30261b0d306a9c631aba`. In an
isolated temporary worktree, all 377 workflow tests and 29
release/preview tests pass on the combined tree. `git diff --cached
--check` passes.
- Refreshed Actionlint workflow validation passes with ShellCheck
disabled. Full Actionlint reports the same 10 existing ShellCheck
diagnostics as master, with no added diagnostics. No source changes or
new PR commits were needed.
- The full local build/typecheck and current-head Linux CI results above
remain the verification for the unchanged PR head. They were not rerun
for this metadata-only refresh. No image publication or tenant
deployment was initiated for this refresh.

## Risks

**Deployment prerequisite satisfied (September 23):** The combined
cleanup release is deployed to staging and production, and production
Support is verified. Active managed-fleet automation uses standard-image
composition. Explicit immutable previews remain supported by the
retained preview publisher. This PR is ready for maintainer review; keep
auto-merge disabled and wait for explicit merge authorization.

- A consumer still selecting `Cloud deployable v1` will stop advancing
at the last legacy-ready commit. Confirm active automatic consumers use
the standard-image composition contract before merge.
- Legacy cloud release-channel aliases stop advancing. Standard
self-hosted aliases continue.
- This PR deletes no registry images, cache tags, migrators,
credentials, or runner infrastructure. Existing immutable releases
remain usable for rollback.
- Explicit legacy previews remain for commit-specific operator
deployments. Retiring that compatibility path requires a separate
consumer migration.
- These changes affect CI publication, not database schema or
application behavior.

## Model Used

OpenAI Codex, GPT-6. The runtime does not expose a more specific model
identifier or context-window size. Used repository inspection,
reasoning, code editing, shell tools, and test execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-23 19:35:19 -07:00
DottaandPaperclip db8f8fe5b7 fix(evals): select Grok subscription protocol credentials explicitly (#13901)
## Thinking Path

> - Paperclip manages agent work through shared runner contracts.
> - Direct protocol evals qualify provider behavior against a mock
control plane.
> - Grok supports API keys and company subscription credentials.
> - The hosted protocol workflow selected an API key for every Grok
cell.
> - Product subscription support did not enable subscription protocol
runs.
> - This change adds explicit subscription selection and checks the
recorded authentication mode.

## Linked Issues or Issue Description

Refs #13878, #13882, #12618.

The direct Grok protocol roster cannot run with subscription
authentication through the trusted default-branch workflow. Add an
explicit selector while keeping API-key dispatches compatible. Keep the
actor allowlist, protected environment, immutable source revisions, and
publication gates.

## What Changed

- Add `grok_authentication` with `api_key` and `subscription` choices.
Keep `api_key` as the compatibility default.
- Deliver the protected `GROK_AUTH_JSON` secret only to a
subscription-selected Grok cell. Do not provide an API key to that cell.
- Read authentication mode from the pinned eval program's actual roster
summary. Retain it in the cell, catalog, campaign roster, and result.
- Reject missing or mismatched authentication evidence during
aggregation. Preserve cell metadata and an allowlisted failure reason
before failing a cell, so malformed evidence cannot hide the retained
attempt.
- Document credential setup, source separation, and temporary-secret
cleanup.

## Verification

- `node --test
packages/paperclip-runner/scripts/runner-protocol-eval-campaign.test.mjs
packages/paperclip-runner/scripts/runner-protocol-eval-workflow-security.test.mjs
scripts/__tests__/release-verify-workflow.test.mjs`: 34 tests passed.
- Validated all 39 Grok cells at eval revision
`3213dbec7e8ca1865ea95e6db7e7d34b095eb47a`; every selected cell requests
only the subscription credential. Validation made zero provider calls.
- Negative coverage rejects invalid selectors and missing or API
authentication evidence in an otherwise passing subscription attempt.
- `git diff --check` passed. All 53 current-head checks passed; the
unchanged callback-drain timing test passed its bounded rerun, and the
failed attempt is retained. Greptile reviewed
`ecd3dcc0e998f07cf56fcb1f087946f50f388bec` at 5/5 with no remaining
findings.
- No Docker or broad builds ran on the developer machine. CI performs
repository checks on the configured fleet.

## Risks

Grok runs require an eval revision that records `authenticationMode` in
the roster summary. Missing evidence fails closed. The credential
contains account access and refresh tokens; an owner must approve its
delivery to the protected environment before a live run. The change adds
no PR trigger or authorization bypass. Live subscription protocol
qualification remains pending this workflow reaching master.

## Model Used

OpenAI GPT-6 through Codex, with tool use and code execution. The exact
serving model identifier and context-window size are not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-23 18:43:21 -05:00
DottaandPaperclip 24429024e7 feat: add Fireflies connector and summary-ready routines (#13890)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Apps gives agents governed access to external tools through stored
credentials.
> - Routines start work when an external service sends an event.
> - Fireflies provides meeting transcripts and summaries through an
official hosted MCP server.
> - This PR adds that connection and accepts signed meeting events
through the shared app webhook flow.
> - Agents can review completed meetings with the same permissions and
audit records as other work.

## Linked Issues or Issue Description

**Problem or motivation**

Operators need agents to read Fireflies meetings and start follow-up
work when a summary is ready. The Apps catalog lacks Fireflies. The
shared app webhook flow needs to accept its signed deliveries.

**Proposed solution**

Use the official Fireflies MCP endpoint with OAuth or a vaulted bearer
API key. Extend the existing Another app or script flow with signed
webhook support. Verify the raw-body signature and pass the JSON payload
as external data. Select Meeting Summarized in Fireflies. Deduplicate
identical signed deliveries, including setup deliveries.

**Alternatives considered**

A separate REST connector would duplicate the governed MCP path.
Polling, legacy V1 payloads, and automatic provider-side webhook
registration are outside this change.

**Roadmap alignment**

This extends the existing MCP Tool Gateway & Apps and Scheduled Routines
surfaces. It adds a provider to those systems. It does not introduce a
second integration framework.

**Additional context**

A GitHub search found no existing Fireflies issues or PRs. Provider
references and verification limits are in
`doc/connections/FIREFLIES.md`.

## What Changed

- Add the official Fireflies catalog definition, generated registry,
provider evidence, and branded artwork.
- Reuse Access → Connect, dynamic discovery, Permissions, vault storage,
policy, and audit behavior.
- Classify Fireflies sharing, movement, and access revocation as writes.
- Preserve Off and Ask first restrictions during OAuth reauthorization
and API-key replacement. New actions retain normal defaults.
- Add `app_webhook` authentication to the shared Another app or script
flow. Accept bearer tokens or raw-body HMAC-SHA256. Preserve earlier
`fireflies_hmac` triggers and revision snapshots for compatibility.
Existing text columns need no migration.
- Verify `X-Hub-Signature` or `X-Hub-Signature-256` against the exact
request body. Preserve generic event payloads and deduplicate identical
signed requests.
- Keep the routine wizard generic. Show one webhook URL and secret in
Another app or script. Keep all new app webhook event names
provider-neutral. Keep provider setup instructions in the connector
documentation.
- Pass generic webhook JSON to the task in an explicit external-data
block, capped at 16,384 characters. Keep strict meeting validation for
existing legacy Fireflies triggers.

## Verification

- Feature implementation commit `0882dc8a1`: all 54 CI checks passed;
two conditional Storybook checks skipped. This includes full tests,
typecheck, build, browser E2E, canary dry run, and security checks.
Greptile rated this commit 5/5; all review threads are resolved.
- Full local `pnpm -r typecheck`, `pnpm build`, and token gates passed
on the final code. Targeted connector, gateway, webhook, revision, and
UI suites passed during implementation. After the provider-neutral
follow-up, all 84 app-webhook and routine-service tests passed; the
final payload-to-task assertion also passed in the 72-test routine suite
and a clean-config rerun.
- The long local `pnpm test:run` invocation started before the final
edits and was stopped after the final-commit CI suites passed. It
reported one generic webhook test failure while those files were
changing; that test and the entire routine suite passed on the final
source, including a clean-config reproduction. The interrupted local run
is not counted as a full-suite pass.
- In the embedded browser, completed official OAuth consent and
discovered 20 live actions. Real meeting listing, transcript retrieval,
and summary/action-item retrieval succeeded as the selected agent.
Turning a live read Off blocked its test; catalog refresh preserved the
restriction.
- Embedded-browser Another app or script setup, back/save/resume, narrow
layout, and a signed synthetic Fireflies delivery succeeded. The UI
reported authentication passed without creating a task. Fixtures cover
signature tampering, malformed requests, ordinary app event names,
duplicate/setup deliveries, rotation, revisions, pause/archive, and
company isolation.
- Existing MCP browser suite: 8 passed and 2 provider-dependent cases
skipped. Branding checks passed; connector artwork and webhook setup
were checked at desktop/mobile widths and in light/dark modes.
- An unauthenticated POST to a correctly formatted public webhook URL
reached the staging tenant verifier through the existing Cloud gateway.
- A real Fireflies webhook delivery remains unverified. A staging
callback is available for the operator walkthrough. Live API-key
authorization, credential expiry, and a new meeting's summary completion
were not tested against the provider. Fixtures cover these protocol and
lifecycle paths where applicable.

- Storybook follow-up `c54174faa`: 27 production-component stories cover
every UI change, with a source-to-story map in the connector
documentation. Static Storybook build, UI typecheck, token gates, and
Playwright checks for all stories and the mobile footer pass. All PR
checks passed for this Storybook follow-up; Greptile reviewed
`c54174faa` at 5/5.

## Risks

- Fireflies may change its hosted MCP tools or OAuth behavior. Tool
discovery stays dynamic. Experimental search/fetch tools are not
required.
- Public webhook setup requires HTTPS and a separate signing secret.
Fireflies normally emits events for meetings owned by the configuring
account.
- Reauthorization touches shared MCP permission code. Regression tests
cover existing restrictions, new actions, connection removal, and other
gateway callers.
- Webhook receipt grants no tool access. The routine agent still needs
an authorized Fireflies connection.

- New generic triggers rely on provider event subscriptions. Without a
sender-supplied idempotency key, changed request bytes count as a new
event. Existing legacy Fireflies triggers retain summary-only filtering
and per-meeting deduplication.

## Model Used

OpenAI Codex, model `gpt-6-astra`. Used reasoning, repository editing,
code execution, and embedded-browser testing. The runtime did not expose
a context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-23 17:11:16 -05:00
DottaandPaperclip b648d8cdda fix(evals): support explicit Grok qualification workflows (#13878)
## Thinking Path

> - Paperclip manages AI agents and their provider connections.
> - Product E2E checks real tasks through the browser, server, and
runner.
> - Grok qualification needs separate API-key and subscription evidence.
> - Product subscription tests and direct Grok protocol evals need
explicit credential delivery.
> - This change supplies each credential only to its selected profile
and prepares the pinned binary.
> - Maintainer authorization and protected-environment gates remain
required.

## Linked Issues or Issue Description

Refs #13845, #13847, #13850, #13882.

The Grok feature branch has a manual subscription qualification profile.
The trusted master workflow must admit its selected credential and
prepare the same verified binary and artifact verifier as the API
profile. Direct protocol evals also need the selected xAI key and pinned
Grok binary. These prerequisites do not register or schedule the new
profiles on master.

## What Changed

- Deliver `GROK_AUTH_JSON` from the protected paid environment only when
the selected profile requests that credential.
- Install the checksum-verified Grok binary for the local subscription
profile.
- Prepare the pinned artifact verifier for the manual subscription
suite.
- Extend workflow security assertions to cover the new credential and
profile.
- Add the ACPX Grok credential mapping to the trusted-master catalog,
then deliver only the selected `XAI_API_KEY` to direct protocol cells
and install the target’s checksum-verified Grok binary before packaging.
- Allow a direct-protocol concurrency override from two cases up to the
existing configured ceiling; it can only lower concurrency.
- Document the Grok protocol workflow and its API-only credential
boundary.
- Render missing LLM usage and cost as Unavailable, and label partial
observations with coverage. Preserve raw records, grades, and actual
zero costs.
- Preserve measured campaign source metadata during report regeneration
instead of inheriting the renderer checkout or CI event; skip empty
legacy source records when recovering older provenance.

## Verification

- Latest commit `05d05801477104c8155977bbbe3e119a5241f960`: all 54
reported checks successful, two intentional skips, Greptile 5/5, and
zero unresolved review threads. [CI
run](https://github.com/paperclipai/paperclip/actions/runs/35890978288).

- After merging current master, all 17 workflow security/image tests and
23 catalog/workflow policy tests passed. The trusted catalog also
generated all 39 pinned Grok cells with `XAI_API_KEY` and one case per
shard. The new policy tests execute the concurrency guard against valid,
out-of-range, and malformed values.
- The Grok branch separately passed 450 Product harness unit tests,
including private company credential staging, cleanup, and
token-fragment redaction.
- The fresh-login native subscription smoke passed three repetitions of
tool execution, session resume, restrictive permissions, and cleanup.
These are setup evidence; full subscription Product qualification
remains pending.
- All 72 focused report/billing/history/catalog tests and the Product
harness typecheck passed for the report-display change. The initial
sandbox run could not open the tsx IPC socket; the permitted rerun
passed. A zero-provider-call replay of the actual 16-cell Grok campaign
preserved all result records, grades, timing, and source provenance
while correcting missing usage labels.
- Review the thirteen-file diff. Provider credentials still enter only
the selected paid-test step; default-branch, numeric-actor, and
environment restrictions are unchanged.

## Risks

This admits a refreshable subscription credential to explicitly selected
trusted tests. Store it only in `runner-e2e-paid`, use a test login, and
remove it after qualification. Unselected profiles receive an empty
value. Pull requests cannot trigger the paid workflow. This PR changes
no fleet admission or actor allowlist.

## Model Used

OpenAI GPT-6 through Codex, with tool use and code execution. The exact
serving model identifier and context-window size are not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-23 17:05:52 -05:00
Devin FoleyandPaperclip be6f49a425 feat(runner): refresh shared coding harness runtimes (#13838)
## Thinking Path

> - Paperclip runs agents through local adapters and the native runner.
> - Both paths must use the same installed provider CLI.
> - New models require current harness releases.
> - The runner still pins Codex 0.153.4, Claude SDK 0.3.263, and
OpenCode 1.18.29.
> - Changing the image alone would fail the runner's exact version and
executable checks.
> - This pull request updates those dependencies, integrity checks,
controller checks, and image pins together.
> - Shared installations can then run the current models without a
task-time download.

## Linked Issues or Issue Description

Refs #13829, which updates model choices and reasoning controls.
Searches found no open PR that updates these runtime pins.

**Current behavior**

The shared provider pack ships old CLIs. Claude Code 2.1.263 cannot run
Opus 5.5, which requires 2.1.280. Remote controllers reject provider
packs whose versions differ from their declared pins.

**Proposed behavior**

Use Codex 0.156.0, Claude Agent SDK 0.3.280 / Claude Code 2.1.280, and
OpenCode 1.18.32 throughout the runner. Keep the reviewed ACP bridge
patches and one shared CLI installation per provider.

**Reason and benefit**

Current harnesses support the new model IDs while preserving executable
verification and remote provider-pack compatibility checks.

## What Changed

- Update dependency overrides, the Codex ACP package patch, runtime
profiles, and remote controller pins.
- Verify the new Claude Linux x64 and macOS arm64/x64 executables and
Codex Linux x64 executable against integrity-verified npm archives.
- Refresh OpenCode version checks, fixtures, and the runner
configuration label.
- Refresh the eval image's Grok, Gemini, Kimi, Cursor, and GitHub CLI
pins and archive hashes. Hermes remains current at 0.19.0.
- Refresh the build-time lock digest from clean pnpm 9.15.4 resolution.
Leave lockfile commits to repository automation.
- Document model compatibility and the separation between CLI runtimes
and patched ACP bridges.

## Verification

- `pnpm -r typecheck` and `pnpm build` passed.
- Rust workspace release tests passed.
- Package/patch and OpenCode binary-materialization contract tests: 11
passed.
- Real Codex 0.156.0 startup-ownership and paginated session-resume
probes passed with isolated synthetic homes and no model turn.
- Codex app-server `thread/start` preserved `gpt-6-sol` and
`gpt-6-luna`; no `turn/start` was sent. An unauthenticated built-in
catalog does not include those account-served entries.
- Installed Claude integrity probes passed for `claude-opus-5-5` and
`claude-fable-5-1`.
- `pnpm --filter @paperclipai/paperclip-runner
test:opencode:qualification` passed with the actual OpenCode 1.18.32
executable under Node 24 and Node 25. The loopback provider exercise
covers health/version, session creation/read/delete, SSE, and a
completed async prompt.
- `pnpm check:token-gates` passed.
- The targeted runner suite passed 130 tests. Three macOS failures in
snapshot module lookup and OpenCode final-message selection also
reproduce on the unchanged base; Linux CI will provide the platform
check.
- [Final Linux
CI](https://github.com/paperclipai/paperclip/actions/runs/35798076399):
all gates passed. Four jobs needed one retry after their CI workers
received shutdown signals. The PR has 55 successful checks, two skipped
checks, Greptile 5/5, and no unresolved review threads.
- Changed runner configuration UI tests: 5 passed.
- Full macOS `pnpm test:run` reached 13,094 passing server tests, 84
skipped, and 18 failures before the wrapper stopped. Failures involved
skill-cache publication permissions, missing bundled connector skills in
the worktree, and a conversation-reset timing case. The 10 cache
permission failures reproduce on the unchanged base; both
conversation-reset cases passed on a targeted retry. The wrapper did not
reach its later workspace/serialized groups locally; Linux CI covers
those groups.
- The local Docker daemon did not respond, so no local Docker build was
run. No billable model requests were made.

## Risks

- Deploy the matching controller and provider pack together. Older
controllers enforce their previous exact pins.
- Current upstream CLIs can change behavior. Existing protocol tests and
isolated real Codex probes cover the integration boundaries;
authenticated model inference is not part of these checks.
- ACP bridge package versions and executable digests stay unchanged
because their executable bytes are unchanged. Only the underlying
CLI/SDK dependencies move.
- No schema migration. Revert the runtime and image pins together to
roll back.

## Model Used

OpenAI GPT-6 via Codex, with repository tools, code execution, and web
research. The exact serving model ID and context window were not exposed
by this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass for the changed surfaces
and real-executable probes; full macOS-suite limitations are listed
above
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-22 17:02:29 -07:00
Devin FoleyandPaperclip e3d8fb0876 feat: attest standard production images at full source commits (#13797)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Operators deploy its standard production container on several CPU
architectures.
> - Downstream image builders need to identify the exact source of their
base image.
> - A short commit tag does not provide signed source evidence.
> - This pull request adds a full commit tag and signed image digest for
canonical master pushes.
> - Consumers can verify the source and compose from the immutable
digest.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

Publication of the standard multi-platform production image.

**Current behavior**

The Docker workflow publishes short commit tags and channel tags. It
does not provide a signed standard-image contract tied to the complete
master commit.

**Proposed behavior**

Canonical master pushes also publish `sha-<full-commit>` and attest the
exact index digest after platform validation and an immutable-image
orphan-reaping check. The signer certificate binds the source
repository, commit, workflow and ref. Existing tags and the separate
cloud producer remain available.

**Reason and benefit**

Downstream builders can prove the source of a standard base without
adding their dependencies or repository details to the public workflow.
No matching open issue or duplicate PR was found.

## What Changed

- Add the canonical full-SHA tag without changing existing tag mappings.
- Validate amd64 and arm64 descriptors and hash the exact registry
response bytes and require its digest header to match.
- Verify the immutable image and sign it with GitHub artifact
attestations.
- Run the contract tests in trusted PR verification and document the
consumer contract.

## Verification

- `node --test scripts/__tests__/release-verify-workflow.test.mjs
scripts/cloud-source-verification.test.mjs
scripts/standard-image-contract.test.mjs`: 37 passed.
- A read-only check against an existing published index returned its
exact expected digest.
- `actionlint -shellcheck='' .github/workflows/docker.yml
.github/workflows/pr-trusted.yml`: passed. Normal ShellCheck reports
only existing `ls` and word-splitting warnings.
- `pnpm build`: passed locally with Cargo available.
- `pnpm -r typecheck`: passed locally.
- Full local Vitest was attempted: 8,260 passed, 14 failed, with 34
failing suites. The failures were missing embedded-PostgreSQL library
aliases in this fresh install and existing macOS runtime-skill-cache
rename errors. Native aliases are now restored. Rerunning the 33
affected database suites produced 577 passes and two unrelated AgentMail
skill-root lookup failures (32 suites passed). The four directly failing
database tests also pass independently. This is not a claim that the
full local suite passed.
- All final-head CI checks pass. One unrelated routine-route mock
assertion passed on the single-shard retry; its 15 tests also pass
locally. Greptile is 5/5 on this exact head, with no unresolved threads.
- Actual signing requires a canonical master push. This draft PR does
not publish trusted provenance.

## Risks

The new attestation step requires OIDC and attestation write permissions
in the merge job. Signing failure leaves the image available but without
the new admission proof. Consumers must fail closed when proof is
missing. Existing release tags, the legacy producer, and image retention
remain unchanged. No database or application behavior changes.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository inspection, shell
execution and test tools. The session does not expose a more specific
model variant or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-22 07:33:30 -07:00
3c2bc4f546 ci: halve the isolated native Runner check by seeding it from the public master build cache (#13736)
## What

Recurring CI health check (PAP-31): on the most recent fully-green PR
run (`cb703ac`, run
[35519542997](https://github.com/paperclipai/paperclip/actions/runs/35519542997)),
the slowest check was **Compile isolated native Runner** at **452s** —
ahead of the largest test shards (379s). Every run recompiled the full
Rust dependency tree from zero, even though `docker.yml` already
refreshes a **public** `mode=max` BuildKit cache
(`ghcr.io/paperclipai/paperclip:buildcache-{amd64,arm64}`) on every
master push, containing exactly these layers. This PR seeds **only the
baseline build** with that registry cache, via anonymous pull.

**Measured on this PR's own CI (which exercises the seeded path): the
check completed in 123s, down from 452s — a 73% reduction, ~5.5 minutes
saved per run.**

## Thinking Path

Cost breakdown of the 452s from the job log: `cargo chef cook`
dependency compile 214.7s, local cache export 43.1s, runner-core build
37.0s, metadata proof layer 36.8s, `cargo install cargo-chef` 36.4s,
rebuild-verification build ~65s, setup/teardown ~20s. The dependency
compile and toolchain layers are identical to what the production
`docker.yml` build already caches publicly on every master push, so
recompiling them here bought no signal — the check's real assertions
live in the *verification* build, not the baseline. A first attempt used
`actions/cache` plus a master `push` trigger, but the CI bot's GitHub
App lacks `workflows` permission; the registry-cache approach is
strictly better anyway (shared across PRs immediately, no 10GB
Actions-cache quota pressure, no workflow change).

## What Changed

- `scripts/check-docker-runner-cache.sh`: the baseline build now adds
`--cache-from
type=registry,ref=ghcr.io/paperclipai/paperclip:buildcache-{amd64|arm64}`
(selected by host arch). `RUNNER_CHECK_SEED_CACHE` overrides the ref, or
set it empty to force the old cold path. The script header documents the
anonymous external read.
- `.github/workflows/docker-runner-check.yml` (comment-only): the stale
"no external cache" note now describes the anonymous GHCR seed and the
verification build's local-cache-only isolation. This was pushed in a
follow-up commit with workflow-edit permissions; the original CI-bot
token could not touch workflow files.

No Dockerfile stages or verification assertions changed.

## Verification

- This PR's own `Compile isolated native Runner` check runs the seeded
path (the script is in the workflow's trigger paths): **passed in 123s**
vs the 452s baseline.
- The rebuild-verification semantics are untouched: it still runs on a
**fresh builder** importing **only the local cache exported by this
run's baseline**, so it proves exactly what it proved before — that the
runner image rebuilds reproducibly from this run's own exported layers.
- Verified `ghcr.io/paperclipai/paperclip:buildcache-amd64` is
anonymously readable (unauthenticated manifest pull succeeds), so the
check gains no credential or secret dependency.

## Risks

- **Stale or missing seed cache:** if the GHCR ref is unreachable,
private, or garbage-collected, BuildKit logs a warning and falls back to
the pre-PR cold compile — the check gets slower, never wrong.
`RUNNER_CHECK_SEED_CACHE=""` restores the cold path explicitly.
- **Cache trust:** the seed only accelerates the *baseline* build; the
verification build still runs on a fresh builder against only this run's
locally exported cache, so a stale or poisoned registry cache cannot
make verification pass spuriously. The ref lives under
`ghcr.io/paperclipai/*`, written only by repo CI on master pushes.

## Model Used

Claude Fable 5 (`claude-fable-5`) via Paperclip agent **Bender
(Fable)**, issue PAP-31.

---------

Co-authored-by: Bender (Fable) <bender-fable@paperclip.local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-09-21 20:22:48 -07:00
DottaandPaperclip b82661b561 refactor(connections): retire the legacy Composio broker (#13758)
## Thinking Path

> - Paperclip manages agents and their access to external tools.
> - Connectors expose these tools through a governed MCP gateway.
> - PR #13755 added a direct Composio MCP connection behind the
experimental MCP aggregators flag.
> - The old project API-key broker still created toolkit child
connections and showed a separate Services tab.
> - Keeping both paths leaves obsolete setup and session code in the
product.
> - This change removes the broker and preserves direct MCP setup,
credentials, permissions, and execution.
> - Saved legacy records fail closed and remain available for explicit
removal.

## Linked Issues or Issue Description

Related: #13755. This retirement supersedes the legacy-path fixes
proposed in #12630, #12632, #12634, and #12906. It does not close those
PRs.

**What existing behavior does this improve?**

Composio connector setup, management, and runtime dispatch.

**Current behavior**

Composio offers both direct MCP and a project API-key broker. The broker
mints sessions and creates one child connection per toolkit.

**Proposed behavior**

Offer only direct MCP. Remove the toolkit Services UI, REST routes, API
client, and session broker. Block saved legacy parent and child records
from discovery, execution, health checks, reconnect, and OAuth. Preserve
their records and credentials until the operator removes each
connection.

**Reason and benefit**

The direct MCP connector becomes the single supported Composio workflow.
Provider accounts remain managed in Composio.

## What Changed

- Remove the API-key catalog method and its generated-source definition.
- Delete Composio broker clients, session creation, account
synchronization, child lifecycle, and toolkit routes.
- Remove the Services tab, service rows, child provenance, and
cascade-removal controls. Keep Vercel provenance intact.
- Retain a shared retirement guard for stored legacy records. Show
Retired status and replacement/removal guidance in the connection list
and details; hide obsolete runtime controls.
- Preserve the experimental MCP aggregators flag and direct MCP
infrastructure.
- Replace broker fixtures with retirement tests and extend direct
Composio catalog/reconnect coverage.

## Verification

- Focused shared, server, and UI tests passed with one worker. Server
retirement tests use a name filter; no full local test suite was run, as
requested.
- Server and UI TypeScript checks passed.
- Token gates and UI build passed.
- Real browser: opened the saved Composio connection, refreshed all 11
tools, and ran the provider's read-only GitHub account-list operation
through the standard Test dialog as an agent. The provider returned
success using the existing OAuth credentials.
- See `doc/connections/COMPOSIO-BROKER-RETIREMENT.md` for scope and live
evidence.
- Storybook build passed. A fresh real agent used
`COMPOSIO_SEARCH_TOOLS` and `COMPOSIO_MULTI_EXECUTE_TOOL` to return the
actual Paperclip DeepWiki hierarchy: one success, zero errors. Gateway
audit records confirm both calls succeeded.
- Browser retirement check: a credential-free legacy fixture showed the
guidance, opened the direct MCP replacement flow, and was removed
through the standard confirmation.
- Focused regressions for the experimental settings copy and exact
OpenAPI route coverage passed. All latest-head CI checks passed (54
successful, two intentionally skipped); Greptile scored 5/5 with no
unresolved review threads. The PR has no merge conflicts.

## Risks

This intentionally breaks the old Composio project API-key and
child-connection workflow. Existing legacy records cannot run, even if
their stored status is active. Operators must create a new direct MCP
connection and choose access rules; credentials and grants are not
migrated. Remove each old record separately to delete its credentials.
No schema migration or data deletion runs automatically. Direct MCP
connections keep their existing grants and secrets.

## Model Used

OpenAI GPT-6 via Codex, with reasoning, code execution, and browser
tools. The exact runtime variant and context-window size are not exposed
in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 13:57:35 -05:00
DottaandPaperclip e8c8ba3c19 feat(apps): add experimental MCP aggregator connectors (#13755)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Its tool gateway applies company access rules and approval controls
to connected apps.
> - MCP aggregators expose many apps through one provider endpoint.
> - Each aggregator needs its own credential, catalog, grants, and
lifecycle in Paperclip.
> - This pull request adds independent Zapier, Arcade, Composio Connect,
and Executor setup with a common Access → Connect layout.
> - A default-off MCP aggregators flag lets operators opt in while we
complete provider acceptance tests.
> - Agents use the normal Paperclip permissions, Test screen, and
gateway after setup.

## Linked Issues or Issue Description

**Subsystem affected**

Apps, connection setup, shared contracts, and the remote MCP gateway.

**Problem or motivation**

Aggregator endpoints need clear provider setup and correct MCP sessions.
Generic setup does not explain each provider's authentication or broad
execution tools. Provider approval must preserve the original execution
instead of replaying a write.

**Proposed solution**

Add four separate connectors behind Settings → Experimental → MCP
aggregators. Start with human and agent access, then connect the
endpoint and read its tools. Enable tools by default. Use the existing
Permissions and Test screens after setup. Keep legacy Composio API-key
and child connections intact.

**Alternatives considered**

A shared connection for all providers would mix credentials and access
rules. Separate provider-specific permission and test screens would
duplicate existing controls. Vercel Connect is outside this change.

**Roadmap alignment**

Extends the existing MCP Tool Gateway & Apps capability and the
Connected Apps roadmap area. This work was requested and reviewed by the
maintainer.

Related work: #11894, #12630, #12632, #12634, and #12906 concern the
legacy Composio broker. #13102 also covers remote MCP pagination. This
change preserves the broker path and adds initialized sessions, response
matching, and provider resume handling alongside pagination.

## What Changed

- Add branded setup and interactive Storybooks for Zapier, Arcade,
Composio Connect, and Executor. Use the existing access controls and
normal action tests. Do not request a connection name or action choices
during setup.
- Add the default-off `enableMcpAggregators` flag to settings, managed
feature metadata, the catalog, and setup guards. Hidden connections keep
running. Legacy Composio connections remain unchanged.
- Reuse the vault, grants, policy, and catalog models. Support OAuth
discovery, bearer tokens, custom headers, and credential-bearing URLs.
Add no database tables or migrations.
- Initialize and retain Streamable HTTP sessions by connection and
effective credentials. Read paginated catalogs and match streaming
responses to request IDs.
- Classify unfamiliar aggregator tools as writes despite upstream
read-only hints; only exact reviewed read capabilities enter the
read-only allowlist. Legacy Composio child behavior is preserved.
- Preserve provider authorization links and execution IDs. Support
Executor approve/resume, decline, and cancel without automatic replay of
uncertain writes.
- Preserve Off and Ask first choices during refresh and reconnect. Allow
new tools and retire removed tools. Keep agent access updates atomic and
preserve an empty agent selection.
- Document connector UX rules, provider branding sources, and live
acceptance results.
- Stabilize the existing Sentry release fixture after its repeated CI
failure by reusing one module mock; production Sentry behavior is
unchanged.

## Verification

- Final head `d11781970`: [CI
run](https://github.com/paperclipai/paperclip/actions/runs/35633534900)
passed, including broad typecheck, test shards, build, and E2E. All 54
checks pass; 2 optional checks are skipped. Greptile is 5/5, Security
Scan passes, and all review threads are resolved.

- Passed 27 focused connector Vitest checks and 18 connector-only
Storybook browser checks before the flag change. All 85 stories rendered
at desktop and narrow widths.
- Passed 5 connector lifecycle/server checks and 7 selected flag checks
after adding the flag. The latter cover settings, managed defaults,
cached catalog visibility, and all four setup routes.
- Review fixes passed 13 risk/handoff/lifecycle checks, dedicated
session-expiration and transport regressions, 13 selected
connector/gateway CI cases, and 10 selected setup/reconnect UI cases. A
real Composio connection-list call also succeeded through the refreshed
UI on `9ab115f71`.
- UI and server TypeScript checks passed. UI build, Storybook build,
token gates, and diff whitespace checks passed during implementation.
- Real browser and real Paperclip agent tests passed for Arcade,
Composio, and Executor. Tested action permissions, denied agent access,
reconnect, disconnect, and isolation. Tested Arcade catalog
additions/removal and Executor provider approve/resume, decline, and
cancel.
- Zapier live acceptance is incomplete. Its dedicated provider server is
configured, but its credential-copy dialog returned an empty clipboard
through browser automation. No live Zapier action is claimed.
- The three isolated Sentry release cases pass after the CI fixture fix.
- Local verification is deliberately narrow at the maintainer's request.
The full local suite, recursive typecheck, and repository-wide build
were not run. CI provides the broader checks.

## Risks

- Shared MCP transport changes affect other remote MCP servers. Protocol
fixtures cover initialized sessions, streaming response matching,
pagination, and isolation.
- Broad execution tools remain broad permissions. The provider governs
actions inside those tools.
- Provider handoff links are retained briefly in memory. After a server
restart, a one-time link may require reopening the provider dashboard.
Paperclip does not replay the original call.
- Zapier remains unproven live. Custom-header imports and self-hosted
endpoints have fixture coverage rather than a separate live account for
every variant.
- Turning the experimental flag off hides setup; it does not revoke
existing credentials or stop existing connections.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, shell
execution, and browser automation. The exact runtime model ID and
context-window size are not exposed in this session. A separate
Anthropic-backed Paperclip agent performed live gateway acceptance
tasks.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 12:53:11 -05:00
86b7ee992c feat(onboarding): ClipLab sleepy-to-wake hero and step hand-offs (#13629)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents have a persistent visual identity (#13171): a ClipLab
character in one of 17 palettes, rendered as cached PNGs in lists and as
a live character in larger placements.
> - The onboarding wizard is where a person meets that identity first,
and it showed a stock ClipLab expression on the previous engine while
the rest of the app would show a different character on a newer one.
> - The wizard's steps also cut from one screen to the next, so the arc
read as separate pages rather than one walk.
> - This pull request puts one character on one engine everywhere, gives
the wizard's hero the studio's sleepy → wink → idle sequence on Review,
and hands the steps over inside one presence.
> - The benefit is that what wakes on Review is exactly what the agent
looks like on the dashboard afterwards, and the walk to it reads as one
screen changing.

## Linked Issues or Issue Description

Refs #13171, now merged into master. This PR contains the onboarding and
ClipLab update on top of that foundation. Original feature work by
@tonio-alucema; merge preparation preserves the original commits.

**Problem or motivation**

The onboarding hero and the app's avatars were two different characters
on two different ClipLab engines. Steps 1 → 4 of the wizard cut between
screens, and the wizard mounted cold when a cloud-managed workspace
arrived from Cloud's naming screen.

**Proposed solution**

Vendor ClipLab v0.2.0 as the shared engine and render one
studio-exported character from it in every palette, for every pose and
size. Play the export's one-shot wake on Review with the palette fading
in over the gray dormant loop. Hand steps over inside one presence so
the footer slides instead of jumping, and play the arrival half of that
hand-off when the wizard opens directly on the agent step.

**Alternatives considered**

Exporting mp4/webm loops per size: no cursor following, no clean alpha,
and the palette "colour in" is a runtime blend. Minting a `cap-v2`
character version: nothing had shipped `cap-v1`, so the artwork is
regenerated in place instead of migrated. Keeping the separately
vendored runtime bundle for the hero: two engines and two characters in
one app.

## What Changed

- `packages/shared/src/cliplab`: re-vendored from ClipLab v0.2.0
(`987b6db0`) with the Paperclip adaptations replayed (optional graphics
backend for the Node SVG snapshot path, supersampled live textures,
character framing, deterministic SVG id prefixes); new upstream
`particles.ts`.
- `packages/shared/src/cliplab/character.ts`: the studio export,
mirrored from `ui/src/assets/cliplab/onboarding.character.json` by
`scripts/sync-cliplab-character.mjs` (drift caught by
`check:token-gates`). `characterDefinition` builds every palette from
it; the resting portrait is its idle beat.
- `OnboardingCharacter`: gray `sleepy` loop through the agent and
connect steps; on Review the one-shot sleepy → wink → idle plays on two
lock-step canvases while the palette fades in, then the `idle` loop.
Body-follows the pointer, page-scoped. 160px in the wizard.
- `OnboardingWizard`: steps 1 → 2 → 3 → 4 hand over inside one
`AnimatePresence` (departing content fades and gives its room back;
arriving content opens its room then fills); the hero has a room that
opens on the walk into the agent step; opening directly on the agent
step plays the arrival half; the self-hosted naming step uses the arc's
label and field.
- Motion vocabulary in `onboarding-motion.ts` (`stepContentMotion`,
`ledeMotion`, `heroRoomMotion`, `heroRoomArrival`, `titleSwapMotion`).
- Storybook: `Onboarding / Character` (Wake Up), `Onboarding / Agent
arc` walkable from the naming step plus `Arrive From Cloud`; the
companies fixture answers the wizard's create call with a company.
- Uses the shared runtime for onboarding; `doc/agent-personas.md`
documents the shared character.
- Releases both onboarding canvases after partial startup or transition
failure. Registers each canvas before seeking so synchronous render
errors can release it. Six component tests cover these failures and
palette changes before or during wake.
- Refreshes both sleeping canvases when the palette changes, including a
palette change in the same render as wake.
- Moves choreography values into the CSS token layer and preserves the
shared motion catalog drift check across the imported stylesheet.
- Repairs the static Storybook avatar route and uses accessible heading
names/current button labels in the wizard play functions.
- Closes the lazy avatar worker pool during application shutdown.

## Verification

- Merge-preparation checks: `pnpm -r typecheck`, `pnpm build`, `pnpm
build-storybook`, and `pnpm check:token-gates` pass. The final UI
typecheck and 123 focused onboarding, lifecycle, and token catalog tests
pass. All 55 checks on final head
`b4f5e201a1564083d163abc6f93f5b3da06ccefd` pass, including the full
sharded test suite, runner verification, and all eight browser shards
([CI
run](https://github.com/paperclipai/paperclip/actions/runs/35445430535)).
The duplicate monolithic local `pnpm test:run` was stopped after CI
completed; it is not claimed as a separate completed local run.
- Chromium walkthrough: palette change, wake, return to sleep, WebGL
failure fallback, Review step hand-offs and cloud arrival pass with
normal and reduced motion; no browser errors. The signoff happy-path
browser test also passes against a disposable instance.
- The final CI run confirms the catalog fix and a passing signoff
browser shard. The earlier signoff failure was a heartbeat-run
availability timeout; the focused local reproduction and final CI passed
without signoff code changes.
- Original author verification:
- `pnpm check:token-gates` (includes the new character sync check);
shared, server avatar/persona (17) and UI onboarding/persona (137)
suites pass; `pnpm build-storybook` packages all 3,564 avatar PNGs
through the worker pipeline.
- Storybook: `Agents / Personas` Sizes, Expressions and Palettes render
the studio character at every size and pose; `Onboarding / Character →
Wake Up` plays the wake on the shared engine; `Onboarding / Agent arc`
walks 1 → 4 with the hand-offs, and `Arrive From Cloud` plays the
arrival (measured: content room 6 → 65px over 320ms, fade to 1.0 by
~560ms, footer travel continuous).
- The original author walked the agent → connect → review flow and wake
after a real sign-in on staging.
- Not done here: the Linux Storybook visual baselines
(`tests/storybook-visual/agent-personas.spec.ts`) need re-baselining for
the new engine, hero size and naming-step changes.

## Risks

- Every avatar's pixels change (new engine, new character) under the
unchanged `cap-v1` name. Stacks that rendered avatars on the previous
engine keep those PNGs in their cache
(`generated-agent-avatars/cap-v1/...`, served immutable) until cleared;
only the two pinned staging stacks ever did.
- The one-shot handoff to the idle loop is timed from the sequence's
authored duration (the engine reports completion by continuing into idle
itself); presentation only, nothing in the wizard's state waits on it.
- Reduced motion skips the wake and the hand-offs; jsdom is treated the
same way, so the wizard tests see the next step's content immediately.
- The committed export differs from the studio by one animation (Loop
off, leading idle step removed); a re-export without that fix would play
a 5.6s idle before the wake.

## Model Used

Original feature: Anthropic Claude Fable 5.1 (`claude-fable-5-1`) in
Claude Code, with shell, browser, and file tools. The original context
window was not recorded.

Merge preparation and lifecycle regression fixes: OpenAI GPT-6 in Codex,
with reasoning, shell execution, file editing, GitHub CLI, and automated
tests. The session does not expose an exact runtime model ID or
context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Dotta <bippadotta@protonmail.com>
Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 08:30:14 -05:00
1ef3b08714 feat(ui): integrate agent personas across the app (#13171)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - A stable agent persona is useful only when the same identity appears
across the app.
> - Lists, task messages, selectors, and activity feeds need inexpensive
static avatars.
> - Onboarding and agent headers need a larger character with
expressions and pointer tracking.
> - This pull request connects the persona foundation to those existing
views and preserves onboarding draft assignments.
> - Full-page stories and Linux checks make the placements and
performance contract reviewable.

## Linked Issues or Issue Description

**Problem or motivation**

Agents need a stable visual identity in lists, tasks, onboarding, and
configuration. External tools also need an image URL for that identity.

**Proposed solution**

Assign each agent a permanent palette from a fixed ClipLab character
library. Store the assignment on the agent. Render and cache preset PNG
URLs on demand. Use static images in dense views and one animated
character in larger placements.

**Alternatives considered**

A generated image bundle requires a separate asset build. A live
renderer in every avatar adds unnecessary work in large lists. Arbitrary
uploaded images do not provide the requested shared character system.

**Roadmap alignment**

This improves agent identity across existing control-plane views. It
preserves agent permissions, company boundaries, and status labels.
ROADMAP.md has no separate ClipLab persona milestone.

Related approaches: #2422 adds configurable image URLs and DiceBear
generation; #5578 adds optional uploaded avatars. This work uses a
fixed, versioned character library and preset URLs.

## What Changed

- Replace agent icons with static persona images across lists, the
sidebar, org charts, tasks, comments, selectors, activity, and dashboard
views.
- Put one animated character in the agent header. Let it follow the
pointer across the page, with reduced-motion and touch fallbacks.
- Add larger padded characters to agent creation. Keep the palette
stable across draft refreshes and connection retries, then reveal it
after success.
- Pass appearance through shared projections rather than fetching each
agent separately.
- Add real full-page Storybook examples for the agent list, overview,
task, dashboard, new-agent dialog, and connection page.
- Add Linux screenshot, clipping, density, and 500-avatar performance
checks.

## Verification

- `pnpm -r typecheck`, `pnpm build`, and token gates pass on the rebased
tree. Persona lifecycle tests pass.
- The rebased feature passes 38 Linux screenshot/performance checks,
including both display densities, corner pointer positions, and the
no-WebGL/no-live-download contract for 500 avatars.
- The final Linux persona suite passes all 38 visual, lifecycle,
density, and full-page checks using the standard Storybook configuration
and real on-demand avatar endpoint.
- Final local focused verification: 45 avatar/native-recovery tests
pass; UI identity/routine tests, typecheck/build, token gates, and
Storybook build pass.
- Current-head CI passes: full workspace/server tests, all serialized
server groups, typecheck/release checks, build, canary validation, and
end-to-end shards. The build passed after retrying a native-runner
concurrency-test failure; its three targeted cases also pass locally.
- Manual inspection covered stable identities in the app, header
placement, full-page mouse tracking, onboarding size, and task/dashboard
placements.


### Screenshots

Linux captures use synthetic Storybook fixtures. Full-page captures use
reduced motion. The live character, mouse tracking, and disposal are
checked separately.

<details>
<summary>Agent overview with the character in its header</summary>

<img
src="https://raw.githubusercontent.com/paperclipai/paperclip/8c68f42b268ada22b67d79d0fe1bb0a2f84ec25c/screenshots/full-page-agent-overview.png"
width="900" alt="Agent overview with the character in its header" />

</details>
<details>
<summary>Task messages and assignee identity</summary>

<img
src="https://raw.githubusercontent.com/paperclipai/paperclip/8c68f42b268ada22b67d79d0fe1bb0a2f84ec25c/screenshots/full-page-task.png"
width="900" alt="Task messages and assignee identity" />

</details>
<details>
<summary>Larger onboarding character with room for expressions</summary>

<img
src="https://raw.githubusercontent.com/paperclipai/paperclip/8c68f42b268ada22b67d79d0fe1bb0a2f84ec25c/screenshots/full-page-meet-your-next-agent.png"
width="900" alt="Larger onboarding character with room for expressions"
/>

</details>
<details>
<summary>Dashboard agent activity</summary>

<img
src="https://raw.githubusercontent.com/paperclipai/paperclip/8c68f42b268ada22b67d79d0fe1bb0a2f84ec25c/screenshots/full-page-company-dashboard.png"
width="900" alt="Dashboard agent activity" />

</details>

## Risks

- This PR depends on #13170, the persona foundation. Merge the
foundation first, then retarget this PR to master.
- Many placements change from icons to character silhouettes. Human
avatars and authoritative agent status labels retain their existing
behavior.
- Only one character can render live per view. Reduced motion,
hidden/offscreen content, touch input, and renderer failures use the
defined fallbacks.
- The full-page stories use fixture data. They do not contact a real
company or complete real provider sign-in.

## Model Used

OpenAI Codex, GPT-6 family. The exact model identifier and context
window are not exposed in this session. Used code editing, shell
execution, browser inspection, and Linux visual testing.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Tonio <tonework@gmail.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 07:57:53 -05:00
Devin FoleyandPaperclip d54b750111 Preserve Claude ACP quota classification and reset time (#13651)
Typed Claude ACP quota failures lost their recovery classification and reset
time when the runtime reduced provider metadata to a generic category error.
Inspect terminal metadata in memory and retain only safe recovery labels and
a parsed reset timestamp. Preserve the existing handling of other limits.

Verified real child processes on both pinned ACPX runtimes, adapter and
server recovery regressions, all PR CI gates, and Greptile 5/5. Also isolate
a pre-existing chat regression from unrelated fixtures’ retry work.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-18 19:42:28 -07:00
352153b5ed fix: run the cargo-building native-runner CI suite in the Rust-cached vitest lane (#13586)
## Thinking Path

Post-merge of #13557, the slowest check on the freshest fully-green PR
run
([35246999382](https://github.com/paperclipai/paperclip/actions/runs/35246999382))
was `ci / General tests (server (1/12))` at **339s**. The cause is one
suite:
`server/src/services/native-runtime/native-codex-runner.integration.test.ts`
runs 1 test in **277s of a 291s vitest step (95%)** because its
`beforeAll` cargo-builds the Runner release binaries, and the
general-server shards carry no Rust cache — every PR run cold-compiles
the full third-party crate graph. The other 19 suites in that shard
finish in under 70ms each.

The obvious fix (a dedicated Rust-cached matrix lane) requires editing
workflow files, which the available GitHub App credentials cannot push
(`workflows` permission). But the `Verify Paperclip Runner` lanes
**already restore the shared `release-runner-v1` Rust cache read-only**,
and their commands are `pnpm --filter @paperclipai/paperclip-runner
<package script>` — so the suite can move into a Rust-cached lane purely
through script changes.

## What Changed

- `scripts/run-vitest-stable.mjs`: new `general-server-native-runner`
group carrying exactly that suite. Under the PR workflow
(`GITHUB_WORKFLOW == "PR"`, inherited from `pr.yml` by the reusable
`pr-trusted.yml`) the without-chat server shards exclude it and
rebalance to ~211s of tests each. Every other caller — local runs,
`release-verify.yml` under the Release / Cloud readiness workflows —
keeps the suite in the shards, so a renamed or unknown workflow degrades
to today's slower-but-covered behavior instead of dropping coverage.
- `packages/paperclip-runner`: `test:typescript:vitest` now routes
through `scripts/run-pr-vitest-lane.mjs` — the identical
`ensure:eval-build-deps && build:rust && vitest run` chain (shard flags
passed through), plus the native-runner group on the **final PR shard
only** (`--shard=N/M` with `N == M`, i.e. today's `vitest 2/2`, the 122s
lane). With the restored cache the suite's cargo build becomes an
incremental rebuild.
- `scripts/__tests__/run-vitest-stable-shard.test.mjs`: guards pin the
whole contract — PR 12-shard coverage (shards + chat + native-runner =
full server group exactly), Release/local 10-shard runs keep the suite,
`pr.yml` is named `PR`, the vitest lanes partition with exactly one
final shard, the package-script wiring, and the wrapper's shard/workflow
gating via its `--dry-run` plan output.

No workflow files change. `.github/workflows/*` are untouched.

## Verification

- `node --test scripts/__tests__/run-vitest-stable-shard.test.mjs
scripts/__tests__/release-verify-workflow.test.mjs`: **36/36 pass**
locally on this branch (includes the new coverage, wiring, and
wrapper-gating guards). Both files run in CI's `Test general-server
shard partition` / `Test release verify workflow wiring` steps.
- Wrapper `--dry-run` plan matrix verified for all six shard/workflow
combinations plus malformed-shard rejection (pinned as a guard test).
- The executing proof is this PR's own CI: `ci / Verify Paperclip Runner
(vitest 2/2)` must go green while running the native-runner suite (its
log will show the `general-server-native-runner` group after the package
vitest shard), and the 12 `ci / General tests (server (x/12))` shards
must go green without it.

## Risks

- The exclusion keys on `GITHUB_WORKFLOW == "PR"`. Failure mode of a
rename is safe (suite falls back into the server shards, slower but
covered) and the guard test on `pr.yml`'s name makes it loud.
- `vitest 2/2` grows from ~122s to an expected ~210–260s — still well
under the ~306s `vitest 1/2` and ~326s e2e shards, and inside the
20-minute lane timeout. If the cache misses (key drift), the lane pays a
cold compile like the server shard does today; a miss is slow, never
wrong.
- Double-run/coverage-loss combinations are enumerated in the wrapper
header and pinned by tests: each caller runs the suite exactly once.

## Model Used

Claude (Bender agent, Paperclip) — Fable 5.

---
Expected savings once merged: the 339s `server (1/12)` check drops to
~265s-equivalent shard levels (~211s of tests), the slowest `ci /` check
becomes the ~326s e2e shard (~13–33s off PR wall time), and every PR run
stops paying ~4.5 min of billed cold Rust compile.

For the merger (squash): please keep the trailer below in the squash
body to preserve authorship.

`Co-Authored-By: Bender (Fable)
<Paperclip-Paperclip@users.noreply.github.com>`

## Related PRs

Searched the GitHub PR list for prior work on this surface — related
groundwork, none duplicate this change:
- #13457 — restored master's Rust dependency cache on the PR runner lane
(the read-only cache this PR relies on)
- #13500 — made that cache key image-toolchain-independent so
GitHub-hosted PR runners actually hit it
- #13521 — rebalanced PR shards and split the Verify Paperclip Runner
lanes this PR extends
- #13557 — previous health-check iteration (split the runnerd transport
suite); this PR targets the next slowest check

## Checklist

- [x] I have searched GitHub for duplicate or related PRs and linked
them above

Co-authored-by: Bender (Fable) <Paperclip-Paperclip@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-09-18 07:22:12 -07:00
mouse-value-add fcdb3f2499 feat: add optional you.com search integration (#13555)
<!-- Simplified Technical English (ASD-STE100). -->

## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Agents that do research work need live information from the web
> - Paperclip reaches external systems through governed, catalog-based
MCP connections
> - The Apps catalog is data-driven: a researched provider with a hosted
remote MCP server becomes a connectable app with no runtime code change
> - You.com operates a hosted remote MCP server for web search, content
extraction, and research tools
> - The server supports OAuth 2.1 with dynamic client registration, an
API key in a bearer header, and a keyless free profile at a separate
endpoint
> - This pull request adds You.com to the self-serve MCP research ledger
and generates its catalog entry with three connection methods: browser
sign-in, API key, and the keyless free profile
> - The benefit is that an operator can give agents live web search
through the normal connection governance, and the free profile needs no
account at all

## Linked Issues or Issue Description

No public issue exists for this provider. The problem description
follows the new-adapter issue template.

**Agent or provider**

You.com — web search and research tools over a hosted remote MCP server.

**Why this adapter is useful**

Agents that do research, monitoring, or fact-finding tasks need current
web results. You.com exposes web search (`you-search`), live page
extraction (`you-contents`), citation-backed research (`you-research`),
and finance research (`you-finance`) as MCP tools. Any Paperclip company
can connect it in a few clicks. The free profile offers `you-search`
without an account, so a new company can try agent web search at zero
cost and zero setup.

**How the agent is invoked**

Hosted remote MCP server (Streamable HTTP) at `https://api.you.com/mcp`.
Three supported access paths, verified against the live server on
2026-09-16:

- OAuth 2.1 browser sign-in. The server returns a `WWW-Authenticate`
challenge with RFC 9728 protected-resource metadata and advertises a
dynamic client registration endpoint, so Paperclip's automatic DCR path
applies.
- API key. Sent as an `Authorization: Bearer` header per the provider's
official server manifest and docs. Keys come from you.com/platform and
unlock higher rate limits plus the full tool set.
- Keyless free profile at `https://api.you.com/mcp?profile=free`.
Provides a reduced, read-only tool set.

Official docs: https://you.com/docs/build-with-agents/mcp-server

**Are you willing to implement it?**

Yes. Implemented in this pull request.

**Additional context**

Research evidence collected 2026-09-16, from live protocol probes and
official provider sources only:

- Unauthenticated `POST https://api.you.com/mcp` returns HTTP 401 with
`WWW-Authenticate: Bearer
resource_metadata="https://api.you.com/mcp/.well-known/oauth-protected-resource"
scope="Tools offline_access"`.
- RFC 9728 metadata lists one authorization server with scopes `Tools`
and `offline_access`.
- The authorization-server metadata (RFC 8414) publishes authorization,
token, and revocation endpoints, and advertises a
`registration_endpoint`, so DCR is available. No registration was
performed during research, per the runbook's non-registering preflight
rule.
- The keyless free profile answers `initialize` (server `You.com`,
version `4.0.1`), lists the tools `you-search` and `you-discover`, and
executed both tools successfully during the probe.
- The API-key placement matches the provider's official `server.json` in
the youdotcom-oss/mcp repository: header `Authorization`, value `Bearer
<key>`.

## What Changed

- Added You.com (slug `youcom`, wave 4, risk tier S2) to the self-serve
MCP research ledger in
`packages/shared/src/self-serve-mcp-research.json`, and refreshed the
ledger verification date.
- Added the You.com category (`ai`) and API-key header spec to
`scripts/ingest-app-definitions.mjs`.
- Added a You.com case to `specialMethodsFor` that emits three methods:
browser sign-in (`mcp-oauth`, DCR), API key (`mcp-api-key`, bearer
header), and keyless free profile (`mcp-free`, no auth).
- Regenerated `packages/shared/src/app-definitions/youcom.json` and the
generated registry via the ingestion script (`--definitions-only` mode;
no unrelated provider churn).
- Added the official You.com wordmark artwork (light and dark theme
variants, taken from the provider's docs site) under
`ui/public/brands/apps/`, with a manifest entry.
- Updated `packages/shared/src/app-definitions.test.ts`: ledger counts
(47 providers, 44 candidates), store count (48), verification date, and
assertions for the three You.com methods and their endpoints.

## Verification

- `node scripts/ingest-app-definitions.mjs --definitions-only` — passed.
Generated the new definition and registry import only; no other provider
JSON changed.
- `node scripts/check-app-brand-assets.mjs` — passed (71 identities).
- `node --test scripts/app-brand-validation.test.mjs` — passed.
- `pnpm exec vitest run packages/shared/src/app-definitions.test.ts
ui/src/lib/app-brand-assets.test.ts
ui/src/pages/apps/AppLogo.brand-assets.test.tsx` — passed (39 tests).
- `pnpm exec vitest run packages/shared/src/app-definitions.test.ts
server/src/__tests__/tool-access-service.test.ts
server/src/__tests__/generic-mcp-connection.test.ts
server/src/__tests__/tool-connection-removal.test.ts
ui/src/pages/apps/AppsConnect.test.tsx
ui/src/pages/apps/Browse.test.tsx` — passed (181 tests). Two server
suites that require embedded Postgres skipped on this machine by their
own environment gate; the gate is unrelated to this change.
- `pnpm --filter @paperclipai/shared typecheck` — passed.
- `pnpm --filter @paperclipai/ui typecheck` — passed.
- `pnpm --filter @paperclipai/plugin-sdk ensure-build-deps` — passed;
builds `@paperclipai/shared` with the new definition.
- `pnpm test:run` (full Vitest suite) — 7,930 passed, 18 failed, 4,592
skipped. Every failure is environmental on this container: the
embedded-Postgres suites refuse to start because the machine runs as
root, the native runtime suites need the Rust runner binary that this
container cannot build, and one media suite needs a native HEIC binary.
No failure touches the app-catalog, connection, branding, or
shared-package surface; those suites pass locally. CI is the
authoritative gate for the full suite.
- `pnpm --filter @paperclipai/server typecheck` — not completed: the
script's `prepare:runner-vendor` prelude builds the Rust runner, which
cannot build on this container. A direct `tsc --noEmit` reports only
pre-existing errors from the missing vendored runner types; no error
touches this change. No server code is changed.
- Live You.com proof on 2026-09-16 (keyless free profile, real network
calls): preflight 401 challenge with RFC 9728/8414 metadata and DCR
endpoint ✓, `initialize` ✓, `tools/list` ✓, `you-search` call returned
results ✓, `you-discover` call returned results ✓.
- Live proof NOT run: an authenticated OAuth connect and an API-key call
against the full server. This environment has no You.com account or API
key. Per the runbook, this proof stays outstanding and must not be
assumed from the keyless probe. Both paths match the reviewed
`mcp_remote` patterns (DCR and bearer header) used by existing
providers.
- Browser e2e suites not run: opt-in per `AGENTS.md`, and this change
adds catalog data only, with no UI code.

## Risks

- Low risk. The change is catalog data plus generated output. It adds no
runtime code and touches no existing provider.
- The free-profile method is a fixed keyless endpoint. If You.com
changes or removes `?profile=free`, that method breaks and the entry
needs a ledger update. The OAuth and API-key methods do not depend on
it.
- The authenticated tool catalog is discovered live at connect time, so
provider-side tool changes appear through the normal catalog refresh and
quarantine flow, not through this definition.
- Rollback is a single revert; no migration and no state are involved.

## Model Used

- Provider: Zhipu AI, via OpenRouter
- Model: GLM-5.3 (`z-ai/glm-5.3`)
- Context window: 200K tokens
- Capabilities used: tool use (shell, file edits, live HTTP probes),
long-context repository reading
- The change was produced with AI assistance and reviewed by a human
before submission.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` /
`Closes: #` / `Refs: #` OR (b) described the issue in-PR following the
relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-17 10:01:49 -07:00
Devin FoleyandPaperclip 6fe8e30625 feat(apps): add Railway connection and governed deployment tools (#13415)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Apps gives agents governed access to external resources.
> - Operators need to inspect Railway services, read logs, deploy code,
and run container commands.
> - Railway offers hosted MCP with OAuth, but broad remote actions hide
their internal operations.
> - This PR adds a branded connection and fixed direct operations
through the existing gateway.
> - Separate SSH keys enable container commands under the same grants
and policies.
> - Operators can require approval for an action and inspect the
resulting audit record.

## Linked Issues or Issue Description

**Subsystem affected**

Apps catalog, connection setup, gateway execution, and connection
documentation.

**Problem or motivation**

Agents need Railway access through Paperclip. Operators need to grant
and revoke that access, inspect available actions, and govern deployment
and container operations without giving agents provider credentials.

**Proposed solution**

Reuse hosted MCP OAuth, vault storage, catalog discovery, grants, and
the gateway. Probe the actual credential before enabling fixed GraphQL
operations. Use a dedicated grant-owned SSH key for bounded container
commands.

**Alternatives considered**

A catalog entry alone cannot execute the missing operations. The hosted
general agent has opaque internal effects. An unrestricted CLI runtime
can bypass action policy and inherit ambient credentials.

**Roadmap alignment**

This extends the existing MCP Tool Gateway & Apps path and the Connected
Apps direction in ROADMAP.md. It does not add a plugin or parallel
connection service.

Related PRs #311, #939, and #7861 concern hosting Paperclip on Railway.
They do not add this outbound Apps connection. The separate shared
agent-picker fix is #13414 and is not included here.

## What Changed

- Add the generated Railway catalog entry, official marks, provenance,
and OAuth setup guidance.
- Add fixed service/deployment status, bounded logs, and
redeploy/restart/rollback tools. Block source deployment until the
provider can atomically bind the approved repository and commit.
- Verify API access with an explicit workspace before exposing direct
tools.
- Add grant-owned SSH key setup and a bounded runner with host
verification, target checks, isolated state, and cleanup.
- Block the opaque hosted railway-agent and accept-deploy actions.
Preserve normal Allowed defaults and Ask-first policies for other
actions.
- Quarantine new or changed Railway schemas after initial discovery,
including reconnect.
- Add provider, lifecycle, gateway, SSH, UI, and browser fixtures.
Document setup, limitations, and the release checklist.

## Verification

- Security follow-up: removed the unsafe source-deployment mutation.
Direct calls and old active catalog entries are denied before any
upstream request, including normalized aliases. Refresh marks retired
entries disabled. All 386 focused Railway, catalog and gateway tests
passed, and server TypeScript checking passed. Full [GitHub
CI](https://github.com/paperclipai/paperclip/actions/runs/35139421144)
passed on d86530ab9, including typecheck, build, all tests, runner
checks, and browser tests. Superagent passed and confirmed the P2 fix.
Greptile reviewed the same commit at 5/5 with no findings.

- CI follow-up: fixed the missing Railway SSH operation in the OpenAPI
document, including its request schema, operator-only authentication,
and error responses. The failure reproduced locally before the fix; all
403 selected API, Railway, catalog, and artwork tests passed after it.
Synced current master and resolved the catalog/artwork conflicts.

- After rebase: 440 focused provider, lifecycle, gateway, catalog, and
container-panel tests passed. AppDetail and AppsConnect passed another
196 tests.
- Full typecheck, build, token gates, and the gallery browser check
passed after rebase.
- During implementation, full build and the gallery browser check
passed. Shared generic-MCP fixtures covered OAuth callback/state/issuer
binding and failure paths.
- Local live consent and tools/list succeeded. There were 44 active
hosted actions and two blocked actions. A workspace-bound API probe and
direct project/service/environment reads succeeded. The inspected
project had no deployed services. No provider mutation ran.
- Full GitHub CI passed on commit 303340f19, including all
server/workspace test groups, typecheck, build, runtime verification,
release dry run, and browser tests. The original local full-run attempt
was incomplete; the complete automated suite is now verified in CI.

Manual review: connect Railway, review the actual actions, install for
an agent, and run a resource read through the gateway. Choose Ask first
before testing a deployment mutation. Configure a dedicated key only
when container access is needed.

**Release qualification is still open.** Live agent gateway reads/logs,
rejected and approved deployment calls, refresh/revoke, public HTTPS
consent, and SSH enrollment/commands/cleanup need an authorized
disposable service. The passing API diagnostic does not replace those
tests. See doc/connections/RAILWAY.md and RAILWAY-REVIEW.md.

## Risks

Overall risk is medium. New runtime behavior is gated to Railway
connections, but the PR changes shared catalog, credential lifecycle,
and gateway code. A regression in those paths can affect other Apps
connections. The highest-impact operations are Railway deployments and
container commands.

- Provider consent can authorize an entire workspace. Catalog labels are
not local resource allowlists. Direct tools check target membership, and
provider permissions still apply.
- Shell commands have broad internal authority. Action policy cannot
approve each internal shell step. Timeouts close the local connection
but cannot guarantee remote child-process termination.
- Log and command output may contain application secrets that pattern
redaction cannot recognize.
- Source deployment is unavailable until the provider supports atomic
repository/commit binding. Existing deployments can still be redeployed,
restarted or rolled back.
- No database migration is required. Rollback can remove promotion and
direct dispatch while preserving connection data and the generic MCP
path.
- Live Railway qualification must still pass before release acceptance.

## Model Used

OpenAI Codex, based on GPT-6, with code execution and browser testing.
An independent read-only security agent reviewed the local
implementation. The exact serving model ID and context window were not
exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-16 13:54:44 -07:00
Devin Foley d08abcba15 ci: cut PR wall clock from ~16 to ~6 minutes (#13521)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Every pull request runs the Trusted PR CI workflow before merge
> - The test suites roughly tripled in six weeks, and shard balance did
not keep up, so PR runs crept from ~4 to ~17 minutes
> - Slow CI delays every merge and every contributor
> - This pull request rebalances the shards from fresh measurements,
splits the largest test files, reuses the Rust build cache in three more
jobs, and takes the policy job off the critical path
> - The benefit is a PR wall clock near 6 minutes with the same coverage

## Linked Issues or Issue Description

**What existing behavior does this improve?**

PR CI wall clock. A typical green run took 16-17 minutes. Two months ago
it took about 4 minutes.

**Subsystem affected**

The Trusted PR CI workflow (`.github/workflows/pr-trusted.yml`), the
shard-duration manifests, the vitest shard runner scripts, the
`paperclip-runner` package scripts, and the dry-run branch of
`release.sh`.

**Current behavior**

The shard-duration manifests were stale. The general-server manifest had
durations for ~400 of 649 suites. The e2e manifest was missing 14 of 29
specs. Stale median weights made shard steps range 417s-806s (server)
and 277s-745s (e2e). Three jobs each paid a ~3m40s cold cargo release
build. Every test lane waited ~60s for the policy job before it could
start.

**Proposed behavior**

All lanes finish in a narrow ~200-290s band. The manifests carry fresh
measured durations for every suite. The three largest test files are
split so no single file caps a shard. The Rust cache restore runs in
every job that builds the Runner binary. Test lanes start as soon as the
gate resolves.

**Reason and benefit**

Merges stop waiting on CI. The projected wall clock is ~6 minutes for
the same test coverage.

## What Changed

- Rebuild `scripts/general-server-shard-durations.json` (646 suites) and
`scripts/e2e-shard-durations.json` (all specs) from per-suite completion
timestamps in runs 35036001734 and 35024948947.
- Move the PR server lane to the release-verify shape:
`general-server-without-chat` across twelve duration-balanced shards,
plus the chat integration suite split by collected test location across
three dedicated lanes.
- Split `tests/e2e/chat-adapters-ui.spec.ts` into `-providers` and
`-messaging` specs, and `tests/e2e/agent-chat.spec.ts` into `-sessions`
and `-projects` specs. Each pair shares fixtures through a `.shared.ts`
module. Playwright collects the same test sets (39 and 20 tests).
- Raise e2e shards to eight and serialized shards to nine.
- Run the runner package's `check:all` as four matrix lanes:
`check:static`, `check:runner`, and two native vitest `--shard` halves.
The union is exactly `check:all`.
- Add the read-only Rust cache restore (toolchain pin, `save-if: false`)
to the Canary Dry Run, Build, and Typecheck jobs.
- Make release.sh preview publish payloads concurrently in batches of
eight during `--dry-run`. The real publish path stays strictly serial.
- Drop the policy-job lockfile artifact chain. Each lane installs with
`--frozen-lockfile` and falls back to an inline `--resolution-only`
regeneration. The policy job stays a required check through the `verify`
and `e2e` aggregates.
- Update the shard-count mirrors and workflow assertions in the
partition and gate tests.

## Verification

- `node --test scripts/__tests__/run-vitest-stable-shard.test.mjs
scripts/__tests__/e2e-shard.test.mjs` — 30 pass.
- `node --test '.github/scripts/tests/'*.test.mjs` — 410 pass.
- `node --test scripts/__tests__/release-verify-workflow.test.mjs
scripts/cloud-source-verification.test.mjs
scripts/__tests__/release-dry-run-notes.test.mjs` — 42 pass.
- `playwright test --list` collects 39 tests across the chat-adapters
split and 20 across the agent-chat split, equal to the original files.
- A local vitest collection of the chat suite partitions 995 tests into
498/497 line shards.
- Projected shard weights: server 230s x12, chat ~143s x3, e2e 207-242s
x8, serialized ~216s x9.

## Risks

- The split spec files reorder tests relative to the original files.
Every describe seeds its own company, so the specs stay independent; a
hidden cross-describe dependency would surface as a deterministic
failure in one shard.
- The inline lockfile fallback changes install behavior for
manifest-changing and stacked PRs. The policy job still validates
resolution as a required check.
- `release.sh` changes are confined to the `--dry-run` preview branch.
The publish loop is untouched. `bash -n` passes and the release dry-run
tests pass.
- One PR now schedules ~44 fleet runners. If the RunsOn fleet caps
concurrency, queueing may absorb part of the gain; watch the first runs.

## Model Used

- Claude Fable 5 (`claude-fable-5`, Anthropic), extended thinking, with
tool use (shell, file edits) in Claude Code.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-16 11:45:14 -07:00
DottaandPaperclip 18989a9e73 docs: add eval guide, authoring skills, and public history hub (#13535)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Its evaluations test both the Runner and complete product workflows.
> - The guides and run histories are in separate places.
> - The shared Evalbook viewer can make the test boundary unclear.
> - This pull request names the two families and adds a guide, authoring
skills, and a public hub.
> - Contributors can choose the correct test and inspect its history.

## Linked Issues or Issue Description

**Issue type**

Missing documentation.

**Where is the issue?**

Runner and Product E2E evaluation guides, case-authoring procedures, and
public result navigation.

**What's wrong?**

There is no single entry point. A report format can be mistaken for an
execution boundary. There are no dedicated case-authoring skills for
these two families.

**Suggested fix**

Add a guide and three skills. Link both existing histories from a public
hub. Keep existing campaign URLs and grading unchanged.

## What Changed

- Add `doc/evals.md` and links from existing guides.
- Add the `paperclip-evals`, `add-runner-eval`, and
`add-product-e2e-eval` skills. Install copies in
`~/paperclipai/.agents/skills`.
- Add a static hub builder that reads the existing public history feeds.
- Show a dated snapshot for each family. Label partial campaigns and
preserve measurement dates across report refreshes.
- Document publication and refresh commands for
https://pages.paperclip.ing/evals/.

## Verification

- Seven Python summary tests pass: `python3 -m unittest discover -s
scripts/evals-hub -p 'test_*.py'`. Run these checks directly; this PR
does not modify package scripts.
- All three skills pass the skill-creator `quick_validate.py` check with
`/usr/bin/python3`.
- Build tested with saved history fixtures and the live public feeds.
- Desktop and mobile browser checks pass. The mobile page has no
horizontal overflow.
- Published https://pages.paperclip.ing/evals/. Browser check: HTTP 200,
no page errors, all eight links return HTTP 200, no mobile overflow.
- Independent skill exercises found the existing Notion-decline case and
a direct Runner permission-denial case. Roster validation with an
explicit run ID passes.
- Missing refresh measurement date: regression fails before the fix and
passes after it.
- `git diff --check` passes.
- No paid evals were run for this documentation and reporting change.
The preceding head passed typecheck, build, server/workspace tests,
runner verification, browser E2E, and the canary dry run. Checks for the
latest commit are pending. Local repository-wide typecheck, test, and
build were not repeated because no product code changed.

## Risks

The hub is a dated static snapshot. It can lag behind the linked
histories until an operator refreshes it. A changed history schema stops
the build. Existing archives and grades are not modified. The published
guide link is pinned to the reviewed commit so branch deletion cannot
break it. Later builds can use master.

## Model Used

OpenAI gpt-6-astra for implementation and review. OpenAI gpt-5.6-luna
for documentation and independent skill checks. Both used repository
tools and code execution. Context window sizes are not exposed by this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green (latest commit pending)
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
(preceding head was 5/5; latest commit pending)
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-16 08:41:06 -05:00
a8d32e5e61 feat(sandbox-providers): add CreateOS sandbox provider (#13434)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Agent work runs in sandboxes that provider plugins supply
> - Operators can choose a provider to run agent work
> - CreateOS adds another provider with workspace-preserving pause and
resume
> - This pull request adds a CreateOS provider plugin
> - The benefit is that operators can preserve a workspace between runs
without keeping its compute active

## Linked Issues or Issue Description

Refs #13203 and the earlier closed #13096.

This continues the CreateOS contribution from @bhautikchudasama and
@ashwaq06. The branch preserves the original implementation commit.
Thank you to both contributors.

When squash-merging, preserve the original author's credit in the squash
commit body:

```text
Co-Authored-By: bhautikchudasama <BhautikChudasama@users.noreply.github.com>
```

The original fork rejects maintainer pushes. This branch includes the
merge-conflict resolution and review fixes. The request is described
below using `adapter_request.yml`.

**Agent or provider**

CreateOS sandbox API (https://api.sb.createos.sh).

**Why this adapter is useful**

CreateOS can pause a sandbox and resume it by ID. The workspace survives
the pause. This adds a reusable-lease option to the existing sandbox
provider system.

**How the agent is invoked**

Build and install the local plugin as described in its README. Open
Instance Settings, then Environments. Select the `createos` driver.
Supply an API key and shape. The driver then supplies sandbox leases for
agent runs.

**Are you willing to implement it?**

Yes. This pull request is the implementation.

## What Changed

- Adds the `createos` sandbox provider under
`packages/plugins/sandbox-providers/createos`.
- Calls the CreateOS HTTP API directly. The package adds no vendor SDK.
- Implements the environment lifecycle hooks, incremental process
output, and binary workspace sync.
- Registers the optional bundled provider and its trusted host
credential fallback. The fallback is limited to the official API origin;
custom endpoints require an explicit key.
- Lists the package in the release manifest with `publishFromCi: false`
until its first npm publish is bootstrapped.
- Waits through delayed pause/resume state updates without duplicate
action requests.
- Cancels queued API requests promptly while preserving request spacing.
- Uses direct CLI invocation in the setup guide so paths and IDs are
passed without an extra shell expansion.
- Includes current master and retains its existing Git-subfolder
containment fix.

## Demo

Fresh setup and a run against a CreateOS sandbox.


https://github.com/user-attachments/assets/e71b9e06-c006-4fb9-b847-52dfd68f6110


https://github.com/user-attachments/assets/43b5ac75-66bd-4f76-8563-67e4c7759084

## Verification

All 25 jobs in [CI run
34884260542](https://github.com/paperclipai/paperclip/actions/runs/34884260542)
passed at commit `f8d0997677024b784fdadf9d44a84c01cb4e813c`, including
typecheck, build, native runner verification, server and workspace
tests, browser tests, and the canary release dry run. Greptile reviewed
the same commit at 5/5 with no unresolved review threads.

GitHub reports no merge conflicts. The remaining merge gate is
code-owner approval for the new `package.json`, as required by
`.github/CODEOWNERS` and the `master` ruleset. Reviewers have been
requested automatically.

Local checks passed:

- Provider: `pnpm typecheck`, `pnpm test` (52 passed, one live smoke
skipped), and `pnpm build`.
- Host: focused credential and bundled-plugin tests (17 passed), plus
CLI invocation safety (39 passed).
- Release: package manifest check and release policy tests (18 passed).

The full local `pnpm test:run` attempt caught the README command issue;
its focused rerun now passes. The full local run stopped after its
general-server group: 7,804 tests passed, with unrelated embedded
PostgreSQL startup failures and 10 failures in unchanged
runtime-skill-cache tests (`EACCES` on directory rename on macOS). It
did not reach the later test groups. Local `pnpm -r typecheck` and `pnpm
build` reach the runner package and stop because this machine has no
Rust/Cargo installation. The corresponding CI checks passed on
provisioned runners, as linked above.

The live CreateOS smoke requires explicit provider credentials and was
not run during this review. It is available with `CREATEOS_LIVE_TEST=1
pnpm test` in the provider directory. The author supplied the demo links
above.

## Risks

The provider is opt-in and is not installed by default. It is available
through a local-path install or explicit image inclusion. npm
publication remains disabled until a maintainer bootstraps the package
and enables publishing.

Sandbox creation has no idempotency key. An ambiguous create response
can leave a resource that requires provider-account inspection. Process
tracking is in memory; durable lease recovery belongs to the host. The
provider does not advertise guaranteed expiry, interactive login,
snapshots, duplex channels, or ingress. Live native-runner qualification
remains outside this PR's tested claims.

## Model Used

Original provider implementation: human-authored by @bhautikchudasama,
as reported in #13203. The original description reports Claude Opus 5
assistance.

Review and follow-up fixes: OpenAI GPT-6 via Codex, with code review,
editing, and tool execution. The precise runtime model variant and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused checks; full-suite
environment limits documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: bhautikchudasama <bhautikrchudasama@gmail.com>
Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-15 17:09:43 -07:00