Commit Graph
4468 Commits
Author SHA1 Message Date
DottaandPaperclip 924f07be8c feat(chat): simplify Slack onboarding and account linking (#13638)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Chat connections let people start and continue that work from Slack.
> - Setup mixed app creation, credentials, URL verification, account
linking, and testing on the same screens.
> - People also needed a safe way to link their own Slack identity after
the first operator finished setup.
> - This pull request gives each step a clear place and keeps membership
approval separate from identity linking.
> - It also makes connection details easier to use and fixes misleading
callback health behind HTTPS proxies.

## Linked Issues or Issue Description

**Subsystem affected**
Cross-cutting: chat routes and services, shared contracts, and the Apps
board UI.

**Problem or motivation**
Slack onboarding made users find settings without enough guidance. A
second user needed operator help to link their account. Activity stopped
at 100 records, and TLS termination could mark working callbacks as
stale.

**Proposed solution**
Use six setup steps with editable app names, a generated manifest,
credential guidance, URL verification, account linking, and an optional
message test. Send each Slack user a private, expiring confirmation
link. Require company membership or an approved access request before
linking. Add cursor pagination and tolerate the internal HTTP hop in
callback diagnostics.

**Roadmap alignment**
This improves the existing connected-app surface and supports CEO Chat
without changing the task-and-comments model. The maintainer requested
and reviewed the flow during a live Slack test drive.

**Additional context**
Related work: #7, #3349, #13000, and #13620. Those cover broader chat
capabilities, older webhook paths, or plugins. This PR improves the
existing native connector's setup and account-linking flow. HTTPS
documentation was published separately in
paperclipai/paperclip-docs#128.

## What Changed

- Split Slack onboarding into six clickable sidebar steps. Keep
secondary and primary actions on one row.
- Generate the Slack creation link and read-only manifest from editable
app, bot, and command names. Add credential prefix validation and direct
instructions.
- Add live account-link status and an optional mention-based message
test.
- Add private, single-use Slack account invitations and membership
access requests. Retain cloud authentication/bootstrap checks and
enforce the chat rollout flag in all identity APIs. Default new Slack
connections to linked users only.
- Put Settings, Access, Conversations, and Activity in the sidebar.
Simplify conversation rows and remove active header badges.
- Add 25-item activity pages, stable timestamp/ID cursors, and replay
safety across pages. Preserve the legacy array API for clients without
pagination parameters.
- Fix false callback warnings when HTTPS terminates at a proxy. Keep
host, port, and path drift detection.
- Document the setup flow, pagination, callback diagnostics, and shared
wizard footer rule.

## Verification

- Passed: `pnpm -r typecheck`, `pnpm build`, and `pnpm
check:token-gates`.
- Passed: focused Slack callback and pagination integration tests; UI
clipboard, wizard, pagination, and activity tests; OpenAPI route tests.
The final access-gate fix also passes 27 focused tests covering cloud
authentication/bootstrap, nonmember invitations, token validity, and the
server-enforced rollout flag.
- Passed: all 1,002 chat integration tests, 6,356 UI tests, and all 11
provider browser scenarios (including mobile light/dark navigation).
After rebase, the identity route, sidebar, and 25 clipboard tests pass.
- The full local `pnpm test:run` was attempted. The first run found 14
Slack fixtures that needed explicit guest access; those are fixed and
the complete chat suite passes. Unrelated embedded PostgreSQL
startup/resource failures and timeouts prevented a clean full local run.
All CI checks pass on `2d858b036`, including the full chat, server,
workspace, build, typecheck, and browser suites.
- Live test drive: Slack app creation, credential setup, URL
verification, private account confirmation, mention messages, and thread
replies. Verified the callback warning clears for the existing proxied
connection.
- Review: create a Slack connection, follow the six steps, link a second
user's account, and browse older activity with Next and Previous.

## Risks

- Identity invitations carry a temporary capability. Tokens are hashed,
expire after 15 minutes, work once, and require explicit confirmation by
a company member. Access requests do not grant membership.
- New Slack connections reject unlinked people by default. Existing
connection settings remain intact.
- Activity is a live ledger. Updated action rows can move forward in
time. Older pages do not poll.
- Proxy tolerance affects health display only. Slack signature checks
and proxy authentication settings remain unchanged.
- No database migration or package-lock changes.

## Model Used

OpenAI Codex, based on GPT-6, with reasoning, repository tools, code
execution, and browser verification. The runtime does not expose an
exact model build ID or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (targeted suites; full
local-run limitations documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.918.0-canary.7
2026-09-18 17:23:53 -05:00
Devin FoleyandPaperclip 4b8dc416da Fix default isolation for projects without workspace configuration (#13636)
Require a company-scoped configured workspace before applying the operator default for Git worktree isolation. Projects that only have a plain managed directory retain their existing behavior. Explicit isolation requests still require a valid checkout.

Add policy and heartbeat integration regressions and document the default. The regression fails before the fix. An isolated checkout passes 541 relevant tests, and the server TypeScript check passes. Local repo-wide typecheck and build require the missing Rust toolchain; all CI lanes passed, and Greptile reviewed the refreshed head at 5/5 with no comments.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
canary/v2026.918.0-canary.6
2026-09-18 14:49:45 -07:00
scotttongandClaude Opus 5 8f69e0af9d fix(ui): confirm every copy action (#13603)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Much of the work means moving an opaque value somewhere else: an
agent id, a callback URL, a webhook secret, an error payload for a bug
report
> - Copy buttons all over the app do that, and most already say whether
it worked — an icon that flips to a check, a label that reads "Copied"
> - Eleven did not. They passed the promise to `.catch(() => {})` and
showed nothing at all
> - A copy that says nothing looks exactly like a copy the browser
refused, and the only way to tell is to paste somewhere and look
> - The clipboard really does refuse: over plain HTTP on a non-secure
host the async Clipboard API is unavailable, and the fallback can still
fail
> - This pull request gives every copy action an affordance, through two
shared hooks
> - The benefit is that a person can tell a copy from a failure without
leaving the page

## Linked Issues or Issue Description

**What happened?**

Eleven copy buttons gave no feedback. They called `copyTextToClipboard`
and discarded the result:

```
onClick={() => void copyTextToClipboard(robotEmail).catch(() => {})}
```

Nothing changes on screen, whether the write succeeded or failed. This
is the same surface where most other copy buttons do show a check or a
"Copied" label, so the silent ones read as broken.

**Expected behavior**

Every copy action says whether the clipboard took the value. A rejected
write says so rather than staying silent or claiming success.

**Steps to reproduce**

1. Open the connector setup flow for an app that shows a sharing email
or a callback URL.
2. Press **Copy** next to that value.
3. Before this change nothing on screen changes. Compare with the copy
button on a login panel, which flips to a check.
4. To see the failure case, open Paperclip over plain HTTP on a
non-localhost host, where the async Clipboard API is not available.

**Paperclip version or commit**

`45586170e` on `master`.

**Deployment mode**

Local.

**Agent adapter(s) involved**

Not adapter-specific (core bug).

**Additional context**

Commit `1fa36be35` moved every copy through `lib/clipboard.ts`, so the
call sites are findable. Of 62 call sites, 51 already showed feedback
and 11 did not.

## What Changed

- New `ui/src/lib/use-copy-action.ts` with two hooks:
- `useCopyAction` for a control that stays on screen. It returns
`copied` and `failed` for the inline swap the rest of the app already
uses, and resets itself.
- `useCopyToast` for a menu item, which unmounts with its menu before an
inline state could be read, so its confirmation goes to the toast
viewport.
- Both wait for the write to resolve before reporting success, and both
report a rejection as a failure. `AdapterLoginChrome.test.tsx` already
held that line for one button; the helper now makes it the default.
- The eleven silent sites now use one of the two: the connector setup
flow's sharing-email and callback-URL buttons, the runner inspector's
copy-value and copy-path buttons, the agent bubble and skill studio menu
items, annotation "Copy link" on both its hosts, both secret-error
detail buttons, the chat webhook secret, and the motion tweak panel's
export box.
- The chat webhook secret follows its own file's convention: a label
that reads "Webhook secret copied", matching the manifest buttons beside
it.
- The 51 sites that already had feedback are untouched. Converting them
would be a large diff with no visible change.
- `clipboard-usage.test.ts` gains a static check, sibling to the one
that already keeps copies on the shared helper. It fails if a new copy
site ships with no affordance.

## Verification

- `npx vitest run ui/src/lib/use-copy-action.test.tsx
ui/src/lib/clipboard-usage.test.ts ui/src/lib/clipboard.test.ts
ui/src/components/AdapterLoginChrome.test.tsx` — 18 tests pass.
- The new hook tests cover the three cases that matter: no confirmation
while the write is still in flight, a failure state on a rejected write,
and a return to rest after the reset delay. The toast hook is covered
for both tones.
- `npx vitest run ui/src/pages/apps/AppsConnect.test.tsx
ui/src/components/task-chat ui/src/components/RunnerInspector.test.tsx
ui/src/components/DocumentAnnotation` — the suites over the touched
components pass.
- `npx tsc -b ui` — clean.

## Risks

Low. Each change is additive at its own call site and the copy path
itself is unchanged. The static check is the only part that touches
future contributors: it is one assertion, and its pattern list is easy
to extend or drop.

## Model Used

Claude Opus 5 (`claude-opus-5`), 1M context, extended thinking, with
tool use and code execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [ ] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
canary/v2026.918.0-canary.5
2026-09-18 14:32:19 -07:00
scotttongandClaude Opus 5 34355e5109 fix(task-chat): selecting an option no longer jumps to the next question (#13602)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - An agent that needs a decision asks in the task composer, which
renders a question set one question per page
> - Single-choice questions use radio options, and the composer also
ships a "Next" button and pagination arrows
> - Selecting a radio option also set a pending advance, waited for a
short confirm animation, and turned the page on its own
> - That takes the page away from the reader while they are still
reading it, and a misclick costs them the question
> - Nothing on the option says that clicking it will navigate
> - This pull request makes selection answer the question and nothing
else
> - The benefit is that moving between questions is always something the
reader chose to do

## Linked Issues or Issue Description

**What happened?**

In the task composer, selecting an option with a radio button jumps to
the next question. `QuestionForm.toggleOption()` set `pendingAdvance`
for any single-select question that was not on the last page, and an
effect then read `--motion-question-confirm`, waited that long, and
called `setPage(page + 1)`. With reduced motion it advanced immediately,
with no pause at all.

The result is that a click meant to answer a question also navigates
away from it. There is no way to read the rest of the page after
choosing, and no way to undo the jump other than pressing the back
arrow.

**Expected behavior**

Selecting an option records the answer and stays on the question. Moving
to the next question stays an explicit act.

**Steps to reproduce**

1. Get an agent to ask a question set with two or more single-choice
questions.
2. Open the question in the task composer.
3. Click one of the radio options.
4. Before this change the composer moves to the next question on its
own.

**Paperclip version or commit**

`45586170e` on `master`.

**Deployment mode**

Local.

**Agent adapter(s) involved**

Not adapter-specific (core bug).

**Additional context**

Every route forward already shipped beside the options, so removing the
shortcut traps no one:

- a footer button that reads "Next" on any page but the last, and the
submit label on the last;
- previous and next pagination arrows;
- "Skip" on questions that are not required.

The same question sets render outside the composer in
`IssueThreadInteractionCard`, which never had the auto-advance. This
change brings the two surfaces to the same behavior.

## What Changed

- `QuestionForm` no longer sets a pending advance when a single-choice
option is selected. The effect that turned the page is removed with it.
- Selecting an option still clears a submit error about a missing
answer, which is the case that error is about.
- The confirm-before-advance animation existed only to soften the jump.
Its `--motion-question-confirm` token, its keyframes, its class, and the
now-unused `confirming` prop on the option button are removed.
- Tests that asserted the jump now assert the opposite: selection leaves
the page number unchanged, "Next" advances, and number-key selection
stays on the same question.

## Verification

- `npx vitest run ui/src/components/task-chat` — the composer,
interaction card, protocol card, and motion token suites pass.
- `npx vitest run ui/src/components/task-chat/TaskChatComposer.test.tsx`
— 87 tests pass, including four new tests for the stationary behavior.
- `npx tsc -b ui` — clean.
- Keyboard: options keep `role="radio"` inside a `radiogroup`, and a
test asserts that number-key selection selects without navigating.

## Risks

Low. One deliberate behavior is removed. It costs a reader one extra
click per question on multi-question sets, and the explicit control for
that click already existed and is already tested.

## Model Used

Claude Opus 5 (`claude-opus-5`), 1M context, extended thinking, with
tool use and code execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [ ] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-18 14:18:39 -07:00
scotttongandClaude Opus 5 378f11994b fix(connections): stop the sign-in screen from adding a phantom step (#13601)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Agents reach other products through connections, which people create
in the connector setup wizard
> - That wizard shows a stepper: a row of dots, a label under each, and
"Step 1 of 2" in the header
> - When a curated OAuth app hands off to the provider, a waiting screen
replaces the wizard while Paperclip prepares sign-in
> - That waiting screen drew its own stepper with three hardcoded
labels, so the stepper grew a dot at the exact moment the person pressed
Connect
> - A stepper that grows mid-flow tells the person they missed a screen
> - This pull request makes the waiting screen show the step model of
the flow that opened it
> - The benefit is that the stepper keeps the same shape from the first
screen to the handoff

## Linked Issues or Issue Description

**What happened?**

The connector setup wizard shows a two-step stepper for a curated OAuth
app: two dots, labelled "Access" and "Sign in", with "Step 1 of 2" and
then "Step 2 of 2" in the header. When you press Connect, a third dot
appears, labelled "Ready".

The extra dot comes from the OAuth waiting screen.
`OAuthConnectStateScreen` rendered its own header with
`labels={["Access", "Sign in", "Ready"]}`, a constant that did not
depend on the wizard that opened it.

"Ready" is also not a step the wizard can reach. The success screen
hides the step header, so no one ever sees a third step become active.

**Expected behavior**

The stepper keeps the same number of steps from the first screen through
the provider handoff. Waiting for browser sign-in is part of the last
step, not a step of its own.

**Steps to reproduce**

1. Open **Apps → Connect** and pick a curated app that signs in through
OAuth, for example Notion.
2. Look at the stepper. It shows two dots and reads "Step 1 of 2".
3. Complete the Access step and press the button that starts sign-in.
4. Look at the stepper on the "Preparing secure sign-in" screen. Before
this change it shows three dots and adds a "Ready" label.

**Paperclip version or commit**

`45586170e` on `master`.

**Deployment mode**

Local.

**Agent adapter(s) involved**

Not adapter-specific (core bug).

**Additional context**

The reporter suggested removing the stepper, on the condition that no
connector flow has three or more steps. One does.
`ConnectionSetupFlow.tsx` keeps `STEP_LABELS = ["Pick app", "Access",
"Add your key"]` for the generic path, which a person walks when they
paste an MCP endpoint instead of choosing a curated app. The two-step
counts apply only after an app is selected. The condition fails, so this
pull request keeps the stepper and repairs the phantom step alone.

## What Changed

- `OAuthConnectStateScreen` takes a `steps` prop: the labels and active
index of the flow that opened it.
- The curated OAuth path passes its own two-step model, so the count
does not change at the handoff.
- The generic pasted-endpoint path passes the three-step model it was
already walking, so its count does not change either.
- The default for hosts without a wizard of their own, such as the
paste-a-config tab, is `["Access", "Sign in"]`. The invented "Ready"
step is gone.
- `AppsConnect.test.tsx` gains a test that reads the stepper before and
after the handoff and fails if the two differ.

## Verification

- `npx vitest run ui/src/pages/apps/AppsConnect.test.tsx
ui/src/features/connections ui/src/pages/tools/PasteConfigTab.test.tsx`
— 185 tests pass.
- `npx vitest run ui/src/pages/apps/generic-mcp-connect.test.ts` — 27
tests pass.
- `npx tsc -b ui` — clean.
- The new test fails on the previous code. With the old three-label
constant restored it reports `["Access", "Sign in", "Ready"]` where it
expects two labels.

## Risks

Low. The change is limited to which labels the waiting screen draws. No
connection logic, no network call, and no navigation changes.

## Model Used

Claude Opus 5 (`claude-opus-5`), 1M context, extended thinking, with
tool use and code execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [ ] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
canary/v2026.918.0-canary.4
2026-09-18 14:11:35 -07:00
352153b5ed fix: run the cargo-building native-runner CI suite in the Rust-cached vitest lane (#13586)
## Thinking Path

Post-merge of #13557, the slowest check on the freshest fully-green PR
run
([35246999382](https://github.com/paperclipai/paperclip/actions/runs/35246999382))
was `ci / General tests (server (1/12))` at **339s**. The cause is one
suite:
`server/src/services/native-runtime/native-codex-runner.integration.test.ts`
runs 1 test in **277s of a 291s vitest step (95%)** because its
`beforeAll` cargo-builds the Runner release binaries, and the
general-server shards carry no Rust cache — every PR run cold-compiles
the full third-party crate graph. The other 19 suites in that shard
finish in under 70ms each.

The obvious fix (a dedicated Rust-cached matrix lane) requires editing
workflow files, which the available GitHub App credentials cannot push
(`workflows` permission). But the `Verify Paperclip Runner` lanes
**already restore the shared `release-runner-v1` Rust cache read-only**,
and their commands are `pnpm --filter @paperclipai/paperclip-runner
<package script>` — so the suite can move into a Rust-cached lane purely
through script changes.

## What Changed

- `scripts/run-vitest-stable.mjs`: new `general-server-native-runner`
group carrying exactly that suite. Under the PR workflow
(`GITHUB_WORKFLOW == "PR"`, inherited from `pr.yml` by the reusable
`pr-trusted.yml`) the without-chat server shards exclude it and
rebalance to ~211s of tests each. Every other caller — local runs,
`release-verify.yml` under the Release / Cloud readiness workflows —
keeps the suite in the shards, so a renamed or unknown workflow degrades
to today's slower-but-covered behavior instead of dropping coverage.
- `packages/paperclip-runner`: `test:typescript:vitest` now routes
through `scripts/run-pr-vitest-lane.mjs` — the identical
`ensure:eval-build-deps && build:rust && vitest run` chain (shard flags
passed through), plus the native-runner group on the **final PR shard
only** (`--shard=N/M` with `N == M`, i.e. today's `vitest 2/2`, the 122s
lane). With the restored cache the suite's cargo build becomes an
incremental rebuild.
- `scripts/__tests__/run-vitest-stable-shard.test.mjs`: guards pin the
whole contract — PR 12-shard coverage (shards + chat + native-runner =
full server group exactly), Release/local 10-shard runs keep the suite,
`pr.yml` is named `PR`, the vitest lanes partition with exactly one
final shard, the package-script wiring, and the wrapper's shard/workflow
gating via its `--dry-run` plan output.

No workflow files change. `.github/workflows/*` are untouched.

## Verification

- `node --test scripts/__tests__/run-vitest-stable-shard.test.mjs
scripts/__tests__/release-verify-workflow.test.mjs`: **36/36 pass**
locally on this branch (includes the new coverage, wiring, and
wrapper-gating guards). Both files run in CI's `Test general-server
shard partition` / `Test release verify workflow wiring` steps.
- Wrapper `--dry-run` plan matrix verified for all six shard/workflow
combinations plus malformed-shard rejection (pinned as a guard test).
- The executing proof is this PR's own CI: `ci / Verify Paperclip Runner
(vitest 2/2)` must go green while running the native-runner suite (its
log will show the `general-server-native-runner` group after the package
vitest shard), and the 12 `ci / General tests (server (x/12))` shards
must go green without it.

## Risks

- The exclusion keys on `GITHUB_WORKFLOW == "PR"`. Failure mode of a
rename is safe (suite falls back into the server shards, slower but
covered) and the guard test on `pr.yml`'s name makes it loud.
- `vitest 2/2` grows from ~122s to an expected ~210–260s — still well
under the ~306s `vitest 1/2` and ~326s e2e shards, and inside the
20-minute lane timeout. If the cache misses (key drift), the lane pays a
cold compile like the server shard does today; a miss is slow, never
wrong.
- Double-run/coverage-loss combinations are enumerated in the wrapper
header and pinned by tests: each caller runs the suite exactly once.

## Model Used

Claude (Bender agent, Paperclip) — Fable 5.

---
Expected savings once merged: the 339s `server (1/12)` check drops to
~265s-equivalent shard levels (~211s of tests), the slowest `ci /` check
becomes the ~326s e2e shard (~13–33s off PR wall time), and every PR run
stops paying ~4.5 min of billed cold Rust compile.

For the merger (squash): please keep the trailer below in the squash
body to preserve authorship.

`Co-Authored-By: Bender (Fable)
<Paperclip-Paperclip@users.noreply.github.com>`

## Related PRs

Searched the GitHub PR list for prior work on this surface — related
groundwork, none duplicate this change:
- #13457 — restored master's Rust dependency cache on the PR runner lane
(the read-only cache this PR relies on)
- #13500 — made that cache key image-toolchain-independent so
GitHub-hosted PR runners actually hit it
- #13521 — rebalanced PR shards and split the Verify Paperclip Runner
lanes this PR extends
- #13557 — previous health-check iteration (split the runnerd transport
suite); this PR targets the next slowest check

## Checklist

- [x] I have searched GitHub for duplicate or related PRs and linked
them above

Co-authored-by: Bender (Fable) <Paperclip-Paperclip@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
canary/v2026.918.0-canary.3
2026-09-18 07:22:12 -07:00
Devin FoleyandBender de9414dd8b fix: return empty read instead of past-EOF range when log reader is caught up (#13592)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The server streams agent run logs to the UI. It reads the log in
byte ranges from a local file, or from an S3 mirror after a pod restart.
> - The range math in `readS3Range()` clamps the range end up to the
range start. A fully caught-up reader then asks S3 for the range
`bytes=total-total`.
> - S3 rejects a range that starts at the end of the object. It returns
a 416 `InvalidRange` error. The API turns this into a 500 error, and the
log poller repeats it.
> - This pull request removes the clamp. A caught-up or past-EOF reader
now gets an empty read, and the server does not send an invalid range to
S3.
> - The benefit is that log polling after a pod restart does not cause
repeated 500 errors.

## Linked Issues or Issue Description

No public GitHub issue exists for this bug. The description follows the
bug report template.

**What happened?**

The run-log API returns a 500 error when a client polls a run log that
lives on the S3 mirror and the client is fully caught up (`offset ===
total`). The cause is in `server/src/services/run-log-store.ts`. The
function `readS3Range()` computes `end = Math.max(start, Math.min(start
+ limitBytes - 1, total - 1))`. When `offset === total`, the
`Math.max(start, …)` clamp forces `end` up to `start`. The `start > end`
empty-read guard does not operate, and the code sends `Range:
bytes=total-total` to S3. S3 rejects a range that starts at or past the
end of the object with a 416 `InvalidRange` error. The error monitor
records this error many times, only in the staging environment, because
only the S3 fallback path is sensitive to it. The local-file path has
the same math, but Node file streams accept past-EOF reads. The function
`readFileRange()` in
`server/src/services/workspace-operation-log-store.ts` has the same
latent math.

**Expected behavior**

A caught-up reader gets an empty read: `{ content: "", nextOffset:
undefined }`. The server does not send an invalid range request to S3.
The poller sees no contract change.

**Steps to reproduce**

1. Start a run and let it write a run log.
2. Let the log upload to the S3 mirror, and remove the local file (this
occurs when the pod restarts).
3. Poll the run-log read endpoint until the client offset is equal to
the log size.
4. Poll one more time. The server sends `bytes=total-total` to S3, S3
returns 416 `InvalidRange`, and the API returns a 500 error.

**Relevant logs or output**

```
InvalidRange: Invalid range
    at readS3Range (server/src/services/run-log-store.ts)
```

## What Changed

- `server/src/services/run-log-store.ts` — remove the up-clamp in
`readLocalRange()` and `readS3Range()`. A caught-up or past-EOF reader
gets an empty read.
- `server/src/services/workspace-operation-log-store.ts` — apply the
same fix to the shared math in `readFileRange()`.
- `server/src/services/run-log-store.test.ts` — the in-memory S3 mock
now rejects past-EOF ranges with `InvalidRange`, the same as real S3.
Add two regression tests for caught-up readers on the S3 path and on the
local path.

## Verification

- Run `pnpm vitest run server/src/services/run-log-store.test.ts`. All
17 tests pass.
- Revert only the source fix, and the new regression test fails with the
exact caught-up scenario. This shows the test covers the bug.
- Run the suites that use the workspace operation log store
(`workspace-runtime-control-recovery`,
`workspace-operations-reconciliation`). All 15 tests pass.
- Run `tsc --noEmit` on the server package. It reports no errors.

## Risks

- Low risk. The change only affects the empty and caught-up boundary of
range reads. Normal in-range reads give byte-identical results.
- Behavior change: a read with `limitBytes <= 0` now returns an empty
chunk instead of one byte. No caller passes a non-positive limit (the
default is 256000).
- Caught-up local reads keep the `nextOffset: undefined` semantics, so
pollers see no contract change.

## Model Used

- Claude Fable 5 (Anthropic, model ID `claude-fable-5`), with extended
thinking and tool use, run through the Claude Agent SDK.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Bender (Fable) <noreply@paperclip.ing>
canary/v2026.918.0-canary.2
2026-09-18 07:21:11 -07:00
Devin FoleyandPaperclip 45586170e1 fix(ui): reveal task history while retries wait to start (#13597)
## Thinking Path

> - Paperclip lets people manage AI agents and inspect their work.
> - Task conversations combine comments with run transcripts.
> - The first reveal waits for the relevant history to load.
> - Scheduled retries have not started, so the log reader does not
hydrate them.
> - Waiting for those retries keeps the whole conversation hidden after
its header loads.
> - This change reveals available history while a retry waits to start.

## Linked Issues or Issue Description

**What happened?**

A task header loads, but the conversation stays behind the loading
overlay while a linked run has `scheduled_retry` status. The readiness
check expects that run in `hydratedRunIds`, although the log reader does
not read its log. A retry delay can therefore become a task-loading
delay.

**Expected behavior**

Show existing comments and transcripts once their data is ready. A retry
that has not started must not block the conversation.

**Steps to reproduce**

1. Open a task with existing comments and a linked run waiting in
`scheduled_retry`.
2. Let the comments and other run transcripts finish loading.
3. Observe that the header loads but the conversation stays hidden until
the retry changes status.

**Paperclip version or commit**

Reproduced on master at `3f1d897a7` with regression tests.

**Deployment mode**

Web UI, built from source. This is a client readiness bug.

Searched public issues and PRs for scheduled retries, conversation
loading, and history loading. No duplicate found. Related: #10255
changes the server response for active runs with no log; it does not
address this scheduled-retry readiness gate. This focused bug fix does
not add a roadmap feature.

## What Changed

- Exclude `scheduled_retry` runs from the initial task-history readiness
gate.
- Extend transcript mocks to expose per-run hydration state.
- Add six regression cases across legacy and native runs. Existing
history loads during retry waits, while unhydrated running and completed
runs still block the first reveal.
- Document the exception in a code comment. No user command or
configuration changes need documentation.

## Verification

- All six new cases fail before the production fix.
- `pnpm --filter @paperclipai/ui exec vitest run
src/components/TaskChatThread.test.tsx
src/components/transcript/useLiveRunTranscripts.test.tsx
src/components/transcript/useNativeRunTranscripts.test.tsx`: 164 passed.
- `pnpm --filter @paperclipai/ui typecheck`: passed.
- `pnpm --filter @paperclipai/ui build`: passed.
- `pnpm check:token-gates` and `git diff --check`: passed.
- `pnpm -r typecheck` and `pnpm build`: attempted; both stop in the
native runner package because the local machine has no Rust `cargo`
executable. UI checks pass separately.
- `pnpm test:run`: started locally, then stopped before completion after
full CI passed on the same commit. The local full-suite result is
incomplete.
- CI on `46daa8b06`: all 53 checks passed; two optional Storybook jobs
were skipped. This includes the full test matrix, build, typecheck,
native runner checks, browser tests, and canary dry run.
- Greptile: 5/5 on `46daa8b06`, with no review threads or actionable
findings. The branch has no merge conflicts with master.

## Risks

Low risk. The exception applies only to runs waiting in
`scheduled_retry`. Running and completed runs retain their existing
readiness checks. Retry scheduling, transcript fetching, and server
behavior stay the same. No migration is needed.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository inspection, code
execution, and test tools. The exact backend revision and context-window
size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.918.0-canary.1 nightly/v2026.918.0-nightly.0
2026-09-17 22:07:27 -07:00
scotttong 3f1d897a7c copy(connectors): say "organization" in the connector setup flow (#13589) 2026-09-17 19:37:14 -07:00
Devin FoleyandPaperclip 5442f2d869 fix: repair managed Git launchers in sandbox projects (#13588)
## Thinking Path

> - Paperclip runs agents in local and remote execution environments.
> - Managed GitHub launchers select credentials for each Git operation.
> - Remote launchers are written inside the project checkout as
extensionless CommonJS scripts.
> - An ES module project makes Node interpret those launchers as ESM, so
they crash before credential resolution.
> - When the launcher can start, empty identity variables also override
valid repository and command-line Git configuration.
> - This change gives the launchers their own CommonJS scope and clears
empty identity overrides while preserving managed credential isolation.

## Linked Issues or Issue Description

**What happened?**

In a repository with `"type": "module"`, the managed `git` and `gh`
launchers fail immediately with `ReferenceError: require is not defined
in ES module scope`. The launchers use CommonJS but inherited the
enclosing project's module type.

Sandbox agents also report empty `GIT_AUTHOR_NAME` and
`GIT_COMMITTER_NAME` variables and try to unset them for each command.
With no managed identity available, the Git launcher recreated those
empty values. `git commit` failed with `fatal: empty ident name`, even
with explicit `user.name` and `user.email` configuration.

**Expected behavior**

Managed `git` and `gh` start in both ES module and CommonJS projects.
Local commits with an explicitly configured identity work without manual
environment cleanup. Managed credentials and captured identity continue
to take precedence. Missing identity does not silently select the host
user's details.

**Steps to reproduce**

1. Create a sandbox project whose `package.json` contains `"type":
"module"`.
2. Stage the managed GitHub launchers and run `git --version` or `gh
--version`. Before this fix, the launcher fails at its first
`require()`.
3. In a CommonJS project with no available managed identity, configure
repository `user.name` and `user.email`, or supply them with `git -c`.
4. Run `git commit --allow-empty -m test`. Before this fix, both
identity configuration forms fail with empty identity.

**Paperclip version or commit**

Reproduced from master commit `165b10bd9`.

**Deployment mode**

Sandbox execution. The shared launcher is also used for managed local
and SSH execution.

Related work: #13094 introduced the local-operation fallback; #13053
changes launcher discovery on Windows. Neither fixes empty identity
overrides. Related identity work in #8945 and #8946 configures worktree
authorship and does not remove these environment overrides.

## What Changed

- Stage `package.json` with `"type": "commonjs"` in the launcher
directory before the Node scripts. Keep the project's package
configuration unchanged.
- Leave inherited author and committer variables unset in the real Git
process. When credentials are absent, require explicit Git identity
configuration with `user.useConfigOnly`.
- Clear empty identity merge overrides in staged shell profiles after
environment merging. Preserve nonempty captured identity values.
- Exercise real Git commits with repository and command-line identity,
broker failures, and managed-user switching. Verify startup in ES module
and CommonJS projects, shell cleanup, and captured identity
preservation.
- Document launcher module scope and local identity behavior in the
execution GitHub identity contract.

## Verification

- Confirmed both new local-commit regression cases fail before the fix
with `fatal: empty ident name`.
- Confirmed the new ES module project regression fails before the fix
with `require is not defined in ES module scope`.
- Focused launcher and shell tests: 28 passed.
- `pnpm exec vitest run --project @paperclipai/adapter-utils --exclude
'**/dist/**'`: 1,216 passed, 11 skipped across 58 files.
- `pnpm --filter @paperclipai/adapter-utils typecheck` and `pnpm
--filter @paperclipai/adapter-utils build`: passed.
- `pnpm -r typecheck` and `pnpm build`: attempted; both stop in the
unchanged native runner because Cargo is not installed on this machine.
- Full `pnpm test:run`: started locally; stopped the duplicate run after
the complete CI suite passed. No local full-suite success is claimed.
- CI on `99ea8050e`: all 53 checks passed (2 skipped), including full
tests, typecheck, build, native runner checks, and browser checks.
- Greptile reviewed `99ea8050e`: 5/5 with no findings or unresolved
comments. GitHub reports no merge conflicts with master.
- No live sandbox or GitHub push probe performed.

## Risks

- The new package scope is confined to the run-specific launcher
directory. It does not change the project's module type, launcher names,
or credential selection.
- Without a managed identity, an explicitly configured repository author
can now create local commits. GitHub access remains subject to the
existing credential broker. Global/system Git configuration, ambient
credentials, and SSH identity remain isolated.
- Managed identity still wins over repository settings. Missing local
identity still fails instead of guessing host details.
- New or resumed executions must stage the updated launcher and shell
profiles. Existing processes retain their prior files and environment
until refreshed. No database migration or sandbox image rebuild is
required.
- Revert this change to restore the prior behavior.

## Model Used

- OpenAI GPT-6 via Codex, with code inspection, implementation, and
local test execution. The hosted model variant and context window were
not exposed.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub references)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (the affected adapter-utils
package)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.918.0-canary.0
2026-09-17 18:40:22 -07:00
DottaandPaperclip 84fe89906d fix: complete native agent review handoffs (#13581)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Native execution uses durable runs, issue locks, wake requests, and
typed tool authority
> - A child can finish with a native agent review request while its
original assignee stays responsible for the work
> - The reviewer then needs a bounded execution path that can inspect
the child, record one decision, and finish safely
> - Before this change, assignee-only gates rejected the reviewer or
left the parent waiting after the child review ended
> - This pull request adds typed reviewer admission, scoped reviewer
tools, durable wake and recovery handling, and parent continuation
evidence
> - The benefit is that native review handoffs complete without changing
child ownership or granting broad mutation access

## Linked Issues or Issue Description

Refs: #13314
Refs: #13574

**What happened?**

A native child run could report `needs_review` for an agent reviewer.
The reviewer wake then failed assignee and execution-lock checks. The
child remained in review and the parent remained waiting.

**Expected behavior**

The named reviewer should receive one durable wake. The reviewer should
inspect the child and resolve the exact review card. The child assignee
should stay unchanged. The parent should receive the recorded review
outcome after the child reaches its terminal state.

**Steps to reproduce**

1. Run a native task with a different named agent reviewer.
2. Keep the child assigned to its original worker.
3. Let the worker finish with a native completion review request.
4. Start the durable reviewer wake.
5. Resolve the review and finish the reviewer run.
6. Observe the child and parent state.

**Paperclip version or commit**

Base: `e926b1301`. PR head: `b31ad9ab8`. Live reviewer verification
source: `eea171aae`.

**Deployment mode**

Built from source.

**Installation method**

Built from source (pnpm build).

**Agent adapter(s) involved**

Not adapter-specific (core bug).

**Access context**

Both.

**Database mode**

Embedded PostgreSQL in the isolated live test fixtures.

## What Changed

- Add server-validated native review assignment facts.
- Admit only the exact company, issue, source run, decision, revision,
addressee, and resolver policy.
- Give reviewer runs a narrow set of Paperclip read and resolve tools.
File and shell access follow the configured agent and environment
policy, so reviewers can run tests.
- Separate server-owned reviewer instructions from untrusted persisted
review data. Escape the data boundary; retain server-enforced
authorization.
- Keep the child assignee unchanged. Atomically claim the reviewer run,
wake request, and issue execution lock. A competing lock prevents
provider startup.
- Require the exact running reviewer session and current issue lock to
resolve its assigned card. Reject missing, unrelated, or terminal
reviewer runs.
- Add durable reviewer wake, lock, stale-card, and abandoned-run
recovery handling.
- Prevent duplicate native wake dispatches during deferred admission and
recovery.
- Carry accepted or rejected child review outcomes into parent task
context and continuation evidence.
- Add focused server, runner, and native protocol coverage.
- Preserve upstream continuation rules. Add child review decisions as
separate evidence, while keeping real human answers in their own field.
- Return actionable completion validation feedback to both providers.
Permit a corrected completion after rejection. Keep strict terminal
acknowledgment validation.
- Apply exclusive shared-workspace locks to sandbox environments. Local
and SSH folders can run concurrently, including when old settings
request serialization.
- Repair test timing, native event parsing, and the review artifact
assertion. Allow a valid reject, correct, and accept review sequence.
Check the accepted card against its reviewer run and decision. Keep
polling within the existing deadline when review acceptance precedes the
parent wake projection; report a specific missing-continuation error at
timeout.
- Apply the ACPX pending-call limit to reserved finish/block calls, with
capacity-release and cancellation tests.

## Verification

- `pnpm build`: passed on `eea171aae`.
- `pnpm -r typecheck`: passed on `eea171aae`.
- `pnpm test:e2e:runner:unit`: 359 tests passed in 30 files on
`b31ad9ab8`; runner E2E typecheck also passed.
- `pnpm check:token-gates`: passed.
- Focused DB review, reviewer authority, and prompt-boundary checks: 31
tests passed. They cover invalid reviewer runs, competing locks, atomic
admission, duplicate claims, and valid resolution.
- Heartbeat, workspace, and recovery checks: 30 tests passed.
- ACPX sidecar suite: 27 tests passed. Moving the capacity guard back
below reserved handling makes both new regression cases fail.
- Four focused live continuation checks passed on their first attempt at
`f15f55e0a`: answer updates scope (6/6 each on Codex and Claude) and
question tool guidance (12/12 each). These cases do not use the reviewer
prompt path changed afterward.
- Fresh Codex and Claude review-handoff checks passed all 29 native
checks each on their first attempt at `eea171aae`. Both runs received
the expected fixed prompt and completed cleanup. Only the six selected
live flows were tested; no full paid provider catalog run.
- The final commit only extracts the existing test-harness timeout
diagnostic into a shared helper and adds positive and negative coverage.
Removing the accepted-review guard makes two regression assertions fail;
restoring it passes all six timeout tests. Production runtime code,
prompts, deadlines, and grading criteria are unchanged by this final
commit.
- Deadline regressions: a valid continuation delayed 20 seconds succeeds
within its 30-second unit-test deadline; an absent wake returns a
specific candidate-failure diagnostic at that same deadline. Both
assertions failed before the fix. Production E2E deadlines remain
unchanged.
- Historical native failures remain recorded: Docker availability
failures; a valid reject/correct/accept sequence that the first-card
grader misread; and a test that rejected the gap between accepted child
review and parent wake projection. No failed result was regraded. The
latest tests use a protected reference to the pinned Docker image and
the unchanged artifact oracle and time limits.
- Full repository verification runs in GitHub CI. Local verification
uses the focused suites above, full build, and full typecheck. An
unchanged Codex shutdown timing test failed once in CI, passed in
isolation, and its full shard passed on the final commit without changes
to that test or its causal code path. The original failure is retained
in the verification record. Greptile reviewed `b31ad9ab8` at 5/5 with no
outstanding actionable findings. All review threads are resolved. All
current-head CI gates passed, including the isolated native runner
Docker build (55 successful checks; two skipped by the workflow).

## Risks

- Reviewer admission depends on exact persisted decision and interaction
bindings. A stale or changed card is rejected.
- Paperclip control-plane tools are limited to inspection and review
resolution. This is not a filesystem permission boundary; provider file
and shell access retain the configured policy.
- Deferred wake recovery changes dispatch receipt coalescing. A
scheduler regression could delay a continuation if the receipt state is
wrong.
- Parent review outcomes are evidence for the model. They do not grant
tool authority or change issue ownership.
- This change does not address legacy lease-hold handoff behavior.

> Roadmap review: native execution, review gates, and durable recovery
are existing roadmap capabilities. This PR completes a narrow
reliability path for those capabilities.

## Model Used

OpenAI `gpt-6-astra` with reasoning, tool use, and code execution.
OpenAI `gpt-5.6-luna` assisted with bounded implementation, review, and
journal work. Context window size is not exposed by this session. Live
test subjects use `gpt-5.6-sol` and `claude-sonnet-5`; they are not the
PR authors.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.917.0-canary.6
2026-09-17 15:52:19 -05:00
Nicky LeachandPaperclip e926b13017 fix: keep sandbox termination progressing after bridge loss (#13287)
## Thinking Path

> - Paperclip must stop remote execution after losing its controller.
> - A bridge can remain blocked while the sandbox still incurs costs or
performs actions.
> - Waiting forever for that bridge prevents provider termination.
> - A temporary provider outage can also exhaust cleanup attempts
permanently.
> - This pull request bounds bridge drain and persists cleanup retries
with backoff.
> - Cleanup ends only after provider confirmation, without a user
accepting uncertain side effects.

## Linked Issues or Issue Description

Builds on merged #13285. Related #13254 added exact provider termination
receipts; merged #13272 adds explicit user retry. Merged #13352 stops
active sandbox startup before waiting for setup. This PR preserves that
immediate cancellation path and extends bounded teardown to ordinary
release and destroy. Cleanup continues automatically after repeated
provider failures. Refs #12953 for provider failures blocking execution.

**What happened?**
Daytona release waits for in-flight bridge activity before stop/delete.
A dead bridge can prevent that wait from finishing. The host also stops
cleanup after five failed attempts.

**Expected behavior**
Provider termination proceeds after a bounded bridge drain. Cleanup
retries survive service restarts and provider outages.

**Steps to reproduce**
Start a sandbox command whose bridge promise never resolves, then
release its lease. Separately, persist a pending-cleanup lease with five
failed attempts and recover the provider.

**Deployment mode**
Hosted Paperclip with a Daytona provider; rebased onto master at
`728f7185f` on September 14.

## What Changed

- Bound bridge drain and provider lifecycle calls. Prefer stop for
reusable sandboxes, with delete fallback.
- Persist cleanup attempt identity, renewable in-flight deadline, and
cooldown. Fence completion writes against superseded attempts.
- Preserve scoped explicit Retry and its activity log. Explicit Retry
can skip cooldown, but cannot take over a live cleanup attempt.
- Continue cleanup after five failures with slower retries and an
operator warning.
- Exclude leases in cooldown before paging so they do not starve due
work.
- Add hung-bridge, restart, provider-recovery, and concurrent-cleanup
regressions.

## Verification

- Rebased onto master at `728f7185f`. The outstanding diff contains only
cleanup changes; the merged controller-ownership prerequisite is
excluded.
- Daytona plugin suite: 160 passed, including immediate startup
cancellation, graceful release, hung activity, and teardown regressions.
- `pnpm exec vitest run
server/src/__tests__/heartbeat-pending-cleanup-sweep.test.ts`: 31
passed. Two added integration cases verify explicit Retry during
cooldown and while another cleanup owns the lease. They also verify run
scoping and the activity log.
- Targeted cleanup and cancellation cases in
`environment-runtime.test.ts`: 20 passed.
- Earlier live disposable Daytona test: provider stop ended background
work, resume preserved files without restarting the old process, a new
command succeeded, and the sandbox was deleted. This verifies provider
behavior; it was not repeated for this rebase.
- Latest-head CI and automated review are pending. Broad local tests,
typecheck, and build were not rerun for this focused rebase; CI supplies
those checks.

## Risks

- Timing out bridge drain permits provider termination; it never
supplies a stop receipt.
- A crashed cleanup attempt remains protected for 15 minutes, then
becomes eligible again. Repeated failures retry every 30 minutes after
escalation.
- The existing counter saturates at the escalation threshold; the new
attempt identity and deadline prevent overlapping claims.
- No schema, UI, telemetry, lockfile, or workflow change.

## Model Used

OpenAI GPT-6 through Codex, using reasoning, repository inspection, code
execution, and test tools. The precise backend revision and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.917.0-canary.5
2026-09-17 10:01:57 -07:00
mouse-value-add fcdb3f2499 feat: add optional you.com search integration (#13555)
<!-- Simplified Technical English (ASD-STE100). -->

## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Agents that do research work need live information from the web
> - Paperclip reaches external systems through governed, catalog-based
MCP connections
> - The Apps catalog is data-driven: a researched provider with a hosted
remote MCP server becomes a connectable app with no runtime code change
> - You.com operates a hosted remote MCP server for web search, content
extraction, and research tools
> - The server supports OAuth 2.1 with dynamic client registration, an
API key in a bearer header, and a keyless free profile at a separate
endpoint
> - This pull request adds You.com to the self-serve MCP research ledger
and generates its catalog entry with three connection methods: browser
sign-in, API key, and the keyless free profile
> - The benefit is that an operator can give agents live web search
through the normal connection governance, and the free profile needs no
account at all

## Linked Issues or Issue Description

No public issue exists for this provider. The problem description
follows the new-adapter issue template.

**Agent or provider**

You.com — web search and research tools over a hosted remote MCP server.

**Why this adapter is useful**

Agents that do research, monitoring, or fact-finding tasks need current
web results. You.com exposes web search (`you-search`), live page
extraction (`you-contents`), citation-backed research (`you-research`),
and finance research (`you-finance`) as MCP tools. Any Paperclip company
can connect it in a few clicks. The free profile offers `you-search`
without an account, so a new company can try agent web search at zero
cost and zero setup.

**How the agent is invoked**

Hosted remote MCP server (Streamable HTTP) at `https://api.you.com/mcp`.
Three supported access paths, verified against the live server on
2026-09-16:

- OAuth 2.1 browser sign-in. The server returns a `WWW-Authenticate`
challenge with RFC 9728 protected-resource metadata and advertises a
dynamic client registration endpoint, so Paperclip's automatic DCR path
applies.
- API key. Sent as an `Authorization: Bearer` header per the provider's
official server manifest and docs. Keys come from you.com/platform and
unlock higher rate limits plus the full tool set.
- Keyless free profile at `https://api.you.com/mcp?profile=free`.
Provides a reduced, read-only tool set.

Official docs: https://you.com/docs/build-with-agents/mcp-server

**Are you willing to implement it?**

Yes. Implemented in this pull request.

**Additional context**

Research evidence collected 2026-09-16, from live protocol probes and
official provider sources only:

- Unauthenticated `POST https://api.you.com/mcp` returns HTTP 401 with
`WWW-Authenticate: Bearer
resource_metadata="https://api.you.com/mcp/.well-known/oauth-protected-resource"
scope="Tools offline_access"`.
- RFC 9728 metadata lists one authorization server with scopes `Tools`
and `offline_access`.
- The authorization-server metadata (RFC 8414) publishes authorization,
token, and revocation endpoints, and advertises a
`registration_endpoint`, so DCR is available. No registration was
performed during research, per the runbook's non-registering preflight
rule.
- The keyless free profile answers `initialize` (server `You.com`,
version `4.0.1`), lists the tools `you-search` and `you-discover`, and
executed both tools successfully during the probe.
- The API-key placement matches the provider's official `server.json` in
the youdotcom-oss/mcp repository: header `Authorization`, value `Bearer
<key>`.

## What Changed

- Added You.com (slug `youcom`, wave 4, risk tier S2) to the self-serve
MCP research ledger in
`packages/shared/src/self-serve-mcp-research.json`, and refreshed the
ledger verification date.
- Added the You.com category (`ai`) and API-key header spec to
`scripts/ingest-app-definitions.mjs`.
- Added a You.com case to `specialMethodsFor` that emits three methods:
browser sign-in (`mcp-oauth`, DCR), API key (`mcp-api-key`, bearer
header), and keyless free profile (`mcp-free`, no auth).
- Regenerated `packages/shared/src/app-definitions/youcom.json` and the
generated registry via the ingestion script (`--definitions-only` mode;
no unrelated provider churn).
- Added the official You.com wordmark artwork (light and dark theme
variants, taken from the provider's docs site) under
`ui/public/brands/apps/`, with a manifest entry.
- Updated `packages/shared/src/app-definitions.test.ts`: ledger counts
(47 providers, 44 candidates), store count (48), verification date, and
assertions for the three You.com methods and their endpoints.

## Verification

- `node scripts/ingest-app-definitions.mjs --definitions-only` — passed.
Generated the new definition and registry import only; no other provider
JSON changed.
- `node scripts/check-app-brand-assets.mjs` — passed (71 identities).
- `node --test scripts/app-brand-validation.test.mjs` — passed.
- `pnpm exec vitest run packages/shared/src/app-definitions.test.ts
ui/src/lib/app-brand-assets.test.ts
ui/src/pages/apps/AppLogo.brand-assets.test.tsx` — passed (39 tests).
- `pnpm exec vitest run packages/shared/src/app-definitions.test.ts
server/src/__tests__/tool-access-service.test.ts
server/src/__tests__/generic-mcp-connection.test.ts
server/src/__tests__/tool-connection-removal.test.ts
ui/src/pages/apps/AppsConnect.test.tsx
ui/src/pages/apps/Browse.test.tsx` — passed (181 tests). Two server
suites that require embedded Postgres skipped on this machine by their
own environment gate; the gate is unrelated to this change.
- `pnpm --filter @paperclipai/shared typecheck` — passed.
- `pnpm --filter @paperclipai/ui typecheck` — passed.
- `pnpm --filter @paperclipai/plugin-sdk ensure-build-deps` — passed;
builds `@paperclipai/shared` with the new definition.
- `pnpm test:run` (full Vitest suite) — 7,930 passed, 18 failed, 4,592
skipped. Every failure is environmental on this container: the
embedded-Postgres suites refuse to start because the machine runs as
root, the native runtime suites need the Rust runner binary that this
container cannot build, and one media suite needs a native HEIC binary.
No failure touches the app-catalog, connection, branding, or
shared-package surface; those suites pass locally. CI is the
authoritative gate for the full suite.
- `pnpm --filter @paperclipai/server typecheck` — not completed: the
script's `prepare:runner-vendor` prelude builds the Rust runner, which
cannot build on this container. A direct `tsc --noEmit` reports only
pre-existing errors from the missing vendored runner types; no error
touches this change. No server code is changed.
- Live You.com proof on 2026-09-16 (keyless free profile, real network
calls): preflight 401 challenge with RFC 9728/8414 metadata and DCR
endpoint ✓, `initialize` ✓, `tools/list` ✓, `you-search` call returned
results ✓, `you-discover` call returned results ✓.
- Live proof NOT run: an authenticated OAuth connect and an API-key call
against the full server. This environment has no You.com account or API
key. Per the runbook, this proof stays outstanding and must not be
assumed from the keyless probe. Both paths match the reviewed
`mcp_remote` patterns (DCR and bearer header) used by existing
providers.
- Browser e2e suites not run: opt-in per `AGENTS.md`, and this change
adds catalog data only, with no UI code.

## Risks

- Low risk. The change is catalog data plus generated output. It adds no
runtime code and touches no existing provider.
- The free-profile method is a fixed keyless endpoint. If You.com
changes or removes `?profile=free`, that method breaks and the entry
needs a ledger update. The OAuth and API-key methods do not depend on
it.
- The authenticated tool catalog is discovered live at connect time, so
provider-side tool changes appear through the normal catalog refresh and
quarantine flow, not through this definition.
- Rollback is a single revert; no migration and no state are involved.

## Model Used

- Provider: Zhipu AI, via OpenRouter
- Model: GLM-5.3 (`z-ai/glm-5.3`)
- Context window: 200K tokens
- Capabilities used: tool use (shell, file edits, live HTTP probes),
long-context repository reading
- The change was produced with AI assistance and reviewed by a human
before submission.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` /
`Closes: #` / `Refs: #` OR (b) described the issue in-PR following the
relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
canary/v2026.917.0-canary.4
2026-09-17 10:01:49 -07:00
Devin FoleyandBender ec40bd8bf6 ci(runner): split retained-settlement suite out of runnerd-codex-transport.test.ts (#13557)
Moves the 19-case 'settles only retained control authority' family (~156s) plus its two private helpers from runnerd-codex-transport.test.ts (8,895 lines, 398s sequential) into a new runnerd-codex-transport-settlement.test.ts so vitest can schedule the two files onto separate workers. Pure code move, no test-logic changes; vitest collects the identical 182 test names.

Measured on the PR's own CI run: 'ci / Verify Paperclip Runner (vitest 1/2)' dropped from 531s to 340s and the end-to-end PR workflow from ~545s to ~415s.

Co-Authored-By: Bender (Fable) <Paperclip-Paperclip@users.noreply.github.com>
2026-09-17 09:09:57 -07:00
DottaandPaperclip e26d787928 Shorten continuation prompts and verify question tool guidance (#13574)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents must continue tasks using user answers without losing earlier
requirements or approval gates.
> - The wake prompt mixed human decisions with prior tool evidence and
repeated detailed question instructions.
> - Those instructions belong with the question tool, with a short
routing hint in the wake.
> - The Runner evals need to prove that answers, approvals, and
completed work survive later turns.
> - This PR shortens the prompts, separates authenticated answers, and
adds continuation tests with useful screenshots.

## Linked Issues or Issue Description

Refs #13517. This is a follow-up to the merged onboarding skill and
Runner E2E work. Related #13539 covers responses received while a run is
active; this PR preserves its cases and adds continuation coverage.
Existing continuation/recovery and question PRs were searched; none
covers this prompt/documentation and eval change.

**What existing behavior does this improve?**

The instructions sent when an agent continues a task, the native
human-input tool documentation, and the evidence captured by Runner
full-stack E2E.

**Current behavior**

The wake repeats a long question-tool guide. Human answers appear
alongside untrusted prior results. Screenshot capture can finish at DOM
load while the task still shows a spinner, even when backend behavior
checks pass.

**Proposed behavior**

Keep earlier requirements unless the user changes them. Treat
clarification as distinct from approval. Give authenticated human
responses a scoped field. Keep tool and agent results as evidence. Put
detailed question behavior in the tool descriptor and retain one routing
sentence in the native wake. Wait for the correct task and loaded
conversation before taking screenshots.

**Reason and benefit**

Reduce repeated prompt text and make authority boundaries clear. Test
that real question cards, later answers, approval gates, and completed
child tasks still work. Make screenshots useful for human review.

## What Changed

- Shorten shared continuation instructions for legacy and native
runners. Separate authenticated user responses from tool results and
agent summaries.
- Remove the detailed question guide from native wake prompts. Keep its
behavior in the canonical `request_human_input` descriptor and existing
payload schema. Regenerate semantic contracts and fixture hashes.
- Add five continuation cases across four local profiles. Add a
dedicated choice-then-text case for native Codex and native Claude. All
22 cells join the shared full E2E campaign.
- Cover revised scope, clarification without approval, hostile
instructions in a handoff file, and reuse of a completed child after
restart. Keep production instructions and fixed user facts.
- Capture continuation screenshots only when the intended task and
conversation have rendered. Add provider-free browser regressions for
loaders and wrong-task capture.
- Preserve current master’s extra tool and onboarding cases. The default
campaign now contains 166 cells; 35 manual everyday cells remain
separate.

## Verification

- `pnpm -r typecheck`: passed after replay on current master.
- `pnpm test:e2e:runner:unit`: 340 passed. Harness typecheck passed.
- `PAPERCLIP_PLAYWRIGHT_CHANNEL=chrome pnpm
test:e2e:runner:browser-support`: 4 passed. These tests failed against
immediate screenshot capture and passed after the fix.
- Focused continuation and native-input tests: 36 passed locally. The
tool-authority suite could not initialize embedded PostgreSQL locally,
including one isolated retry; its 17 assertions did not run locally. The
full remote server shards passed on this PR commit.
- `pnpm build`: passed after replay on current master. `pnpm test:run`
was attempted locally but hit the same embedded PostgreSQL
initialization failure; the remaining local run was stopped after
complete remote CI passed. This is not claimed as a full local test
pass.
- [Full PR
CI](https://github.com/paperclipai/paperclip/actions/runs/35232755685):
passed on `6a22128c14f4552d0613a6d9a25955db4a1ed02f`. All
server/chat/workspace/serialized shards, browser shards, Runner checks,
typecheck, build, canary and policy checks passed. The isolated native
Runner build and security checks also passed: 57 successful checks, with
two expected Storybook skips.
- Greptile reviewed the exact PR head at 5/5, with no findings or
unresolved review threads. The PR has no merge conflicts.
- [Live question-docs
report](https://pages.paperclip.ing/runner-e2e-question-docs-35227647794/):
3/3 passed at source `83dd132f2` before replay on master. Native Codex
and Claude each asked a choice, waited, asked a text question, and saved
both answers. Claude also passed a completed-child restart case. All
three native turns are checked for absence of the old question block.
- [Earlier continuation
report](https://pages.paperclip.ing/runner-e2e-continuation-35154943615/):
all five continuation cases passed on native Claude. The report retains
campaign and revision provenance and separately shows two unresolved
onboarding behavior failures.
- [Before/after prompt
report](https://pages.paperclip.ing/runner-prompt-comparison-20260917/):
full text, current recorded Claude inputs, and reproducible
reference-token counts. The controlled wake comparison removes 401
reference tokens; the net counted input reduction is 339 after charging
the larger tool description. These are text-size estimates, not measured
billing savings.

## Risks

- Prompt wording affects model behavior. Live results cover the stated
cases, not every provider or conversation. Legacy profiles are
registered but were not rerun for this change.
- The optional continuation field changes prompt data only; there is no
database migration or new production API.
- Authenticated answer projection excludes generated summaries and
agent-resolved interactions. It preserves the answer’s question or
approval scope.
- The screenshot guard can expose UI loading failures that earlier runs
hid. Backend grading alone no longer makes those captures valid.
- The two prior onboarding failures remain separate product issues: work
before acceptance and a missing saved plan. This PR does not claim the
entire onboarding suite passes.

## Model Used

OpenAI Codex, GPT-6, with reasoning, repository tools, code execution,
and browser verification. The exact deployed model identifier and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass — targeted tests above; the
full local database-startup limit is documented
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.917.0-canary.3
2026-09-17 09:32:54 -05:00
327ab2fe38 fix(grok-local): do not pin empty GROK_HOME over host login (#13570)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Local adapters such as `grok_local` invoke a host CLI (`grok`) for
each heartbeat
> - Grok authenticates from `GROK_HOME/auth.json` when that env var is
set, otherwise from `~/.grok`
> - Recent work (#12469, #12618) isolated subscription credentials into
a company-scoped Grok home filled only by sandbox device login
> - Local_trusted instances have only a Local environment, so that login
never runs, the company home stays empty, and execute still sets
`GROK_HOME` to it
> - This pull request stops pinning `GROK_HOME` on local subscription
runs unless the company home already has usable auth, a managed AI
connection supplied a home, or the run is remote/sandbox
> - The benefit is that `grok login` on the host works again for local
Grok agents, without leaking host credentials into sandboxes

## Linked Issues or Issue Description

Fixes: #13568

Related PRs (predecessors, not duplicates):

- Refs #12469
- Refs #12618
- Refs #12696

I searched GitHub for `GROK_HOME`, `grok login`, `not signed in`, and
`device-code`. No existing PR restores host-login fallback for local
`grok_local` runs.

## What Changed

- Local subscription execute no longer sets `GROK_HOME` when the company
Grok home has no usable `auth.json`
- Remote/sandbox runs and managed AI connections still pin `GROK_HOME`
so they cannot fall through to the host login
- Local runs still pin `GROK_HOME` once a company home has a usable
credential (completed device login)
- Adapter configuration notes document the host-login vs company-home
split
- Subscription detection respects an explicit empty `XAI_API_KEY` that
clears an inherited host key. This keeps a valid company login selected.
- Tests cover a real child process reading fixture host credentials,
custom host homes, malformed company credentials, API-key overrides,
managed connections, and empty remote homes.
- Original fix by @hawikk. The follow-up preserves the contributor
commit and adds independent regression coverage.

## Verification

- `pnpm exec vitest run packages/adapters/grok-local`: 131 tests pass in
12 files.
- `pnpm --filter @paperclipai/adapter-grok-local typecheck`: passes.
- The new subprocess host-login regression fails against `master` and
passes with this fix. It uses disposable fixture credentials and makes
no provider request.
- The explicit-empty-key regression fails against the contributed commit
and passes with the follow-up.
- `pnpm -r typecheck` and `pnpm build`: pass locally.
- `pnpm test:run`: attempted locally, then stopped after embedded
PostgreSQL startup failures. A focused retry of
`ai-legacy-compatibility.test.ts` reproduced the same startup failure
after five attempts.
- All CI checks pass on `f4a380fec`: 54 successful checks and two
intentional Storybook skips. This includes all general tests, serialized
server suites, runner checks, browser shards, typecheck, build, and
canary packaging.
- Greptile reviewed `f4a380fec` at 5/5 with no actionable findings.
GitHub reports no merge conflicts.
- Grok CLI 1.0.13 is installed on the verification host, but it has no
signed-in account. A live authenticated inference run was not performed.

## Risks

- Low. Behavior changes only local subscription runs whose company Grok
home has no usable `auth.json`.
- Remote/sandbox isolation is unchanged: those runs still pin
`GROK_HOME` and never use host `~/.grok`.
- Managed AI connections still pin even with an empty home (fail closed
rather than using the host account).
- Operators who previously copied `auth.json` into the company home keep
the pinned-home path.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- Provider: xAI Grok
- Model: Grok 4.6 (`grok-4.6`)
- Tool use: yes (repository search, local tests, GitHub issue/PR)
- Human-authored: no — AI-assisted implementation
- Follow-up review, code, and tests: OpenAI GPT-6 in Codex, with
reasoning, repository tools, and code execution. Exact backend model ID
and context window are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Dotta <bippadotta@protonmail.com>
canary/v2026.917.0-canary.2
2026-09-17 09:12:40 -05:00
5d9b20ccf0 fix(ui): keep task composer available while pause state loads (#13562)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Task messages enter through the shared task composer.
> - The task page waits for a separate tree-control query to find active
pause holds.
> - The old loading guard disables the composer until that request
settles. A slow or stalled request blocks messages on both desktop and
mobile.
> - This PR allows sends while that query is pending. The server still
rejects paused board messages before saving a comment or waking an
agent.
> - Query failures and known pause holds still block the composer.
> - Regression tests cover pending sends, successful responses, query
failures, and late root or inherited pauses.

## Linked Issues or Issue Description

Fixes: #13561

Related: #13569 adds pending/error coverage for the same fix. This PR
now covers those cases and the late-pause transitions.

The report overstates two details: `isPending` clears after a successful
response, and the shared loading guard also affects mobile. The
reproduction holds the request pending. It does not prove why the
reporter's request remained unresolved.

## What Changed

- Preserve the original two-line fix that removes the pending-state
composer block.
- Add eight page regression cases using a real QueryClient and a
controlled API promise. Cover desktop and mobile, submission before the
query completes, successful resolution, rejected resolution, and late
root/inherited pause responses.
- Repair an existing server CI failure in a separate commit. Reuse the
shared unique-violation helper so Drizzle-wrapped duplicate inserts
become retryable document conflicts. Add a deterministic regression and
preserve unrelated database errors.
- Repair an existing Inbox test race in a separate commit. Wait for
workspace metadata, which resolves independently of the task list.

## Verification

- Red: restore the pre-fix `IssueDetail.tsx` and run the new `composer
tree control` cases. Both desktop and mobile pending-send cases fail
with `Checking task status…`. The other six cases pass.
- Green: restore the original PR fix. All 343 tests in IssueDetail,
TaskChatThread, and TaskChatComposer pass.
- The server comment/reopen and artifact-review suites pass (179 tests).
They include POST and PATCH pause checks that return 409 before any
comment, task mutation, or wakeup.
- The document error regression fails before the shared-helper fix and
passes after it. The document, artifact-review, and database-error
suites pass (31 tests). The handler matches only the issue-document key
constraint; revision and other constraint errors retain their original
identity.
- The Inbox suite passes (27 tests).
- `pnpm check:token-gates` passes.
- `pnpm -r typecheck` and `pnpm build` pass locally. The full CI matrix
passes on `5af1ed0b43269247aaba406cf4fd4d7fe1a22e75`: 54 successful
checks and two opt-in Storybook skips, including all general/serialized
tests, runner checks, release checks, and eight browser shards.
- One server shard initially hit an unrelated `EADDRINUSE` on test port
52000. Its single rerun passed without code changes.
- Greptile completed successfully on that exact commit with 5/5 and no
outstanding findings or review threads.
- The serial local `pnpm test:run` was started, then stopped after the
complete parallel CI matrix passed. It is not claimed as a completed
local full-suite run. The focused local suites above did finish
successfully.

For a manual reproduction, delay the task's `/tree-control-state`
response, open the task, and enter a message. Send should remain
available during the delay. Resolve the response with an active pause
hold and confirm that the pause takeover replaces the composer. Reject
the request and confirm that the error blocks sends.

## Risks

A user can attempt a send before the pause response arrives. The server
remains authoritative and returns 409 for a paused task before saving or
waking work. The known pause takeover and query-error block remain.
There are no schema, API, or styling changes.

The document change restores the existing conflict/retry behavior for
wrapped database errors. It does not retry unrelated database failures.
The Inbox change affects test synchronization only.

## Model Used

- Original fix: Anthropic Claude Sonnet 4.6 (`claude-sonnet-4-6`), 200k
context, tool use and code editing, as reported by the author.
- Review, regression tests, and CI repairs: OpenAI GPT-6 (`gpt-6-astra`)
through Codex, with reasoning, tool use, and code execution. The session
does not expose its context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: austinpilz <austinpilz@users.noreply.github.com>
Co-authored-by: Dotta <bippadotta@protonmail.com>
2026-09-17 08:40:40 -05:00
Devin FoleyandPaperclip 165b10bd98 fix: enable GitHub Actions MCP toolset (#13553)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The GitHub connector lets agents use repository tools through MCP.
> - GitHub excludes Actions from its default MCP toolsets.
> - Approval of Actions permissions therefore does not make workflow
tools appear in Paperclip.
> - This pull request adds Actions to the requested toolsets for
discovery and execution.
> - Users can refresh existing connections and use workflow tools under
the existing access rules.

## Linked Issues or Issue Description

**What happened?**

GitHub Actions tools remain absent after the GitHub App receives Actions
read/write access and the user refreshes actions in Paperclip. Paperclip
does not request the Actions MCP toolset.

**Expected behavior**

Authorized GitHub connections expose workflow tools, including
`actions_run_trigger` with `method: "run_workflow"`, so agents can
dispatch an existing release workflow.

**Steps to reproduce**

1. Connect GitHub to Paperclip with access to a repository that has a
dispatchable workflow.
2. Grant the GitHub App Actions read/write permission and approve the
installation update.
3. Refresh the connection's actions in Paperclip.
4. Observe that the workflow tools are absent.

**Paperclip version or commit**

Base commit: `fae698031`.

**Deployment mode**

Hosted instance with a managed GitHub connection. The same missing
header affects PAT connections.

No matching public issue or pull request was found in the duplicate
search.

## What Changed

- Send `X-MCP-Toolsets: default,actions` through the shared GitHub MCP
header helper. This covers discovery, refresh, and execution for
existing and new managed or PAT connections, including legacy rows
identified through `transportConfig`.
- Test catalog refresh, tool risk classification, and workflow dispatch
through a mock MCP server.
- Document tool names, workflow arguments, required GitHub permissions,
and the refresh step.

## Verification

- Passed both affected test suites: `pnpm exec vitest run
server/src/__tests__/tool-access-service.test.ts
server/src/__tests__/tool-gateway.test.ts` (391 tests).
- Passed `pnpm check:token-gates` and `git diff --check`.
- Live provider check: the default catalog returned 45 tools.
`default,actions` returned 49 tools, with no tools removed. The four
added tools were `actions_get`, `actions_list`, `actions_run_trigger`,
and `get_job_logs`.
- Live `actions_get` / `get_workflow` call succeeded. No workflow was
dispatched during live verification.
- Passed `pnpm -r typecheck` and `pnpm build` with the existing Rust
toolchain added to PATH.
- Rechecked server typecheck and build after the legacy-connection fix;
both passed.
- The full local test run has reported three skills-cache failures in
`company-skills-service.test.ts`. All three reproduce on the untouched
base commit (`fae698031`) on this macOS host: runtime-cache directory
renames fail with `EACCES`. The full run remains in progress.
- Greptile: 5/5 on `7d391e3c7`, with no unresolved review threads.
- After deployment, use **Refresh actions** on an existing GitHub
connection and verify the workflow tools appear.

## Risks

- Refreshed GitHub catalogs expose more tools. Existing access,
approval, and quarantine rules still apply. `actions_run_trigger` keeps
GitHub's destructive classification because it also supports
cancellation and log deletion.
- GitHub still enforces token and installation permissions. Dispatch
requires Actions write permission and a workflow with
`workflow_dispatch`.
- No database migration or saved connection edit is required.

## Model Used

OpenAI GPT-6 through Codex, with code execution and tool use. The exact
serving model ID and context window are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.917.0-canary.0 nightly/v2026.917.0-nightly.0
2026-09-16 17:10:34 -07:00
Devin FoleyandPaperclip fae6980310 revert(apps): restore Google connector visibility (#13552)
## Thinking Path

> - Paperclip helps people manage AI agents for work.
> - The Connectors catalog lists services that agents can use.
> - PR #13551 temporarily hid Google connectors.
> - We now want to restore their catalog visibility.
> - This PR reverts that change and restores the previous catalog
behavior.

## Linked Issues or Issue Description

Refs: #13551

Revert the temporary removal of Google connectors from the UI.

## What Changed

- Restore Gmail and eight Google Workspace entries to the catalog.
- Restore the matching branding flags and original catalog and service
tests.
- Remove the temporary-hiding documentation note.

This is an exact revert of commit
`cf1e873ab24277d55ffd3ab06074f77014dc4015`.

## Verification

- Passed: 507 catalog, UI, and connection service tests.
- Passed: `pnpm check:token-gates` and `node
scripts/check-app-brand-assets.mjs`.
- Passed: `pnpm --filter @paperclipai/ui... build` and `pnpm --filter
@paperclipai/ui... typecheck`.
- Full local build and typecheck stop at the Rust runner because `cargo`
is not installed.
- Full local Vitest was not repeated because the unchanged base has
confirmed macOS skill-cache permission failures. The full CI suites
passed.
- Passed: all GitHub CI gates; Greptile 5/5 on commit
`4e3dddef0ebfef1f99001e7735822ed4cba852ab`, with no review threads.
- Reviewer check: open Connectors and confirm that Gmail and Google
Workspace entries appear again.

## Risks

Low risk. This restores the previous catalog visibility and setup entry
points. Connector implementations and saved connection data are
retained.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, tool use, and code
execution. The exact deployment ID and context window size are not
exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.916.1-canary.6
2026-09-16 15:35:16 -07:00
Devin FoleyandPaperclip cf1e873ab2 fix(apps): temporarily hide Google connectors (#13551)
## Thinking Path

> - Paperclip helps people manage AI agents for work.
> - The Connectors catalog lists services that agents can use.
> - We need to temporarily remove Google connectors from the UI.
> - The catalog already separates visibility from retained definitions.
> - This PR uses that setting so Google can return with a small change.

## Linked Issues or Issue Description

**What existing behavior does this improve?**
The Connectors catalog and its setup entry points.

**Current behavior**
The catalog shows Gmail and eight Google Workspace connectors.

**Proposed behavior**
Temporarily hide those nine entries. Keep their definitions and existing
connections.

**Reason and benefit**
Make the temporary UI removal easy to reverse.

**Breaking changes**
Fresh catalog setup no longer offers Google. Saved connections keep the
existing management and reconnect paths.

## What Changed

- Add the nine Google connector slugs to the existing hidden list.
- Match the branding manifest visibility flags.
- Update existing catalog and service tests. Keep backend Google
connection coverage and document how to restore visibility.

## Verification

- Passed: 507 targeted tests covering catalog definitions, URL matching,
setup routing, connector UI, branding, and the connection service.
- Passed: `pnpm --filter @paperclipai/ui... build` and `pnpm --filter
@paperclipai/ui... typecheck`.
- Passed: `pnpm check:token-gates` and `node
scripts/check-app-brand-assets.mjs`.
- Full local build and typecheck stop at the Rust runner because `cargo`
is not installed.
- Stopped the full local Vitest run after skill-cache permission
failures. Three failures in `company-skills-service.test.ts` also
reproduce on the unchanged base branch. The final connector service
suite passes all 319 tests.
- Greptile: 5/5 on the current commit, with no open review threads. CI
is retrying one unrelated preview-server readiness timeout. That test
file passes all seven tests locally.
- Reviewer check: open Connectors in a company with no Google
connections. Gmail and Google Workspace entries should be absent.
Existing saved connections remain manageable.

## Risks

Low risk. This uses the existing catalog visibility mechanism. No
connector implementation, credential, or database schema is removed.
Restoring visibility requires updating both the hidden list and branding
manifest.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, tool use, and code
execution. The exact deployment ID and context window size are not
exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.916.1-canary.5
2026-09-16 15:04:46 -07:00
Devin FoleyandPaperclip 6fe8e30625 feat(apps): add Railway connection and governed deployment tools (#13415)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Apps gives agents governed access to external resources.
> - Operators need to inspect Railway services, read logs, deploy code,
and run container commands.
> - Railway offers hosted MCP with OAuth, but broad remote actions hide
their internal operations.
> - This PR adds a branded connection and fixed direct operations
through the existing gateway.
> - Separate SSH keys enable container commands under the same grants
and policies.
> - Operators can require approval for an action and inspect the
resulting audit record.

## Linked Issues or Issue Description

**Subsystem affected**

Apps catalog, connection setup, gateway execution, and connection
documentation.

**Problem or motivation**

Agents need Railway access through Paperclip. Operators need to grant
and revoke that access, inspect available actions, and govern deployment
and container operations without giving agents provider credentials.

**Proposed solution**

Reuse hosted MCP OAuth, vault storage, catalog discovery, grants, and
the gateway. Probe the actual credential before enabling fixed GraphQL
operations. Use a dedicated grant-owned SSH key for bounded container
commands.

**Alternatives considered**

A catalog entry alone cannot execute the missing operations. The hosted
general agent has opaque internal effects. An unrestricted CLI runtime
can bypass action policy and inherit ambient credentials.

**Roadmap alignment**

This extends the existing MCP Tool Gateway & Apps path and the Connected
Apps direction in ROADMAP.md. It does not add a plugin or parallel
connection service.

Related PRs #311, #939, and #7861 concern hosting Paperclip on Railway.
They do not add this outbound Apps connection. The separate shared
agent-picker fix is #13414 and is not included here.

## What Changed

- Add the generated Railway catalog entry, official marks, provenance,
and OAuth setup guidance.
- Add fixed service/deployment status, bounded logs, and
redeploy/restart/rollback tools. Block source deployment until the
provider can atomically bind the approved repository and commit.
- Verify API access with an explicit workspace before exposing direct
tools.
- Add grant-owned SSH key setup and a bounded runner with host
verification, target checks, isolated state, and cleanup.
- Block the opaque hosted railway-agent and accept-deploy actions.
Preserve normal Allowed defaults and Ask-first policies for other
actions.
- Quarantine new or changed Railway schemas after initial discovery,
including reconnect.
- Add provider, lifecycle, gateway, SSH, UI, and browser fixtures.
Document setup, limitations, and the release checklist.

## Verification

- Security follow-up: removed the unsafe source-deployment mutation.
Direct calls and old active catalog entries are denied before any
upstream request, including normalized aliases. Refresh marks retired
entries disabled. All 386 focused Railway, catalog and gateway tests
passed, and server TypeScript checking passed. Full [GitHub
CI](https://github.com/paperclipai/paperclip/actions/runs/35139421144)
passed on d86530ab9, including typecheck, build, all tests, runner
checks, and browser tests. Superagent passed and confirmed the P2 fix.
Greptile reviewed the same commit at 5/5 with no findings.

- CI follow-up: fixed the missing Railway SSH operation in the OpenAPI
document, including its request schema, operator-only authentication,
and error responses. The failure reproduced locally before the fix; all
403 selected API, Railway, catalog, and artwork tests passed after it.
Synced current master and resolved the catalog/artwork conflicts.

- After rebase: 440 focused provider, lifecycle, gateway, catalog, and
container-panel tests passed. AppDetail and AppsConnect passed another
196 tests.
- Full typecheck, build, token gates, and the gallery browser check
passed after rebase.
- During implementation, full build and the gallery browser check
passed. Shared generic-MCP fixtures covered OAuth callback/state/issuer
binding and failure paths.
- Local live consent and tools/list succeeded. There were 44 active
hosted actions and two blocked actions. A workspace-bound API probe and
direct project/service/environment reads succeeded. The inspected
project had no deployed services. No provider mutation ran.
- Full GitHub CI passed on commit 303340f19, including all
server/workspace test groups, typecheck, build, runtime verification,
release dry run, and browser tests. The original local full-run attempt
was incomplete; the complete automated suite is now verified in CI.

Manual review: connect Railway, review the actual actions, install for
an agent, and run a resource read through the gateway. Choose Ask first
before testing a deployment mutation. Configure a dedicated key only
when container access is needed.

**Release qualification is still open.** Live agent gateway reads/logs,
rejected and approved deployment calls, refresh/revoke, public HTTPS
consent, and SSH enrollment/commands/cleanup need an authorized
disposable service. The passing API diagnostic does not replace those
tests. See doc/connections/RAILWAY.md and RAILWAY-REVIEW.md.

## Risks

Overall risk is medium. New runtime behavior is gated to Railway
connections, but the PR changes shared catalog, credential lifecycle,
and gateway code. A regression in those paths can affect other Apps
connections. The highest-impact operations are Railway deployments and
container commands.

- Provider consent can authorize an entire workspace. Catalog labels are
not local resource allowlists. Direct tools check target membership, and
provider permissions still apply.
- Shell commands have broad internal authority. Action policy cannot
approve each internal shell step. Timeouts close the local connection
but cannot guarantee remote child-process termination.
- Log and command output may contain application secrets that pattern
redaction cannot recognize.
- Source deployment is unavailable until the provider supports atomic
repository/commit binding. Existing deployments can still be redeployed,
restarted or rolled back.
- No database migration is required. Rollback can remove promotion and
direct dispatch while preserving connection data and the generic MCP
path.
- Live Railway qualification must still pass before release acceptance.

## Model Used

OpenAI Codex, based on GPT-6, with code execution and browser testing.
An independent read-only security agent reviewed the local
implementation. The exact serving model ID and context window were not
exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.916.1-canary.4
2026-09-16 13:54:44 -07:00
DottaandPaperclip d0b67bfe71 feat: queue approvals and answers during active runs (#13539)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Users guide running agents through messages, questions, and approval
cards.
> - Messages already wait in a queue when an agent is running.
> - Card responses did not appear in that queue. Some question answers
also steered a later run without a user click.
> - A fast approval could invalidate the agent's review handoff and
cause it to stop its own run.
> - This pull request gives card responses the same queue controls and
preserves the exact response during delivery.
> - Users can wait for completion or explicitly send the response with
Interrupt or Steer.

## Linked Issues or Issue Description

Refs #13517, which is merged. This PR targets master and adds queued
interaction responses on top of the onboarding changes. Related
continuation work: #10519 and #12866.

**What happened?**

Accepting a proposal while its source run was active left a saved
response outside the message queue. The agent could then lose its review
path, reassign the task, and cancel itself. Answers to older questions
could also steer another active turn without a click.

**Expected behavior**

Save the response immediately. Queue its continuation behind the active
run. Deliver it after completion, or when the user explicitly chooses
Interrupt or Steer. Preserve approval revisions and answer choices.

**Steps to reproduce**

1. Let an agent publish a confirmation card while its run is still
active.
2. Accept the card before the agent finishes its review handoff.
3. Inspect the message queue and the task's next run.

**Paperclip version or commit**

Reproduced on da8a3876c with the onboarding changes from #13517.

**Deployment mode**

Local development from source. The fix covers legacy adapters and native
Runner turns.

## What Changed

- Project resolved cards into the existing queue as immutable responses.
Keep answers and exact approval revisions.
- Require an explicit click to steer a response into a compatible native
turn. Use Interrupt when a fresh session is required.
- Preserve typed response context through interruption, cleanup waits,
and normal queue promotion. Keep the direct answer channel for a
provider blocked on its original question request.
- Accept the source run's review handoff after its card resolves. Reject
stale agent reassignment that would orphan a queued response.
- Add deterministic regression tests and an `accept-while-running` case
to the first-task suite. Require recorded timestamp overlap before that
case can pass.
- Keep the first-task skill name out of user-facing messages.

## Verification

- Red-green: the original route failed the queue regression; the changed
route passes it.
- Focused server/UI tests: 139 passed, including 64 queue-route tests.
- Runner harness unit tests: 314 passed.
- Server, UI, and Runner E2E typechecks passed. UI token gates passed.
- Full repository typecheck and build passed. Server typecheck passed
again after review fixes.
- Review regressions: 165 queue/reopen route tests, 53 wake admission
tests, and 18 run identity tests passed. Approval acknowledgement
recovery and both message/approval arrival orders are covered.
- Full local test run: 12,401 passed; three new admission regressions
ran against a cached pre-fix module. A fresh run of that entire suite
passed (53 tests). The complete CI suite passed on the final commit.
- Previous-head CI at `c28e2ef12`: 32 checks passed and 2 optional
Storybook checks skipped. Every server/workspace/browser shard, Runner
verification, build, typecheck/release registry, canary, policy, and
security check passed. Greptile: 5/5, no unresolved threads. Earlier
interrupted CI workers were replaced by this fresh complete run.
- After integrating the updated parent: 314 harness tests, 119
queue/admission tests, 44 onboarding/question-delivery tests, and 13
native recovery tests passed locally. Full repository typecheck and
build passed.
- Clarified the skill wording preference: routine replies describe the
action without announcing the internal skill; direct questions and
permission/security/execution disclosures remain truthful.
- The paid `accept-while-running` scenario is registered for all four
local first-task profiles. It has not been run against a model in this
change.

- Rebased onto the merged parent at `11921075a`; the resulting tree
exactly matches the locally verified integration tree. Final-head CI on
`b53054807` passed: 54 successful checks, 2 optional Storybook checks
skipped, no failed checks. Every new server/browser shard, aggregate
verify/e2e gate, Runner, typecheck, build, canary, and security check
passed on the first attempt. Greptile reviewed this exact head at 5/5
with no unresolved threads.

## Risks

- Responses now wait instead of implicitly steering another active turn.
A provider blocked on the original question still receives its answer
directly.
- Approval receipts cannot be edited, discarded, or reordered as
comments. This preserves the recorded decision.
- Interruption must still prove that the prior execution stopped. The
tests cover cleanup waits and duplicate delivery.
- The new paid overlap case can be unexercised if the model finishes
before the click lands. It cannot pass without evidence of overlap.
- No database migration is required. This repairs the existing approvals
and execution controls; it does not implement the roadmap's work-stream
queues.

## Model Used

OpenAI GPT-6 through Codex. The exact deployed model ID and
context-window size were not exposed in this session. Capabilities used:
agentic reasoning, repository inspection, code editing, terminal
commands, and test execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.916.1-canary.3
2026-09-16 14:44:36 -05:00
DottaandPaperclip 11921075a4 Add first-task onboarding skill and Runner E2E coverage (#13517)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The first task helps a new user define and approve useful work.
> - That workflow needs reusable instructions and tests against the
production experience.
> - Native Codex and Claude must load the assigned skill, including
after resume.
> - Maintainers need recorded conversations and precise failed checks to
judge regressions.
> - This pull request adds the first-task skill and a suite in the
shared Runner E2E harness.
> - It keeps behavior results separate from informational quality scores
and incomplete recordings.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

The first onboarding task and the Runner E2E report used to review it.

**Current behavior**

Onboarding embeds its policy in a hidden brief. Native Codex drops the
skill-instructions setting at the Rust boundary. The shared E2E harness
has no onboarding suite or full conversation view.

**Proposed behavior**

Assign and invoke `/first-task` for the onboarding task. Send selected
Codex skills as structured protocol inputs. Run twelve scenarios across
legacy Codex, legacy Claude, native Codex, and native ACPX Claude.
Include all 48 cells in full campaigns. Show recorded chat, question and
approval cards, exact checks, instructions, and billing in the shared
dashboard.

**Reason and benefit**

Measure the real onboarding experience before changing prompts.
Distinguish infrastructure failures, behavior failures, and unexercised
journey steps.

**Breaking changes**

No database migration or production API change. First-task instructions
now live in an assigned skill. The user-edited persona is preserved; the
skill includes the maintainer-approved proposal-mode mapping and
saved-plan requirement.

Related: #11043 is earlier onboarding work. #13422 already fixes native
Claude model pinning, context delivery, and read permissions on master;
this branch includes those fixes through its base. The new Claude
recovery test supplements them.

## What Changed

- Extract and assign the first-task skill while retaining the production
greeting and opening question.
- Carry the Codex skill-instructions flag through thread start and
resume. Resolve explicit task skill references only against assigned
skills and send native skill inputs.
- Invoke an unambiguously selected assigned skill through Claude ACPX’s
native slash-command parser on initial and resumed turns, retaining the
entire task/wake envelope as its argument. Do not carry that invocation
into ordinary tasks.
- Restore the saved single-task proposal modes: confirmation card, or
saved plan with revision-targeted checkbox approval. Explicit plan
requests also require a saved plan.
- Add first-response and complete-journey cases with fixed user facts,
acceptance checkpoints, durable outcome checks, and accounting for child
runs.
- Fail the eval when choice questions have fewer than two real options.
Recognize planning documents without treating them as completed work.
- Add optional, bounded quality judging as explicit post-processing.
- Render full conversations and static interaction cards in the shared
report. Conversations start folded. Show original and regraded results
and incomplete journeys distinctly.
- Keep credential-persistence scanning outside the first-task behavioral
suite; retain public evidence redaction.
- Refresh generated capability references after the API-reference edits.
- Correct shared native question guidance and tool schemas: choices need
at least two meaningful options; open-ended questions use canonical text
fields with the required compatibility payload. Verify both formats
through real tool-authority persistence.
- Disable announcements automatically for every isolated Runner E2E
process and label the gallery environment/provider/target explicitly.
- Remove CI races in the GitHub connection browser test and native
session recovery test by waiting for the actual async work before
asserting its results.

## Verification

- `pnpm exec vitest run
server/src/services/onboarding-first-task-assets.test.ts
server/src/__tests__/issue-onboarding-first-task-routes.test.ts`: 19
passed.
- `pnpm --dir packages/paperclip-runner exec vitest run
src/drivers/acpx/runtime-host.test.ts
src/drivers/acpx/native-skill-prompt.test.ts
src/cli/acpx-runtime-sidecar.test.ts`: 70 passed. Native command
forwarding and the 1 MiB input boundary both failed before their fixes
and passed afterward. Coverage includes changed skills on reopen,
approval context, and an ordinary subsequent task.
- Runner E2E unit suite: 306 passed. Harness typecheck passed. The 64
first-task fixture and grader tests also pass.
- Full repository typecheck and build passed locally. Server typecheck
and Runner build passed again after the native-command change.
- Full GitHub Actions CI passed on `23e56447b`: all
server/workspace/browser shards, Runner verification, typecheck/release
registry, build, canary, policy, and Docker checks. Greptile reviewed
this exact head at 5/5 with no unresolved threads. The earlier broad
local run had database startup/timing failures that passed isolated
retries; the complete remote suite is green.
- Merge verification against current master: 312 harness tests and 13
native recovery tests passed. Regenerated semantic contracts and fixture
hashes pass their consistency check. Full local typecheck and build also
passed on the stacked queue branch. After merging the latest master and
preserving the GitHub setup timing regression in the split browser
suite, both focused GitHub browser tests passed. Three CI timing/startup
flakes passed local verification and one remote retry; all latest-head
checks are green.
- Real pinned Claude SDK and Claude ACP JSON-RPC probes against a local
mock API confirmed that `/skill-name` expands the assigned skill body
before the model request and retains the task arguments. A prose mention
does not. The probes made no paid model calls. The ACP probe used the
current first-task skill body and retained the wake arguments.
- [Full 48-case campaign and
report](https://pages.paperclip.ing/runner-e2e-first-task-35053063880/):
44 passed after three interrupted Codex cases completed in targeted
reruns. Original results, regrades, and all 51 executions remain in the
report provenance.
- [Claude campaign after the shared-question
fix](https://pages.paperclip.ing/runner-e2e-first-task-claude-35099525201/):
10/12 passed with zero single-option failures. All 12 recorded the
current assigned skill and corrected guidance. The failures exposed
skipped skill invocation and a missing saved plan. This PR adds native
command invocation and explicit saved-plan instructions; the subsequent
report below still shows behavior failures.
- [Fresh 12-case Claude
report](https://pages.paperclip.ing/runner-e2e-first-task-claude-35102737804/)
at `78452129e`: 10/12 pass after correcting two false proposal-matcher
failures. The recordings said “Here is the task I will create and
run/complete” in approval cards; the old matcher missed that word order.
Regression tests failed before the fix and pass after it. Original
results and offline regrade provenance remain linked. No agent rerun was
needed. Zero single-option-question failures; two behavior failures
remain: direct work before acceptance on a plain first message, and an
explicit plan request without a saved plan. Neither check was relaxed.
The follow-up `82087ac7e` fixes command-prefix size accounting;
`94aefb1f3` fixes only that proposal matcher.
- Report browser checks confirm folded conversations, rendered cards,
explicit Local/Daytona labels, and no page errors. The published-object
audit scanned 1,306 text files across 2,154 objects with no
credential-format findings or prohibited files. Image pixels and unknown
token formats are outside that scan.

## Risks

- Model behavior is nondeterministic. One campaign is evidence, not a
guarantee. The two remaining Claude behavior failures are visible in the
report and require further product work; this PR does not claim all
onboarding scenarios pass.
- The suite checks persisted Paperclip effects. It cannot prove the
absence of arbitrary external effects.
- Historical recordings can miss later journey steps. These remain
incomplete, never passes.
- Native profiles switch runtime after the production onboarding wizard
because it does not yet expose a native option.
- Quality scores are informational and cannot override behavioral
failures.

## Model Used

OpenAI Codex, GPT-6, with reasoning, repository tools, and code
execution. The exact deployed model identifier and context-window size
are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.916.1-canary.2
2026-09-16 14:24:32 -05:00
Devin Foley dcb04a8062 fix(claude-local): read a macOS isolated login from its suffixed Keychain item (#13519)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Connecting a Claude subscription during onboarding uses an isolated
login: the wizard points `claude` at a per-connection
`CLAUDE_CONFIG_DIR` and then verifies the credential before saving the
connection
> - The verifier reads `.credentials.json` from that directory — but on
macOS, Claude Code does not write a credentials file at all: it stores
the OAuth credential for a custom config dir in a per-directory Keychain
item named `Claude Code-credentials-<first 8 hex chars of sha256(dir)>`
> - So on macOS the connect step can never verify a successful sign-in,
and onboarding dead-ends at "Could not verify the local subscription"
(Linux works because the CLI falls back to writing the file there, which
is why the Docker-based smokes pass)
> - This pull request teaches the credential readers to consult the
login home's own suffixed Keychain item when the file is missing
> - The benefit is that macOS self-hosted users can connect a Claude
subscription during onboarding, while the standing isolation invariant —
an isolated login must never fall through to the machine-level operator
login — is preserved, because only the per-directory suffixed item is
ever read

## Linked Issues or Issue Description

No existing issue. Description follows the bug-report template:

**What happened?**
On macOS, connecting a Claude subscription during onboarding (or from
Connections) always fails with "Could not verify the local subscription.
Run the sign-in command shown for this connection, finish signing in,
then try Connect again" — even after `claude auth login` completes
successfully in the isolated `CLAUDE_CONFIG_DIR`.

**Expected behavior**
After finishing the browser sign-in for the printed command, clicking
Connect verifies the subscription and saves the connection.

**Steps to reproduce**
1. On macOS, run onboarding on a fresh instance and reach "Connect a
model" → Claude → Subscription.
2. Run the printed `export CLAUDE_CONFIG_DIR=… && claude auth login`
command in a terminal on the same machine and complete the browser
sign-in.
3. Return and click Connect. Verification fails every time. Inspecting
the isolated directory shows `.claude.json` with a fully populated
`oauthAccount` but no `.credentials.json`; `security
find-generic-password -s "Claude Code-credentials-<suffix>"` shows the
credential landed in the Keychain, where the verifier never looks.

**Paperclip version or commit**
Reproduced on `2026.915.0-canary.11` (`dffc2b3ca`) with Claude Code
2.1.231.

**Deployment mode**
Self-hosted, authenticated instance on macOS.

**Installation method**
`npx paperclipai onboard` (also affects any macOS install; Linux is
unaffected).

## What Changed

- `packages/adapters/claude-local/src/server/quota.ts`:
- New exported helper `readIsolatedClaudeKeychainToken(loginHome)` —
computes the suffixed service name (`Claude Code-credentials-` + first 8
hex chars of `sha256(loginHome)`) and reads only that item via
`/usr/bin/security`; returns null off macOS
- `readClaudeToken` with a custom `CLAUDE_CONFIG_DIR` now consults that
directory's suffixed item after the file reads miss (previously it
refused the Keychain entirely for custom homes). The unsuffixed operator
item is still gated behind the explicit `allowKeychain` opt-in with no
custom home, unchanged
- `server/src/services/local-ai-credentials.ts`: for anthropic isolated
logins, fall back to the suffixed Keychain item after the hardened
credentials-file reads miss. The file path is untouched and still
preferred; the hardened file reader (`readLocalAiCredentialFile` with
its uid/mode/symlink checks) is not bypassed
- Tests: adapter keychain suite extended (suffixed lookup for custom
homes, no unsuffixed fallback when the suffixed item is absent,
off-macOS null); server verifier suite extended (keychain fallback when
the file is missing, file preferred over keychain, absent-login failure
still never touches the ambient reader)

Security note: the suffix binds each Keychain item to exactly one auth
home, so reading it can only surface the login performed inside that
home. The account-isolation invariant the old code enforced by refusing
the Keychain outright ("never substitute the server operator's login for
a user's isolated login") is preserved — the unsuffixed item is never
consulted for an isolated login, and a new test pins that.

The suffix derivation was confirmed against a live login on macOS: a
real `claude auth login` into an isolated home left no credentials file,
wrote the full `oauthAccount` to `.claude.json`, and created a Keychain
item whose suffix equals the first 8 sha256 hex chars of the exact
`CLAUDE_CONFIG_DIR` string; reading it back with the same `security`
invocation returned the live token, which the new code path then
verifies via the existing quota probe.

## Verification

- `pnpm exec vitest run src/server/quota-keychain.test.ts`
(claude-local): 10 tests pass; full claude-local suite: 287 passed, 1
skipped
- `pnpm exec vitest run src/__tests__/local-ai-credentials.test.ts`
(server): 11 tests pass
- Reverting only the verifier change makes the two new server tests fail
— the suite reproduces the live bug
- End-to-end on macOS: a dev server built from this branch, fresh data
dir, full onboarding walk with a real `claude auth login` into the
printed isolated dir — the connect step verifies and saves the
connection

## Risks

- Low. The change is additive and fail-closed: when the suffixed item is
absent (Linux, older Claude Code versions, no login performed), behavior
is byte-identical to today — the file reads run first and the failure
message is unchanged
- The `security` call runs with the existing 10s timeout and swallowed
errors, matching the established unsuffixed-item code path
- No migrations, no API surface changes

## Model Used

Claude Fable 5 (`claude-fable-5`), extended thinking with tool use
(Claude Code).

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
canary/v2026.916.1-canary.1
2026-09-16 11:45:28 -07:00
Devin Foley d08abcba15 ci: cut PR wall clock from ~16 to ~6 minutes (#13521)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Every pull request runs the Trusted PR CI workflow before merge
> - The test suites roughly tripled in six weeks, and shard balance did
not keep up, so PR runs crept from ~4 to ~17 minutes
> - Slow CI delays every merge and every contributor
> - This pull request rebalances the shards from fresh measurements,
splits the largest test files, reuses the Rust build cache in three more
jobs, and takes the policy job off the critical path
> - The benefit is a PR wall clock near 6 minutes with the same coverage

## Linked Issues or Issue Description

**What existing behavior does this improve?**

PR CI wall clock. A typical green run took 16-17 minutes. Two months ago
it took about 4 minutes.

**Subsystem affected**

The Trusted PR CI workflow (`.github/workflows/pr-trusted.yml`), the
shard-duration manifests, the vitest shard runner scripts, the
`paperclip-runner` package scripts, and the dry-run branch of
`release.sh`.

**Current behavior**

The shard-duration manifests were stale. The general-server manifest had
durations for ~400 of 649 suites. The e2e manifest was missing 14 of 29
specs. Stale median weights made shard steps range 417s-806s (server)
and 277s-745s (e2e). Three jobs each paid a ~3m40s cold cargo release
build. Every test lane waited ~60s for the policy job before it could
start.

**Proposed behavior**

All lanes finish in a narrow ~200-290s band. The manifests carry fresh
measured durations for every suite. The three largest test files are
split so no single file caps a shard. The Rust cache restore runs in
every job that builds the Runner binary. Test lanes start as soon as the
gate resolves.

**Reason and benefit**

Merges stop waiting on CI. The projected wall clock is ~6 minutes for
the same test coverage.

## What Changed

- Rebuild `scripts/general-server-shard-durations.json` (646 suites) and
`scripts/e2e-shard-durations.json` (all specs) from per-suite completion
timestamps in runs 35036001734 and 35024948947.
- Move the PR server lane to the release-verify shape:
`general-server-without-chat` across twelve duration-balanced shards,
plus the chat integration suite split by collected test location across
three dedicated lanes.
- Split `tests/e2e/chat-adapters-ui.spec.ts` into `-providers` and
`-messaging` specs, and `tests/e2e/agent-chat.spec.ts` into `-sessions`
and `-projects` specs. Each pair shares fixtures through a `.shared.ts`
module. Playwright collects the same test sets (39 and 20 tests).
- Raise e2e shards to eight and serialized shards to nine.
- Run the runner package's `check:all` as four matrix lanes:
`check:static`, `check:runner`, and two native vitest `--shard` halves.
The union is exactly `check:all`.
- Add the read-only Rust cache restore (toolchain pin, `save-if: false`)
to the Canary Dry Run, Build, and Typecheck jobs.
- Make release.sh preview publish payloads concurrently in batches of
eight during `--dry-run`. The real publish path stays strictly serial.
- Drop the policy-job lockfile artifact chain. Each lane installs with
`--frozen-lockfile` and falls back to an inline `--resolution-only`
regeneration. The policy job stays a required check through the `verify`
and `e2e` aggregates.
- Update the shard-count mirrors and workflow assertions in the
partition and gate tests.

## Verification

- `node --test scripts/__tests__/run-vitest-stable-shard.test.mjs
scripts/__tests__/e2e-shard.test.mjs` — 30 pass.
- `node --test '.github/scripts/tests/'*.test.mjs` — 410 pass.
- `node --test scripts/__tests__/release-verify-workflow.test.mjs
scripts/cloud-source-verification.test.mjs
scripts/__tests__/release-dry-run-notes.test.mjs` — 42 pass.
- `playwright test --list` collects 39 tests across the chat-adapters
split and 20 across the agent-chat split, equal to the original files.
- A local vitest collection of the chat suite partitions 995 tests into
498/497 line shards.
- Projected shard weights: server 230s x12, chat ~143s x3, e2e 207-242s
x8, serialized ~216s x9.

## Risks

- The split spec files reorder tests relative to the original files.
Every describe seeds its own company, so the specs stay independent; a
hidden cross-describe dependency would surface as a deterministic
failure in one shard.
- The inline lockfile fallback changes install behavior for
manifest-changing and stacked PRs. The policy job still validates
resolution as a required check.
- `release.sh` changes are confined to the `--dry-run` preview branch.
The publish loop is untouched. `bash -n` passes and the release dry-run
tests pass.
- One PR now schedules ~44 fleet runners. If the RunsOn fleet caps
concurrency, queueing may absorb part of the gain; watch the first runs.

## Model Used

- Claude Fable 5 (`claude-fable-5`, Anthropic), extended thinking, with
tool use (shell, file edits) in Claude Code.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-16 11:45:14 -07:00
Devin Foleyandgithub-actions[bot] c49336acdf docs(release): canonicalize v2026.916.0 release notes (#13548)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Stable v2026.916.0 shipped today; its release notes were maintained
during the beta window at `releases/beta/v2026.916.0-beta.0.md`
> - The release workflow's canonicalize step renames the file to its
canonical stable path once the promotion completes, pushing the rename
to a branch for a human-opened pull request (bot-opened PRs trigger no
CI)
> - This pull request is that rename:
`releases/beta/v2026.916.0-beta.0.md` → `releases/v2026.916.0.md`
> - The benefit is that the published GitHub Release body and the
in-repo notes file agree on the canonical path

## Linked Issues or Issue Description

Routine post-release canonicalization for the `v2026.916.0` stable,
generated by the release workflow's `canonicalize_stable_notes` job (run
35126040769). Content is byte-identical to the notes published on the
GitHub Release.

**What happened?** Stable promotion completed; the notes file needs its
canonical name.
**Expected behavior** `releases/v2026.916.0.md` exists on master
matching the published release body.
**Steps to reproduce** n/a — mechanical rename.

## What Changed

- Rename `releases/beta/v2026.916.0-beta.0.md` to
`releases/v2026.916.0.md` (no content change)

## Verification

- Diff is a pure rename; content matches the published GitHub Release
body for v2026.916.0

## Risks

- None — docs-only rename.

## Model Used

None for the change itself (generated by the release workflow); PR
opened via Claude Fable 5 (`claude-fable-5`, Claude Code).

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
canary/v2026.916.1-canary.0
2026-09-16 11:32:21 -07:00
DottaandPaperclip 9fd2e50310 feat: create company skills from runner tasks (#13538)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The Runner gives agents tools to change company resources.
> - Users need agents to save reusable skills during a task.
> - A saved skill needs a visible result that users can inspect and
edit.
> - This pull request adds `create_skill` and a task feed card linked to
Skill Studio.
> - Users can open the saved skill from the task and edit the same
resource.

## Linked Issues or Issue Description

**Subsystem affected**

Runner tools, company skill storage, task feed, and Skill Studio.

**Problem or motivation**

The Runner has no dedicated tool to create a company skill. A user
cannot follow a creation result from the task feed to the saved skill.

**Proposed solution**

Add a company-scoped `create_skill` tool. Save the skill with the
existing company policy. Add one creation card to the task. Open a named
sidebar tab from that card. Let the user open the same skill in Skill
Studio.

**Alternatives considered**

An agent can write a local file, but that file is not a company skill. A
second document copy in the task would become stale after a Studio edit.
The sidebar therefore reads the saved skill directly.

**Roadmap alignment**

This extends the shipped Skills Manager, Skill Studio, and Skills Store
milestone. The maintainer requested and approved this scope. Search
found no duplicate `create_skill` PR or issue. Related UI validation
work: #8715. This PR does not change that validation display.

## What Changed

- Add the real Runner tool, its contract, and its mock implementation.
- Validate the complete SKILL.md and derive company, task, agent, and
run identity from authentication.
- Apply the existing company skill policy. Do not assign the skill to an
agent.
- Make keyed retries return one skill and one creation event. Reject
conflicting retries.
- Make concurrent file creation safe. Never replace an existing
published skill during creation.
- Add a creation card, a named sidebar tab, and an Open in Skill Studio
action.
- Show saved Studio edits when the user returns to the task.
- Add storage, policy, mode, retry, UI, and Product E2E tests. Document
the tool.
- Fix deleted-name reuse, onboarding panel persistence, immediate feed
refresh, and mock validation parity from review.
- Serialize Studio file edits and renames with skill deletion and
recreation. Reject stale editor requests before they can change a
replacement skill.
- Generate the standalone mock parser and validator from the production
contract. Use portable UUIDs so the browser scenario bundle builds.

## Verification

- All latest-head PR checks pass on `145dd76a5`, including all server
shards, browser E2E, Runner verification, build, typecheck, and release
dry run. Greptile: 5/5 with no open findings. An interrupted CI runner
was retried successfully.
- `pnpm -r typecheck`: passed.
- `pnpm build`: passed.
- `pnpm check:token-gates`: passed.
- Review regressions: 73 storage tests, 6 real API tests, 63 UI tests,
and 61 semantic runtime tests passed. Parser synchronization passed.
- CI exposed existing fire-and-forget Sentry test races. Reproduced the
resumption race locally, then synchronized the related sweep and
finalizer assertions on the actual report; all 27 tests across the three
affected files pass.
- Runner scenario browser build and strict content-security-policy
check: passed.
- Runner suite: 2,012 tests passed; 10 skipped.
- `pnpm test:run`: the general-server batch had 12,416 passes and two
failures. The old tool-count assertion was fixed; all 16 authority tests
then passed. The chat webhook test had a socket error; it passed four
isolated reruns.
- Both workspace test groups passed. The isolated route suites
completed. Two socket failures in the initial route batches passed on
individual reruns; all remaining 61 files passed.
- Product E2E `create-skill-studio`: passed with local Codex and local
ACPX Claude.
- Manual browser test: submit a task, observe the real tool call and
creation card, open the sidebar, edit in Studio, save, and return. The
task reached Done. The saved second revision and sidebar tab survived a
server restart.
- The new companion headless Runner Eval passed. Companion coverage PR:
https://github.com/paperclipai/paperclip-evals/pull/23. Daytona was not
run because no immutable runner image was configured.

## Risks

- Database writes and local file writes cannot share one transaction.
Recovery accepts only an exact file-for-file retry after a database
rollback. Conflicting files remain untouched.
- The sidebar displays the current skill. The feed card remains the
historical creation receipt.
- No database migration, dependency, or workflow change is included.
- Remote Daytona behavior still needs a run with a configured immutable
image.

## Model Used

OpenAI GPT-6 (`gpt-6-astra`) handled design, integration, review, and
browser verification. OpenAI `gpt-5.6-luna` assisted with bounded
implementation and eval work. Both used code execution and tool access.
The host did not expose the context window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-16 13:01:58 -05:00
Devin Foley 5b6e54fba8 fix(cli): accept prompt defaults in onboarding wizard validators (#13520)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The CLI onboarding wizard's custom-setup path collects server,
database, storage, and secrets configuration through `@clack/prompts`
text prompts, most of which show a sensible default
> - `@clack/prompts` runs each prompt's `validate` callback on the raw
typed value, and only substitutes `defaultValue` after validation passes
— so pressing Enter to accept a shown default hands the validator an
empty string
> - Eleven of the wizard's validators reject the empty string, which
makes their own displayed defaults unacceptable: accepting "Embedded
PostgreSQL port: 54329" fails with "Port must be an integer between 1
and 65535", "Backup directory" fails with "Backup directory is
required", and so on through every prompt in the custom path
> - This pull request makes each affected validator accept empty input
when a default exists, while keeping all real validation for typed input
> - The benefit is that the custom-setup wizard is walkable by pressing
Enter through the defaults, as the UI clearly intends

## Linked Issues or Issue Description

No existing issue. Description follows the bug-report template:

**What happened?**
In `paperclipai onboard` custom setup, pressing Enter to accept a
prompt's displayed default fails validation on eleven prompts. Confirmed
live on the "Embedded PostgreSQL port" prompt (default 54329 → "Port
must be an integer between 1 and 65535") and the "Backup directory"
prompt (populated default path → "Backup directory is required"). The
only way through is to retype every default by hand.

**Expected behavior**
Pressing Enter accepts the displayed default, as in every standard
`@clack/prompts` flow.

**Steps to reproduce**
1. `npx paperclipai onboard` → choose Custom setup.
2. At "Embedded PostgreSQL port", press Enter to accept the shown
default.
3. Validation rejects it. Same for the backup directory, backup
interval/retention, server port, bind host, storage directory, S3
bucket/region, and secrets key-file prompts.

**Paperclip version or commit**
Reproduced on `2026.915.0-canary.11`; the same validators exist in the
latest stable (`v2026.831.1`) — long-standing, not a recent regression.

**Deployment mode**
Any (the bug is in the CLI wizard).

**Installation method**
`npx paperclipai onboard`.

## What Changed

- Audited all 15 `validate:` callbacks across
`cli/src/prompts/{database,server,storage,llm,secrets}.ts`. Eleven
rejected the empty string while displaying a default; two were already
fine (hostname CSV prompts, where empty parses to `[]`); one is an
intentionally required password with no default (left alone); one
(PostgreSQL connection string) and one (public base URL) are
conditionally required — empty now passes only when a saved default
exists, so fresh setups still enforce the field
- Fix pattern: allow empty/undefined input at the top of each affected
validator; every check for non-empty typed input is unchanged. One
deliberate exception: the bind-host validator validates `(val ||
defaultHost)` so accepting the default still runs the loopback safety
check rather than bypassing it
- New `cli/src/__tests__/prompt-default-accept.test.ts` (repo-convention
vitest + clack mock reproducing real submit semantics — validate raw
`""`, then substitute the default): drives the four prompt modules
end-to-end and unit-exercises each captured validator (empty accepted,
garbage still rejected, required-when-fresh still rejected)

## Verification

- `pnpm exec vitest run src/__tests__/prompt-default-accept.test.ts` in
`cli/`: 13/13 pass; with the prompt fixes stashed: 12/13 fail — the
suite reproduces the bug
- Full cli suite: 483 passed; 15 failures are pre-existing/environmental
on the clean tree too (macOS `/var` realpath mismatch in
`worktree.test.ts`, one parallel-run flake that passes in isolation)
- `pnpm run typecheck` in `cli/`: zero errors in `cli/src` (the 228
pre-existing errors in `../server/src` from missing prebuilt workspace
dist are identical with and without the change)

## Risks

- Low. Typed input validates exactly as before; whitespace-only input is
still rejected (clack substitutes the default only for truly-empty
input). The two conditionally-required prompts still hard-require a
value on fresh setups
- No migrations, no API surface changes; CLI-only

## Model Used

Claude Fable 5 (`claude-fable-5`), extended thinking with tool use
(Claude Code); implementation drafted by a subagent on the same model
and reviewed before commit.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
canary/v2026.916.0-canary.4
2026-09-16 10:23:32 -07:00
Devin Foleyandgithub-actions[bot] de9ac61fb5 docs(release): curate stable notes for v2026.916.0 (#13522)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The release system publishes canary, nightly, beta, and stable
channels; a stable promotion publishes its release notes as the GitHub
Release body
> - At beta publish time the workflow auto-drafts a raw commit-log
skeleton on this branch (`releases/beta/v2026.916.0-beta.0.md`) for
humans to edit
> - The skeleton is a 1,800-line commit dump; stable `v2026.916.0`
cannot ship user-facing notes in that form
> - This pull request replaces the skeleton with the curated changelog
for the `v2026.916.0` promotion: overview, breaking changes, highlights,
fixes, upgrade guide, and contributor credits for the 483-commit range
since `v2026.831.1`
> - The benefit is that the stable promotion reads finished, accurate
notes from master and publishes them as the release body

## Linked Issues or Issue Description

Refs #13403, #13247, #13248, #13268, #13256, #13038, #13299 — the
headline features this changelog describes.

The notes-drafting flow itself (skeleton branch at beta publish, stable
promotion reading the file from master, post-ship canonicalization) is
the standing release process; this PR is the curation step it expects.

## What Changed

- Replaced the auto-drafted skeleton in
`releases/beta/v2026.916.0-beta.0.md` with the curated release notes for
stable `v2026.916.0` (promoted from beta `2026.916.0-beta.0`, source
`dffc2b3ca`)
- Also removes the orphaned `releases/beta/v2026.915.0-beta.0.md`: that
beta's promotion was abandoned before publishing, and this changelog
supersedes it
- Sections: overview, Breaking Changes (5), Highlights (Connections
train leads), Improvements, Fixes, Upgrade Guide (migrations
`0231`–`0279`, new env vars, removed API surface), Contributors

## Verification

- Every PR link and claim was checked against the actual commit range
`dbf052577..667c79ded` (the same range the skeleton header names)
- Migration list enumerated from `git log --diff-filter=A` over that
range; only `0231` and `0236` discard data, called out as such
- New environment variables verified in code (`server/src/config.ts`):
announcement flags (default on), split Sentry DSNs with legacy fallback,
token-broker allowed hosts, cwd `.env` opt-out
- Contributor list built from commit authors plus co-author trailers;
core maintainers excluded per release-notes convention
- Docs-only change: no code paths, no tests to add

## Risks

- Low risk: documentation only. The stable promotion reads this file
from master at dispatch time; if wording needs a follow-up after merge,
the promotion pins the master revision it resolved, so edits must land
before the stable dispatch
- The file is renamed to `releases/v2026.916.0.md` by the post-ship
canonicalization step, as with previous promotions

## Model Used

Claude Fable 5 (`claude-fable-5`), extended thinking with tool use
(Claude Code). The commit-range analysis and draft were produced by a
subagent on the same model and human-review-style checked against the
range before commit.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
canary/v2026.916.0-canary.3
2026-09-16 10:03:01 -07:00
DottaandPaperclip 18989a9e73 docs: add eval guide, authoring skills, and public history hub (#13535)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Its evaluations test both the Runner and complete product workflows.
> - The guides and run histories are in separate places.
> - The shared Evalbook viewer can make the test boundary unclear.
> - This pull request names the two families and adds a guide, authoring
skills, and a public hub.
> - Contributors can choose the correct test and inspect its history.

## Linked Issues or Issue Description

**Issue type**

Missing documentation.

**Where is the issue?**

Runner and Product E2E evaluation guides, case-authoring procedures, and
public result navigation.

**What's wrong?**

There is no single entry point. A report format can be mistaken for an
execution boundary. There are no dedicated case-authoring skills for
these two families.

**Suggested fix**

Add a guide and three skills. Link both existing histories from a public
hub. Keep existing campaign URLs and grading unchanged.

## What Changed

- Add `doc/evals.md` and links from existing guides.
- Add the `paperclip-evals`, `add-runner-eval`, and
`add-product-e2e-eval` skills. Install copies in
`~/paperclipai/.agents/skills`.
- Add a static hub builder that reads the existing public history feeds.
- Show a dated snapshot for each family. Label partial campaigns and
preserve measurement dates across report refreshes.
- Document publication and refresh commands for
https://pages.paperclip.ing/evals/.

## Verification

- Seven Python summary tests pass: `python3 -m unittest discover -s
scripts/evals-hub -p 'test_*.py'`. Run these checks directly; this PR
does not modify package scripts.
- All three skills pass the skill-creator `quick_validate.py` check with
`/usr/bin/python3`.
- Build tested with saved history fixtures and the live public feeds.
- Desktop and mobile browser checks pass. The mobile page has no
horizontal overflow.
- Published https://pages.paperclip.ing/evals/. Browser check: HTTP 200,
no page errors, all eight links return HTTP 200, no mobile overflow.
- Independent skill exercises found the existing Notion-decline case and
a direct Runner permission-denial case. Roster validation with an
explicit run ID passes.
- Missing refresh measurement date: regression fails before the fix and
passes after it.
- `git diff --check` passes.
- No paid evals were run for this documentation and reporting change.
The preceding head passed typecheck, build, server/workspace tests,
runner verification, browser E2E, and the canary dry run. Checks for the
latest commit are pending. Local repository-wide typecheck, test, and
build were not repeated because no product code changed.

## Risks

The hub is a dated static snapshot. It can lag behind the linked
histories until an operator refreshes it. A changed history schema stops
the build. Existing archives and grades are not modified. The published
guide link is pinned to the reviewed commit so branch deletion cannot
break it. Later builds can use master.

## Model Used

OpenAI gpt-6-astra for implementation and review. OpenAI gpt-5.6-luna
for documentation and independent skill checks. Both used repository
tools and code execution. Context window sizes are not exposed by this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green (latest commit pending)
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
(preceding head was 5/5; latest commit pending)
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.916.0-canary.2
2026-09-16 08:41:06 -05:00
4577d10029 fix: prepare everyday artifact and Codex sandbox prerequisites in CI (#13516)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The runner E2E workflow executes paid everyday workflow stories on
disposable CI hosts
> - The everyday artifact oracle requires a pinned Python image and
fails closed when it is absent
> - Fresh CI hosts did not prepare this image before the paid cell, so
project stories failed during preflight
> - This pull request prepares and verifies the pinned image before the
affected everyday cells
> - The benefit is reliable artifact isolation checks on fresh trusted
CI hosts

## Linked Issues or Issue Description

**What happened?**

Fresh trusted CI runners did not have the pinned Python artifact oracle
image.

**Expected behavior:**

The workflow prepares and verifies the pinned image before an everyday
project story starts.

**Steps to reproduce:**

Run an everyday project story on a fresh CI host without the image
cached. The `everyday-artifact.py --preflight` check fails before task
creation.

**Paperclip version or commit:**

`master` at `bd51f157e`.

**Deployment mode:**

Other: GitHub Actions trusted paid workflow.

## What Changed

- Add a matrix-gated CI step for everyday project and recovery cells.
- Check Docker, pull the fixed digest with bounded timeouts, and verify
the exact repo digest.
- Apply the existing Codex sandbox preparation to both native Codex
profiles, including the mini profile.
- Add workflow security assertions for ordering, condition, digest,
timeouts, and secret isolation.
- Document that CI prepares the pinned oracle image.
- Check provisioning eligibility against every catalog cell, and scope
the Daytona registry inspection assertion to the Daytona image job.

## Verification

- `pnpm test:e2e:runner:unit` — 24 files and 222 tests passed.
- `pnpm test:e2e:runner:typecheck` — passed.
- `pnpm exec vitest run --config tests/runner-e2e/vitest.config.ts
workflow-security.test.ts` — 10 tests passed.
- Python artifact oracle calibration — 12/12 passed.
- `git diff --check` — passed.

## Risks

Low risk. The image step runs only for everyday cells that execute the
artifact preflight. The Codex setup now covers both native Codex
profiles. It uses a fixed public image digest and has no provider
credentials.

## Model Used

OpenAI gpt-5.6-luna (implementation subagent) and gpt-6-astra (review
fixes and orchestration), using code execution and repository tools.
Context window size is not exposed by this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Paperclip <paperclip@paperclip.ing>
2026-09-16 07:49:48 -05:00
Nicky Leach ae06329971 test(server): hoist the route module graph in the issue ownership authz suite (#13524) canary/v2026.916.0-canary.0 nightly/v2026.916.0-nightly.1 2026-09-15 22:53:44 -07:00
Nicky LeachandPaperclip e1f245a660 fix(server): recover sandbox leases stranded active after a restart (#13515)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The server starts and stops provider sandboxes through environment
leases
> - A restart can leave a terminal run with an `active` lease
> - The normal orphan recovery path cannot find that lease after the run
ends
> - This pull request adds a bounded sweep that changes the stranded
lease to `pending_cleanup`
> - The existing cleanup sweep then stops the provider sandbox on the
same heartbeat tick
> - The benefit is that stranded sandboxes stop and do not continue to
create provider cost

## Linked Issues or Issue Description

**What happened?**

A restart can occur after the server writes a terminal run status but
before it releases the related environment lease. The lease then stays
`active`, and later recovery does not select it. A second path skips the
lease when its environment row does not exist.

**Expected behavior**

The heartbeat recovery path must find an `active` lease that no live run
can release. It must move that lease to `pending_cleanup`, and the
cleanup sweep must stop the provider sandbox.

**Steps to reproduce**

1. Start a run that owns a provider sandbox lease.
2. End the run and stop the server between the run-status write and the
lease-release write.
3. Restart the server and allow the heartbeat recovery sweep to run.
4. Confirm that the lease reaches `pending_cleanup` and the provider
sandbox receives a stop request.

**Paperclip version or commit**

This pull request targets the current `master` branch at the base commit
used for review.

**Deployment mode**

The change applies to local development and server deployments.

**Installation method**

Built from source with the repository test commands.

**Agent adapter(s) involved**

Not adapter-specific. The change applies to core heartbeat recovery.

**Database mode**

The change uses the existing database tables. It adds no migration.

## What Changed

- Add `sweepOrphanedActiveLeases()` to heartbeat recovery.
- Select only stale `active` leases that have no live run owner.
- Skip leases with a different live lease for the same provider
resource.
- Preserve retained leases and write a failure reason for recovered
leases.
- Limit each sweep to 20 rows.
- Run the recovery sweep before the pending-cleanup sweep.
- Add focused tests for the recovery guards and same-tick cleanup.

## Verification

- `pnpm vitest run
server/src/__tests__/heartbeat-orphaned-active-lease-sweep.test.ts` — 10
tests pass.
- `pnpm vitest run
server/src/__tests__/heartbeat-pending-cleanup-sweep.test.ts` — 22 tests
pass.
- `pnpm --filter @paperclipai/server typecheck` — exits 0.
- The complete CI suite remains the final check for the repository.

## Risks

The sweep changes only stale `active` leases that no live run can
release. The stale threshold, live-resource guard, retained-lease guard,
and page limit reduce false recovery. The change adds no endpoint,
schema change, or migration.

## Model Used

OpenAI Codex, GPT-5, tool-enabled coding agent with repository
inspection, GitHub CLI, and test execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-15 21:52:02 -07:00
Nicky LeachandClaude Opus 5 bd51f157e9 fix(ci): make the Runner Rust cache key independent of the image toolchain (#13500)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Every change goes through the pull request CI workflow, and its
`Verify Paperclip Runner` lane builds and tests the native Rust runner
> - https://github.com/paperclipai/paperclip/pull/13457 made that lane
restore master's prebuilt Rust dependency cache instead of recompiling
313 crates
> - Measured over 26 runs since it went live, the cache hits on the
RunsOn fleet and misses on every GitHub-hosted runner
> - The cause is the cache key: `rust-cache` hashes every installed
toolchain, and each runner image ships a different stable Rust next to
the pinned one
> - This pull request removes the extra toolchains before the key is
computed, in the reader and the writer
> - The benefit is that the saving the cache already delivers, 4.8
minutes per run, reaches the 73% of runs that currently miss it

## Linked Issues or Issue Description

No public GitHub issue exists for this. The problem follows the
enhancement issue template below.

- Refs https://github.com/paperclipai/paperclip/pull/13457 — added the
cache this pull request repairs
- Refs https://github.com/paperclipai/paperclip/pull/13459 — activated
it
- Refs https://github.com/paperclipai/paperclip/pull/13194 — created the
`release-runner-v1` entry on master

I searched this repository for other pull requests touching this cache
and found no duplicate and nothing in flight.

**What existing behavior does this improve?**

The Rust dependency cache added in #13457 misses on GitHub-hosted
runners, so most pull requests still recompile the whole dependency
tree.

**Subsystem affected**

Cross-cutting (multiple of the above). The change touches CI workflow
configuration only. It does not change product code.

**Current behavior**

The cache works, but only on one runner class. Across 26 successful
`Verify Paperclip Runner` jobs since #13459 merged:

| Runner | Runs | Before | After | Change | Cache |
|---|---|---|---|---|---|
| RunsOn fleet | 7 | 16.4m | 11.6m | −4.8m | 4 of 4 full hit |
| ubuntu-latest | 19 | 14.9m | 14.7m | −0.2m | 0 of 6 hit |
| All | 26 | 15.0m | 13.9m | −1.1m | |

Only 27% of runs reach the fleet, so the fleet-wide saving is 1.1
minutes rather than the 4.8 minutes the cache delivers where it lands.

The two runners compute different keys:

```
fleet:         v0-rust-release-runner-v1-Linux-x64-4bb3b8ea-a95b0328
ubuntu-latest: v0-rust-release-runner-v1-Linux-x64-9fdc73e3-a95b0328
```

The lockfile half agrees. The environment half does not. `rust-cache`
logs why, under `Environment considered`:

| Runner | Toolchains it found |
|---|---|
| fleet | 1.97.1 and **1.98.0** |
| ubuntu-latest | 1.97.1 and **1.98.1** |

`rust-cache` hashes every installed toolchain, not only the active one.
Both images carry the pinned 1.97.1. Each also ships its own stable
Rust, and those differ by a patch version. The post-merge writer runs on
a RunsOn image, so the fleet agrees with it and GitHub-hosted runners
cannot.

Pinning `RUSTUP_TOOLCHAIN` in #13457 was necessary but not sufficient.
It fixes which toolchain builds the code. It does not change which
toolchains exist on the image.

**Proposed behavior**

Remove every toolchain except the pin, before the cache step, in both
the reader and the `release-runner-v1` writer. The key then depends on
the pinned compiler and the lockfile alone, not on what the image
happens to carry.

**Reason and benefit**

The cache already proves its value where it lands: release compile drops
from 5m07s to 1m30s, debug from 1m48s to 13s, and the job from 16.4m to
11.6m. This change extends that to the other 73% of runs.

Expected fleet-wide mean: about 11.5m, against 13.9m today and 15.0m
before #13457.

It also removes a standing fragility. The fleet hits today only because
two RunsOn images happen to agree. If either image updates its stable
Rust on its own, the hit rate drops to zero with no code change.

**Breaking changes**

None. The change only affects cache key computation. A miss reproduces
the current behavior.

**Additional context**

The typecheck writer in `release-verify.yml` keeps its current step on
purpose. It restores and saves on the same post-merge image, so its key
never disagrees with itself.

## What Changed

- Added a toolchain normalization block to `Select the pinned Runner
Rust toolchain` in `.github/workflows/pr-trusted.yml`, before the cache
restore. It keeps the pinned toolchain and uninstalls the rest.
- Added the identical block to the same step in the
`verify_paperclip_runner` job of `.github/workflows/release-verify.yml`,
which writes `release-runner-v1`. The reader and the writer must agree,
or the key matches nothing.
- Made the block tolerant. If a toolchain cannot be removed it prints a
notice and continues, so a pull request loses the cache rather than the
run.
- Extended `.github/scripts/tests/pr-runner-rust-cache.test.mjs` with
two tests: the block is byte-identical in both workflows, and it runs
before the cache step in each.

## Verification

Run the workflow shape tests:

```bash
node --test '.github/scripts/tests/*.test.mjs' ./scripts/__tests__/e2e-shard.test.mjs ./scripts/__tests__/release-verify-workflow.test.mjs ./scripts/__tests__/run-vitest-stable-shard.test.mjs ./scripts/cloud-source-verification.test.mjs
```

Result: 471 pass, 0 fail.

I ran the step body against a stub `rustup` to confirm the logic, rather
than only checking syntax. Three paths, all exit 0:

| Case | Result |
|---|---|
| Pin plus an extra stable toolchain | Uninstalls only the extra,
exports `RUSTUP_TOOLCHAIN=1.97.1-x86_64-unknown-linux-gnu` |
| Pin only, already normalized | No uninstall calls, no error |
| No `rustup` on `PATH` | Prints the notice and continues |

I also mutation-tested the new parity assertion. Each mutation edits
only the writer, then the reader and writer disagree:

| Mutation to `release-verify.yml` | Result |
|---|---|
| One word changed in the shared comment | Caught |
| `uninstall` changed to `remove` | Caught |
| Trailing `rustup toolchain list` deleted | Caught |

After this merges, confirm the repair in CI. Take a `Verify Paperclip
Runner` job that ran on `ubuntu-latest` and check the restore step for
`full match: true`. Under `Environment considered`, `Rust Versions` must
list only 1.97.1. The job should finish near 11.5m rather than 14.7m.

## Risks

- **Master must republish the cache once.** This changes the key, so the
existing `release-runner-v1` entry no longer matches.
`cloud-readiness.yml` runs `release-verify.yml` on every push to master,
so the first push after this lands writes the new entry. Pull requests
merged before that miss the cache, which is exactly what most of them do
today. No run breaks.
- **The reader and the writer must stay in step.** If they diverge, no
run hits the cache. The new parity test fails on any difference in the
block, including a comment.
- **Uninstalling the image toolchain is deliberate.** Everything in this
job builds through the pinned 1.97.1, resolved from
`packages/paperclip-runner/rust-toolchain.toml` and `RUSTUP_TOOLCHAIN`.
Nothing in the job uses the image default.
- **Low blast radius.** The change only affects cache key inputs. A miss
compiles from scratch, as today.
- **Image drift is now handled.** The key no longer depends on the
image's own Rust version, so a future image update cannot silently
disable the cache.

## Model Used

Claude Opus 5, provider Anthropic, exact model ID `claude-opus-5`, 1M
context window. Adaptive thinking was on. I used tool use throughout:
the `gh` CLI and the GitHub API to pull 26 post-merge job records and 10
full job logs, log parsing in Python to isolate the per-phase timings
and the two cache keys, a stub `rustup` on `PATH` to exercise the new
step, and local `node --test` runs to verify and mutation-test the
guards. Run through Claude Code.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Note on the unchecked boxes. The CI and Greptile boxes stay unchecked
until those checks finish. On documentation: no document describes the
Runner cache keys, so there is nothing to update.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 17:24:22 -07:00
a8d32e5e61 feat(sandbox-providers): add CreateOS sandbox provider (#13434)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Agent work runs in sandboxes that provider plugins supply
> - Operators can choose a provider to run agent work
> - CreateOS adds another provider with workspace-preserving pause and
resume
> - This pull request adds a CreateOS provider plugin
> - The benefit is that operators can preserve a workspace between runs
without keeping its compute active

## Linked Issues or Issue Description

Refs #13203 and the earlier closed #13096.

This continues the CreateOS contribution from @bhautikchudasama and
@ashwaq06. The branch preserves the original implementation commit.
Thank you to both contributors.

When squash-merging, preserve the original author's credit in the squash
commit body:

```text
Co-Authored-By: bhautikchudasama <BhautikChudasama@users.noreply.github.com>
```

The original fork rejects maintainer pushes. This branch includes the
merge-conflict resolution and review fixes. The request is described
below using `adapter_request.yml`.

**Agent or provider**

CreateOS sandbox API (https://api.sb.createos.sh).

**Why this adapter is useful**

CreateOS can pause a sandbox and resume it by ID. The workspace survives
the pause. This adds a reusable-lease option to the existing sandbox
provider system.

**How the agent is invoked**

Build and install the local plugin as described in its README. Open
Instance Settings, then Environments. Select the `createos` driver.
Supply an API key and shape. The driver then supplies sandbox leases for
agent runs.

**Are you willing to implement it?**

Yes. This pull request is the implementation.

## What Changed

- Adds the `createos` sandbox provider under
`packages/plugins/sandbox-providers/createos`.
- Calls the CreateOS HTTP API directly. The package adds no vendor SDK.
- Implements the environment lifecycle hooks, incremental process
output, and binary workspace sync.
- Registers the optional bundled provider and its trusted host
credential fallback. The fallback is limited to the official API origin;
custom endpoints require an explicit key.
- Lists the package in the release manifest with `publishFromCi: false`
until its first npm publish is bootstrapped.
- Waits through delayed pause/resume state updates without duplicate
action requests.
- Cancels queued API requests promptly while preserving request spacing.
- Uses direct CLI invocation in the setup guide so paths and IDs are
passed without an extra shell expansion.
- Includes current master and retains its existing Git-subfolder
containment fix.

## Demo

Fresh setup and a run against a CreateOS sandbox.


https://github.com/user-attachments/assets/e71b9e06-c006-4fb9-b847-52dfd68f6110


https://github.com/user-attachments/assets/43b5ac75-66bd-4f76-8563-67e4c7759084

## Verification

All 25 jobs in [CI run
34884260542](https://github.com/paperclipai/paperclip/actions/runs/34884260542)
passed at commit `f8d0997677024b784fdadf9d44a84c01cb4e813c`, including
typecheck, build, native runner verification, server and workspace
tests, browser tests, and the canary release dry run. Greptile reviewed
the same commit at 5/5 with no unresolved review threads.

GitHub reports no merge conflicts. The remaining merge gate is
code-owner approval for the new `package.json`, as required by
`.github/CODEOWNERS` and the `master` ruleset. Reviewers have been
requested automatically.

Local checks passed:

- Provider: `pnpm typecheck`, `pnpm test` (52 passed, one live smoke
skipped), and `pnpm build`.
- Host: focused credential and bundled-plugin tests (17 passed), plus
CLI invocation safety (39 passed).
- Release: package manifest check and release policy tests (18 passed).

The full local `pnpm test:run` attempt caught the README command issue;
its focused rerun now passes. The full local run stopped after its
general-server group: 7,804 tests passed, with unrelated embedded
PostgreSQL startup failures and 10 failures in unchanged
runtime-skill-cache tests (`EACCES` on directory rename on macOS). It
did not reach the later test groups. Local `pnpm -r typecheck` and `pnpm
build` reach the runner package and stop because this machine has no
Rust/Cargo installation. The corresponding CI checks passed on
provisioned runners, as linked above.

The live CreateOS smoke requires explicit provider credentials and was
not run during this review. It is available with `CREATEOS_LIVE_TEST=1
pnpm test` in the provider directory. The author supplied the demo links
above.

## Risks

The provider is opt-in and is not installed by default. It is available
through a local-path install or explicit image inclusion. npm
publication remains disabled until a maintainer bootstraps the package
and enables publishing.

Sandbox creation has no idempotency key. An ambiguous create response
can leave a resource that requires provider-account inspection. Process
tracking is in memory; durable lease recovery belongs to the host. The
provider does not advertise guaranteed expiry, interactive login,
snapshots, duplex channels, or ingress. Live native-runner qualification
remains outside this PR's tested claims.

## Model Used

Original provider implementation: human-authored by @bhautikchudasama,
as reported in #13203. The original description reports Claude Opus 5
assistance.

Review and follow-up fixes: OpenAI GPT-6 via Codex, with code review,
editing, and tool execution. The precise runtime model variant and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused checks; full-suite
environment limits documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: bhautikchudasama <bhautikrchudasama@gmail.com>
Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-15 17:09:43 -07:00
scotttongandClaude Fable 5.1 dffc2b3ca1 fix(claude-local): skip expired credentials file when reading the Claude token (#13505)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Agents on the Claude Local adapter run on a local Claude Code
subscription. A user connects that subscription from Agent → Harness /
Runtime → "Connect account".
> - The server verifies the login with `readClaudeToken` in
`packages/adapters/claude-local/src/server/quota.ts`. It reads
`~/.claude/.credentials.json` first and consults the macOS Keychain only
when no file is present.
> - On macOS the Claude CLI refreshes the Keychain item, not the file. A
leftover credentials file keeps an expired token forever, and the reader
ignores `claudeAiOauth.expiresAt`.
> - The stale file shadows the live Keychain login. The usage check
fails and the user sees "Could not verify the local subscription"
although `claude auth status` reports a valid login. Running `claude
auth login` again does not help.
> - This pull request skips a credentials file whose token has expired,
so the reader falls through to the Keychain or returns null.
> - The benefit is that a valid local Claude login connects on the first
try, and a dead token is never sent upstream.

## Linked Issues or Issue Description

No public issue exists for this bug. Description follows
`bug_report.yml`.

### What happened?

Agent → Harness / Runtime → "Connect account" → Claude (Subscription) →
Connect failed with:

> Could not verify the local subscription. Run claude auth login in a
terminal on the machine running Paperclip, then try Connect again.

`claude auth status` on the same machine reported `loggedIn: true`,
`authMethod: claude.ai`, `subscriptionType: max`. The Keychain item
`Claude Code-credentials` held a fresh token. `POST
/api/companies/:id/ai-connections/local/check` returned
`{"status":"sign_in_required"}`.

A leftover `~/.claude/.credentials.json` (written weeks earlier) held an
access token that expired the same day it was written. `readClaudeToken`
returned that token. `fetchClaudeQuota` got a non-OK response from
`/api/oauth/usage`, and the route threw the generic 422.

### Expected behavior

An expired credentials file must not block a valid login. The reader
skips the dead token and falls through to the Keychain. Connect
succeeds.

### Steps to reproduce

1. On macOS, sign in with `claude auth login` (credentials land in the
Keychain).
2. Place a `~/.claude/.credentials.json` with
`claudeAiOauth.accessToken` set and `claudeAiOauth.expiresAt` in the
past.
3. Open an agent → Harness / Runtime → Connect account → Claude
(Subscription) → Connect.
4. Before this change: the "Could not verify the local subscription"
error appears. After: the connection is created.

### Agent adapter(s) involved

Claude Code (`@paperclipai/adapter-claude-local`)

### Operating system

macOS (Keychain-backed credentials). On Linux the file is the live
store; an expired file token now returns null instead of a failing
request, so the user-facing message is unchanged.

## What Changed

- `packages/adapters/claude-local/src/server/quota.ts`:
`parseClaudeCredential` now returns the token plus `expiresAt` (epoch
ms) when the file records one. `readClaudeTokenFromFile` returns `null`
for a token whose `expiresAt` is in the past, so `readClaudeToken` moves
on to the next candidate (second file name, then Keychain when
`allowKeychain` is set). Files without an `expiresAt` keep the old
behavior. `parseClaudeCredentialToken` (used for the Keychain payload)
is unchanged in behavior.
- `packages/adapters/claude-local/src/server/quota-keychain.test.ts`:
three new cases — expired file falls through to Keychain; expired file
with no Keychain access returns `null`; a file with no expiry is still
accepted.

## Verification

- `pnpm --filter @paperclipai/adapter-claude-local exec vitest run
src/server/quota-keychain.test.ts` → 7 passed (4 existing + 3 new).
- `pnpm --filter @paperclipai/adapter-claude-local exec tsc --noEmit` →
clean.
- `pnpm --filter @paperclipai/server exec vitest run
src/__tests__/local-ai-credentials.test.ts` → 9 passed.
- Manual, on the affected machine: with the stale file in place, `POST
/api/companies/:id/ai-connections/local/check` (provider `anthropic`,
method `subscription`) returned `sign_in_required`; with the stale file
removed it returned `ready`. This change makes the first case behave
like the second without touching the file.

## Risks

- Low risk. The only behavior change is for a credentials file that
carries a numeric `expiresAt` in the past. Such a token is already
rejected upstream, so the change removes a guaranteed failure rather
than a working path.
- Clock skew: a machine clock that runs ahead of real time could treat a
token as expired slightly early. The fall-through then reads the
Keychain (macOS) or returns null, which triggers the same "sign in"
message the user already sees for an expired token.
- Keychain payloads are not expiry-checked in this PR. The CLI refreshes
that item itself, and `getQuotaWindows` already falls back to the CLI
`/usage` probe when the OAuth call fails.

## Model Used

- Claude — `claude-fable-5-1` (Claude Fable 5.1) via Claude Code, with
extended thinking and tool use (shell, file edit). Root cause found by
reproducing the server's read path against the local credential file and
Keychain.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [ ] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
v2026.916.0 beta/v2026.916.0-beta.0 nightly/v2026.916.0-nightly.0 canary/v2026.915.0-canary.11
2026-09-15 16:48:07 -07:00
Devin Foley bfb4ceabbb perf(release): wait for npm registry visibility of all packages concurrently (#13495)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The release system publishes ~33 public npm packages per release
across the canary, nightly, beta, and stable channels
> - `scripts/release.sh` publishes them strictly sequentially: publish
one package, poll npm until that version is registry-visible, then start
the next
> - npm accepts a publish in seconds, but registry packument propagation
can lag minutes per package (the CI budget was raised to 30 minutes per
package after two aborted releases), so the total publish time is the
sum of every package's lag — about two hours on a bad npm day, paid by
every channel run including every canary on every master push
> - This pull request keeps the publishes sequential but runs all the
registry visibility polls concurrently once every publish is accepted
> - The benefit is that the wall-clock cost of npm propagation drops
from the sum of all packages' lag to the single slowest package's lag,
with every existing safety property preserved

## Linked Issues or Issue Description

No existing issue. Description follows the enhancement template:

**What existing behavior does this improve?**
The npm publish step of `scripts/release.sh` (Step 5), used by every
release channel.

**Subsystem affected**
Release tooling (`scripts/release.sh`, `scripts/release-lib.sh`).

**Current behavior**
Packages publish one at a time, and after each publish the script polls
npm until that package's version is visible in the registry packument
before publishing the next. With per-package propagation lag of minutes
(observed up to ~15 minutes; per-package CI budget is 30 minutes), the
full 33-package set takes up to ~2 hours of mostly idle waiting.

**Proposed behavior**
Phase 1 publishes every package sequentially exactly as today (a
rejected publish still aborts the batch immediately with exact
attribution). Phase 2 then polls registry visibility for all packages
concurrently. Each package keeps its own `NPM_PUBLISH_VERIFY_ATTEMPTS` ×
`NPM_PUBLISH_VERIFY_DELAY_SECONDS` budget, and any version that never
becomes visible still hard-fails the release, now naming every
straggler.

**Reason and benefit**
Total publish wait becomes the slowest single package's lag instead of
the sum of all lags — typically minutes instead of hours. This shortens
every canary, nightly, beta, and stable run and reduces exposure to job
timeouts during npm slowdowns.

**Breaking changes**
None. Dry-run output is byte-identical in structure, dist-tags are still
applied at publish time (`--tag`), the Sigstore TLOG duplicate-recovery
path is untouched, and the later dist-tag integrity check
(`wait_for_release_registry_state`) is unchanged.

## What Changed

- `scripts/release-lib.sh`: replaced `publish_package_to_npm_and_wait`
with `wait_for_npm_package_versions`, which takes the package tuple list
and polls every package's visibility in background subshells, each
reusing the existing `wait_for_npm_package_version` poll (same
per-package budget), then reports per-package success or fails naming
all stragglers
- `scripts/release.sh` Step 5: the publish loop calls
`publish_package_to_npm` only (sequential, fail-fast on a rejected
publish), followed by one call to `wait_for_npm_package_versions` for
the whole set; Step 6's recap line updated to match
- `scripts/release-lib.test.mjs`: the registry-visibility and
workflow-budget tests now drive the new function (same assertions on
`npm view` counts, virtual sleeps, and the fail-closed message, which
now names the straggler); a new cross-visibility test proves concurrency
— two fake packages that each become visible only after the other has
been polled can only converge when polled in parallel, so the test fails
if the waits ever serialize again

Safety analysis for the ordering change: nothing in the publish loop
resolves sibling packages from the registry.
`prepare-bundled-package.mjs` bundles and patches from the local
workspace tree, and the TLOG duplicate-recovery path only queries the
package it just published. The only consumer of the "visible before next
publish" invariant was the release script's own final verification,
which still runs against the full set.

## Verification

- `node --test scripts/release-lib.test.mjs` — 15 tests pass, including
the new concurrency proof and the existing 15-minute-20-second budget
tolerance test against the new function
- `npm run test:release-registry` — full lane, 140 tests pass
- `bash -n` on both scripts; `shellcheck` reports no findings beyond the
three pre-existing ones on master (verified by comparing counts against
`origin/master` copies)
- A real-release exercise happens on the next master push: every canary
run executes this exact path

## Risks

- Low. The failure mode most worth watching is a release where some
packages become visible and others never do: previously the run stopped
at the first invisible package with later packages unpublished; now all
packages are accepted before visibility is enforced, and the run fails
naming every straggler. Recovery is identical in both worlds (the next
attempt derives a new version number), and the accepted-but-lagging
packages carry the correct dist-tag either way.
- Publish jobs run the source commit's copy of `release.sh`, so this
change takes effect for a given channel only once its source commit
includes this merge — promoted nightlies/betas cut from older commits
keep the old sequential behavior until their trains catch up.

## Model Used

Claude Fable 5 (`claude-fable-5`), extended thinking with tool use
(Claude Code).

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
canary/v2026.915.0-canary.10
2026-09-15 16:09:51 -07:00
DottaandPaperclip 669bd0e7a9 fix(ui): show progress and verify subscriptions during onboarding (#13499)
Automatically verify detected Claude and Codex subscriptions during initial onboarding. Show connection progress, support retries, and ignore stale results after navigation. Prefer the personal default subscription while preserving the later create-agent chooser.

Allow Enter to advance from the agent name field. Add regression tests and production-component Storybook coverage.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-15 17:39:25 -05:00
DottaandPaperclip 9adeabb590 fix(connections): unblock personal MCP auth discovery (#13497)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Connections let people give agents access to external tools.
> - A personal connection needs the current user's authorization.
> - A new MCP URL must be probed before Paperclip can discover its
sign-in method.
> - Requiring a personal grant before that probe prevents sign-in from
starting.
> - This pull request permits the initial probe for a creator-owned
draft with no credentials.
> - People can complete personal setup while later requests retain
authorization checks.

## Linked Issues or Issue Description

**What happened?**

Connecting an unknown MCP URL with "Just me" failed with HTTP 502 and
"This connection needs the current user's authorization". Paperclip
checked for a personal grant before contacting the provider. The health
wrapper also changed the expected authorization error into a server
error.

**Expected behavior**

Discover OAuth and start browser sign-in. Create a personal grant after
consent. For a public endpoint, discover its tools and create the empty
personal grant after a successful probe. Keep missing authorization on
later health checks as HTTP 422.

**Steps to reproduce**

1. Add an unknown remote MCP URL with no saved credentials.
2. Select "Just me".
3. Check the link. Before this fix, the request fails before sign-in or
tool discovery.

**Paperclip version or commit**

The three original regressions fail against `6cfe4acff` with the service
fix removed and pass with it restored.

**Deployment mode**

The defect was reported in production and reproduced in local server
tests with isolated PostgreSQL.

Related work: Refs #11831. Refs #11144. Searches found no duplicate fix.

## What Changed

- Allow an initial credential-free probe only for the creating user's
personal draft with unknown authentication and no supplied credentials.
- Leave OAuth grant creation to the callback. Create an empty personal
grant only after a public probe succeeds.
- Preserve the personal grant for URLs that already contain a
credential.
- Make empty personal grant creation conflict-safe without overwriting a
concurrent grant or duplicating its creation audit.
- Commit the empty grant and audit atomically. Retain the established
public draft identity after a later catalog failure, so a failed retry
cannot remove a successful retry's grant.
- Run the following catalog/default-profile step in its own transaction,
so failures discard partial catalog, profile, binding, and audit changes
without deleting the established identity.
- Verify archived personal connections retain their owner: another user
is rejected before probing, while the original owner can resume setup.
- Preserve `user_authorization_required` and HTTP 422 in health
failures.
- Cover the connect and OAuth callback routes, real loopback HTTP,
credential-bearing URLs, and later health checks.
- Document personal setup and the test fixtures.

## Verification

- Red/green: the original three tests failed with the exact reported
error before the fix and passed after it.
- The credential-bearing personal URL regression also failed before its
guard was added.
- Focused suite: `pnpm --filter @paperclipai/server exec vitest run
src/__tests__/generic-mcp-connection.test.ts
src/__tests__/tool-access-service.test.ts` passed all 385 tests across
the two suites. An earlier run had a socket hang-up in an existing
agent-permissions test; the unchanged suite passed on rerun.
- The concurrent rollback regression failed before its fix because the
successful retry's grant was deleted. It now verifies the grant and
draft survive and a later normal health check succeeds.
- Database fault injection during profile-entry insertion reproduced
partial catalog writes before the transaction fix. The regression now
verifies unchanged catalog rows, no partial profile/bindings, a retained
grant, and successful retry.
- CI's first serialized-server shard 3 attempt failed an existing
peer-agent mutation test (the real run-context guard ran despite the
test's mock). The test passed in isolation and all 108 tests in that
suite passed unchanged locally. The single failed-shard rerun passed
without code changes.
- Final-commit CI: all 32 applicable checks passed on
`639f037987352cab6084c4ebfa5dbf7b0aed6046`, including all 385 affected
tests, the full test matrix, browser suite, build, typecheck, release
checks, and security checks. The two Storybook-only checks were not
applicable and skipped. Greptile is 5/5 with all review threads
resolved. [Successful CI
run](https://github.com/paperclipai/paperclip/actions/runs/35027478353).
- `pnpm -r typecheck` passed.
- `pnpm smoke:mcp-fixtures -- --require-paperclip` passed.
- `pnpm build` passed.
- Full local `pnpm test:run` was attempted: its general-server group
finished with 12,372 passed, 4 failed, and 70 skipped tests. The run
started before review edits; its two MCP failures used the old cached
service (including an insert without the new conflict clause). All 385
focused tests pass on the final code. The other failures were existing
workspace-cleanup and runtime-port tests; their unchanged suites passed
on rerun (66 passed, and 25 passed/3 skipped). The local command stopped
before later groups. The final-commit CI matrix is the full-suite merge
gate; this local run is not claimed as green.

## Risks

- The initial probe must not become a general authorization bypass. It
is restricted to the creating user's draft. Normal health checks retain
authorization enforcement.
- Public endpoints get a personal grant with no secrets only after they
answer successfully. Credential-bearing URLs keep their existing grant.
- The concurrent-probe regression seeds the catalog and default profile
to isolate grant creation. Existing first-time catalog/profile creation
races are outside this change; this does not claim to make the entire
setup flow concurrency-safe.
- No database migration or UI change is required. OAuth tests use a
simulated provider; the public endpoint test uses real loopback HTTP.

## Model Used

OpenAI Codex, GPT-6-based assistant for regression tests and PR
preparation; a GPT-5-based Codex assistant assisted with the initial
implementation. Exact runtime model IDs and context-window sizes are not
exposed in this session. Both used reasoning, repository tools, and code
execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.915.0-canary.9
2026-09-15 17:11:59 -05:00
Devin FoleyandPaperclip 544c3476a8 feat(server): wrap a bare Cloud UI snippet body in a <script> element (#13496)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - A Cloud-managed instance injects an operator-owned HTML snippet
before `</body>` through `injectCloudUiSnippet`.
> - The Cloud control plane delivers that snippet to each instance as an
environment variable through a provider API.
> - The provider edge firewall now base64-decodes the request payload
and blocks any value whose decoded form contains a `<script` marker.
> - A working snippet needs a script tag, so every delivery is now
blocked and the operator cannot ship the snippet at all.
> - This pull request treats a resolved value that does not start with
`<` as a bare script body and wraps it in a `<script>` element at
injection time.
> - The benefit is that the operator can deliver a tag-free body that
the firewall passes, and the instance restores the script element on the
page.

## Linked Issues or Issue Description

No public issue exists. The problem is described below.

Related PRs (searched the PR list; none duplicate this change):
- Refs #13168 — added `injectCloudUiSnippet`, the mechanism this
extends.
- Refs #13245 — added the base64 `_B64` path on the assumption that
base64 clears provider WAFs. That assumption no longer holds; this PR is
the successor.
- Refs #13441 — the in-product feedback approach that the Cloud-owned
snippet replaced (closed).

**What happened?**

`injectCloudUiSnippet` injects `PAPERCLIP_CLOUD_UI_SNIPPET` (or the
base64 `_B64` form) verbatim before `</body>`. A working value must
therefore contain a `<script>` tag. The Cloud control plane delivers
this value as an environment variable through a provider API that sits
behind an edge firewall. The firewall now base64-decodes the payload and
rejects any value whose decoded form contains `<script`. The delivery
request fails, so the snippet cannot reach the instance.

**Expected behavior**

The operator can deliver the snippet through the provider API, and the
instance runs it.

**Steps to reproduce**

1. Build a snippet that contains a `<script>` element.
2. Deliver it to a Cloud instance through the provider variable API, in
plain or base64 form.
3. The firewall rejects the request. The instance never receives the
snippet.

## What Changed

- `injectCloudUiSnippet` now wraps a resolved value that does not start
with `<` in a `<script>...</script>` element. A value that already looks
like markup is injected byte-for-byte, so existing full-`<script>`
snippets are unchanged.
- Added two unit tests: one for the wrap path (plain and base64), one
that confirms `$&`-style replacement tokens in the body survive the
wrap.

## Verification

- `pnpm exec vitest run src/__tests__/cloud-ui-snippet.test.ts` — 17
pass.
- `pnpm exec vitest run src/__tests__/static-index-html.test.ts` — 2
pass.
- Confirmed the changed file has no type errors. The full `pnpm run
typecheck` needs the Rust runner toolchain (`cargo`), which is absent on
this machine; CI runs it in full.
- Verified out of band that a tag-free body clears the provider firewall
and wraps into valid, executable standalone JavaScript with no premature
`</script>` close.

## Risks

Low risk. The change adds a branch that only affects values that do not
start with `<` — previously injected as inert text, never as a running
script. Values that start with `<` keep their exact bytes. The content
is trusted operator HTML, consistent with the existing contract. Roll
back by reverting this commit.

## Model Used

Claude — `claude-fable-5` (Fable 5), extended thinking, with tool use
and code execution in Claude Code. A human author reviewed and verified
the change before submission.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [ ] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-15 15:09:19 -07:00
DottaandPaperclip 9cfa7fd2d1 fix(ui): share compact live and saved activity across runners (#13421)
## Thinking Path

> - Paperclip helps people manage AI agents and inspect their work.
> - Task threads show live activity and saved transcripts from several
runners.
> - Legacy runs used a separate activity renderer that expanded tool
cards as commands arrived.
> - A compact live row makes ongoing work easier to follow.
> - This PR shares the native runner activity group across live and
saved legacy turns.
> - Both paths now use the same labels, icon gutter, animation, and
history controls.

## Linked Issues or Issue Description

Related: Refs #13255, which introduced rolling native-runner activity
groups.

**What happened?**

During a legacy CLI run, each new command added an expanded tool card
under an activity heading. The original legacy parity story rendered a
completed turn, so it did not exercise this live path.

**Expected behavior**

Show one current activity line per commentary group. Roll that line
forward when a new activity starts. Keep tool icons aligned on the left,
use friendly labels, and show history only when expanded.

**Steps to reproduce**

1. Start a task with the Codex local adapter using the CLI engine.
2. Ask it to read two files in separate tool calls and run a test
command.
3. Watch the activity feed while the run is active, then expand its
history.

**Paperclip version or commit**

The live behavior was reproduced on
bdee5ebb21b2d09280e49c88bc329435da1b07c8. This update includes both live
and saved rendering fixes.

**Deployment mode**

Built from source with the local test-drive command and a real Codex CLI
agent.

## What Changed

- Use `TaskChatRunnerActivityGroup` for both saved legacy phases and
`TaskChatLiveTail`.
- Align the legacy Working spinner with the shared activity icon gutter.
- Cover streamed reasoning updates, successive commands, image labels,
hidden details, and explicit history expansion in tests.
- Feed raw legacy transcript events through the real adapter and live
renderer in Storybook. Add live, expanded, narrow, light, and completed
stories with the final reply preserved.

## Verification

- Passed 195 targeted tests covering the live tail, status pill, shared
activity group, native turn, and full task thread.
- Passed UI typechecking, token gates, UI build, and Storybook build
before submission.
- Observed a real Codex CLI run while it read separate files and ran
tests. The current activity stayed at one 32-pixel row without an
accumulated tool list; the spinner and activity icon centers aligned.
- Compared native and legacy Storybooks: identical rolling animation,
fixed height, persistent expansion, truncated long labels, and correct
light and completed states.
- Full repository `pnpm -r typecheck` and `pnpm build` passed after
replaying the branch on current master, including the Rust runner. The
full `pnpm test:run` suite is still running.

## Risks

- Legacy activity now starts collapsed. Users can expand each group and
each row to inspect the same transcript details.
- Live runtime request cards must retain their timeline positions.
Existing task-thread and native-runner regression tests cover this
boundary.
- No API, database, or transcript format changes.

## Model Used

- OpenAI GPT-6 through Codex for the live-path fix, tests, and browser
verification. Capabilities used: reasoning, tool use, local code
execution, and image inspection. The exact deployment ID and
context-window size are not exposed in this session.
- The original saved-turn change recorded OpenAI Codex with GPT-5.6; its
exact variant and context size were not recorded.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-15 16:40:32 -05:00
DottaandPaperclip d49f168381 fix: publish sandbox files on legacy and native runners (#13493)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents must publish generated files so users can inspect their
results after a sandbox stops.
> - Legacy sandbox bridges blocked attachment listing and could not
carry multipart binary uploads through the queue transport.
> - The native runner has a separate verified file registration path
that needs the same durable result.
> - This pull request repairs legacy binary transport, makes native
download receipts explicit, and reveals new outputs in the task
Artifacts tab.
> - Users can open generated files from either runner without a
transport flag change.

## Linked Issues or Issue Description

Refs #13355 for the existing native file publication path. Related
filename fixes: #2615 and #4788. Related sandbox persistence work:
#13376. This change repairs attachment delivery through the existing
API; it does not add workspace persistence.

**What happened?**

The upload helper first lists task attachments to avoid duplicates. Both
legacy bridge allowlists rejected that GET request with 403. A direct
multipart upload also failed: the queue bridge accepted only JSON,
excluded attachment uploads, and converted bytes to UTF-8 text. Enabling
HTTP/2 alone did not fix the missing listing route. These failures
occurred before attachment storage.

**Expected behavior**

Both runners can publish a workspace file, register its work product,
bind it to a response, and return a working download. The file stays
accessible after sandbox deletion. A new output opens the task Artifacts
tab. The agent receives accurate errors and decides how to retry or
report a failure.

**Steps to reproduce**

1. Run a legacy agent in Daytona with the duplex bridge disabled.
2. Invoke the bundled upload helper with Bash on a PNG or PDF.
3. Repeat with the duplex bridge enabled.
4. Register the same file through the native runner with generic API
tools disabled.
5. Retry registration, delete the sandbox, and compare the downloaded
bytes with the original file.

**Paperclip version or commit**

The failing baseline was `f2c5e54dc`. This branch is rebased onto
`6cfe4acff`.

**Deployment mode**

Source checkout with a local API and real isolated Daytona sandboxes.

## What Changed

- Allow authenticated attachment listing, upload, and content download
through both legacy bridge transports.
- Add optional base64 body encoding to queue envelopes. Preserve the
existing UTF-8 contract when the encoding field is absent. Decode binary
bodies before forwarding them.
- Preserve multipart headers. Bound raw bytes, encoded envelopes, and
in-flight reservations. Retain timeout and uncertain-write behavior.
- Preserve helper deduplication and return structured uncertain-write
failures. Document explicit Bash invocation in live skills.
- Add attachment IDs and content/download paths to native registration
receipts. Reuse verified local and remote file reads, attachment
storage, work-product registration, and response binding.
- Preserve Unicode upload filenames and provide a valid
Content-Disposition header.
- Open the task Artifacts tab when new stored outputs arrive, including
a closed desktop panel or mobile drawer. Deduplicate upload and
registration events by object ID. Preserve manual selection on
refetches, edits, and panel remounts.
- Remove task artifact filters, the company Artifacts footer link, and
the unassigned group heading and timestamp.

### Screenshot

![Generated images and a document in the task Artifacts
tab](https://pages.paperclip.ing/sandbox-file-delivery-2026-09-15/artifacts-tab.png)

This is the local display fixture. The image was generated separately
and published through the attachment and work-product APIs.

## Verification

- Post-rebase `pnpm -r typecheck` and `pnpm build` pass.
- The post-rebase local `pnpm test:run` passed 12,369 tests before one
existing conversation reset test timed out; all 33 tests in that suite
pass when rerun with isolated test configuration. The aggregate command
stopped before its remaining groups. GitHub runs the complete suite in
separate shards.
- All [GitHub verification
checks](https://github.com/paperclipai/paperclip/actions/runs/35017893350)
pass on `b66ac276dd3d5fc738a22ecea783400106a494d4`: 32 successful checks
and two configured skips. The native-session recovery assertion
initially raced its fire-and-forget Sentry report; all 13 tests pass
locally, and the same-commit CI rerun passes all 170 suites (3,079
tests).
- Live post-rebase Daytona: all three file-delivery tests pass. They
cover the real Bash helper with the queue bridge, the helper with
HTTP/2, and native `register_deliverable` with generic API tools
disabled.
- Daytona cases cover PNG/PDF bytes, spaced and Unicode names, duplicate
registration, response binding, authorization controls, and
byte-for-byte downloads after sandbox deletion.
- Local focused coverage includes transfer bounds, malformed encoding,
interrupted transfers, remote path containment, and native file
verification. The attachment route suite passes all 32 tests, including
an eight-case filename-header matrix for Unicode and special characters,
inline and forced downloads, and full and partial responses.
- Browser verification confirms image previews, persisted downloads,
automatic Artifacts selection, and preserved manual selection after
edits and reloads. Desktop/mobile component coverage passes. The latest
UI cleanup passes its 10 affected tests and token gates.
- Coverage limit: the Daytona tests call the real helper and native
registration path directly. They do not replay a complete model-led
image-generation task through the browser.

Live command (requires a configured Daytona credential):

```sh
PAPERCLIP_FILE_DELIVERY_DAYTONA=1 pnpm exec vitest run server/src/__tests__/file-delivery-bridges.test.ts
```

## Risks

- Binary queue bodies use more memory because base64 adds encoding
overhead. Transfer and process limits must remain aligned.
- An interrupted write can have an unknown result. The bridge reports
this state and preserves stable retry identities.
- New artifacts intentionally change the active task tab. Existing
history and repeated updates must not take focus again.
- Transport flag defaults, server authorization, frozen skill snapshots,
and completion policies remain unchanged. No schema migration is
required.

## Model Used

OpenAI Codex, based on GPT-6, with reasoning, repository tools, code
execution, and browser testing. The runtime does not expose a more
specific model ID or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.915.0-canary.8
2026-09-15 15:46:57 -05:00
DottaandPaperclip 6cfe4acff7 fix: preserve README images in npm package (#13488)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Its CLI is published as the paperclipai npm package, with the root
README shown on the package page
> - The root README uses repository-relative image paths so images
render correctly on GitHub
> - npm resolves those paths under the package repository directory,
which is cli, so the image requests point to missing cli/doc/assets
files
> - This pull request prepares the generated npm README by converting
only image src and srcset asset paths to stable raw GitHub URLs
> - The benefit is that the same source README remains correct on GitHub
and the published npm README displays its images

## Linked Issues or Issue Description

**Where is the issue?**

The issue is in the root README image assets and the npm packaging step
in scripts/build-npm.sh. The affected public page is
https://www.npmjs.com/package/paperclipai.

**What's wrong?**

The npm build copies README.md into cli/ before publishing. npm resolves
relative image paths beneath the package repository directory, so
doc/assets/banner.jpg becomes cli/doc/assets/banner.jpg. Those files do
not exist, and the images render as broken on npm.

**Suggested fix**

Keep the root README paths relative for GitHub. Rewrite
repository-relative image paths only in the generated npm README copy to
absolute raw.githubusercontent.com URLs.

## What Changed

- Added a small npm README preparation script that rewrites relative
image src and srcset asset paths.
- Updated scripts/build-npm.sh to use the preparation step when
generating the npm package README.
- Added a regression test for src, srcset, immutable refs, absolute
URLs, and non-image Markdown links.
- Pinned release-build image URLs to the source commit, while preserving
tarball builds by passing their known source refs.

## Verification

- Passed: node --test scripts/prepare-npm-readme.test.mjs
- Passed: bash -n scripts/build-npm.sh scripts/e2e-install-lifecycle.sh
scripts/e2e-update-migrations.sh
- Passed: focused README and E2E migration harness tests
- Passed: git diff origin/master...HEAD --check
- Generated README asset URLs were checked against raw GitHub and all
seven returned HTTP 200.

## Risks

Low risk. The change affects only the temporary README generated for npm
packaging. It does not change the GitHub README or runtime code. The
generated npm README depends on the public raw GitHub asset URLs
remaining available.

## Model Used

OpenAI Codex, GPT-5. Tool-enabled repository inspection, code execution,
browser verification, and git/GitHub operations were used.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with Fixes: # / Closes: #
/ Refs: # OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub #NNN / github.com/paperclipai/paperclip URLs)
- [x] My branch name describes the change (e.g. docs/... or fix/...) and
contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented risks above
- [ ] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-15 13:34:30 -05:00
Devin Foleyandgithub-actions[bot] 081006bee4 docs(release): curate stable notes for v2026.915.0 (#13487)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The release system publishes canary, nightly, beta, and stable
channels; a stable promotion publishes its release notes as the GitHub
Release body
> - At beta publish time the workflow auto-drafts a raw commit-log
skeleton on this branch (`releases/beta/v2026.915.0-beta.0.md`) for
humans to edit
> - The skeleton is a 1,800-line commit dump; stable `v2026.915.0`
cannot ship user-facing notes in that form
> - This pull request replaces the skeleton with the curated changelog
for the `v2026.915.0` promotion: overview, breaking changes, highlights,
fixes, upgrade guide, and contributor credits for the 483-commit range
since `v2026.831.1`
> - The benefit is that the stable promotion reads finished, accurate
notes from master and publishes them as the release body

## Linked Issues or Issue Description

Refs #13403, #13247, #13248, #13268, #13256, #13038, #13299 — the
headline features this changelog describes.

The notes-drafting flow itself (skeleton branch at beta publish, stable
promotion reading the file from master, post-ship canonicalization) is
the standing release process; this PR is the curation step it expects.

## What Changed

- Replaced the auto-drafted skeleton in
`releases/beta/v2026.915.0-beta.0.md` with the curated release notes for
stable `v2026.915.0` (promoted from beta `2026.915.0-beta.0`, source
`667c79ded`)
- Sections: overview, Breaking Changes (5), Highlights (Connections
train leads), Improvements, Fixes, Upgrade Guide (migrations
`0231`–`0279`, new env vars, removed API surface), Contributors

## Verification

- Every PR link and claim was checked against the actual commit range
`dbf052577..667c79ded` (the same range the skeleton header names)
- Migration list enumerated from `git log --diff-filter=A` over that
range; only `0231` and `0236` discard data, called out as such
- New environment variables verified in code (`server/src/config.ts`):
announcement flags (default on), split Sentry DSNs with legacy fallback,
token-broker allowed hosts, cwd `.env` opt-out
- Contributor list built from commit authors plus co-author trailers;
core maintainers excluded per release-notes convention
- Docs-only change: no code paths, no tests to add

## Risks

- Low risk: documentation only. The stable promotion reads this file
from master at dispatch time; if wording needs a follow-up after merge,
the promotion pins the master revision it resolved, so edits must land
before the stable dispatch
- The file is renamed to `releases/v2026.915.0.md` by the post-ship
canonicalization step, as with previous promotions

## Model Used

Claude Fable 5 (`claude-fable-5`), extended thinking with tool use
(Claude Code). The commit-range analysis and draft were produced by a
subagent on the same model and human-review-style checked against the
range before commit.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-09-15 11:31:44 -07:00
DottaandPaperclip f4cdc7b231 fix: recover transient workspace bootstrap scans (#13481)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The control plane prepares task workspaces before it starts an
agent.
> - Workspace preparation reads Git state so it can preserve edits and
exclude private files.
> - A failed scan was treated as a non-Git folder and lost its actual
failure code.
> - The resulting generic setup failure could not recover, even when the
cause was temporary.
> - This pull request keeps the cause and uses the existing bounded
retry schedule before provider startup.
> - Tasks can recover without human intervention, while permanent
failures and exhausted retries stop with useful guidance.

## Linked Issues or Issue Description

**What happened?**

A Git scan error during managed repository preparation became
`Configured repository folder is not a Git checkout`, followed by
generic `setup_failed`. The agent never started. Generic recovery could
not distinguish a temporary timeout from a bad workspace configuration.

**Expected behavior**

Keep the closed scan error code. Retry temporary timeouts and queue
saturation under the existing shared budget. Preserve edits, exclusions,
ownership, and pause gates. Stop permanent failures and exhausted
retries with a specific explanation. Do not replay historical generic
setup failures.

**Steps to reproduce**

1. Configure a task project with a local Git source that must be copied
into its managed repositories.
2. Make the ignored-file scan exceed its timeout before the agent
starts.
3. Before this fix, the snapshot returns null and the run ends as
non-retryable `setup_failed`.
4. Use the disposable browser fixture in
`tests/e2e/workspace-bootstrap/README.md` to inject real timeouts and
test the full recovery path.

**Paperclip version or commit**

Reproduced against `4510bf7c9e2fcbeb043445850928b5dcb79908ca`.

**Deployment mode**

Built from source. The defect is in core workspace setup, not a specific
model provider.

Related work: Refs #13442 (managed repository preparation), Refs #11572
(bounded Git scheduler), Refs #12997 (separate adapter startup retry
work), Refs #13469 (separate terminal-workspace scan performance work).

## What Changed

- Return the non-Git fallback only for repository discovery. Propagate
failed scans of a confirmed repository.
- Replace full ignored status output with an ignored-only directory
listing. Preserve NUL-delimited paths and exclusions.
- Preserve typed, sanitized scan errors through workspace preparation
and persist pre-provider failure details.
- Retry only timeouts and queue saturation, using the existing durable
two-retry budget and issue gates. Prevent generic recovery from adding
another budget.
- Show workspace-specific failure copy and actionable exhausted-recovery
notices.
- Add red-green unit tests, real-database restart and retry-boundary
tests, and opt-in browser acceptance fixtures with real Git subprocess
timeouts.
- Document the recovery contract and browser verification procedure.

## Verification

- Red: injected scan failures returned null instead of rejecting; setup
lost the timeout code; task-thread and recovery notices had generic
copy.
- Green: 119 focused adapter/backend tests, 20 recovery-boundary tests,
and 136 task-thread tests.
- `pnpm -r typecheck` — passed.
- `pnpm build` — passed on the final production code.
- `pnpm check:token-gates` — passed.
- The initial local `pnpm test:run` overlapped source edits and was
interrupted after two late-added assertions saw pre-fix behavior; it is
not counted as a green full run. A fresh final-head run passed all 229
tests across the six affected adapter/backend/UI suites. The clean
latest-head CI full test matrix passed: all five general-server shards,
all five serialized-server shards, and all three general-workspace
shards.
- Latest-head CI also passed all three browser shards and their
aggregate gate, typecheck and release registry, build, runner
verification, canary dry run, policy, Docker context integrity, and
security gates. Greptile: 5/5, with the review thread resolved.
- Browser: created a task in a disposable instance. A real Git timeout
scheduled recovery, the next run completed through the run-scoped API
without manual Retry, and Done survived reload. The deterministic
process worker checked preserved source edits and excluded private
files; no model calls were made.
- `WORKSPACE_BOOTSTRAP_TEST_URL=<disposable-instance-url> pnpm exec
playwright test --config
tests/e2e/workspace-bootstrap/playwright.config.ts` — 2 passed (3.6
minutes). The persistent case made exactly three failed attempts, never
started the worker, showed the cause-specific notice, stayed stopped for
another scheduler tick, and retained Blocked after reload.
- Extra red-green coverage: 50 recovery tests passed after fixing an
exhausted-bootstrap classification that incorrectly implied unknown
provider actions. Missing or uncertain evidence still retains the safety
hold.
- Verified the documented Git executable override during repository
seeding.

## Risks

- A confirmed repository scan failure now fails closed instead of
falling back to directory sync. This prevents unfiltered copying but
makes previously hidden errors visible.
- Temporary host problems can create up to two additional setup
attempts, 30 seconds apart. Permanent scan errors do not auto-retry.
Generic recovery cannot reset this budget.
- The durable retry path still enforces ownership, pause, and work
eligibility. Integration tests cover restart, duplicate promotion,
pause, exhaustion, and non-retryable categories.
- No schema migration, new runtime setting, new retry budget, production
deployment, or historical task replay.

## Model Used

OpenAI Codex, GPT-5-based coding agent, with reasoning, repository
tools, shell execution, and browser testing. The exact deployment model
ID and context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.915.0-canary.6
2026-09-15 12:37:21 -05:00
cceeb0aa66 test(runner): add everyday workflow evaluation harness (#13474)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The native runner must support project work, delegation, hiring, and
service access.
> - Browser tests exposed lost connection access, rejected helper
events, and stalled recovery.
> - Some eval failures also came from incorrect fixtures and decision
controls.
> - This pull request fixes those paths and adds eight everyday workflow
stories.
> - The tests retain observed failures and verify delivered files
independently.
> - The benefit is repeatable evidence for common user tasks and their
remaining gaps.

## Linked Issues or Issue Description

Related work: #13404 contains earlier workflow fixes. #13300 and #13470
changed the CI contracts used by the harness security tests. Merged
companion:
[paperclip-evals#22](https://github.com/paperclipai/paperclip-evals/pull/22).

**What happened?**

Native ACPX sessions did not receive the assigned connection gateway.
Codex helper events could arrive before their spawn receipt and fail
thread validation. A parent continuation could take a shared workspace
before its child retried. A failed native continuation could leave the
task status without a clear recovery blocker. The eval harness also
confused tool approvals with new connection requests and could reject a
valid delegated download.

**Expected behavior**

Keep assigned gateway access and its approval checks. Verify helper
lineage before accepting helper progress. Let a waiting child proceed
before automatic parent recovery. Preserve a failed task's recovery
ownership. Grade the actual requested workflow and its delivered files.

**Steps to reproduce**

Run the everyday workflow suite with the native Codex and Claude
profiles. Exercise service approval, connection refusal, delegated
project work, and teammate reuse. The commands and case requirements are
in `tests/runner-e2e/EVERYDAY-WORKFLOWS.md`. Use `pnpm
test:runner-recovery` for controlled crash and replacement cases.

## What Changed

- Pass the scoped connection gateway binding through the native ACPX
host and sidecar.
- Recognize Codex helper lineage from parent metadata and spawn
receipts. Verify early helper events with `thread/read`. Keep helper
events separate from root completion authority.
- Guide agents to use persistent hiring, child tasks, dependency
records, and a blocked handoff while waiting for a child.
- Defer automatic parent recovery while a child has an active execution
path in the same shared workspace. Allow parent recovery when the child
needs review.
- Record Blocked status and recovery evidence when a failed native
continuation needs reconciliation, including existing active or
escalated incidents. Preserve their owner and retry budget.
- Add eight browser-driven workflow cases. Use real decision controls,
explicit child feedback delivery, managed hiring credentials, and
independent ZIP checks inside a bounded Docker sandbox. Verify sandbox
availability before task creation. Record screenshot SHA-256 at capture.
- Keep runner crash probes in controlled recovery tests. Preserve the
original failure when cleanup also fails.
- Display missing accounting and replay revisions as unavailable. Align
harness security assertions with the approved CI changes.
- Make the channel-rejection browser fixture bind its file after the
send captures its payload. This prevents live refresh from removing the
file before the simulated race.

## Verification

- Full workspace `pnpm -r typecheck` passed after merging current
master.
- Runner E2E typecheck passed. Harness unit tests passed: 216/216.
- Wake-queue database tests passed: 55/55. The two added
existing-incident tests failed before the fix and pass after it.
- Docker artifact calibration passed: 12/12. Host-file and host-loopback
isolation tests failed before the fix and pass after it. Read-only
delivery and output limits are also verified.
- Full `pnpm build` passed. Targeted recovery tests passed: 83/83.
- The channel-rejection browser test passed five consecutive runs after
fixing the fixture race found in CI.
- Local general-server (12,351 tests), UI (6,250), CLI (485), and
workspace package groups passed. The monolithic run stopped at an
unchanged lock-heartbeat fixture race; the isolated workspace group
passed on rerun (shared: 747/747). A separate local serialized run
passed 97 files before two socket errors in the unchanged issue-list
route suite; that suite passed 15/15 on isolated rerun. These local full
commands did not finish uninterrupted; the complete CI matrix below
covers the remaining suites.
- Final head `0fb293733fe307be7e6667ae8f1364077d0c6455`: **34 successful
checks, 2 expected skips**, including every server/workspace shard,
browser shard, native runner verification, build, and typecheck. [Final
CI
run](https://github.com/paperclipai/paperclip/actions/runs/34989136700).
- Greptile reviewed this exact head at **5/5**; all review threads are
resolved. Both Superagent security checks are successful.
- ACPX credential-boundary tests passed: 118/118. Superagent accepted
the runner/sidecar versus provider-environment trace and cleared its
finding.
- The latest paid local campaign on source
`f6a2fdf7ac2af859826a2ae627ff4125a5478529` passed 22/24 cases: Sol 8/8,
Claude 7/8, Mini 7/8. These results predate the merge with current
master.
- The two remaining failures are in `hire-reuse`: Claude exceeded the
attempt deadline during final review; Mini made invalid deliverable tool
calls and remained Blocked.
- Six Daytona cases were not run because the matching immutable runner
image was unavailable. This PR does not claim new remote model results.

## Risks

The changes affect connection admission, helper identity, and recovery
scheduling. Assigned gateway grants and user approval still govern
service calls. The workspace admission gate still exists; the broader
folder-sync design is separate work. Provider behavior can still cause
the two recorded hiring failures. No database migration is required.
Paid cases are opt-in and have bounded attempt deadlines. Project
stories now require Docker and the documented pinned Python image on the
harness host.

## Model Used

OpenAI `gpt-6-astra` performed implementation, diagnosis, and
substantive review. OpenAI `gpt-5.6-luna` assisted with verification, PR
preparation, and review tracking. Both used repository tools and code
execution. Context-window sizes were not recorded.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused checks and
isolated reruns; full-run limitations are documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: OpenAI GPT-5.6 Luna <noreply@openai.com>
Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-15 11:04:16 -05:00
DottaandPaperclip 4510bf7c9e ci: use code-owner-reviewed master for trusted PR workflow (#13470)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Pull request CI uses a trusted workflow on the AWS runner fleet.
> - The caller used a fixed SHA that also needed runner-group admission.
> - A mainline pin update left CI queued because the group still allowed
older SHAs.
> - This pull request calls the trusted workflow on master, which
requires code-owner review.
> - New merged workflow versions can use the existing master
runner-group entry.

## Linked Issues or Issue Description

Related: #12968. That Dependabot PR proposes another SHA rotation. This
change keeps this first-party workflow on master instead.

**What happened?**

CI run 34975562974 stayed queued because its trusted workflow SHA was
absent from the runner-group allowlist. The fleet itself was healthy.

**Expected behavior**

New reviewed versions of the trusted workflow on master should receive
runner access without a separate SHA allowlist update.

**Steps to reproduce**

Change the caller to a new trusted workflow SHA without adding that SHA
to the restricted runner group. Its jobs remain queued. The master
reference removes that recurring synchronization step.

## What Changed

- Call `paperclipai/paperclip/.github/workflows/pr-trusted.yml@master`.
- Exclude this exact first-party workflow from Dependabot updates.
- Update the existing E2E shard workflow tests for the master caller
contract.
- Document the runner-group entry, required code-owner review, and
old-reference retention.

## Verification

- `actionlint .github/workflows/pr.yml` passed.
- `node --test scripts/__tests__/e2e-shard.test.mjs
.github/scripts/tests/cloud-runner-routing.test.mjs
.github/scripts/tests/pr-runner-rust-cache.test.mjs
.github/scripts/tests/pr-dependency-cache.test.mjs` passed: 35 tests.
- Parsed Dependabot YAML and checked the exact workflow exclusion.
- `git diff --check` passed.
- Live GitHub checks confirmed `.github/**` has code owners, CODEOWNERS
has no errors, and the active master ruleset requires code-owner review.
This covers the trusted workflow, caller, and CODEOWNERS itself.
- The approved organization setting now allows
`paperclipai/paperclip/.github/workflows/pr-trusted.yml@refs/heads/master`.
All previous references and other runner-group settings remain intact.
- Full application typecheck, tests, and build were not repeated locally
for this workflow-only change. This PR's CI and review are pending.

## Risks

New versions of the trusted workflow take effect for new callers after
merge to master. Keep code-owner review and master protection enabled.
Existing administrator pull-request bypasses remain unchanged.
Third-party action pins, runner routing, and infrastructure are
unchanged. Older callers still use their SHA pins; their allowed refs
remain in place.

## Model Used

OpenAI GPT-6 through Codex performed implementation and orchestration;
its exact runtime variant and context-window size were not exposed.
OpenAI `gpt-5.6-luna` with high reasoning inspected workflow assumptions
and applied the approved runner-group setting. Both used code and tool
access.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
canary/v2026.915.0-canary.5
2026-09-15 09:43:24 -05:00
Nicky Leach 5b913e7943 ci: activate restore-only Rust dependency cache (#13459)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Every change to Paperclip goes through the pull request CI workflow
before it merges
> - That workflow is split in two on purpose: `pr.yml` is the caller,
and it pins `pr-trusted.yml` at an immutable SHA
> - The pin means a change to `pr-trusted.yml` on master does nothing
until someone advances the pin
> - https://github.com/paperclipai/paperclip/pull/13457 added a
read-only Rust dependency cache to the `Verify Paperclip Runner` lane,
and it is inert for that reason
> - This pull request advances the pin, which is the second step of that
rollout
> - The benefit is that the saving measured in #13457 starts to apply,
about 3.9 minutes per run and about 4.7 compute-hours each day

## Linked Issues or Issue Description

This pull request is the activation half of a two-step rollout. #13457
merged on 2026-09-15, so this pull request now targets master directly.

- Refs https://github.com/paperclipai/paperclip/pull/13457 — added the
cache step this pull request activates. Merged as `f97a3f886`.
- Refs https://github.com/paperclipai/paperclip/pull/13302 — the
previous activation, and the change that introduced the `# Pin:` comment
convention this pull request follows
- Refs https://github.com/paperclipai/paperclip/pull/13300 — the change
#13302 activated, and the current pin target

I searched this repository for other pull requests that move this pin.
One is open:

- Refs https://github.com/paperclipai/paperclip/pull/12968 — an
automated bump of the same pin. See Risks.

**What existing behavior does this improve?**

The pull request CI lane still recompiles the full Rust dependency tree
on every run, because the cache step added in #13457 is not yet part of
the active CI definition.

**Subsystem affected**

Cross-cutting (multiple of the above). The change touches CI workflow
configuration only. It does not change product code.

**Current behavior**

`.github/workflows/pr.yml` pins `pr-trusted.yml` at `44dde2de`, the
squashed commit of #13300. GitHub reads `pr.yml` from the pull request
and takes every job from `pr-trusted.yml` at that SHA. A change to
`pr-trusted.yml` on master therefore has no effect on any pull request
until the pin advances.

#13457 is the only change to `pr-trusted.yml` since that pin, and it is
currently inert.

**Proposed behavior**

Advance the pin to the commit that carries the cache step, and update
the `# Pin:` comment to name the pull request it activates.

**Reason and benefit**

The saving measured in #13457 begins to apply. Master's own warm-cache
lanes run the same checks in 3.4 minutes against 7.1 minutes cold. The
net saving is about 3.9 minutes per run after the 20 second restore,
across about 73 runs each day.

**Breaking changes**

None. The activated change only adds a cache restore. A cache miss
reproduces today's behavior exactly.

**Additional context**

The last five activations all landed on the same day as the change they
activated: #13302, #12860, #12810, #12509, and #12464. This pull request
follows that convention. #13457 merged today.

## What Changed

- Advanced the `uses:` pin in `.github/workflows/pr.yml` from `44dde2de`
(#13300) to `f97a3f886`, the squashed merge of #13457.
- Updated the `# Pin:` comment to name #13457 and the capability it
activates, matching the convention #13302 introduced.

The diff is the same two lines every previous activation changed.

## Verification

Run the workflow and pin tests:

```bash
node --test ./scripts/__tests__/e2e-shard.test.mjs ./scripts/__tests__/run-vitest-stable-shard.test.mjs ./scripts/__tests__/release-verify-workflow.test.mjs ./scripts/cloud-source-verification.test.mjs '.github/scripts/tests/*.test.mjs'
```

Result: 469 pass, 0 fail. This includes `pr.yml calls the trusted PR
workflow at an immutable SHA`, which reads the pinned workflow out of
git and asserts on its content.

Confirm the pin resolves to a workflow that contains the cache step:

```bash
git show $(grep -oE '[0-9a-f]{40}' .github/workflows/pr.yml):.github/workflows/pr-trusted.yml | grep -c "Restore Runner Rust dependencies (read only)"
```

This prints `1`.

This pull request also verifies itself. GitHub uses the pull request's
own `pr.yml` for `pull_request` events, so this run takes its jobs from
the newly pinned workflow. The `Verify Paperclip Runner` job in this run
is therefore the cached version, running the exact definition this pull
request makes active. Check its log for `Cache restored from key:
v0-rust-release-runner-v1-Linux-x64-...`, confirm cargo prints no
`Compiling` lines for third-party crates, and compare the job duration
against the 15.0 minute baseline recorded in #13457.

## Risks

- **An automated pin bump is open and may race this.** #12968 moves the
same pin. It is a no-op today, because `pr-trusted.yml` is identical
between the two SHAs. If it rebases after #13457 lands, its target moves
to a commit that contains the cache step, and merging it would activate
the change with a stale `# Pin:` comment that still names #13300.
Closing #12968 before merging this avoids the ambiguity.
- **The activated change itself is low risk.** It only adds a read-only
cache restore. A miss reproduces today's behavior. #13457 records the
full risk list.
- **The rollback is one commit.** Restoring the previous pin value
returns CI to the current definition without touching `pr-trusted.yml`.

## Model Used

Claude Opus 5, provider Anthropic, exact model ID `claude-opus-5`, 1M
context window. Adaptive thinking was on. I used tool use throughout:
`git` to confirm the merge strategy, the pin history, and the squashed
merge SHA, the `gh` CLI and the GitHub API to read the repository merge
settings and to find the open automated bump, and local `node --test`
runs to verify the pin resolves and the guard tests pass. Run through
Claude Code.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Note on the unchecked boxes. The CI and Greptile boxes stay unchecked
until those checks finish on this pull request. On tests: the existing
pin guard in `scripts/__tests__/e2e-shard.test.mjs` already covers this
change, so this pull request adds no new test. On documentation: no
document describes the pull request CI pin, so there is nothing to
update.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
canary/v2026.915.0-canary.4
2026-09-15 01:19:14 -07:00
Nicky Leach f97a3f886e perf(ci): restore master's Rust dependency cache on the PR runner lane (#13457)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Every change to Paperclip goes through the pull request CI workflow
before it merges
> - One lane of that workflow, `Verify Paperclip Runner`, builds and
tests the native Rust runner
> - That lane compiles all 313 third-party crates from scratch on every
pull request, in two profiles
> - Master already builds and stores exactly those compiled crates, but
the pull request lane never reads them
> - This pull request restores that existing cache on the pull request
lane, read-only
> - The benefit is about 3.9 minutes less compute per run, or about 4.7
compute-hours each day, with no change to what CI checks

## Linked Issues or Issue Description

No public GitHub issue exists for this. The problem follows the
enhancement issue template below.

Related merged pull requests, found by searching this repository:

- Refs https://github.com/paperclipai/paperclip/pull/13194 — created the
`release-runner-v1` cache that this pull request reads
- Refs https://github.com/paperclipai/paperclip/pull/13300 — established
the restore-only dependency cache pattern that this pull request follows
- Refs https://github.com/paperclipai/paperclip/pull/13326 — split the
post-merge runner job into the `protocol` and `rust` lanes used as the
warm-cache baseline below
- Refs https://github.com/paperclipai/paperclip/pull/13259 — applied the
same caching idea to the post-merge typecheck job

I found no open pull request that duplicates this work.

**What existing behavior does this improve?**

The `Verify Paperclip Runner` job in `.github/workflows/pr-trusted.yml`
recompiles the full Rust dependency tree on every pull request.

**Subsystem affected**

Cross-cutting (multiple of the above). The change touches CI workflow
configuration only. It does not change product code.

**Current behavior**

The job has no Rust cache. Its only cache step restores the pnpm store.
Each run therefore downloads about 280 crates and compiles all 313
third-party crates twice, once in the `dev` profile and once in the
`release` profile.

Measured over 12 successful runs on 2026-09-15, the job takes 15.0
minutes on average. The range is 12.1 to 17.4 minutes.

| Phase | Mean | Share |
|---|---|---|
| `check:eval-kernel` | 0.0m | — |
| `typecheck:typescript` and protocol manifest | 0.1m | 1% |
| `build:rust` — cargo `dev` profile | 1.8m | 12% |
| TypeScript tests (`node --test` and vitest, 143 files) | 5.8m | 39% |
| `check:replay-goldens` | 0.1m | 1% |
| `typecheck:rust` — `cargo fmt` and `cargo check` | 0.9m | 6% |
| `test:rust` compile — cargo `release` profile | 4.9m | 33% |
| `test:rust` run | 0.8m | 5% |
| Parity checks | 0.0m | — |
| `check:api-authority` | 0.5m | 3% |

Cargo reports the two compiles directly. The `dev` profile takes 1m24s
to 1m53s. The `release` profile takes 3m54s to 5m36s.

**Proposed behavior**

The job restores master's existing Rust dependency cache before it runs
the checks.

Master already writes this cache. The post-merge `rust` lane in
`.github/workflows/release-verify.yml` writes `release-runner-v1`. It
also warms both profiles into that entry. The entry is 677MB and lives
on `refs/heads/master`. Pull request branches are allowed to read caches
from the default branch.

The pull request lane restores that entry read-only. Master stays the
only writer. A pull request never saves a branch-scoped copy. A pull
request never evicts the shared entry. This matches the rule the pnpm
store in the same file already follows.

**Reason and benefit**

Master runs these same checks with a warm cache. That gives a direct
measurement of the saving.

| Check set | Cold (pull request today) | Warm (master) | Change |
|---|---|---|---|
| `check:runner` and `check:api-authority` | 7.1m | 3.4m | −3.7m |
| `check:eval-kernel` and `check:protocol` | 7.8m | 7.3m | −0.5m |
| Cache restore step | — | 20–21s | +0.35m |

The net saving is about 3.9 minutes per run, or about 26%. About 73 runs
execute this job each day. The daily saving is therefore about 4.7
compute-hours.

The protocol side changes very little. Vitest dominates that side, not
compilation.

The cache should almost always hit.
`packages/paperclip-runner/runner/Cargo.lock` changed in 7 of the last
1697 commits on master. A miss costs nothing more than today's behavior.

**Breaking changes**

None. The change adds two steps to one CI job. It does not change any
check, any test, or any product code.

**Additional context**

No documentation covers pull request CI caching, so this pull request
updates no documents.

## What Changed

- Added a `Select the pinned Runner Rust toolchain` step to the
`verify_paperclip_runner` job in `.github/workflows/pr-trusted.yml`. The
step exports `RUSTUP_TOOLCHAIN`. `rust-cache` hashes `rustc -vV` into
the cache key, so the key needs the pinned compiler. This lane had no
rustup step before, so `rustc` resolved to each runner image's default
instead of the pinned 1.97.1.
- Made that step tolerate a missing `rustup`. `release-verify.yml` runs
on one post-merge fleet image. The gate in this workflow routes to
either `ubuntu-latest` or the public pull request fleet. A missing
`rustup` now costs the cache. It does not fail the pull request.
- Added a `Restore Runner Rust dependencies (read only)` step that uses
`Swatinem/rust-cache` with `save-if: false`.
- Mirrored every cache key input from the master writer in
`release-verify.yml`: the same action SHA, `workspaces`, `shared-key`,
`cache-workspace-crates`, and `cache-bin`. Neither side sets
`prefix-key`. Any drift causes a silent miss and a full recompile.
- Added `.github/scripts/tests/pr-runner-rust-cache.test.mjs`. It
asserts key-input parity across the two workflow files, the step order,
and the restore-only contract.

## Verification

Run the workflow shape tests:

```bash
node --test '.github/scripts/tests/*.test.mjs'
```

Result: 413 pass, 0 fail.

I also mutation-tested the new assertions. I applied each mutation to
the workflow, ran the new test file, then restored the file. All 7
mutations fail the suite:

| Mutation | Result |
|---|---|
| `shared-key` changed to `pr-runner-v1` | Caught |
| `save-if` changed to `true` | Caught |
| `cache-bin` changed to `true` | Caught |
| `rustup show active-toolchain` line deleted | Caught |
| `RUSTUP_TOOLCHAIN` export line deleted | Caught |
| `working-directory` line deleted | Caught |
| `command -v rustup` guard deleted | Caught |

To confirm the cache key inputs match the master writer, parse both
workflows and compare:

```bash
ruby -ryaml -e 'pr=YAML.safe_load(File.read(".github/workflows/pr-trusted.yml"), aliases: true); rv=YAML.safe_load(File.read(".github/workflows/release-verify.yml"), aliases: true); a=pr["jobs"]["verify_paperclip_runner"]["steps"].find{|s| s["uses"].to_s.include?("rust-cache")}; b=rv["jobs"]["verify_paperclip_runner"]["steps"].find{|s| s["uses"].to_s.include?("rust-cache")}; puts a["uses"]==b["uses"]; %w[workspaces shared-key cache-workspace-crates cache-bin].each{|k| puts "#{k}: #{a["with"][k]==b["with"][k]}"}'
```

Every line prints `true`.

After the pin advances (see Risks), confirm the cache works in CI. The
restore step log must show `Cache restored from key:
v0-rust-release-runner-v1-Linux-x64-...`. Cargo must stop printing
`Compiling` lines for third-party crates. The job should drop from about
15.0 minutes to about 11.0 minutes.

## Risks

Low risk overall. A cache miss produces exactly today's behavior, so the
worst case is no improvement.

- **This change does nothing until a second pull request lands.**
`.github/workflows/pr.yml` pins this reusable workflow by SHA. The file
header describes this two-step rollout. A follow-up pull request must
bump that pin. That follow-up validates itself, because GitHub uses the
pull request's own `pr.yml` for `pull_request` events.
- **A key mismatch would silently remove the benefit.** The runner
images here may ship a different default `rustc` than the post-merge
fleet. The toolchain step pins the compiler to prevent this. The new
test guards the remaining key inputs. If the first runs still miss,
compare `rustup show` output against a master run.
- **A missing `rustup` degrades quietly.** The step prints a GitHub
notice and continues. CI stays green and the run is simply uncached.
- **No cache poisoning path.** Pull requests only read. `save-if: false`
stops any write. GitHub also isolates pull request cache writes from the
default branch.
- **The cache entry can expire.** GitHub evicts unused entries after 7
days and enforces a repository size limit. Master pushes are frequent,
so the entry should stay warm. Eviction only causes a miss.

## Model Used

Claude Opus 5, provider Anthropic, exact model ID `claude-opus-5`, 1M
context window. Adaptive thinking was on. I used tool use throughout:
the `gh` CLI to read 12 job logs and step timings from recent CI runs,
the GitHub Actions cache API to read cache entry keys and sizes, `git
log` to measure `Cargo.lock` churn, and local `node --test` runs to
verify and mutation-test the change. Run through Claude Code.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Note on the unchecked boxes. The branch name
`claude/paperclip-runner-conditional-81a300` carries a tool-generated
suffix, so it does not meet the branch naming rule. I can rename it if
you want. The CI and Greptile boxes stay unchecked until those checks
finish on this pull request.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-09-15 01:16:34 -07:00
Devin FoleyandPaperclip 4cc387f907 fix(ci): remove npm propagation from cloud readiness (#13456)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Paperclip Cloud needs a verified image and matching database
migrator before it can deploy a merge.
> - New npm package versions can take minutes to become downloadable
after the package build finishes.
> - The direct producer now publishes signed archives and a complete
dependency lockfile for each master commit.
> - This pull request makes readiness verify those artifacts and removes
the duplicate automatic npm migrator run.
> - Deployment still requires all source checks, exact image identity,
migration compatibility, and pinned dependencies.

## Linked Issues or Issue Description

Refs: #13455, #13454, #13192

**What existing behavior does this improve?**
The time from a master merge to the `Cloud deployable v1` signal.

**Current behavior**
Readiness polls npm metadata for the new DB and shared versions. An
automatic dispatcher also starts a separate npm-only migrator workflow.
A measured source built its packages at 06:22:41 UTC on 2026-09-15, but
both npm archives were not downloadable until 06:31:56 UTC.

**Proposed behavior**
Wait for the successful exact-source direct producer, verify its signed
manifest and all pinned downloads, and publish readiness only after the
existing source and image jobs pass. Keep manual npm migrators and
branch previews available.

**Reason and benefit**
Remove new-version npm propagation from merge-to-deployable time. The
gain depends on whether image building or source verification finishes
later; it is not a fixed subtraction from every run.

## What Changed

- Require a successful producer from the canonical repository, exact
commit, master ref, expected workflow, and approved event.
- Verify the manifest's GitHub attestation with the hosted GitHub CLI.
Enforce the exact source SHA, master workflow identity, and hosted
runner.
- Download and validate both archives and the complete dependency
lockfile after publication succeeds. Reject invalid signatures,
inaccessible objects, corrupt bytes, and source mismatches.
- Remove automatic npm-only migrator dispatch. Retain manual release and
branch-preview publication.
- Document the cloud feature-switch prerequisite and coordinated
rollback.

## Verification

- `node --test .github/scripts/tests/*.test.mjs`: 405 pass.
- Focused readiness, routing, preview, and artifact tests: 249 pass.
- Workflow lint and `git diff --check`: pass.
- `pnpm test:release-registry`: 139 pass after installing this
worktree's dependencies.
- All latest-head GitHub CI checks passed. Greptile is 5/5 with no
unresolved comments.
- Application source is unchanged. Common-source local typecheck and
build passed. The full local application suite has the documented macOS
read-only-directory rename limitation from #13454 (13 failures in two
unchanged suites); Linux CI is the final application gate.
- Live readiness verification of master
da77a0c28c passed in 6.52 seconds,
including the real GitHub CLI signature policy and all artifact
downloads.
- Cloud consumer resolution with the certificate encoding fix passed in
5.45 seconds with zero npm metadata requests or npm processes. The
consumer is deployed and enabled in staging and production; their live
resolution APIs passed in 2.34 and 2.38 seconds. Both report the
expected fixed harness commit. A fresh tenant deployment follows this
cutover merge.

## Risks

- `Cloud deployable v1` no longer promises npm preview availability.
Enable the cloud direct-artifact consumer in staging and production
before merging this change.
- Artifact storage and GitHub attestations become required services for
new direct releases. Missing or invalid evidence fails explicitly.
- Restore the old npm dispatcher and readiness gate together before
disabling the consumer switch. Retain artifacts referenced by existing
releases.
- This change does not expand AWS runner access. The producer and
readiness bookkeeping use GitHub-hosted runners. Existing PR allowlists
and source verification gates remain enforced.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, and code
execution. The exact serving model ID and context window are not exposed
by this environment.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused checks; full
application host limitation documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-15 01:14:43 -07:00