Commit Graph
2000 Commits
Author SHA1 Message Date
nguyenm7 af7f03e42e fix(apps): correct connector logo identity and geometry 2026-09-30 17:21:20 -07:00
DottaandPaperclip c8f874311c fix(ui): hide Google connectors only on the Connections page (#14774)
## Thinking Path

> - Paperclip helps people manage AI agents for work.
> - The Connections page lists the apps and saved accounts that agents
can use.
> - Google Workspace verification is still pending.
> - Google entries must be temporarily hidden from this page without
removing their implementations.
> - This PR filters the final page rows, including saved Google
accounts, after the page resolves their provider.
> - Definitions, direct setup routes, OAuth profiles, credentials, and
runtime access stay intact.
> - Review instances can keep the prior UI by staying on their pinned
app release.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

Temporary provider visibility on the Connections landing page.

**Current behavior**

The page can show Google Workspace catalog entries and saved accounts
while verification is pending.

**Proposed behavior**

Hide all nine Google Workspace rows on this page. Keep every other
connector and all Google integration code unchanged. Use an existing
release pin for review instances instead of a hostname exception in the
app.

**Reason and benefit**

Pause public discovery without disabling existing runtime tools or
removing the implementation needed for verification and later
re-enablement.

**Breaking changes**

Google accounts are no longer visible on this landing page. Direct setup
and management routes remain available. This is not an access-control
restriction.

Related completed work: #13551 used catalog-level visibility. This
change is deliberately limited to the landing page and also covers saved
account rows. #14740 reduced Google scopes; this change leaves those
scopes unchanged. No duplicate open PR or matching open issue was found.

## What Changed

- Derive the Google app slugs from the existing Workspace profile
registry.
- Filter the combined catalog and saved-account rows only inside
`Browse`.
- Cover all nine Google entries, active/draft/disabled accounts, legacy
connection metadata, mixed-provider rows, and independently identified
non-Google connectors in regression tests.
- Document the display-only hold, pinned review builds, and how to
restore visibility after approval.

## Verification

- Passed: `pnpm exec vitest run ui/src/pages/apps/Browse.test.tsx
ui/src/pages/apps/AppsConnect.test.tsx` (199 tests, including the latest
master changes).
- Passed: `pnpm check:token-gates`.
- Passed: `pnpm build`.
- Passed: `pnpm -r typecheck` and `pnpm build` after merging the latest
master. An earlier overlapping run hit a local runner codesign race;
sequential checks passed.
- Passed again after the final custom-provider fix: `pnpm --filter
@paperclipai/ui typecheck` and `pnpm --filter @paperclipai/ui build`.
- The full local `pnpm test:run` was started, then stopped after the
full remote CI suite passed to avoid continuing duplicate long-running
work on the developer machine. It is not claimed as a completed local
pass.
- All 54 latest-head CI checks passed. Two non-applicable Storybook jobs
were skipped. One serialized server job lost its self-hosted runner
connection; its single retry passed.
- Greptile: 5/5 on `aeda167bf4494feed6ee0de2585960511fb02918`, with no
unresolved review threads.
- Confirmed in the existing review instance that all nine Google entries
still appear after its current release was pinned. No new app release
was deployed to that instance.
- Reviewer steps: open Connections on this branch with Google catalog
entries and saved Google accounts. None should appear. Non-Google
connectors must remain. Direct Google setup routes must still load.

## Risks

- Existing Google accounts cannot be found on this page during the hold.
Their data and runtime access remain unchanged.
- This is a UI-only filter, not an authorization gate. Direct routes and
API access still work by design.
- Review instances must not receive this UI build until the hold is
removed. Their existing release pin excludes fleet app upgrades; an
explicit targeted upgrade must still be avoided.
- No migrations, backend changes, broker changes, or credential changes.

## Model Used

OpenAI Codex (GPT-5-based coding agent), with reasoning, tool use, code
execution, and browser inspection. The exact deployment model ID and
context window are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 18:22:25 -05:00
d6fa1fd1ef feat(ui): streamline account menu profile access and add Invite shortcut (#14480)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Humans work in teams, so Paperclip has a multi-user system with
logins, profiles, and a company sidebar.
> - The account menu in the lower-left corner of the sidebar is the main
entry point for a user's own settings.
> - The account menu spent half of its rows on "View profile" and "Edit
profile". Users open these rows rarely.
> - The account menu had no fast path to invite a new member. The invite
flow is different on a self-hosted instance and on Paperclip Cloud.
> - This pull request removes the two profile rows, makes the header a
link to the profile, adds an "Edit profile" button on the profile page,
and adds an "Invite" row.
> - The benefit is a shorter account menu that keeps profile access and
gives users with the invite permission a one-click path to invite
people.

## Linked Issues or Issue Description

No public GitHub issue exists for this change. The description below
follows the enhancement template.

Refs #14060. That earlier pull request holds the first two commits and
the first Greptile review. It closed when the branch got a new name to
remove an internal ticket id. All Greptile findings from both reviews
are fixed in this branch.

**What existing behavior does this improve?**

The account menu in the sidebar (`SidebarAccountMenu` and its
`.production` variant) and the user profile page at `/u/:userSlug`.

**Subsystem affected**

ui/ — React + Vite board UI

**Current behavior**

The account menu shows the user's picture and name at the top. Below
them, the menu shows "View profile" and "Edit profile" rows, then the
other rows. The header is not a link. The menu has no row to invite
people. On a self-hosted instance, the user must open Company settings,
then Members, then the Invites tab. On Paperclip Cloud, the user must
open the Members page and use the Cloud People link there.

**Proposed behavior**

The account menu does not show "View profile" or "Edit profile". The
picture and name at the top of the menu are a link to the user's own
profile page. The profile page shows an "Edit profile" button when the
viewer looks at their own profile. The account menu shows an "Invite"
row with the same `UserPlus` icon as the company menu. On a self-hosted
instance, the row opens the Members page on the Invites tab, and shows
only to boards that hold the `users:invite` grant (company owners and
admins, instance admins, and local boards). On Paperclip Cloud, the row
opens the People settings for the current stack, and shows only to the
owner or admin of the stack, the same rule the Members page uses.

**Reason and benefit**

Users open their profile rarely, but the two rows took half of the menu.
Inviting people is a common task, but it needed three clicks and a
different path on Cloud. The new menu is shorter, keeps profile access
in one tap on the header, and gives one "Invite" entry point on both
hosting modes to the users who can invite.

**Breaking changes**

None. Routes, API responses, and settings keys do not change. The
profile page and the invite pages keep their current URLs.

## What Changed

- `SidebarAccountMenu.tsx` and `SidebarAccountMenu.production.tsx`:
remove the "View profile" and "Edit profile" rows. Make the picture and
name header a link to the user's profile. Add the "Invite" row after
"Settings" in the streamlined menu and first in the production menu.
- Header structure: `master` added a staging commit SHA link under the
email in the same header. The profile link is now a stretched overlay
behind the header content, so the SHA anchor sits beside the email and
the header has no nested anchors.
- `ui/src/lib/userProfileLinks.ts` (new): build the profile path from
the user id first. The profile endpoint treats the id as the one unique
slug, so two members with the same display name get different links.
Name and email are a fallback only when the session has no id.
- `ui/src/hooks/useCloudInviteUrl.ts` (new): read the Cloud stack
portfolio and return the People settings URL only when the current stack
role is owner or admin. This is the same rule as the Members page.
- `ui/src/hooks/useCompanyInviteAccess.ts` (new): read the current board
access snapshot and report whether the board may invite people to the
selected company on a self-hosted instance. Local boards and instance
admins pass. Other boards need an active owner or admin membership, the
roles that carry `users:invite`. This follows the same client-side gate
pattern as `ToolsAdminGate` and the run ledger. The server stays
authoritative.
- Invite row visibility: the row is hidden when the operator hides
`company.members` or `company.invites`, and until the health check
resolves. On self-hosted instances, the row is hidden until the board
access snapshot loads and when the board cannot invite. On Cloud, the
row is hidden when no People URL can be built. The in-app Invites tab is
never a fallback on Cloud, because it drives a different flow.
- `ui/src/pages/UserProfile.tsx`: add an "Edit profile" button that
links to `/company/settings/instance/profile`. The button shows only on
the viewer's own profile and follows the `instance.profile`
hidden-settings gate.
- Tests: extend `SidebarAccountMenu.test.tsx`; add
`userProfileLinks.test.ts`, `UserProfile.test.tsx`, and
`useCompanyInviteAccess.test.ts`.
- No documentation references the removed menu rows, so no docs change
is needed.

## Verification

Run the focused tests from the `ui/` directory:

```bash
pnpm vitest run src/components/SidebarAccountMenu.test.tsx src/hooks/useCompanyInviteAccess.test.ts src/lib/userProfileLinks.test.ts src/pages/UserProfile.test.tsx
```

- 52 tests pass in these four files. They cover the header link by user
id, the removed rows, the header overlay with no nested anchors, the
self-hosted invite target, the self-hosted permission gate (owner,
admin, instance admin, and local board see the row; an operator does
not, on both menu variants), the Cloud invite target with no
`target="_blank"`, the menu order, the hidden-settings gate on both
variants, the Cloud role gate (a plain member sees no row), the
no-fallback rule when Cloud stack metadata is missing, the staging
commit SHA link from `master`, and the own-profile-only "Edit profile"
button.
- `tsc -b` in `ui/` reports no errors in the changed files. The only
errors are pre-existing `@paperclipai/plugin-sdk/ui` resolution errors
in `PluginOrganizationSwitcher.tsx` from an unbuilt workspace package.

Manual steps:

1. Sign in and open the account menu in the lower-left corner. Confirm
the menu has no "View profile" or "Edit profile" rows.
2. Click your picture or name at the top of the menu. Confirm your
profile page opens and shows an "Edit profile" button.
3. Open another user's profile. Confirm the page shows no "Edit profile"
button.
4. On a self-hosted instance, as a company owner or admin, click
"Invite". Confirm the Members page opens on the Invites tab. As an
operator or viewer, confirm the menu shows no "Invite" row.
5. On Paperclip Cloud, as a stack owner or admin, click "Invite".
Confirm the Cloud People settings page opens in the same tab. As a plain
member, confirm the menu shows no "Invite" row.

## Risks

- Low risk. The change is limited to the UI and touches ten files.
- Users who know the "View profile" and "Edit profile" rows must learn
the new header link. The header has hover and focus styles to show that
it is a link.
- On self-hosted instances, the "Invite" row depends on the board access
snapshot from `/cli-auth/me`, which other gates in the UI already use. A
member with a custom `users:invite` grant but an operator or viewer role
does not see the row. That member can still use the Members page. The
row is a shortcut, not the only path.
- On Paperclip Cloud, the "Invite" row depends on the stack portfolio
query. When that query fails or the role is unknown, the menu hides the
row instead of sending the user to the wrong flow.
- Both menu variants change together, so a behavior difference between
them is not expected.

## Model Used

- Provider: Anthropic. Model: Claude Fable 5.1 (`claude-fable-5-1`).
- Run through Claude Code on the Claude Agent SDK inside a Paperclip
`claude_local` agent, with extended thinking and tool use (file edits,
shell, tests, GitHub API).
- The model wrote the code, the tests, and this description. A human
reviewed the pull request and requested the review fixes.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Bender (Fable) <bender@paperclip.local>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: devinfoley <139239+devinfoley@users.noreply.github.com>
2026-09-30 16:18:59 -07:00
DottaandPaperclip 33f2b3a159 fix: separate GitHub tools and code review bot connections (#14750)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The Connectors catalog lets people give agents tools or connect
agents to conversations.
> - GitHub put these two uses behind one card and an extra choice.
> - People should choose the connection they need from the catalog.
> - This pull request keeps GitHub for tools and adds GitHub Code Review
Bot as a separate card.
> - Each card opens its setup directly. Both use the existing connection
code.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

GitHub connector discovery and setup.

**Current behavior**

With chat connectors enabled, GitHub opens a menu that asks whether to
use tools or create a bot. Saved tools and bots share the same catalog
entry.

**Proposed behavior**

GitHub opens tool account access. GitHub Code Review Bot opens agent
selection. Saved bots and drafts appear under the bot card.
Chat-disabled instances show only GitHub tools.

**Reason and benefit**

The catalog names the two uses and removes an extra setup choice. The
bot keeps the existing GitHub provider, credentials, endpoint IDs, setup
steps, and runtime.

**Additional context**

Related work: https://github.com/paperclipai/paperclip/pull/12843 and
https://github.com/paperclipai/paperclip/pull/14594 established GitHub
account identity. This change preserves that tool flow. No duplicate
catalog split was found.

## What Changed

- Split the generated app definitions into GitHub tools and GitHub Code
Review Bot. Reuse the existing GitHub logo and channel method.
- Open bot setup directly, including old resume and reconnect links.
- Put existing bot endpoints and drafts under the bot card. Hide
duplicate internal chat applications.
- Keep pasted GitHub URLs mapped to the tool connection.
- Add seven Storybook states for the catalog, saved connections,
disabled chat, both setup paths, mobile, and light mode.
- Fix narrow-screen bot rows so the label cannot overlap status and
setup actions.
- Update catalog, route, browser, and API tests, plus the GitHub
connector guide.

## Verification

- [Hosted
Storybook](https://d1p6rlowie26tp.cloudfront.net/storybook/branches/codex~2Fgithub-review-connection/?path=/story/connections-github-and-code-review-bot--catalog):
seven states built from this branch. The deployment passed its
public-file verification.
- All GitHub checks pass on `d13a2cd53561645bb2a15c6f8e75a61a936d6459`.
Two optional Storybook jobs skip under their normal trigger rules; the
manual Storybook deployment passes. The branch has no merge conflicts.
- Greptile: 5/5 on the current head, with no review comments or
unresolved threads.
- `pnpm -r typecheck`, `pnpm build`, `pnpm check:token-gates`, and `pnpm
build-storybook` passed. The final Storybook fixture also passed UI
typecheck and the hosted build.
- Targeted catalog, URL matching, routing, grouping, brand, and chat UI
contract tests passed.
- GitHub provider browser tests: 2 passed. These cover direct tool setup
and the bot setup and management lifecycle with provider responses
mocked.
- Embedded-browser test on an isolated local instance: opened both
cards, selected an agent, saved a bot draft, and resumed the same
endpoint under the bot card after a reload.
- Storybook Tool Setup and Bot Setup assertions pass in the published
preview. Chat Disabled assertions pass locally. Inspected mobile and
light mode, including the draft-row layout and official GitHub marks.
- Local full-suite limitation: `pnpm test:run` was not clean. A
cross-company route assertion failed in the aggregate run and passed in
isolation; a workspace-runtime test reached its 30-second hook timeout.
Some isolated database reruns skipped when the embedded-PostgreSQL
availability probe failed. The local aggregate was stopped after CI
completed. The corresponding full CI suites pass all 360 tool-access
tests and all 162 workspace-runtime tests.
- No live GitHub authorization or installation was performed. The
isolated instance correctly stopped at the cloud enrollment or public
HTTPS prerequisites.

## Risks

- Low scope: catalog presentation and routing change. There is no
database migration or provider credential change.
- Existing GitHub bot URLs now open bot setup directly. The tool route
remains `/apps/connect?source=github`.
- The bot remains behind the existing chat-connectors feature flag.
Existing endpoints retain `provider: github`.
- Channel applications are represented by endpoint rows. Regression
tests cover legacy bot applications, tools, active bots, and drafts
together.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, code execution, and
embedded-browser tools. The exact deployed model ID, context window
size, and reasoning setting are not exposed to this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 17:15:02 -05:00
DottaandPaperclip 018993140f feat: let agents name prompt-only tasks (#14761)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Users create tasks with a title and a description.
> - A required title adds work when the prompt already explains the
request.
> - An agent can name the task once it reads that request.
> - This pull request accepts prompt-only tasks and starts them with a
short prompt slice.
> - A scoped title tool lets the assigned agent replace that slice early
without changing execution state.
> - A live browser eval checks the real agent call, saved title, audit
entry, and preservation of user titles.

## Linked Issues or Issue Description

**Subsystem affected**

Cross-cutting: task creation, shared contracts, database, server, runner
tools, and board UI.

**Problem or motivation**

Users must currently write a title before they can submit a detailed
task prompt. The agent has enough context to write a useful title
itself.

**Proposed solution**

Make the title optional when a description is present. Save the first
120 characters of the normalized prompt as a provisional title. Ask the
assigned agent to call `set_task_title` early. Use an atomic
provisional-title guard to preserve titles supplied or edited by users.
Keep explicit titles supported.

Related: #14543 and #14556 concern empty-title submission. This change
intentionally enables that submission when a prompt is present, instead
of requiring a title.

## What Changed

- Add the `titleNeedsGeneration` field with an idempotent migration.
Keep existing titles unchanged.
- Add `PUT /api/issues/:id/title` and the native and legacy
`set_task_title` tool. Enforce company access, active-run ownership,
shared, bounded retry receipts across native/HTTP calls, and
transactional audit logging. Refresh external-object links after commit,
with the same feature gate and plugin detectors as ordinary title edits.
- Add early naming guidance in Standard, Ask, and Plan task context.
Preserve the description, status, and assignment.
- Allow prompt-only root and child task creation, plus draft restoration
in the New Task dialog. Keep user titles supported.
- Add an opt-in Product E2E suite for prompt-only Standard and Ask
tasks, plus an explicit-title control. It checks actual provider calls
within the first five tools, persisted state, audit attribution, and the
reloaded UI.
- Preserve a closed vocabulary of API key maintenance phrases in
declared prose while rejecting opaque credential suffixes. Add one
bounded naming retry after wording is rejected, without treating the
rejected call as a saved title.
- Repair the native cleanup receipt check exposed during full
verification: accept matching input digests, retain legacy input checks,
and reject conflicting receipts.

## Verification

- Live Product E2E on `f43478473800e3a46b85c5ee79677efdb15108e7`: **3/3
passed** with native Codex `gpt-5.4-mini`, first attempts only,
automatic retries disabled. Standard and Ask each saved “Rotate expired
API key” on their first tool call, with matching persisted state and a
single same-run audit entry. The explicit-title control retained its
user title with zero title writes. All three verified the reloaded
browser UI.
- Campaign: `local-2026-09-30T21-30-11-021Z`. Earlier failed campaigns
are retained separately; they exposed credential-prose handling and
prompted the naming recovery fix. No failed result was regraded or
deleted.
- Reproduce with `pnpm test:e2e:runner -- --id
task-titles.runner-codex-mini.local.prompt-title-standard --id
task-titles.runner-codex-mini.local.prompt-title-ask --id
task-titles.runner-codex-mini.local.preserve-explicit-title
--max-automatic-retries 0` and an authorized provider key.
- Full `pnpm -r typecheck` and `pnpm build` passed on the latest commit.
The runner build used the configured external eval source tree.
- Product E2E unit suite: **61 files, 818 tests passed**; E2E typecheck
and UI token gates passed.
- Title API/native regressions cover prompt-only and explicit child
creation, user edits, ownership/company isolation, external reference
refresh, cross-surface retry replay, and the 64-key limit without
receipt eviction. All passed. Prompt-context coverage: **44 tests
passed**.
- Rust credential regressions: **35 tests passed**, including benign
maintenance qualifiers and opaque credential rejection in every declared
prose field. Catalog/report reconciliation: **28 tests passed**. Native
recovery: **560 tests passed**.
- Broad local `pnpm test:run`: **14,555 tests passed** in the general
server group; two suites failed to initialize embedded PostgreSQL and
the existing 40,000-file Git streaming stress test exceeded its
300-second macOS timeout. All three suites then passed in isolation (**5
tests passed**) without code or timeout changes. The original full local
command exited nonzero and is not being represented as a clean full run.
- Latest-head GitHub checks are green: **53 passed, 4 skipped, zero
failed or pending**, including all test shards and the canary packaging
dry run. Greptile reviewed the same commit at **5/5**, with zero
unresolved review threads.

## Risks

- The additive database field must reach the server and UI together. The
migration uses `IF NOT EXISTS` and defaults existing tasks to a final
title.
- Title generation depends on the assigned agent running. Tasks without
a run keep their provisional title.
- Live qualification covers the native Codex path in Standard and Ask
modes. API/legacy and Plan behavior have deterministic coverage.
- The credential-prose exception validates the entire suffix against a
closed maintenance vocabulary. Unknown suffixes, assignments, quoted
values, credential prefixes, and diagnostics retain strict checks.

## Model Used

OpenAI Codex, based on GPT-6, with reasoning, tool use, and code
execution. The exact deployment ID and context window are not exposed in
this session. The live eval uses the native Codex `gpt-5.4-mini`
profile.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 16:48:24 -05:00
DottaandPaperclip ad55d0a281 fix(connections): repair personal credentials and request write access (#14739)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents use Apps through a gateway that checks identity, company
access, and action policies.
> - Personal pasted credentials can point to company secrets. Setup can
show success while the gateway rejects every call.
> - Several OAuth methods also omit the scopes needed for their
supported write actions.
> - This pull request gives setup, health checks, and invocation the
same credential rules. Owners repair existing connections by
reconnecting.
> - New connections request reviewed permissions for their supported
actions. Read-only choices remain available under Advanced.
> - Agents can use the connections people give them, while existing
consent, identity boundaries, and action restrictions remain enforced.

## Linked Issues or Issue Description

Refs #14009 and #14008. This addresses the personal-credential defect.
The separate GitHub organization-identity selection defect is outside
this change.

Related work: #13942 fixed part of new personal-key setup. #14200
independently fixes legacy personal reconnect and protects managed-agent
profile credentials during removal. This PR covers that ownership
invariant across key and secret-URL setup, reconnect, health, discovery,
and invocation, and keeps owner reconnect as the repair path. #14059
tracks requested versus provider-asserted OAuth scopes; it remains
separate work. I searched open PRs and issues for Zapier, Airtable
scopes, connector writes, and personal credential failures.

**What happened?**

A Zapier secret URL saved through personal setup can become a company
secret referenced by a user grant. Health checks bypass the gateway's
ownership check, so the connection appears healthy but calls fail with
`grant_credential_invalid`. Custom-header paths can also receive a
duplicate `credentials.` prefix. Omitted OAuth scopes make write access
depend on provider defaults.

**Expected behavior**

Personal invocation credentials belong to the selected user. Setup,
health, and actual calls enforce the same rule. New connections request
documented permissions for supported read and write actions. Existing
tokens gain no permissions without provider consent.

**Steps to reproduce**

1. Connect Zapier or a generic secret URL with the personal identity.
2. Allow an agent to use the connection and complete setup.
3. Invoke a tool through a run-scoped gateway. The legacy layout fails
ownership validation despite successful setup.

**Paperclip version or commit**

The implementation started from
`44736c9c7c67b7b646ead9d51721db10f5b83835` and was rebased onto master
at `94e8dec56`.

**Deployment mode**

Built from source. Regression tests use isolated PostgreSQL fixtures and
controlled MCP transports.

## What Changed

- Share credential writing, ownership validation, and canonical paths
across initial setup, resume, reconnect, rotation, health, discovery,
and gateway calls. Keep OAuth client-registration secrets separate from
invocation credentials.
- Existing personal connections with company-scoped credentials require
owner reconnect with a fresh key or secret URL. Reconnect creates a
correctly owned value and updates the existing grant and declarations.
There is no automatic ownership backfill or new startup hook.
- Preserve PostgreSQL timestamp precision when reconnect checks whether
a grant changed. Previously, converting the timestamp to a JavaScript
Date could reject reconnect with a false concurrent-change error.
- Protect credentials used by other grants, connections, bindings,
managed-agent profiles, routine triggers, or secret proposals from
connection removal.
- Review all 117 tool methods, including 84 OAuth methods. Record
explicit scopes or documented provider-default exceptions with official
evidence. Add Airtable's seven scopes, Hugging Face repository/job
scopes, and other documented MCP permissions.
- Prefer available write-capable methods. Put explicit read-only choices
under Advanced. Explain pasted-key permissions and offer reconnect for
missing OAuth consent. Preserve existing grants, policies, Google
availability gates, and curated scope allowlists.
- Reconnect generic secret URLs and custom headers using their stored
credential fields. Refresh the catalog after setup, correct reconnect
feedback and error guidance, and let Cancel exit invalid setup while
Save & exit retains draft-saving behavior.
- Apply ownership checks to the new GitHub repository/skill connection
picker. Align the permission audit with the Google scope reductions
merged on master.
- Add run-scoped gateway, ownership, owner-reconnect, OAuth URL,
insufficient-scope, UI, and catalog-wide regression coverage. Update the
connector playbook and permission audit.

## Verification

Latest commit `97bc0b86e0eae0ec892e4ac44beff1a66164b20e` passes all
CI/status gates (55 completed check runs, no failures or pending checks)
and has a completed Greptile review at **5/5 with no outstanding
findings**. GitHub reports the PR as mergeable/CLEAN.

- **Embedded browser:** used the actual server and built UI from this
worktree, a fresh isolated database, and local HTTP MCP fixtures.
Completed personal bearer-key, secret-URL, and custom-header setup;
reproduced the legacy ownership failure; reconnected through the owner’s
form; and completed writes afterward. Read-back was verified for
bearer-key and secret-URL connections. Public organization-wide setup
appeared immediately in Browse without reload. Zapier URL
validation/Cancel and Google’s enrollment gate were also exercised.
- **Persistence and invocation:** verified user ownership, canonical
`credentials.authorization` / `remote.url` / `headers.X-Api-Key`
declarations, and unchanged connection/grant identity. The old company
secrets retain their ownership. Separate HTTP calls through an actual
run-scoped gateway session completed a write and read-back.
- **Backend coverage:** the final gateway suite passes all 82 cases,
including catalog Zapier and generic inline reconnect. It checks
company/user isolation, canonical declarations, same-endpoint URL
validation, fresh credentials, retained restrictions, and real gateway
read/write execution using fixture transport. A timestamp with
PostgreSQL microseconds covers the former false reconnect conflict.
- **Local checks:** 368 catalog, gateway, repository, and UI tests
passed before the final extra Zapier case; 49 GitHub skill access tests
also passed. All three Apps browser regressions pass, including
reconnect through the actual form and catalog visibility without reload.
Full `pnpm -r typecheck`, `pnpm build`, server typecheck after the final
patch, and token gates passed. Full tool-access service runs hit varying
15-second Google fixture timeouts; both affected cases and the updated
reconnect assertion pass in isolation (3 tests). The complete test
matrix passes in CI on this head.
- **Verification limits:** no live provider account was available for
Zapier/Airtable/OAuth consent or account-bound write proof. Public
metadata and local fixtures do not establish provider consent. The
original development database clone failed on a pre-existing missing
`tool_connections_transport_check` constraint; browser acceptance used a
fresh isolated database created by the normal CLI onboarding flow.

## Risks

- Existing broken personal connections stay unusable until their owner
reconnects. Health, discovery, and invocation return an actionable
ownership error; startup does not rewrite credential ownership.
- Scope changes affect new authorization requests. Providers may still
require resource selection, account roles, paid plans, or app
verification. Existing consent and action restrictions remain unchanged.
- Shared credentials are retained rather than reassigned or revoked.
Provider-default exceptions and unavailable live checks are documented
in `doc/connections/CONNECTOR-PERMISSION-AUDIT.md`.
- No new endpoint, database table, lockfile change, or CI workflow
change is included.

## Model Used

OpenAI Codex, based on GPT-6, with reasoning, code editing, shell
execution, web research, and browser tools. The exact deployment model
ID and context window were not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 14:32:34 -05:00
DottaandPaperclip d432dc7fa3 Add GitHub-synced skill sources (#14713)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Company skills supply instructions and files to those agents.
> - GitHub imports already exist, but users cannot manage repositories
as skill sources.
> - Repository refresh also needs caller-authorized access and complete
local packages.
> - This pull request adds Sources inside Skills and reuses GitHub
connections from Apps.
> - Installed snapshots let agents use skills without fetching GitHub
during a run.
> - Manual refresh preserves skill identity and leaves failed imports on
their last good version.

## Linked Issues or Issue Description

**Subsystem affected**

Cross-cutting: skills UI, server, database, shared contracts, and
runtime materialization.

**Problem or motivation**

Users keep skills in GitHub repositories. They need a clear way to
select, import, and refresh those skills. Existing imports do not expose
repository management or consistently preserve supporting files.

**Proposed solution**

Add company-scoped skill sources. Browse repositories from all
accessible GitHub connections, or paste a public repository or branch
URL. Select whole skill packages, inspect included files and reference
warnings, and install complete, immutable snapshots. Refresh each source
manually.

**Alternatives considered**

Project repository settings hide the workflow from Skills. A second
GitHub connector would duplicate credentials and grants. Upstream
editing and PR creation are separate work.

**Roadmap alignment**

This implements the Skills Manager direction in ROADMAP.md. The
maintainer requested this scope and reviewed the component and full-app
journey stories before implementation.

Related reports: Refs #10285, Refs #10949, Refs #13464. Related work:
#14356, #13656, #9268.

## What Changed

- Add source and entry records, an idempotent migration, company-scoped
APIs, and legacy GitHub import adoption.
- Reuse current caller grants and credential refresh. Combine and
deduplicate repository inventories across accessible connections. Pasted
public URLs also prefer the active user’s authorized connections. Tokens
stay in the Git child environment, never argv or disk.
- Fetch a shallow Git snapshot at one immutable commit. Scan the full
local tree, including hidden and nested folders. Read Git objects
without checkout or archive transformations and enforce nested package
boundaries.
- Bound Git downloads to 128 MiB and three minutes. Cancel active
process groups and remove incomplete downloads. Preserve cancellation
and deadlines while progress drains; close stalled HTTP progress streams
after 30 seconds. Reuse caller-scoped temporary snapshots for
preview/import after reauthorization.
- Index repository package boundaries once and cap expanded work at
1,000 packages, 10,000 files, and 100 MiB, including repeated copies of
shared blobs. Bound path depth and the shared path index. Discovery
keeps audited manifests without retaining all package bodies.
- Resolve moving branches before fetching so unchanged discovery reuses
caller-scoped snapshots. Limit active scans, scan frequency, and new
downloads per caller and company; quotas apply before metadata reads and
across connections, and cached scans do not consume the download quota.
- Store complete versions with script content, binary bytes, and
executable modes. Preserve these through copies, runtime caches, and
runner packaging.
- Stage downloads before publication. Use source leases, revision
checks, and transactional activity records. Keep installed versions
after failures, upstream deletion, deselection, and disconnect.
- Add the approved import flow, Sources page, selection tree,
provenance, read-only Studio behavior, and saved return from GitHub
setup.
- Add package manifests, commit-pinned file previews, and separate
runtime requirements and reference warnings. Supporting files are
included together; nested skills remain independently selectable.
Preview requests reauthorize the caller and re-audit package content.
- Show installed skills as compact links beneath each source. Repository
titles open GitHub. Keep Refresh, Select skills, and Disconnect source
in a three-dot menu. Source rows omit the branch, imported count, and
refresh timestamp; action alignment and repository titles work at narrow
widths.
- Stream discovery metadata over an opt-in NDJSON response. Show
measured Git download progress and real package/file counts, animate
newly checked skills, support cancellation, and require a complete scan
before selection. Keep the existing JSON API.
- Retain component stories and add a separate full-app journey story
group. Include fixed progress states and interactive scan,
large-repository, interruption, and saving stories.
- Update Skills documentation and product contracts. Suppress private
GitHub skill references in telemetry. Privacy review requested for the
telemetry changes.

## Verification

- Local repository typecheck, full build, token gates, and Storybook
build passed during this work. Focused transport, authorization,
scanner, persistence, route, and UI tests pass. The final UI refinement
passes all eight focused UI tests, UI typecheck/build, and token gates.
The scanner resource and repeated-discovery fixes pass 132 focused
scanner, transport, authorization, source-service, route, and rate-limit
tests, plus server typecheck/build. Full-suite verification comes from
CI; the older full local Vitest run was stopped after unrelated chat
failures and a font-test failure, all of which passed in fresh focused
runs. At commit `1098d5996`, all 54 active checks pass; two optional
Storybook jobs are skipped. CI covers repository typecheck, build, the
full test suites, browser shards, and the canary dry run. Greptile is
5/5 with no open findings; the security scan also passes.
- Adversarial scanner tests verify repeated-blob byte accounting with
and without declared sizes, package/file/path caps, one-time repository
indexing, metadata-only discovery audits, and nested package boundaries.
Additional tests cover branch movement, snapshot reuse, caller/company
quotas, isolation across connections, active-lease cleanup, quota
recovery, and rejection before any metadata API call.
- Real Git tests verify hidden paths, exact binary bytes, executable
modes, export-ignore preservation, symlink/submodule reporting, pinned
commits, caller-scoped cache reuse, cancellation, cleanup, and
credential isolation. Regression tests hold both download slots with
permanently blocked progress callbacks, verify timeout/cancellation
cleanup and retry, and exercise HTTP backpressure cancellation. Access
tests cover automatic public-URL connection selection and revoked
grants. Database tests verify company and grant audiences.
- Live isolated browser test: the public `anthropics/skills` scan now
completes and discovers all 20 skills without connecting an account.
Imported canvas-design with all 83 files, opened it from Sources, and
verified the installed binary-font preview/download control. Package
previews also expose the complete file inventory before import.
Cancelled an active Git download and retried successfully to all 20
discovered skills; the browser displayed measured download progress. The
current audits reject four other packages; eligible selections remain
importable.
- Browser checks verify the simplified source rows at desktop and narrow
widths, keyboard navigation into the actions menu, Refresh from the
menu, selection, and fixture disconnect with installed skills retained.
Storybook includes a menu-open checkpoint and a 320px layout.
- Storybook includes receiving/preparing download checkpoints and a
timed full-app import journey, plus cancellation, retry,
large-repository, and saving states. Streaming tests cover split UTF-8
frames, incomplete streams, late responses, cross-company requests, HTTP
errors, and JSON compatibility.
- Earlier live acceptance on this PR imported `stitch-skill` with
`DESIGN.md`, assigned it to an agent, disconnected its source, and ran a
successful Studio test that read both installed files. An editable copy
changed independently. Both Skills variants, mobile selection, and
return from GitHub setup were exercised.
- Private access, revoked credentials, OAuth success return,
binary/script preservation, concurrent refresh, transaction rollback,
version pins, and legacy adoption have automated coverage. A real
private-repository OAuth grant was not created during this test.

## Risks

- The migration groups recognizable legacy imports without provider
calls. Their first successful refresh completes the local package
snapshot.
- Reference checks are advisory. They cover Markdown links and explicit
relative resource paths, not arbitrary runtime dependency graphs.
Preview text is capped at 64 KiB; imported bytes remain complete.
- Git must be installed on the server. Shallow fetches still download
the branch snapshot, including files outside selected packages.
Downloads have size/time/concurrency limits. Temporary caches are
bounded and caller-scoped. GitHub API quota still applies to repository
metadata and the connection picker; content no longer uses per-file API
requests. Failed scans retain installed content.
- Sources depend on the current caller's GitHub access. A saved
connection does not grant access to another person's token.
- GitHub script support and immediate manual refresh are explicit
maintainer-approved requirements. The operator trusts the selected
repository and accepts upstream script and executable-mode changes on
refresh. Static audits are not a sandbox or a guarantee of safe code;
agents may later invoke installed helpers under their runtime
permissions. Import and refresh do not execute scripts, hooks, package
installation, or builds. Raw URL and skills.sh imports keep their prior
script restrictions.
- Source originals remain read-only. Refresh affects subsequent unpinned
runs; explicit pins and active runs retain their versions.
- The telemetry change removes source-managed GitHub identifiers from
skill-reference events. It introduces no event or field. Please review
the privacy boundary.

## Model Used

OpenAI Codex, based on GPT-6, with reasoning, code execution, and
browser tools. The exact serving model ID and context-window size are
not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 13:32:58 -05:00
DottaandPaperclip 3c561642b4 fix(chat): resolve approvals and preserve unanswered questions (#14613)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents ask for decisions and optional details through cards in chat.
> - A clear approval in a message can leave the matching card pending.
> - An unanswered question can also block an unrelated later reply.
> - Decisions need a saved source message, while optional questions need
to remain answerable in history.
> - This pull request records conversational decisions and lets users
move on from questions and answer them later.

## Linked Issues or Issue Description

**What happened?**

Native Claude and Codex could act on approval in chat while the original
approval card stayed pending. Pending question forms stayed above the
composer, were absent from history, and could suppress later chat
replies. A late native question answer could wait for a finished run to
reconnect.

**Expected behavior**

The active agent records a clear approval or refusal against the exact
card and user message. Ambiguous replies do not grant consent. Users can
send another message without answering a question. The question remains
pending in history and can be reopened and answered later. The saved
answer reaches the agent.

**Steps to reproduce**

1. Ask an agent to propose work with a confirmation card, then approve
it in chat.
2. Check that the original card records that approval before work
starts.
3. Ask an interactive question, send an unrelated message, and reload.
4. Open the unanswered question from history and submit an answer.

Related work: #14408 added completion delivery. #14607 tests completion
reporting turns. Neither records conversational answers on approval
cards.

## What Changed

- Add a confirmation endpoint backed by a user comment, with schema
validation, OpenAPI discovery, and native Plan-mode access. Ask mode
remains read-only.
- Check company, active run, actor, current session, message provenance,
revision, and resolver policy. Save the decision and audit in one
transaction. Retries do not repeat effects. Emit resolution telemetry
after commit.
- Give fresh and resumed chat turns the actual pending confirmation
identities. Teach agents to save clear conversational decisions before
acting and to clarify ambiguity.
- Keep unanswered Agent Chat questions as compact history entries. A
newer user message closes the old form. Question cards never contribute
to composer pending counts or navigation, including after dismissing a
fresh form. The history card is the sole reminder; clicking it restores
that exact form and draft.
- Preserve Agent Chat questions when later messages or questions arrive.
Historical ordinary inputs no longer gate later chat replies.
Current-run requests, task execution, and governed approvals keep their
gates. Remove the special acknowledgement-publication proof helpers that
this rule replaces.
- Route answers to finished native runs through durable fresh-wake
delivery, with existing idempotency and source-question context. Settle
late replies against contiguous completed conversation turns and freeze
their history replay; failed, unhandled, and newly arriving messages
remain actionable.
- Add real-component Storybook scenarios, database and UI regressions,
and a three-turn native Claude/Codex E2E case. Capture distinct,
UI-ready screenshots and report the individual assertions.

## Verification

- Focused decision/publication/UI regressions after merging master: 288
passed; subsequent UI draft, failed-send, and conversation checks: 199
passed.
- Native question and durable delivery regressions: 106 passed,
including all four terminal run states and exactly-once late delivery.
Seven targeted regressions fail against the original implementation and
pass with the fix.
- Latest conversation/decision/native-delivery regressions after the
master merge: 121 passed. Covers completed progress, missing or failed
intervening turns, new messages during a late reply, stale sessions, and
frozen retry/replay boundaries. Four new assertions fail before the
ordering fix.
- E2E support suite after the master merge: 792 passed. Negative
controls reject expired cards, wrong questions/answers, stale or missing
replies, unrelated clarification forms, and unexpected tasks.
- The embedded-browser walkthrough caught one additional defect:
dismissing a fresh question still showed a composer badge. Both Cancel
and close-button regressions failed before the fix. The fix at
`65f2ade12` passes 170 chat-thread tests and 792 E2E support tests.
After merging master, 232 chat-thread/confirmation tests, server/UI
typechecks, and token gates pass. The preview and two-provider live E2E
pass at `e5512a206`; Greptile is 5/5 with zero unresolved threads at
that commit. All 55 checks are now successful at `e5512a206` (four
conditional checks skipped), including the aggregate verification gate
and clean-install canary test. The first attempt was interrupted by
simultaneous CI worker shutdowns; one failed-job rerun passed without
code changes.
- [Published
Storybook](https://d1p6rlowie26tp.cloudfront.net/storybook/branches/codex~2Fchat-approval-resolution/?path=/story/chat-comments-agent-chat-unanswered-questions--moved-on):
nine real-component scenarios. Manually exercised move on, reopen,
preserve draft, answer later, answer one of multiple questions, and a
custom mobile answer in the embedded browser. Retested fresh Cancel and
close-button dismissal in the updated build, then reopened and submitted
the preserved Green selection and inspected its answered receipt. Static
preview has no live model/backend; its callbacks are fixture responses.
- [First live
campaign](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36714504406-1/)
reproduced the late-answer completion-state defect on both providers
despite correct saved answers and acknowledgements. It also exposed a
valid imperative clarification rejected by the old oracle. Both issues
are fixed with regression controls; this failing run is retained as
evidence.
- [Four-cell
qualification](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36717804064-1/)
passed 4/4 at `2bf8a1009`: unanswered-question return and ambiguous
confirmation, each on native Claude and Codex. Inspected saved state,
source-message decisions, visible cards, and agent replies. Both
late-answer chats settled to waiting; no unrequested tasks were created.
[Final branch
rerun](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36719666238-1/)
passed 2/2 at `142630720`: the same unanswered-question journey after
merging master, plus an additional screenshot and browser assertion for
the actual late-answer acknowledgement.
- [Composer-reminder
E2E](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36727006818-1/)
passed 2/2 at `5b62c52d9`: native Claude and Codex, three turns each,
with explicit no-badge assertions before and after reload. Inspected
saved pending/answered state, both screenshots with a clear composer,
and actual Blue acknowledgements; all five behavioral matchers passed
per provider and neither created tasks. Cost coverage is partial; this
is bounded workflow qualification.
- [Fresh-dismissal
E2E](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36742773318-1/)
passed 2/2 at `e5512a206`: native Claude and Codex, including fresh
Cancel, clear composer, reopen, unrelated message, reload, late Blue
answer, and actual agent acknowledgement. All five behavioral matchers
pass per provider. Inspected the fresh-dismissal screenshots and saved
pending/answered identity; neither created tasks. Cost coverage is
partial (4/6 runs).
- Prior evidence remains available in [the earlier
campaign](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36642252725-1/).
Its early loading screenshot and overwritten final capture prompted the
UI-ready, distinct screenshot fixes.

## Risks

- The model interprets intent. The server verifies permission and
provenance; it does not infer consent from text. Ambiguous and unrelated
replies are not approvals.
- Historical questions can accumulate. They remain visible, pending, and
answerable; no automatic answer or expiry is invented.
- The change to completion gates is scoped to Agent Chat and ordinary
historical inputs. Current-turn and governed approvals retain their
existing controls.
- Live qualification is limited to the selected stories. Broader native
onboarding finalization remains separate work.
- No database migration. Telemetry adds no fields or values; the
contract and README document the commit boundary. Privacy review was
requested on the PR.

## Model Used

OpenAI Codex, GPT-6 family, with reasoning, repository tools, code
execution, and browser-test orchestration. The exact model ID and
context-window size are not exposed to this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 11:46:46 -05:00
DottaandPaperclip 0e5830887b perf: reveal task content sooner and parallelize issue reads (#14727)
## Thinking Path

> - Paperclip helps people manage AI agents and their work.
> - A task page must show saved replies quickly so a person can read the
work.
> - The title could appear while the conversation waited for unrelated
metadata and transcripts.
> - Thread requests also waited for enriched task details, while the
server read several independent fields in sequence.
> - This change shows saved content as soon as it is ready, starts
thread reads earlier, and runs independent server reads together.
> - Native event history takes priority over legacy log fallback, and
mentioned tasks load on intent.
> - Content-free timing spans make the remaining server delays visible
without recording task content.

## Linked Issues or Issue Description

**What happened?**

Task titles and properties appeared quickly, but saved conversation
content stayed hidden for several more seconds while metadata and run
transcripts loaded.

**Expected behavior**

Saved replies and the task description should be readable without
waiting for supporting history. Returning to a cached task should show
content within a frame or two.

**Steps to reproduce**

1. Open a task with saved comments and completed runs.
2. Delay the task activity and runs responses by five seconds in the
browser.
3. Observe whether saved content remains hidden until those responses
finish.
4. Navigate away and return to the task to check cached navigation.

**Paperclip version or commit**

The change was developed from `1b48e73e0` and rebased onto `44736c9c7`.

**Deployment mode**

Built from source, tested in an authenticated staging deployment and
with local response replay.

Related: #14667 overlaps the transcript reveal behavior and adds
separate retry UX. This PR also changes navigation prefetch, parent
metadata gates, native log fallback, server read scheduling, and timing
spans. #12647 proposes a separate SQL predicate optimization in the runs
service. #13597 and #13095 are earlier loading fixes.

## What Changed

- Start activity and runs when the task page mounts, alongside task
details and comments, using the route reference for shared query keys.
Hover/focus prefetch does not start full history reads.
- Reveal saved comments and descriptions while metadata and transcripts
load. Preserve strict waits for linked-comment navigation and tasks with
only runtime content.
- For settled native runs, fetch legacy logs only when event history is
empty or fails. Preserve live-log subscriptions for queued and running
native runs. Fetch mentioned-task details on hover or focus.
- Run independent issue-detail enrichment and run metadata reads in
parallel while preserving recovery dependencies.
- Add `Server-Timing` phases and opt-in OpenTelemetry spans, plus
regression tests and observability documentation.

## Verification

- **313 tests passed** across the initial seven focused component,
cache, timing, and scroll suites. After review fixes, **313 tests
passed** across five task-page, cache, prefetch, and live-transcript
suites (`pnpm exec vitest run` with `--maxWorkers=1`; these sets
overlap).
- `pnpm -r typecheck` and `pnpm build` passed locally before the final
UI-only review fixes. UI typecheck/build and `pnpm check:token-gates`
passed after those fixes. The final commit also passes full typecheck
and build in CI.
- The full local Vitest run was attempted. Several unrelated embedded
PostgreSQL fixtures failed to start, and parallel test workers hit
timeouts. Focused reruns passed. A later local full-suite rerun was
stopped after the complete CI suite passed; it is not claimed as a local
full-suite pass.
- Final commit `d1e147145`: **54 checks passed, 2 skipped**, including
all server/workspace test shards, all eight browser E2E shards, full
build/typecheck, Runner checks, release registry, and canary
clean-install verification. [CI
run](https://github.com/paperclipai/paperclip/actions/runs/36734642034).
Greptile **5/5** after two reviews; both findings fixed and all review
threads resolved. No merge conflicts.
- Live browser tests on an existing task with two saved replies: median
full reload to visible content fell from **1.92 s** (3 samples) to
**1.40 s** (5 samples). Cached return fell from **421 ms** (1 sample) to
**29 ms** (3 samples). These are observed samples, not a performance
guarantee.
- With activity and runs delayed by five seconds, saved content appeared
in **1.38 s** on desktop and **1.33 s** on mobile. The inspected comment
did not move when metadata arrived. Verified history expansion, task
properties, pending-input navigation, dashboard return, and mobile
layout.
- A separate local replay with fixed responses reduced visible-content
time from **5.12 s** to **2.15 s**. This isolates frontend behavior and
is not a live-server benchmark.

## Risks

Progressive history can change the thread after first paint. Existing
anchor behavior is retained and covered by tests and delayed-response
browser checks. Query aliases must stay aligned for invalidation.
Parallel reads can increase short bursts of database work; dependent
recovery operations remain ordered. Full reloads still depend on network
and task-detail latency.

No schema or authorization change. OpenTelemetry remains disabled
without an operator endpoint. The added spans use a closed set of phase
names and carry no task IDs, task content, or exception text.

I checked `ROADMAP.md`; this is a performance fix within the existing
task page.

## Model Used

OpenAI GPT-6 in Codex. The runtime does not expose a more specific model
ID or context-window size. The agent used reasoning, code editing,
terminal tools, and Chrome performance profiling.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 10:26:28 -05:00
DottaandPaperclip 44736c9c7c fix(ui): align runner commentary with chat replies (#14716)
## Thinking Path

> - Paperclip helps people manage AI agents and their work.
> - Agent chat shows run commentary and the agent's reply in one thread.
> - The runner commentary text starts 4px left of the following reply.
> - The runner activity group omits the gutter that the reply bubble
uses.
> - This pull request gives runner commentary the same token-based
gutter.
> - Both text blocks now share one left edge.

## Linked Issues or Issue Description

**What happened?**

In agent chat, the first commentary line of a completed run starts about
4px left of the following agent reply. This is visible in the Codie
conversation.

**Expected behavior**

Runner commentary and ordinary agent reply text should have the same
left edge.

**Steps to reproduce**

1. Open an agent chat with a completed run that has commentary and an
agent reply.
2. Compare the first commentary line with the following reply at the
left edge.
3. Inspect the wrappers. Before this change,
`task-chat-phase-interstitial` has zero left padding and
`task-chat-agent-bubble` has 4px.

## What Changed

- Add the existing `px-1` spacing token to runner commentary.
- Add a regression test that compares the runner commentary gutter with
the ordinary agent reply gutter.

## Verification

- `pnpm check:token-gates`
- `pnpm --filter @paperclipai/ui exec vitest run
src/components/task-chat/TaskChatRunnerActivityGroup.test.tsx
src/components/task-chat/TaskChatTurn.test.tsx
src/components/task-chat/TaskChatBubble.test.tsx`
- `pnpm --filter @paperclipai/ui typecheck`
- `pnpm --filter @paperclipai/ui build`
- A full repository typecheck and build passed earlier on this branch.
All latest-commit CI gates passed, including the server shard after a
transient preview readiness failure was rerun.
- Browser render: runner commentary and reply paragraphs both start at
x=32px after the change.

## Risks

Low risk. This changes only the horizontal padding of runner commentary.
Long commentary lines may wrap 8px sooner. The full local test suite was
stopped after the code changed during its run; the focused UI tests and
latest-commit CI suite passed.

> I checked `ROADMAP.md`. This bug fix does not duplicate planned core
work.

## Model Used

OpenAI GPT-6 in Codex. The runtime did not expose a more specific model
ID or context window size. The agent used reasoning, shell tools, and
read-only browser inspection to diagnose the layout and verify the
change.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 09:48:41 -05:00
DottaandPaperclip 1b48e73e0b feat(ui): add secondary navigation for agent chat (#14706)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agent Chat already provides a persistent conversation with each
agent.
> - Its shortcuts share the primary navigation and do not give chats a
dedicated place.
> - People need to find agents, start a chat, and switch conversations
without moving the page layout.
> - This pull request adds a secondary chat sidebar and a landing page
around the existing chat surface.
> - The same conversation, composer, history, and context panel remain
in use.

## Linked Issues or Issue Description

Refs #13283 and #13420. This extends the existing experimental Agent
Chat navigation after review of the component and page stories. It
supports the CEO Chat roadmap item through the existing task-backed
conversation model.

**Subsystem affected**

The board UI and the company-scoped conversation list API.

**Current behavior**

Chat shortcuts sit inside the primary navigation. There is no dedicated
landing page with a searchable conversation list. A separate landing
header also moves the sidebar when an agent is selected.

**Proposed behavior**

Show a Chat entry in primary navigation. Keep a searchable agent sidebar
beside the chat content. The plus button starts or reopens the current
user's single conversation with that agent. Keep the header and sidebar
in the same positions before and after selection.

**Reason and benefit**

People can find agents and return to persistent conversations without
leaving the chat area or creating duplicate chats.

**Breaking changes**

The experimental chat navigation changes. Explicitly adding a chat now
resolves its conversation immediately. Direct visits to unused agent
chat URLs remain read-only. The existing per-agent routes and message
contracts remain compatible. No database migration is required.

## What Changed

- Add an account- and company-scoped conversation list endpoint with the
existing access checks, feature gate, and OpenAPI entry.
- Add the live secondary sidebar, landing page, avatars, search, loading
states, errors, and retry controls.
- Make the agent picker wait for chat creation and display failures.
Existing agents reopen the same conversation. A dismissed selection
cannot close a reopened picker or navigate over a newer choice.
- Preserve recent-activity ordering and terminated agents’ chat history.
Scope live list refreshes to the current user’s conversation events. A
failed historical-agent lookup leaves healthy chats usable and offers a
focused retry.
- Keep the sidebar and header stable across chat routes. Keep mobile
selection in the navigation drawer.
- Use the production components in Storybook. Prepare the theme and
mobile viewport before mounting the page to avoid the startup flash.
- Update product documentation, the design guide, and navigation tests.
Replace old browser expectations for stars and recent shortcuts with
persistent conversation and layout coverage.

## Verification

- `pnpm -r typecheck`, `pnpm build`, and `pnpm check:token-gates` pass
after rebasing onto master.
- Focused UI tests pass, including 105 sidebar, picker, and live-update
checks after review fixes. The 33 conversation service and route tests
and 10 OpenAPI checks pass, including ownership, feature gating, and
concurrent creation.
- The full local test run passed 14,163 server tests before three
environment or timeout failures. The embedded Postgres startup,
connector socket, and native runner failures all passed direct reruns.
- Browser test-drive verification covers a real provider reply, add and
reopen, persisted history after reload, no-match search recovery, mobile
drawer dismissal, and top-aligned context panels.
- Browser measurements confirm that the sidebar has the same position
and dimensions on the landing page and an agent conversation.
- Storybook builds and its add-and-reopen interaction passes.
- The revised browser regression passes locally against a freshly built
throwaway instance. It covers stable sidebar geometry, add/reopen
uniqueness, drafts, search, history, and terminated-agent history after
reload. The full CI browser suite also passes.
- Latest commit `b323577d9523180104df4000eaceedea2772608c`: all 54
completed checks pass, including the complete server/workspace/browser
suites, aggregate verification, build/typecheck, security scans, and
canary packaging. The two Storybook jobs are skipped by their workflow
conditions. [CI
run](https://github.com/paperclipai/paperclip/actions/runs/36714052050).
- Greptile reviewed this same commit at 5/5 with no remaining actionable
findings; all review threads are resolved.
- Reviewer path: enable Agent Chat, click Chat, use plus to choose an
agent, send a message, switch away, and reopen that agent. One
conversation must remain, with its history intact.

## Risks

- The new sidebar lists persistent conversations instead of starred and
recent shortcuts.
- Chat creation is asynchronous. Errors stay visible in the picker, and
delayed responses cannot navigate into a previous company or account.
- The shell adjustment is limited to chat routes and preserves the
existing conversation implementation.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, code execution, and browser
tools. The session does not expose the exact API model ID or
context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#123` / `Refs #123` OR (b) described the issue in-PR following the
relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 07:35:32 -05:00
DottaandPaperclip d72389bee2 feat: add Browser Use Cloud connector and live task browsers (#14627)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The Apps gateway gives agents governed access to external tools.
> - Browser Use Cloud can run browser work, but a tool result alone does
not let a person watch or take over.
> - A task needs a durable browser session, a visible viewer, and
recorded costs.
> - This pull request adds a Browser Use Cloud v4 connection and
interactive browser tabs on tasks.
> - People can follow the work, interact with the page, and retain the
browser after the agent finishes.

## Linked Issues or Issue Description

**Problem or motivation**

Agents need governed access to Browser Use Cloud. People need to see and
interact with the same browser from the task. A browser must remain
available after a run finishes and appear at the correct point in the
task feed.

**Proposed solution**

Add a native REST connection for the v4 API. Bind each session to its
company, task, agent, and credential grant. Open its interactive viewer
in the task side panel. Record provider costs as financial events. Use
`browser-use-cloud` as the app and connector key. Keep its skill with
the connector and deliver it only with authorized connection tools.

**Alternatives considered**

A v3 MCP connection would expose tools without the v4 lifecycle
integration. An external viewer link would leave the task. A fixed
viewer size would prevent pages from responding to changes in the task
pane.

**Roadmap alignment**

This extends the governed Apps gateway and Connected Apps roadmap. It
uses the existing task, grant, secret, approval, and financial records.
The work was requested by the maintainer. A search found no duplicate
Browser Use connector PR or issue.

## What Changed

- Add the Browser Use Cloud app, brand asset, API-key connection, and
profile settings under the `browser-use-cloud` key.
- Bundle the `browser-use-cloud` skill with the connector. Keep it out
of global `skills/` discovery. Deliver it only with authorized task/run
connection tools. Remove retired connector skill keys from runtime
overlays and preserve unrelated browser skills.
- Expose seven v4 tools through the governed gateway and deliver them to
native and CLI agents.
- Persist sessions, browsers, runs, event and recovery cursors, shutdown
leases, and cumulative cost accounting. Recover uncertain paid starts
without replaying them.
- Enforce task ownership, credential grants, approvals, revoked access,
and budget limits.
- Add interactive task browser tabs and compact chronological feed
entries. Retain the viewer across tab switches and keep visible idle
browsers open.
- Add debounced automatic viewport fitting, standard size presets, and a
viewer ownership lease.
- Add lifecycle, authorization, accounting, viewport, UI, and Storybook
coverage.
- Add an idempotent database migration after the current master
migration. Preserve deployed migration hashes. Migrate pre-release Cloud
connection and financial keys without replacing grants, credentials, or
browser history.
- Document provider behavior, live acceptance results, and the lack of
documented passkey forwarding.

## Verification

- Full workspace typecheck and production build pass on the updated
branch.
- Token gates, brand asset validation, module boundaries, and migration
ordering pass.
- Cloud tests verify global skill exclusion, authorized task/run
delivery, unassigned agents, disabled connections, revocation, adapter
isolation, and secret exclusion. The existing AgentMail connector
assignment test also passes.
- Migration replay runs twice against existing browser work and
financial records. It preserves the records and avoids duplicate costs.
- The focused provider, app catalog, OpenAPI, connection gateway, and
migration regression suites pass. Recovery coverage includes lost
replies, process crashes, provider rejection, and browser arrival
acknowledgement.
- All 54 checks pass on `2974b5f03641ad0cea3c941d8c02579316fa8c92`,
including the full test matrix, browser E2E shards, build, typecheck,
security, and release canary. Two optional Storybook jobs are skipped.
- Greptile is 5/5 on the same commit, with zero unresolved review
threads. The corrected review uses the actual master-to-head diff.
- The local `pnpm test:run` started and was stopped after the full CI
matrix passed. It did not complete locally; the full-suite result above
comes from CI.
- Earlier live acceptance used an isolated company with a capped
provider credential. The agent opened paperclip.ing, the embedded viewer
accepted navigation, and the same browser stayed available after
completion and tab switches.
- The local Storybook build passes. Stories cover the panel, footer,
feed entries, settings, lifecycle failures, and viewport modes with an
offline viewer fixture.

## Risks

- Browser Use charges for hosted work. Provider caps and local budget
checks reduce exposure; reported costs can arrive after work completes.
- Viewer and CDP URLs grant access to the browser. The server validates
and restricts them. They are excluded from agent results and durable
event data.
- Runtime resizing of v4 agent browsers uses a provider option confirmed
by live testing but absent from its published agent schema. Resizing
during a click may invalidate coordinates. Fixed presets remain
available.
- Viewport ownership is process-local and resets on restart. The
lifecycle and accounting records remain in the database.
- The original intermittent embedded-viewer stall has not been fully
diagnosed. A bounded reconnect and active-session recovery cover the
observed failure paths.
- Live tests did not cover every revocation, approval, rate-limit, or
restart case. Deterministic integration tests cover those paths. Passkey
forwarding is not claimed.
- Unknown create outcomes keep the credential available for cleanup.
Run-list absence cannot prove a paid POST was rejected, so recovery
stays pending until it can identify provider work.

## Model Used

OpenAI Codex, GPT-6. Used reasoning, repository search, code execution,
browser interaction, and test tools. The exact serving model ID and
context-window size were not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 07:13:16 -05:00
DottaandPaperclip 2f6fa3b6dc fix: recover provider authentication inside tasks (#14629)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents need a working model provider connection to run a task.
> - A provider can reject a stored credential after the task starts.
> - The failed run must ask the responsible user to repair that
connection.
> - This pull request adds that request directly to the task and reuses
Connections sign-in.
> - The user can choose an API key or subscription, then continue the
task with a fresh session.

## Linked Issues or Issue Description

**What happened?**

A run that ended with `acpx_auth_required` or another known provider
authentication error did not immediately offer an inline way to connect
the provider. A repair form could also lock the user to the failed
account's sign-in method.

**Expected behavior**

Show a provider connection card in the task as soon as the
authentication failure is saved. Allow the responsible user to connect
or repair the provider with any supported sign-in method. Keep the
connection name automatic and resume the task after successful setup.

**Steps to reproduce**

1. Run a task with a supported provider and an expired or invalid
credential.
2. Let the run fail with a provider authentication error.
3. Open the task and attempt to repair the connection.

Related work: Refs #13724 and #13726. This change adds the inline task
repair flow and method choice.

## What Changed

- Classify provider authentication failures and create one connection
request for the current task. A persisted blocked classification
suppresses automatic retries only after the repair card is created;
unsupported providers retain their existing recovery path.
- Mark only the attributed, unchanged credential as needing sign-in.
Preserve credentials that were refreshed after the failed run started.
- Reuse the provider sign-in controls inside the task. Allow API key and
subscription choices for Claude, Codex, and Grok. Keep names hidden and
generate a default from the user, provider, and method.
- Keep the existing account when reconnecting with the same method.
Create and select another account when the method changes. Validate
updates to explicit agent bindings through the normal agent save path.
- Require explicit adoption for legacy agent authentication. Validate in
the agent environment, then commit the binding, connection install,
audit, and card completion in one transaction. Keep failed setup and
account selection visible and retryable.
- Add regression tests and update the specification and Connections
documentation.

## Verification

- Fresh local verification: 199 tests passed across the inline form,
provider method selector, default naming, authentication and recovery
classifiers, run liveness, OpenAPI routes, database adoption/rollback,
and Cursor execution suites. The adoption database suite also passed
against disposable Docker PostgreSQL.
- Full repository `pnpm build` and `pnpm -r typecheck` passed on the
latest commit. Token gates are clean.
- Embedded browser: opened real task cards from seeded authentication
failures; switched Claude from API key to subscription and back;
switched Codex from subscription to API key; confirmed the name field
stays hidden. Provider sign-in was not completed with real credentials.
- The broad local `pnpm test:run` started before review fixes and was
interrupted after the working tree changed; it is not counted as a
passing full run. Fresh focused tests passed. CI supplies the full test
and browser suite results for the current commit.
- CI is green on commit `4b97a4e447045ff3d7516525a187a5d1d21e0d4c`: 54
checks passed and two Storybook checks were skipped by their path rules.
The workspace preview job passed on one rerun after a local-server
startup timeout; its rerun passed 835 tests.
- Greptile is 5/5 on the same commit with no actionable findings and no
unresolved review threads.

## Risks

- Incorrect authentication classification could prompt for a connection
unnecessarily. Tests exclude tool authorization, quota, and unrelated
runtime failures.
- A method change selects the new personal provider default, which also
applies to other agents that use that user's default. Explicit account
bindings use the existing permission and runtime validation path.
- Credential invalidation must not race with refresh or reconnect. The
code compares the saved credential generation and grant update time
under locks.
- No database migration or new credential storage format is required.

## Model Used

OpenAI GPT-6 through Codex. The exact model ID and context window size
were not exposed in this session. Capabilities used: reasoning,
repository editing, shell commands, database tests, and embedded-browser
interaction.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused tests listed
above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 06:39:37 -05:00
Devin Foley f38b5693f6 fix: always enable keyboard shortcuts (#14643)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The web UI has keyboard shortcuts for the inbox, task lists, cases,
and task detail, plus global shortcuts such as `c`, `/`, `?`, `[`, and
`]`
> - Shortcut enablement was an instance-wide General setting until
#14141 moved it to a per-user preference that defaults to off
> - The move did not carry the old instance value over, so every
existing user lost shortcuts on upgrade and had to find a new toggle
under Profile settings
> - A toggle that only turns off a standard, input-safe feature costs a
setting, a database column, two API routes, and a React context for
little benefit
> - This pull request removes both the instance setting and the personal
preference and enables keyboard shortcuts for every signed-in user
> - The benefit is one less thing to configure, no silent loss of
shortcuts on upgrade, and less code to maintain

## Linked Issues or Issue Description

Refs #14141 (the change that introduced the personal preference).

**What existing behavior does this improve?**

Keyboard shortcuts in the web UI stay off unless each user turns them on
in Profile settings.

**Subsystem affected**

Web UI shortcuts, Profile settings, instance general settings, the
`/api/auth/preferences` routes, and the `user` table.

**Current behavior**

Shortcuts default to off per user. #14141 moved the toggle from Instance
settings → General to Profile settings and did not carry the old
instance value over. Users who had shortcuts on lost them after the
upgrade and had to find the new toggle.

**Proposed behavior**

Keyboard shortcuts are always enabled for every signed-in user. There is
no instance setting and no personal preference. Shortcuts already ignore
key presses inside text inputs and modal dialogs, so an opt-out is not
needed.

**Reason and benefit**

Fewer settings, no silent loss of shortcuts on upgrade, and removal of a
database column, two API routes, a query hook, and a React context that
existed only to gate this feature.

**Breaking changes**

`GET` and `PATCH /api/auth/preferences` are removed. `PATCH
/api/instance/settings/general` no longer accepts `keyboardShortcuts`;
that schema is strict, so the key now returns 400.
`instance.general.keyboardShortcuts` is no longer a valid
`PAPERCLIP_HIDDEN_SETTINGS` key; the parser ignores unknown keys with a
warning.

## What Changed

- Removed the Keyboard shortcuts section from Profile settings, the
`useUserPreferences` hook, `queryKeys.auth.preferences`, and
`authApi.getPreferences` / `authApi.updatePreferences`.
- Removed `GeneralSettingsContext`. The inbox, legacy inbox, task list,
legacy task list, cases, and task detail pages no longer gate their key
handlers.
- Removed the `enabled` option from `useKeyboardShortcuts`. The app
shell always registers the global shortcuts.
- Removed `GET` and `PATCH /api/auth/preferences`, their OpenAPI
entries, and the `currentUserPreferencesSchema` /
`updateCurrentUserPreferencesSchema` validators.
- Removed `keyboardShortcuts` from `InstanceGeneralSettings`, the
general settings zod schema, the settings service defaults, and
`HIDEABLE_GENERAL_SECTIONS`.
- Added migration `0289_drop_user_keyboard_shortcuts`, which drops
`user.keyboard_shortcuts`.
- Updated `AGENTS.md`, `doc/SPEC.md`, `doc/SPEC-implementation.md`, and
`docs/deploy/environment-variables.md`.
- Parsed the stored general settings row with
`instanceGeneralSettingsSchema.strip()` in the feedback vote path, so a
retired key left in the row cannot reset the sharing preference to
`prompt` and overwrite the stored choice.
- Kept every bare global shortcut (`c`, `?`, `[`, `]`, `/`) out of open
modal dialogs in `useKeyboardShortcuts`; only `/` had that guard before.
- Updated the affected tests and added a Profile settings test that
asserts the toggle is gone, a hook test for the modal dialog guard, and
a feedback service regression test for the retired-key case.

## Verification

- Typecheck passes for `@paperclipai/shared`, `@paperclipai/db`
(including the migration numbering and safety checks),
`@paperclipai/server`, and `ui`.
- `pnpm exec vitest run
server/src/__tests__/instance-settings-routes.test.ts
server/src/__tests__/openapi-routes.test.ts
server/src/__tests__/auth-routes.test.ts
server/src/__tests__/sentry.test.ts` → 119 passed.
- `pnpm exec vitest run ui/src/components/Layout.test.tsx
ui/src/pages/ProfileSettings.test.tsx ui/src/pages/IssueDetail.test.tsx
ui/src/pages/Inbox.test.tsx ui/src/pages/Cases.test.tsx
ui/src/hooks/useKeyboardShortcuts.test.tsx ui/src/pages/Agents.test.tsx
ui/src/pages/InstanceGeneralSettings.test.tsx` → 286 passed.
- `pnpm exec vitest run packages/shared/src/settings-visibility.test.ts`
→ 16 passed.
- `pnpm exec vitest run ui/src/hooks/useKeyboardShortcuts.test.tsx` → 7
passed.
- `pnpm exec vitest run server/src/__tests__/feedback-service.test.ts`
(embedded Postgres) → the new retired-key test passes with the fix and
fails without it.
- Manual: sign in with no settings changed, open the inbox, press `j`
and `k` to move the selection, press `?` to open the cheatsheet. Open
Settings → Profile and confirm there is no Keyboard shortcuts section.

## Risks

- The migration drops a column. It uses `DROP COLUMN IF EXISTS`, and the
column has no readers after this change. If you roll back to a build
from before this PR after the migration has run, re-add the column
first: `ALTER TABLE "user" ADD COLUMN "keyboard_shortcuts" boolean
DEFAULT false NOT NULL;`. The older build's ORM selects that column when
it loads users.
- Any external client that still sends `keyboardShortcuts` to `PATCH
/api/instance/settings/general` receives a 400. No in-repo client does.
- Stored `instance_settings.general.keyboardShortcuts` values are
stripped on read and ignored.
- Users who never turned the toggle on now get shortcuts. The handlers
skip text inputs, contenteditable regions, and modal dialogs, so typing
is unaffected.

## Model Used

Claude Fable 5.1 (`claude-fable-5-1`) in Claude Code, with extended
thinking and tool use.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-29 21:28:06 -07:00
Vyacheslav ScherbininandClaude Opus 5.5 3b34d45a92 fix(ui): make destructive-foreground readable in the light theme (#14328)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The board UI uses shadcn theme tokens from `ui/src/index.css`.
> - Delete and remove dialogs confirm with a red button: `bg-destructive
text-destructive-foreground`.
> - In the light theme, `--destructive-foreground` has the same value as
`--destructive`. The label is red text on a red fill, so it is not
visible.
> - Operators cannot read the button that deletes data or removes a
connection.
> - This pull request sets the light `--destructive-foreground` to the
near-white value that the dark theme already uses.
> - The benefit is that every destructive confirm button shows a
readable label in both themes.

## Linked Issues or Issue Description

Fixes #14327

I searched open and closed PRs for `destructive-foreground`, "Remove
connection" and related terms. I found no PR that changes this token.

## What Changed

- `ui/src/index.css`: in `:root`, `--destructive-foreground` changes
from `oklch(0.577 0.245 27.325)` (the same as `--destructive`) to
`oklch(0.985 0 0)`. This is the value in `.dark`.
- Added `ui/src/components/DestructiveThemeTokens.test.ts`. The test
converts both tokens from OKLCH to sRGB and computes the WCAG contrast
ratio. `:root` must reach 4.5:1 (AA for normal text). `.dark` must reach
3:1, because it keeps the brighter red chosen in the gallery review
(about 3.6:1 today).

The fix applies to all five buttons that use
`text-destructive-foreground`: remove connection (`Browse.tsx`,
`Connections.tsx`), delete folder (`FolderControls.tsx`), delete
environment (`CompanyEnvironments.tsx`) and remove chat endpoint
(`ChatEndpointDetail.tsx`). No other file uses the token.

## Verification

- `pnpm exec vitest run src/components/DestructiveThemeTokens.test.ts`
in `ui`: 2 tests pass. The light theme contrast is about 4.8:1.
- Before the fix, the test fails for `:root` (1:1, red on red). With a
light gray foreground `oklch(0.88 0 0)` it also fails (3.3:1 < 4.5:1).
- Manual check on a self-hosted instance in the light theme: Apps →
connection → **Remove connection** showed a red button with no visible
label. After the fix, the label is white and readable.


Before and after, light theme (the same classes as the confirm button:
`bg-destructive text-destructive-foreground`):

<img
src="https://gist.githubusercontent.com/gentslava/3aa0d498e2475dd44e12bed11e7d7862/raw/c793570ebd416fe04cc390df8554515b4207590d/remove-connection-before-after.png"
alt="Remove connection button before and after the fix" width="520">

## Risks

- Low risk. The change is one token in the light theme. Only the five
confirm buttons above read `--destructive-foreground`.

## Model Used

- Claude Opus 5.5 (`claude-opus-5-5`) in Claude Code, with tool use
(shell, file edits). It found the cause, wrote the fix and the test, and
wrote this description.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

@cryppadotta @devinfoley a tiny one-line theme fix: the delete buttons
in the light theme are red-on-red right now, so people can't see what
they are about to click. Would be great to get it in when you have a
minute 🙏

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 16:34:26 -07:00
DottaandPaperclip e912f0df53 fix(ui): open text attachments in task tabs (#14297)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agent work often ends with a Markdown or plain-text file.
> - Task attachments currently open outside the task panel.
> - Users need to inspect those files while keeping the task
conversation in view.
> - This pull request opens text attachments in task tabs and adds
rendered, raw, and download controls.
> - The same controls work in the mobile task drawer.

## Linked Issues or Issue Description

**What happened?**
Opening a text attachment did not put its content in a task tab.
Markdown files had no in-task rendered/raw toggle.

**Expected behavior**
Open Markdown and text attachments in one reusable task tab. Show
Markdown as rendered content or raw text. Download the original file.

**Steps to reproduce**
1. Upload a Markdown file and a plain-text file to a task comment.
2. Open each attachment from the task conversation or artifact list.
3. Switch Markdown between Rendered and Raw. Download both files.
4. Repeat at a mobile viewport width.

Related work: #14193 controls artifact tab arrival. This change adds
text attachment content tabs.

## What Changed

- Route text attachment opens from conversation and artifact cards into
task tabs.
- Add a text attachment panel with accessible Rendered, Raw, and
Download controls.
- Preserve ordinary links for other file types.
- Support the selected attachment in the mobile drawer.
- Keep text-tab actions on the current rich artifact cards, including
CSV previews.
- Render attachment image references and diagram source without loading
media URLs.
- Add browser regression tests, component tests, Storybook examples, and
usage documentation.

## Verification

- Full workspace typecheck, production build, Storybook build, and UI
token gates pass locally.
- All 6,960 UI tests pass. The additional media regression passes
against the real Markdown renderer and fails before the fix. CSV
coverage verifies direct downloads and text tabs after preview.
- Both desktop and mobile browser cases pass locally. They check
rendered/raw Markdown, literal plain text, reusable tabs, review
controls, and exact original download bytes. The local fixture used a
separate database port because an existing socket occupied the default
range.
- The full CI test matrix passes on
`82be5efbef26927b237a031725bb3d7fa79f637f`. The duplicate local `pnpm
test:run` was stopped after this CI result; it did not complete locally.
- Greptile is 5/5 on the final commit. All review threads are resolved,
and the security scan passes.
- All 54 final-head checks pass, including the canary dry run. The two
optional Storybook jobs are skipped.

## Risks

- Text attachment links now open in the task panel. Other content types
keep their existing link behavior.
- File display still depends on the existing authenticated attachment
route. There are no API or database changes.
- Raw text is displayed as text, including strings that look like HTML.
Rendered Markdown keeps media references inert.

## Model Used

OpenAI Codex, based on GPT-6, with code execution, browser testing, and
subagent tool use. The runtime does not expose an exact serving model
variant or context-window size. Recovered earlier implementation changes
were reviewed and tested; their exact model metadata is unavailable.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-29 16:40:10 -05:00
DottaandPaperclip 81a52eb740 fix(ui): allow touch scrolling in new-task pickers (#14599)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The new-task composer lets users select an assignee and override its
model.
> - On phones, these selectors open as large sheets outside the task
dialog DOM.
> - The parent dialog's scroll lock cancels touch drags in those sheets.
> - This pull request gives each mobile sheet its own modal scroll
boundary.
> - Users can scroll to an option, select it, and continue editing their
task.

## Linked Issues or Issue Description

**What happened?**

An iPhone Safari user could not drag through the assignee or model list
in the new-task composer. The list stayed at the top. A touch-enabled
Chromium reproduction also showed canceled touchmove events and an
unchanged scroll position.

**Expected behavior**

A finger drag scrolls the list. A tap selects an option. Closing the
picker preserves the task draft and selected values.

**Steps to reproduce**

1. At phone width, open New Task in a company with enough agents to
overflow the picker.
2. Open Assignee and drag upward through the list.
3. Select a Codex agent, open Codex options, choose Custom, and open the
model selector.
4. Repeat the drag with enough models to overflow the available
viewport.

**Paperclip version or commit**

Reproduced on `e5bf9d49a`. The fix is rebased onto `b3eb03fcb`.

**Deployment mode**

Local development in an isolated test instance. The user reported iPhone
Safari. Browser automation uses native Chromium touch input at phone
dimensions; a physical iPhone was not available.

Related: #14250 introduced the large mobile entity picker sheets.
Duplicate search found no existing fix for their touch scroll boundary.

## What Changed

- Use modal Radix popovers for mobile entity sheets. Their lists can
scroll while the background stays locked.
- Suppress opening when Radix restores trigger focus after dismissal.
Escape and outside taps now close the picker without reopening it.
- Add a failing-before/fixed-after regression for a mobile sheet inside
a parent dialog.
- Add a browser test for native touch scrolling in both lists, tap
selection, draft retention, close-button/Escape/outside dismissal, and
desktop mouse/keyboard selection.
- Document the browser test command.

## Verification

- Passed `pnpm -r typecheck`, `pnpm build`, and `pnpm
check:token-gates`.
- Passed 41 focused tests for `InlineEntitySelector` and
`NewIssueDialog`.
- Passed the new browser test against a disposable instance running this
worktree. It uses 22 real fixture agents and a fixed 24-model catalog.
It covers 390×844 and 390×430 phone viewports and a 1280×900 desktop
viewport.
- Inspected the rendered new-task form and both selectors. Selections
returned to the draft with its title intact.
- `pnpm test:run` was attempted and stopped after confirming that local
database suites were skipping because macOS has exhausted its system
semaphore limit (`initdb`: `could not create semaphores: No space left
on device`). It did not complete locally. The complete test matrix
passed in Linux CI, including all 71 tool-gateway tests.
- All 150 browser tests passed in CI, including the new native-touch
regression.
- Greptile reviewed the latest commit at 5/5. Its desktop focus finding
is fixed and the review thread is resolved. All checks on
`2ab22727ca768b4134c9c840bbc8e4c80d7da674` are complete: 53 passed, 2
intentionally skipped Storybook deployment checks, and the Snyk status
passed.

## Risks

Low risk. Mobile selectors now own focus and scroll isolation. This
changes their modal behavior, so the tests cover nested dismissal and
focus return. Desktop selectors remain non-modal. No schema, API, or
dependency changes.

## Model Used

OpenAI Codex (GPT-6), with reasoning, repository inspection, code
execution, and browser testing. The exact backend deployment ID and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-29 13:37:03 -05:00
DottaandPaperclip 29c8fb0b66 fix(composer): match effort options to adapter execution (#14576)
## Thinking Path

> - Paperclip is an open source control plane for AI agents.
> - The issue composer lets operators choose a model and effort for a
task message.
> - The picker must show only effort levels that the selected adapter
can execute.
> - The Codex Runner fix in #14568 exposed similar gaps in other
adapters.
> - Grok hid a supported control. Claude used one range for all models.
Kimi ACP showed a control that it ignored.
> - This pull request aligns the picker with each adapter and adds
matching stories.
> - Operators can see and save supported effort settings without false
controls.

## Linked Issues or Issue Description

**What happened?**

The composer hid Grok effort. It showed an incomplete effort range for
some Claude models and a slider for Claude Haiku. It showed Kimi effort
on the default ACP engine even though that engine drops the value.

**Expected behavior**

The composer should show only effort levels supported by the selected
model and execution engine. Grok effort should reach its adapter as
`reasoningEffort`.

**Steps to reproduce**

1. Select a Grok agent and `grok-4.7`. Observe that the picker has no
effort slider.
2. Select Claude Sonnet 5 or Haiku 4.5. Observe the generic low to high
slider.
3. Select a Kimi agent on ACP. Observe the slider even though ACP
ignores effort.

Related fix: #14568.

## What Changed

- Use Claude and Grok adapter capability helpers in the production
picker and Storybook preview.
- Map Grok composer effort to `reasoningEffort` in per-message
overrides.
- Show Kimi effort only when its agent uses the CLI engine, including
when it runs its default model.
- Show Grok effort when it runs its default model.
- Add design and production stories for Claude, Grok, and Kimi ACP and
CLI cases.
- Document the picker rules and add focused tests.

## Verification

- `pnpm --filter @paperclipai/ui exec vitest run
src/components/task-chat/composer-run-settings.test.ts`
- `pnpm --filter @paperclipai/ui typecheck`
- `pnpm check:token-gates`
- `pnpm -r typecheck`, `pnpm test:run`, and `pnpm build` passed on the
first commit.
- The focused test, UI typecheck, token check, and Storybook build
passed again after the review fixes.
- Browser smoke: production stories show effort for default Grok and CLI
Kimi, and hide it for ACP Kimi and Claude Haiku.

## Risks

- Existing Kimi ACP agents lose a slider that could not change
execution. Their stored effort override remains unchanged until the next
edit.
- Claude and Grok ranges follow their adapter catalogs. A provider can
change model support before the catalog is updated.

## Model Used

OpenAI Codex based on GPT-6. The exact runtime model ID and context
window are not exposed in this session. Reasoning, tool use, and code
execution were used.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-29 10:34:29 -05:00
da887ea3e9 fix(runner): honor Codex effort selected in composer (#14568)
## Thinking Path

> - Paperclip manages AI agents that work on assigned tasks.
> - The task composer lets a person choose an assignee, model, and
effort for the next run.
> - A Paperclip Runner agent can use Codex as its provider.
> - The composer hid Codex effort for that agent because it checked only
the older Codex adapter.
> - The native Runner input also did not carry an effort choice to
Codex.
> - This pull request carries the chosen effort from the composer to
each Codex turn.
> - People can now select a supported effort and get the effort they
selected.

## Linked Issues or Issue Description

Refs #14322

**What happened?**

The composer showed a model but no effort slider when the assignee used
Paperclip Runner with the Codex provider. A task-level model override
also did not reach the native Runner input.

**Expected behavior**

The composer shows effort choices for a known Codex model. The next
native Codex turn uses the selected model and effort.

**Steps to reproduce**

1. Open a task composer.
2. Select an agent that uses Paperclip Runner with the Codex provider.
3. Select a known Codex model such as `gpt-6-astra`.
4. Open the assignee and model picker. The effort slider is missing
before this change.

## What Changed

- Show known Codex effort levels for Paperclip Runner Codex assignees.
- Save the task effort override in the native run input and send it to
Codex on each turn.
- Apply the task's merged model and effort overrides when the native run
starts.
- Apply a task model override for OpenCode Runner without changing the
agent's provider.
- Add Runner effort tests and desktop and mobile Storybook cases.

## Verification

- `pnpm -r typecheck` passed.
- `pnpm build` passed.
- `pnpm build-storybook` passed.
- `pnpm check:token-gates` passed.
- Focused UI, server, Runner contract, and Codex driver tests passed.
- The full CI test matrix, build, typecheck, and canary dry run passed
on the latest head.

## Risks

- Native Runner inputs add an optional Codex effort field to the current
v5 input. Older inputs keep their previous behavior.
- A known model rejects an effort that its catalog does not support.
Unknown models do not show a slider.

> This fixes an existing composer bug. I checked `ROADMAP.md`; it does
not describe this bug as planned work.

## Model Used

OpenAI Codex, GPT-6. The exact deployment ID and context window are not
exposed in this session. The model used reasoning, code execution, and
repository tools.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: OpenAI GPT-6 Astra <noreply@openai.com>
2026-09-29 09:41:05 -05:00
DottaandPaperclip 7636966452 fix(inbox): keep other users’ failed runs out of Mine (#14572)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The Mine inbox shows work that needs the current user.
> - Failed-run rows used the latest run for every agent in the company.
> - A failure from another user therefore appeared in Mine and its
badge.
> - Run list responses also omitted the responsible user needed to
filter these rows.
> - This pull request uses run ownership for personal failure routing.
> - Users see their own failures and can still inspect company failures
in All.

## Linked Issues or Issue Description

**What happened?**

An agent run started for one user failed. Its row and failure badge
appeared in another user's Mine inbox.

**Expected behavior**

Mine and its badge include failed runs for the current responsible user.
Other users' failures remain available in All and run details.

**Steps to reproduce**

1. Use a company with two human users.
2. Create a failed or timed-out run attributed to the first user.
3. Open Mine as the second user. Before this fix, the failed run appears
there and increases the badge.

**Paperclip version or commit**

Reproduced in regression tests on master at `24beb0057`.

**Deployment mode**

Authenticated deployment with multiple users. Tests also cover the local
single-user board.

Related prior work: #933 addressed inbox dismissal and badge
consistency. No duplicate ownership fix was found.

## What Changed

- Return `responsibleUserId` in normal and summary run lists.
- Share one ownership rule across both inbox versions and client/server
badges.
- Select the latest run per agent before applying the ownership filter.
This prevents old failures from resurfacing on shared agents.
- Keep unattributed historical failures in the local board's Mine view.
Hide them from authenticated users with no matching owner.
- Keep company health alerts outside the personal badge, consistent with
the client.
- Document the routing contract and add page, badge, and database
regression coverage.

## Verification

- Red: the new badge cases failed with three company failures instead of
one personal failure; eight Mine page cases failed across both inbox
versions.
- Green: 113 focused tests pass in `ui/src/lib/inbox.test.ts`,
`ui/src/pages/Inbox.test.tsx`,
`server/src/__tests__/heartbeat-list.test.ts`, and
`server/src/__tests__/inbox-dismissals.test.ts`.
- `pnpm check:token-gates` passes.
- Agent calls on behalf of a user have two additional red-to-green API
regressions.
- Full `pnpm -r typecheck` and `pnpm build` pass. Server typecheck also
passes after the agent-call fix.
- All CI test shards and browser tests pass on `243bfa681`. The
duplicate local `pnpm test:run` was stopped after the CI test lanes
completed; it did not finish locally.

## Risks

- Authenticated users no longer receive unattributed legacy failures in
Mine. Those failures remain visible in All.
- The server badge no longer counts company health alerts, matching the
existing client badge.
- No migration, run state, retry behavior, or company access rules
change.

## Model Used

- OpenAI GPT-6 through Codex, with reasoning, terminal execution, and
browser tools. The exact deployment variant and context window size are
not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-29 09:40:21 -05:00
DottaandPaperclip 24beb00575 feat(runner): add rich ACP transport and durable interaction foundation (#14430)
Add shared rich ACP transport, durable questions and permissions, verified provider packaging, and bounded activity and plan presentation. Keep Cursor, Copilot, and Pi pending their separate provider qualification.

Persist interaction settlement before publication, fence failed writes until fresh recovery, and preserve owned-process cleanup. Incorporate reviewed mainline integration with extended harness coverage.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-29 08:56:21 -05:00
DottaandPaperclip 61b3fd57a6 fix(ui): recover gracefully during server restarts (#14560)
Show a clear reconnecting state during server restarts and retry safe access
checks every five seconds. Preserve open drafts, wait for initial startup
readiness, and keep authentication failures separate from temporary outages.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-29 08:31:05 -05:00
DottaandFry 3ca196b0a6 feat(agents): persist agent files across tasks without revision history (#14420)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - An agent needs personal files across tasks and sessions.
> - AGENTS.md is one file in that directory. Supporting files need the
same persistence.
> - The Instructions Editor and agent runs must share one current
directory.
> - Concurrent runs should apply only the files they change. The last
sync of the same file wins.
> - This pull request uses existing file transport and removes temporary
copies after sync.
> - Old instruction-only sessions keep their restore contract. New saves
do not create revision history.

## Linked Issues or Issue Description

Refs #14325. This replaces its revision-oriented design with persistent
agent files. Keep #14325 unmerged.

Transport prerequisite #14416 merged first at
`d172197117a14b80a1eb2d2835a0e7cce2679656`. This PR now targets master
and remains below 100 changed files.

Related work: #4513 and #8798 cover instruction tooling. This change
handles run synchronization, cross-task personal files, browser editing,
and old-session restoration.

## What Changed

- Keep one current directory per company and agent. Point AGENT_HOME at
a temporary working copy for each active run. Keep task files and
provider HOME separate.
- Restore text, binary files, and nested folders through workspace
transport. Exclude remote agent files from task Git snapshots with a
self-ignoring file inside the reserved runtime directory; never write
through repository-controlled Git metadata.
- Collect after the provider and child processes have stopped. Keep
resumable conversation state.
- Apply changed and deleted files under the agent lock. The last sync
wins for the same file. Unrelated concurrent changes survive.
- Remove temporary copies after successful sync, rejected sync, and
staging failure. Register ownership before copying so restart recovery
can remove interrupted preparation. Retry transient synchronization up
to three times. Preserve the original remote lease reference until
deletion succeeds; restart cleanup never acquires a replacement sandbox.
Do not create captured directories or a conflict-review queue for new
runs.
- Keep browser editing, stale-draft protection, and streaming binary
downloads. Keep the instruction entry and text editor limited to 1 MiB.
- Keep historical agent-folder sync failures on their affected runs
instead of repeating them above current saved instructions. Preserve
legacy candidate review and current browser-save errors. Avoid duplicate
quota warnings while retaining separate sync failures when they describe
a different problem.
- Require target-scoped caller grants for peer instruction access, while
preserving self edits, responsible-user checks, and protected-change
consent.
- Treat full storage as a nonblocking run warning, never an agent pause
or run-admission failure. Restore already-over-quota saved folders so
ordinary agent cleanup can recover; warn on each run until cleanup. The
run detail view shows the warning.
- Allow 256 MiB per file, 2 GiB per directory, and 100,000 entries. Hash
large files as streams. Check editor-save quotas with metadata instead
of hashing unrelated files.
- Preserve old native inputs, instruction-only copies, paths, digests,
and pending legacy candidates. Adopt old revision heads once. New writes
do not append history rows.
- Add idempotent migration 0287 and verify upgrades from the preview
tables and receipts.
- Add nine interactive stories under **Agents / Persistent files**,
including automatic incoming edits, stale browser drafts, and
storage-limit diagnostics.

## Verification

- Merge candidate: `4f5390107ec6ffd80a76d1d2e85530e66f21d079`, after
merging current master and the landed transport prerequisite.
Integration required no manual conflict resolution; the feature remains
99 changed files. Full workspace typecheck, production build, token
gates, and 715 focused tests passed on this merge candidate. Fresh
Greptile review is 5/5 with no unresolved findings. All 55 checks
passed, with four conditional skips, including the build, typecheck,
browser E2E, and canary dry run. A single retry recovered four jobs
interrupted by runner shutdowns; no source changes were required.
- Historical-warning UI fix: all 6,834 UI tests across 640 files passed,
including regression coverage for three old failures, legacy preserved
edits, and warnings scoped to the affected run. Full workspace
typecheck, production build, Storybook build, and token gates passed.
Browser-verified Storybook playtests passed for Historical Failures
After Successful Save, Storage Limit, and Full Storage Run Warning.
- Review follow-ups at `4e20c9fb2`: all 18 focused tests passed,
including external Git directories, linked worktrees, symlinks,
hardlinks, and distinct I/O failures alongside storage warnings. Server
and UI typechecks, token gates, and the production build passed.
- Storage warning regressions at `0724f3012`: all 33 directory tests and
all five heartbeat-list tests passed, with no skips in their successful
runs. They cover repeated runs while full, an already-over-quota saved
folder, cleanup, warnings retained after unrelated save failures, and
bounded warnings in large result JSON. Server typecheck passed after the
final warning fixes.
- Full workspace typecheck, production build, and token gates passed
during this follow-up. Product E2E harness: 631 tests passed across 52
files; harness typecheck passed. Earlier native session/context and
directory/legacy collection suites passed 537 tests; Runner
unit/transport suites passed 329 tests.
- **Real E2E at `0724f3012` (before this follow-up):** legacy local
Codex and native Daytona Codex each passed six tasks, one server
restart, seven independent assertions, and cleanup verification. Both
prove browser-to-agent edits, agent-to-browser edits, nested/binary
restoration, per-file last-sync-wins, a successful run after an
oversized save rejection, and cleanup clearing the warning.
- Native local Codex also passed the six-task quota flow before the
final warning-retention fixes. That pass began at `918d1ed02` while the
bounded-result warning fix was being edited, so it is not claimed as
exact-final-head evidence. Its final-head rerun failed during embedded
PostgreSQL bootstrap before any provider run: the macOS host had 87,365
of 87,381 SysV semaphores occupied. No unrelated services or kernel
limits were changed.
- The final-source report intentionally records **2/3 cells passed**,
preserving the blocked native-local attempt:
`tests/runner-e2e/results/agent-files-quota-final-20260928-report/`.
Earlier failed attempts and provenance notes remain under
`tests/runner-e2e/results/agent-files-quota-final-20260928-input/` and
the original campaign directories.
- Daytona used immutable image
`ghcr.io/paperclipai/paperclip-daytona-runner@sha256:5643f0d801417cae3581833a1a3bc6715b325e028602738d2652c44cac5dc6bf`
and its exact Linux runner binary. Controller source is `0724f3012`;
image source is recorded separately.
- Legacy-session compatibility and all three ACP Stop/resume browser
regressions passed on the prior validated feature head
`169fab46d5af21caa2269b4c1b29b69c933a6951`. They assert the same
provider session is retained and interrupted writes are not replayed.
Migration upgrade tests also passed earlier.
- Nine interactive stories are under **Agents / Persistent files**,
including **Full Storage Run Warning**. Its playtest and visual browser
inspection passed; the warning states that runs continue and the editor
remains available.
- Prior-head checks on `4e20c9fb2`: 55 passed, two conditional jobs
skipped, no failures or pending checks. All eight browser E2E shards and
their aggregate passed. Fresh Greptile review is 5/5 with no findings;
all review threads are resolved, the security scan passed, and GitHub
reports no merge conflicts.
- The broad local follow-up test run was interrupted after host
semaphore exhaustion affected isolated PostgreSQL instances. It also
encountered the existing macOS long-path fixture failure and two timeout
failures. This is not a claim that the full local suite passed. Logs are
retained; focused storage/warning tests passed.

## Risks

- A later sync can overwrite an earlier edit to the same file, including
a saved browser edit. There is no text merge or retained version. This
is the intended last-sync-wins policy.
- A save that exceeds a storage limit is rejected and its temporary copy
is discarded. The run itself continues normally, and later runs restore
the last saved files with a warning until cleanup. Transient sync
failures get bounded retries. An I/O failure partway through a sync can
leave some files updated; a failed receipt does not claim whole-folder
success.
- Larger folders increase copy time, network traffic, and temporary disk
usage. Active runs still need working copies. Terminal runs do not
accumulate archives. Operators must provision disk for agents and
configured concurrency; these limits are not company-wide quotas.
- A restored old native session remains instruction-only until a fresh
session starts. Its original conflict fence and existing pending
candidates remain compatible.
- Provider processes close at the collection boundary. Conversation
resume remains available, but warm process reuse is lost.
- Backups must include the instance filesystem and database. External
bundles keep their existing behavior until explicitly moved to managed
storage.

## Model Used

OpenAI Codex, GPT-6 family. The session does not expose a more specific
model ID or context-window size. Reasoning, code execution, and browser
tools assisted this change. Real provider E2E uses `gpt-5.6-sol`.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Fry (Paperclip) <noreply@paperclip.ing>
2026-09-29 08:25:56 -05:00
Devin FoleyandPaperclip faf72cb1d5 fix: preserve company context across browser hot reload (#14482)
## Thinking Path

> - Paperclip uses a shared company context for browser providers and
consumers.
> - Vite can load a new consumer module while an older provider is
mounted.
> - Recreating the context disconnects that consumer from the mounted
provider.
> - The consumer then reports a missing provider even though one is in
its React ancestry.
> - This change preserves the context object across development module
refreshes.
> - Bounded global error diagnostics distinguish development bundles and
otherwise context-free promise rejections.

## Linked Issues or Issue Description

**What happened?**
A refreshed company consumer can throw `useCompany must be used within a
CompanyProvider`. A real Vite and Chromium reproduction confirms that a
retained provider and a refreshed consumer can hold different context
objects. Global promise rejections also lack the bounded document state
already attached to React boundary errors.

**Expected behavior**
A refreshed consumer should read the mounted provider. Error reports
should identify the loaded bundle mode and bounded browser state while
preserving monitoring opt-in, sign-out, and privacy behavior.

**Steps to reproduce**
Run `pnpm test:e2e:browser-context`. The isolated Vite fixture renders
the real CompanyProvider, imports a new timestamped consumer module, and
renders that consumer below the retained provider. The test fails before
the context change and passes after it. The SDK regression invokes its
real unhandled-rejection handler with an undefined reason.

## What Changed

- Keep the React context object in Vite's per-module `hot.data`. Account
values stay in React.
- Add an isolated browser regression with mocked API responses and no
live instance, discovered by the existing Chrome CI shards.
- Add document-state diagnostics to global errors while preserving
earlier boundary snapshots.
- Tag events with development or production bundle mode and the type of
an unhandled rejected value.
- Document the new test command and diagnostic fields.

## Verification

- Company context, browser context, and Sentry suites: 59 passed.
- `pnpm test:e2e:browser-context`: passed in Chromium. The original
context code fails the reproduction.
- Real SDK tests preserve DSN/sign-out behavior and omit request
context, breadcrumbs, and private DOM data.
- `pnpm check:token-gates`: passed.
- `pnpm -r typecheck`: passed.
- `pnpm build`: passed.
- Local `pnpm test:run` exited in the general-server phase: 13,819
passed, 86 skipped, 21 failures in unchanged filesystem and host-port
suites. Cache permission and long-path failures also reproduce on the
unmodified base; seven other failures involve local runtime port
ownership. This is not a full local pass.
- Full Linux CI and review passed at a0b5270662 (53 successful checks;
two conditional checks skipped). Three jobs interrupted by a runner
shutdown passed on rerun. Greptile is 5/5 with no unresolved comments.
The real browser regression also passed in the normal Chrome CI shard
(3.2 seconds).
- Code-owner approval remains required because the dedicated test
command changes `package.json`.
- Scanned the diff and PR text for credentials and private data before
pushing.

## Risks

The context cache applies only to development hot reload. It retains the
context object, not account state. Production continues to create an
ordinary React context. The diagnostic hook adds only bounded state and
type values, preserves boundary snapshots, and returns the original
event if enrichment fails. It does not suppress errors or restore raw
breadcrumbs. The global rejection diagnostics do not identify the
promise's originating operation by themselves.

## Model Used

OpenAI GPT-6 through Codex, with repository inspection, code editing,
and command execution. The runtime does not expose a more specific model
revision or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused suites and
Chromium regression; full local limitations documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 16:15:14 -07:00
0be2afcca6 feat(ui): improve task composer controls and pending input (#14322)
## Thinking Path

> - Paperclip lets operators assign tasks to AI agents and review their
work.
> - The task composer controls the next message and its assigned agent.
> - Operators needed a way to choose that agent's model and effort
without leaving the composer.
> - The old mode selector, upload button, and input cards made the
mobile composer crowded and hid normal messaging during a pending
decision.
> - Harnesses publish different model and effort capabilities, so the
picker must follow the selected agent.
> - This pull request adds one responsive composer flow, keeps pending
cards visible above it, and protects Codex ACP authentication in the
local test path.
> - Operators can choose run settings, send a message, and answer a
pending card as separate actions.

## Linked Issues or Issue Description

**Subsystem affected**

Task composer UI, issue thread interactions, Codex ACP credential
handling, and Storybook.

**Problem or motivation**

The composer did not expose model or effort for the selected agent.
Mobile actions wrapped poorly. Pending questions and confirmations
replaced the composer. A local Codex ACP test could also reuse host
authentication after the managed key was removed.

**Proposed solution**

Put assignee search, model search, exact model IDs, effort, and fast
mode in one picker. Use a mobile dialog. Replace the direct-upload plus
action and separate mode selector with an Add menu and removable Plan or
Ask chips. Place pending interaction cards above the usable composer.
Keep these cards pending after an ordinary message unless their creator
asks for comment superseding. Replace managed ACP auth files atomically
and isolate the test key from host credentials.

**Roadmap alignment**

ROADMAP.md does not list an overlapping composer milestone. This change
improves the existing task and review flows.

## What Changed

- Added the combined assignee, model, and effort picker to both task
composers. Search matches agent name, role, and harness. The server uses
a curated Codex list by default and honors instance-declared models.
Manual IDs remain available.
- Added an effort slider for known model capabilities, a conditional
Codex fast control, and reset. The picker opens in a modal on mobile.
- Added the Add menu for files, supported goals, Plan mode, and Ask
mode. Plan and Ask are exclusive removable chips. Keyboard mode cycling
remains available.
- Adjusted mobile spacing, avatars, wrapping, and Send placement.
Removed the composer divider.
- Moved pending question, confirmation, review, and related cards above
the composer. Ordinary comments now leave question and confirmation
cards pending by default. The onboarding prompt retains explicit comment
superseding.
- Updated the Storybook composer group with responsive states and the
production picker. Added UI, service, route, and browser regression
coverage.
- Isolated Codex ACP API-key authentication, skipped subscription auth
merge and shared-home copy-back for remote API-key runs, and replaced
the managed auth file atomically.

## Verification

- `pnpm -r typecheck` — passed on the final local head.
- `pnpm check:token-gates` — passed on the final local head.
- `pnpm exec vitest run server/src/__tests__/adapter-models.test.ts
ui/src/components/task-chat/ComposerRunSettingsPicker.test.tsx` — 31
tests passed, including role and harness search, declared Codex models,
and filtering general OpenAI models.
- `pnpm exec vitest run
server/src/__tests__/issue-thread-interactions-service.test.ts` — 74
tests passed.
- `pnpm exec vitest run
packages/adapters/codex-local/src/server/acp.test.ts` — 42 tests passed,
including remote API-key copy-back isolation.
- `pnpm test:run` — attempted locally; the embedded PostgreSQL test
database could not initialize on macOS. The isolated
`heartbeat-run-event-sequencing` suite reproduced that environment
failure. GitHub CI runs the full test matrix for this head.
- `pnpm build` — passed on the final head. `pnpm build-storybook` passed
after the last UI change; only server code, tests, and docs changed
afterward.
- Live local test drive — Codex ACP ran a task with a managed API key.
The test agent was restored to its default ACP configuration afterward.
- Review the interactive stories under the top-level Composer group with
`pnpm storybook`. Check a narrow desktop width and mobile Plan, Ask,
picker, and pending-question states.

## Risks

- A pending card stays open when an ordinary comment changes the
discussion. Its creator can set `supersedeOnUserComment: true` when a
new comment should replace it.
- Model and effort overrides persist on the task until reset or changed.
An unlisted manual model ID may fail when the provider runs it.
- Some harness catalogs do not report effort support. The picker hides
effort for those models.
- No database migration is required.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

OpenAI GPT-6 via Codex. This runtime does not expose the exact model ID
or context window to the task. The model used code execution and browser
tools.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: OpenAI Codex <codex@openai.com>
2026-09-28 23:02:39 +00:00
DottaandPaperclip 24c58e479a Improve task artifacts with rich cards and editable stories (#14469)
Render eight artifact card types from real task records and share them with editable Storybook stories. Preserve document review and media/file actions, and load bounded CSV previews on request.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-28 15:38:06 -05:00
DottaandPaperclip 18e8c121d9 fix(runner): include Grok support in public installs with sandbox prerequisites (#14024)
## Thinking Path

> - Paperclip manages agents through a shared native runner.
> - Built-in harness support should ship with Paperclip's public
distribution.
> - Grok already speaks ACP; it does not require a new public bridge
package.
> - Sandbox provisioning owns the native executable and its pinned
version.
> - The runner must verify that prerequisite without downloading it
during npm installation.
> - This change separates built-in launcher identity from external
runtime identity.
> - Clean npm installation and live staging checks verify the
distribution boundary.

## Linked Issues or Issue Description

Refs #13882, #13973, #13977, #13979.

This follow-up now targets master after #13882 was squash-merged. It
replaces the private `@paperclipai/grok-acp` workspace package with
runner-owned assets. Current master is included so the branch also
contains the merged scheduler, complete-event capture, and durable
cleanup fixes.

## What Changed

- Ship Grok launcher and qualification metadata inside the runner's
compiled output and the public server's vendored runner tree.
- Remove the separate Grok npm package and all package-manager install
hooks for this runtime.
- Require the checksum-verified Grok Build 1.0.13 binary at
`/opt/paperclip/providers/grok/1.0.13/grok` in the selected execution
environment. Provision it explicitly in the Daytona image and CI setup.
- Keep native binaries outside the provider pack. Bind the built-in
launcher into the pack manifest.
- Preserve executable leases, descriptor-backed startup, credential
fences, permissions, and exact ACP model admission.
- Use `builtin:grok-acp` and `native:grok` as profile identities.
Historical package-profile sessions fail closed on resume rather than
being silently reinterpreted.
- Resolve built-in assets from the authenticated sidecar location,
including public server npm layouts. Keep the controller path out of
provider environments.
- Add clean npm tarball installation verification to the existing
trusted canary CI job and the admitted manual EC2 verification path. It
stages a unified release version and runs npm lifecycle scripts, then
verifies missing-prerequisite rejection and admission after separate
provisioning without credentials or inference.
- Include the controller-owned provider pack in stamped Cloud images.
Unstamped local images omit the pack and remain usable; remote ACPX
requires full source provenance.
- Correct CLI approval-page metadata for an already authenticated Cloud
board user; approval authorization remains unchanged.
- Honor explicit native-runner enablement in the Cloud agent picker and
direct setup page, keeping the flag disabled by default.
- Allow selecting the execution environment before connecting
credentials. Include Grok in the existing authenticated hello-probe
flow, targeting its pinned native prerequisite for runner setup.
- Recover an existing subscription sign-in conflict through an explicit
cancel-and-retry action, serialized after cancellation succeeds.
- Preserve the selected ACPX harness before normalizing config fields,
so new Grok agents use the Grok default model.
- Keep the credential-free Cloud provider pack root-owned and readable
after runtime UID remapping; verify manifest and referenced asset access
under an unrelated unprivileged UID during image builds.
- Archive prior failover backups alongside explicitly replaced harness
state, preserving evidence while preventing stale backups from blocking
a fresh replacement.
- Update Daytona image content inputs and contract tests for the
built-in assets and explicit provisioner.
- Document and regression-test the shared `approve-all` default for Grok
setup, saved configuration, and native execution. Explicitly saved
restrictions remain unchanged.

## Verification

Current merge-repair head `df09eb3e1a619430ad8419a0ee9aedd486689b05`
incorporates master `f1a394bd30cb56fb9e479f98b9f50176fe921858` after the
base PR was squash-merged. All 12 conflicts came from incoming files
identical to the tested pre-squash base. The final tree exactly matches
a three-way merge using that original base, preserving built-in Grok
distribution and removal of the obsolete private package. All 252
focused runner/UI tests, six npm-isolation tests, and token gates pass.
Fresh exact-head Greptile review is 5/5 with no outstanding findings;
security scans and EC2 native compilation pass. All current-head CI is
green: 56 successful checks/statuses and four intentional skips ([run
36468768035](https://github.com/paperclipai/paperclip/actions/runs/36468768035)).
The repository owner explicitly authorized bypassing code-owner approval
after all checks passed; no CI checks or repository protection settings
are bypassed or changed. The only remaining PR was removed from the
completed stack metadata to permit native auto-merge.

Earlier integration head `78cb306ecc41b5c96577c26c1d89153b0ef865a1`
includes master `3447609d2247e75e55d91493dda91a608364f672` (2026-09-28).
Two master advances during verification overlapped the eval catalog; the
final merge preserves Grok qualification, completion updates, and
bounded API-response reading in all 348 cells. All 77 focused
catalog/eval/workflow tests pass. Both native stack layers (#14397) are
mergeable, and both exact-head Greptile reviews are 5/5 with successful
security scans and no unresolved review threads. All current-head CI is
green: 56 successful checks/statuses and four intentional skips ([CI
attempts](https://github.com/paperclipai/paperclip/actions/runs/36447124691)).
The initial attempt lost two EC2 runners to shutdown signals and stalled
a third shard during dependency preparation; all three passed the
same-commit failed-job-only retry. Trunk code-owner requirements remain
enforced. The review summary’s non-blocking saved-asset offset
classification note concerns code already merged in #14301; those
runtime files are identical to master and outside this stack’s diff.
Historical live evidence below retains its original source revisions.
[Final public npm
verification](https://github.com/paperclipai/paperclip/actions/runs/36445542764)
passed on `76ea70cd4d13786a042af9df82f0fd7a8c85ae30`: 17 public
packages, an executed offline lifecycle sentinel, unchanged consumer
lock, built-in launcher, missing-prerequisite rejection, and verified
separately provisioned binary/command lease. Provisioning and cleanup
require no host privilege elevation; only the positive probe mounts the
temporary native binary read-only. The verifier is unchanged by the
final master merge. All six isolation tests and an offline npm smoke
test pass. The prior head had 56 green CI checks and a 5/5 review after
two unchanged tests timed out and passed a failed-job-only retry ([CI
attempts](https://github.com/paperclipai/paperclip/actions/runs/36444597313)).
All 56 recovery-display/lineage tests pass; re-review cleared the
already-covered missed-retry concern. Earlier EC2 failures remain
retained: [npm lockfile
rejection](https://github.com/paperclipai/paperclip/actions/runs/36436311203),
[missing compiler in the slim
image](https://github.com/paperclipai/paperclip/actions/runs/36440210984),
and the aggregate 15-minute test timeouts in those broad runs. Both
broad attempts passed typecheck, token gates, Product E2E type/unit
checks and build. The focused EC2 lane preserves the existing
trusted-actor and immutable-source gates.


Earlier documentation/test checkpoint
`ff244c4fd78a7ede5a3e00efe09f475f133ef33e` leaves runtime behavior
unchanged. 154 focused tests pass across configuration building, native
provider resolution, permission policy, credentials, UI configuration,
and new-agent setup (including both Grok auth modes); token gates pass.
All fresh CI is green for this head: 56 successful checks/statuses and
two intentional skips ([run
36367065119](https://github.com/paperclipai/paperclip/actions/runs/36367065119)).
Greptile is 5/5 with no new findings. Grok already inherits the shared
`approve-all` default, so unattended setup requires no manual permission
change.

Runtime head `bb5a9307991f1ac567b781970ef11b39d518e19b` fixes a final
staging continuation failure before provider startup: explicit
replacement archived the old harness but left its failover backups
active, which caused `runner_harness_state_mismatch`. The regression
fails before the fix and passes after it; all eight adjacent
recovery-safety cases also pass. Old backups remain inspectable inside
the continuity archive. All fresh CI is green at this head ([run
36360839248](https://github.com/paperclipai/paperclip/actions/runs/36360839248)),
with a 5/5 review. One unrelated Cursor test timed out in the initial
server shard; the same-commit failed-job rerun passed, and both attempts
are retained. Staging deployment is confirmed healthy on this revision.
The controller image is
`ghcr.io/paperclipai/paperclip@sha256:6ad91c487910ccd2596ff7aed0a3a3ea5233d12b51b83cd6e1402237749b9673`.
The final browser-created staging task passed on this exact revision
with API authentication: context read → structured human question →
controller restart → answer submission → same native provider session
resumed → document saved → task Done. The two turns took approximately
119s and 77s. The actual write receipt was applied, and the saved
document has exactly one revision containing the selected answer and
requested marker. Usage and cost were not reported. [Controller image
build](https://github.com/paperclipai/paperclip/actions/runs/36360889243).

- Previous integration head `a44f7dbb6b6f77cd9ed893756ca453307f281e5f`:
all CI green (53 successful checks/statuses, two intentional skips),
including repository typecheck/build/tests, native Runner tests, browser
shards, and canary installation checks. [CI run
36358672529](https://github.com/paperclipai/paperclip/actions/runs/36358672529).
Greptile is 5/5 with no unresolved findings.
- Focused checks cover Grok credentials, executable admission, launcher
assets, provider-pack paths/permissions, workflow contracts, setup
defaults, CLI authorization, and subscription conflict recovery. All 39
protocol definitions validate. Final integration checks pass 124
catalog/evidence/cache tests and nine project-form tests; token gates
pass. Some local dependency checks could not load the stale installed
dependency tree; the corresponding fresh EC2 checks pass.
- Clean public npm installation passed on EC2 at
`8b172ebcf8e02e30662d830c00f3961e3bd459ec` ([run
36164964900](https://github.com/paperclipai/paperclip/actions/runs/36164964900)):
17 unified-version packages, lifecycle scripts enabled, built-in
launcher present, no separate Grok package or npm-downloaded binary,
missing prerequisite rejected, separately provisioned native executable
and command lease verified. No credentials or inference were used.
Subsequent changes preserve this npm asset layout.
- The immutable Daytona prerequisite image is
`ghcr.io/paperclipai/paperclip-daytona-runner@sha256:98957d5be0ac774d086b6402b5849e8e6356fec70fb8c09fca6eb4ed6de918e0`,
built from `5a2db471f3ddabe77f9f80e76ed27f996cb97fba`. The previous
Cloud controller image was
`ghcr.io/paperclipai/paperclip@sha256:fd914e1ab1e45f741e8e078ff452d16f082d7ac05f9b4b3506d3a3c64150d204`,
built from `a44f7dbb6b6f77cd9ed893756ca453307f281e5f`; it is superseded
by the latest image above. Its EC2 build verified provider-pack access
under an unrelated unprivileged UID.
- Browser staging at `40f898bc4cba73c1dff4e6344a3983ba0fb247ef` passed
full Grok onboarding with the correct `grok-4.7` model, saved credential
delivery, and pinned Daytona execution. A browser-created task read
context and asked the structured human question. After a controller
restart, answering the persisted question resumed the same native
provider session, saved the requested document, and completed the task.
Actual tool outcomes and durable state agree: one question and one
document revision. The two successful turns took 42.7s and 63.1s; usage
and cost were not reported.
- Restricted policy returned the expected `approval_required` outcome.
Functional staging tests explicitly selected `approve-all`; controller
authorization and governed approvals remain enforced. Temporary board
CLI access was revoked and verified rejected (HTTP 401), and the
disposable onboarding agent was paused.

Failures remain retained: the pre-fix continuation failure (its task
remains blocked; the passing final task is fresh), the original Cloud
provider-pack permission failure, the expected restricted-policy denial,
the superseded npm staging failure, and an earlier monolithic CI
infrastructure timeout. Browser CI exposed a project alias/form race;
the final stack uses master's stronger draft-preservation fix and all
browser shards pass. Historical full subscription/API protocol and
Product rosters retain their original source revisions and do not
qualify this packaging revision. No local Docker or Rust build was used.

## Risks

The branch includes master’s draft-preservation fix for project URL
aliases. It keeps the same project’s edit form mounted and clears prior
data when the project or company changes.

Custom sandboxes and local execution hosts must provision the pinned
binary before Grok starts. Missing, changed, unsupported-platform, and
symlinked executables fail admission. The new builtin profile cannot
resume sessions created with the former private-package profile.
Existing Claude/Codex npm bridge profiles retain their package pins.
Grok restricted modes preserve the selected policy but cannot
automatically admit Paperclip calls: ACP permission metadata does not
independently bind tool authority, so those calls stop with
`approval_required`. New Grok configurations default to `approve-all`,
including API configurations that omit the mode. Existing explicitly
restricted configurations remain restricted; controller authorization
and governed approvals remain enforced.

## Model Used

OpenAI GPT-6 through Codex, with tool use and code execution. The exact
serving model identifier and context-window size are not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 14:54:44 -05:00
DottaandPaperclip 2f585ef26a fix(ui): preserve newer drafts after repeated receipt cleanup (#14332)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The task composer keeps unsent text across page reloads.
> - A server receipt confirms a submitted message after a reload.
> - Replayed cleanup can apply the old text offset to a newer draft
twice.
> - This pull request reconciles each confirmed attempt once and
preserves both tabs’ unsent intent.
> - The full newer draft stays available for the next send.

## Linked Issues or Issue Description

**What happened?**

Reloading the classic composer while a save was pending could truncate a
newer draft. Replayed receipt cleanup changed `A newer draft written
while delivery was pending.` into `nding.`. The storage helper rejected
the duplicate settlement, but the effect still changed the editor and
its body reference.

**Expected behavior**

A receipt settles its retained submission once. Later cleanup must
preserve the newer draft in both the editor and browser storage.

**Steps to reproduce**

1. Send a comment and hold its HTTP response after the server accepts
it.
2. Type a newer draft, then reload the page.
3. Restore the matching receipt under React StrictMode.
4. Inspect the newer draft after effect replay and unmount.

The deterministic regression reproduces this on current master. The
existing browser test exposed the issue during #14329 verification.
Related draft persistence code came from #13338. A search found no
separate fix for this duplicate settlement.

## What Changed

- Reconcile each confirmed attempt once in memory so duplicate effects
cannot trim its newer draft again.
- Unlock confirmed drafts when storage writes fail. Never acquire
another tab’s pending receipt or assume its text is newer. Keep a
conflicting local draft and its attachments in tab-scoped session
storage while preserving the shared draft unchanged.
- Add component regressions for StrictMode replay, exact editor and
stored bytes, foreign receipts, attachment-only differences, reload
recovery, unavailable storage, and task navigation before autosave. A
brief notice explains when this tab has a separate draft.

## Verification

- RED: the new test received `nding.` instead of the full newer draft
before the fix.
- Review RED: three cases reproduced a locked composer after failed
storage writes or another tab's settlement. A further negative case
preserves a newer retained attempt and its attachments.
- Cross-tab review RED: four deterministic cases reproduced lost local
text, foreign receipt takeover, attachment loss when text matched, and
overwriting newer stored text.
- GREEN: 118 tests across `IssueChatThread`, `composer-draft`, and
`comment-submit-draft`, including same-mounted A → B → A navigation and
a full recovery-storage failure. Two further lifecycle RED tests verify
finishing a recovery returns to the shared draft on a later visit while
continued typing and attachments retain recovery.
- Both existing browser reload cases passed against a fresh server and
database on exact head `ceb80aca77fc8cc0f813c328ba87025b3e1a2222` (42.5
seconds), with classic mode enabled and disabled. The browser flow
verifies one original comment, a preserved draft, and a successful
second send.
- UI typecheck, token gates, and the shipped static UI build passed. The
initial typecheck required the fresh worktree's plugin SDK build; the
retry passed after that dependency built.
- The session-only failure mock also preserves localStorage on platforms
where both share the Storage prototype; this fixes the Linux workspace
test failure.
- All 56 checks passed on `4c30dcccc4fe5b90dcd08dd1d90475ea89e4a376`;
Greptile is 5/5 with zero open review threads. An unrelated runner
baseline scan hit its existing 100 ms deadline once; its isolated test
and the single failed CI job passed on retry without source changes.
This four-file UI fix does not change server or database code.


- Final alternate-staging acceptance passed on deployed source
`d884e1ab046cc76004e35e6091e9e6e2c918c9eb`, with exact health checked
before and after. Three real browser cases used shipped static UI and
real API saves: reload during an accepted-but-unacknowledged save in
both composer variants, plus two classic tabs with different drafts. The
test deliberately held only its own accepted POST acknowledgement and
temporarily withheld its own GET receipt from one tab to reproduce
settlement ordering. Both drafts survived independent reloads, the
shared draft remained intact after the other tab sent, and finishing
recovery returned to the shared draft on a later visit. Exact request
IDs/counts and server comment bytes passed; no provider was invoked.
Screenshots preserve the live drafts before fixture cleanup.
- An ordinary authenticated Chrome/CUA visit independently showed the
exact original and newer comments; an additional unsent draft survived
reload and was then cleared. All three isolated fixtures are complete,
browser contexts are closed, and original user tasks were untouched.

## Risks

- Reconciliation uses the exact request ID and draft key. Conflicting
drafts stay separate; the tab recovery survives same-tab reload through
session storage and does not outlive the browser tab. If recovery
storage is unavailable, the editor keeps its in-memory text and
explicitly warns the user to copy it before leaving.
- Attachment selections remain bounded to 20 per draft; if combining
equal-text snapshots would exceed that limit, each original selection
stays in its respective draft.
- No API, schema, or migration changes. The status notice uses existing
design tokens.

## Model Used

OpenAI Codex, GPT-6, with tool use and code execution. The runtime does
not expose the exact deployment variant or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 14:03:36 -05:00
DottaandPaperclip f1a394bd30 feat(runner): add Grok Build through native ACP (#13882)
## Thinking Path

> - Paperclip manages AI agents and governs their work.
> - Its native runner uses structured provider protocols for sessions
and tools.
> - Grok Build supports ACP over stdio, but the runner did not expose
it.
> - Native execution requires company-scoped credentials, verified
identities, and permission gates.
> - This change adds Grok through ACPX for local and Daytona execution.
> - Subscription login and explicit API-key execution have separate
credential paths.
> - Qualification grades real tool outcomes, durable state, and browser
workflows.

## Linked Issues or Issue Description

Refs #13845, #13847, #13850, #13878, #13901, #13973, #13977, #13979.

Add **Grok Build** to `paperclip_runner` with `provider: "acpx"`,
`acpxAgent: "grok"`, and model `grok-4.7`. Existing legacy Grok agents
keep their adapter. Merge the three companion fixes (#13973, #13977,
#13979) before treating the integrated Product qualification as deployed
behavior.

## What Changed

- Synchronize shared, TypeScript, Rust, server, validation, and UI
provider contracts.
- Run Grok native ACP stdio through ACPX and the authenticated Paperclip
MCP bridge. Verify the pinned executable and exact ACP model identity.
- Prefer company subscription login. Support an explicit company-secret
API key without automatic paid fallback. Fence refresh and copyback to
the same account and remove private runtime credentials after
containment.
- Preserve selected permissions, cancellation, durable session identity,
resume, and restart recovery. Keep unsupported steering and goals
unavailable. Preserve missing usage and cost as unknown.
- Package checksum-verified Grok Build 1.0.13 for Daytona with an
immutable, signed image built on EC2.
- Add deterministic admission, protocol, permissions, identity,
credential, failure, and cleanup checks. Add the maintained 39-case
protocol roster and separate subscription/API Product profiles.
- Fix live-test findings in reasoning events, reloads, idle-owner
retirement, credential-home cleanup, expired-login model discovery,
launcher pinning, and rerun evidence selection.
- Align control-plane state readers with the transport's 64 MiB bound
while retaining identity, ownership, lifecycle, and size rejection
checks.
- Stabilize two asynchronous CI assertions while retaining actual
outcome and filesystem-evidence checks.

## Verification

Current integration head `f114948376056fe0b6b34c1496ae8667b59daa63`
includes master `3447609d2247e75e55d91493dda91a608364f672` (2026-09-28).
Two master advances during verification overlapped the eval catalog; the
final merge preserves Grok qualification, completion updates, and
bounded API-response reading in all 348 cells. All 77 focused
catalog/eval/workflow tests pass. Both native stack layers (#14397) are
mergeable, and both exact-head Greptile reviews are 5/5 with successful
security scans and no unresolved review threads. All current-head CI is
green: 56 successful checks/statuses and four intentional skips ([CI
run](https://github.com/paperclipai/paperclip/actions/runs/36447097232)).
Trunk code-owner requirements remain enforced. The review summary’s
non-blocking saved-asset offset classification note concerns code
already merged in #14301; those runtime files are identical to master
and outside this stack’s diff. Historical live evidence below retains
its original source revisions.


Earlier integration checkpoint:
`24fc9b94ca0afb21ccdc8d26dbb2e4b258ad72cb`. Refreshed against master
`0f14d2612`, preserving Grok qualification alongside the new accounting
and lifecycle suites. All 124 focused catalog, evidence, and
service-worker checks pass. The current base workflow includes the
explicitly selected public-install verification lane; follow-up #14024
supplies its verifier script. CI at that earlier checkpoint was green
(56 successful checks/statuses, four intentional skips), and the review
is 5/5 with no unresolved findings. Prior feature CI at
`fd73f0a9b1ecdf4094685054028df71739ddc3e1` passed ([run
36148259902](https://github.com/paperclipai/paperclip/actions/runs/36148259902));
that is historical evidence, not a current-head result.

Paid Product measurements use frozen integrated source
`2d939a92b21dcaf5c77c88b54d96784d2ddd0699`, which combines the feature
with #13973, #13977, and #13979. That source passed all 52 CI checks and
clean 5/5 review. Later master syncs incorporate upstream changes. Their
checks remain separate from these pinned live measurements.

| Check | Result and source-pinned report |
| --- | --- |
| Subscription protocol roster | [39/39 first attempts; 206
assertions](https://d1p6rlowie26tp.cloudfront.net/runner-protocol-evals/campaigns/gha-36046839612-1/index.html),
runtime `bc6833f7`, evals `92bb4b8c` |
| API protocol roster | [39/39 first attempts; 206
assertions](https://d1p6rlowie26tp.cloudfront.net/runner-protocol-evals/campaigns/gha-35926577007-1/index.html),
runtime `4a1061c8`, evals `3213dbec` |
| Subscription full Product matrix | [16/16 first attempts; 144
assertions; cleanup
passed](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36096908572-1/index.html),
source `2d939a92` |
| Subscription core repetitions | 18/18: tool use, planning approval,
and Stop/resume each passed three times in local and Daytona profiles.
The full matrix contains repetition one; [repeat
two](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36104551060-1/index.html)
and [repeat
three](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36108100404-1/index.html)
each passed 6/6. Total: 28 unique subscription attempts at `2d939a92`. |
| API smoke and question continuation | [4/4 first attempts; cleanup
passed](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36147315401-1/index.html),
both environments at `2d939a92` |
| Historical API Product coverage | [16/16 full
matrix](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35875144860-1/index.html)
and 18/18 core repetitions at `4a1061c8`; retained as measurements of
that revision |
| Native Daytona proof | Three subscription and three API
MCP/permissions/resume runs passed at `bc6833f7`. Three expired-login
admission and fenced refresh checks passed without inference. All test
sandboxes were removed. |
| Inspectable artifacts and UI | Current-source screenshots verify
planning approval, direct Ask completion, question continuation after
controller restart, and two downloadable project revisions. The project
downloads pass 12 and 18 tests; all 40 independent artifact oracle
checks pass. |
| Provider-free checks | 116 eval-validator tests, 39 Grok definitions,
and 359 enabled/external campaign cells pass. Continuation regressions
above 2 MiB and 16 MiB failed before their fixes; 32 focused
recovery/ownership/size checks pass. |

The 32 unique current-source Product attempts have no failures, retries,
or skipped cells, and all cleanup checks pass. Whole-workflow timing,
model identity, image and provider-pack provenance, attempts, and
accounting coverage are retained in the canonical reports. The report
publisher's conservative `complete=false` flag is preserved; independent
audits verify the exact selected source catalog and immutable result
rows.

Pins: Grok Build `1.0.13 (5e9a58528b76)`, ACPX `0.13.1`, ACP model
`grok-4.7`. Linux binary SHA-256:
`edf79521581bb5e6b95abef848491a6a742e860da3e237ebe86a280d30dce4c1`.
Launcher SHA-256:
`f0b698395a3704ed2ffaf84ea19bdb20c36c8a0a70b7c629c7b6ffe144e59e55`.
Image:
`ghcr.io/paperclipai/paperclip-daytona-runner@sha256:76b24edfd850219e949418b19e4ceba690e84d51d199ade426e484953329b5e9`.
Image build source is `4196a4cd`, recorded separately from application
source `2d939a92`; each campaign verifies the image signature and
provider pack.

Original failed campaigns remain available: [continuation
bound](https://github.com/paperclipai/paperclip/actions/runs/36057718059),
[scheduler/event
capture](https://github.com/paperclipai/paperclip/actions/runs/36071063537),
and [startup cleanup plus EC2
interruption](https://github.com/paperclipai/paperclip/actions/runs/36080870743).
They retain their original grades. No Docker or Rust builds ran on the
developer laptop for these follow-ups.

## Risks

Merge packaging follow-up #14024 with this base before public release.
The follow-up replaces the private Grok bridge package with a built-in
launcher and makes the native binary an explicit sandbox prerequisite.

Three separate, reviewed fixes are part of the tested integrated
behavior: #13973 serializes task-run admission; #13977 captures complete
event evidence; #13979 durably reconciles failed Daytona creation. Each
has green CI and clean 5/5 review. Failed-create recovery has 277 plugin
tests, 92 SDK tests, host-runtime recovery tests, and a real Daytona
lost-deletion-receipt proof. The live proof uses a private file for
journal persistence; database durability is covered by host tests.
Worker death before delivery of a failure envelope remains outside that
recovery mechanism.

Subscription fixtures stage an authorized company login; interactive
browser sign-in is not qualified. Local Product profiles ran on EC2
Linux. The temporary subscription credential was removed from the
protected GitHub environment after all subscription audits, with absence
verified. Runtime homes and refresh copyback remain ownership-fenced.

Protocol results remain pinned to their original revisions; they are not
relabeled as tests of the latest feature commit. New binary/model
versions require qualification. Missing token usage and model cost
remain unknown; runtime estimates do not establish a full bill.
Automatic paid Grok scheduling remains disabled pending separate
reviewed enablement. The 64 MiB bound can increase memory use for
verbose sessions, and larger files still fail closed. No automatic
legacy-agent migration occurs.

## Model Used

OpenAI GPT-6 through Codex, with tool use and code execution. The exact
serving model identifier and context-window size are not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 13:54:36 -05:00
DottaandPaperclip d9d2147171 fix(auth): keep Cloud tenants on the Cloud sign-in flow (#14407)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Cloud owns human identity and passes a verified identity to each
tenant.
> - The tenant can report no session while the Cloud session is still
valid.
> - The access gate and direct `/auth` route then show the instance
password form.
> - This pull request sends those users through the configured Cloud
entry endpoint.
> - Cloud can renew the tenant session or show its login page, then
return to the original task.

## Linked Issues or Issue Description

**What happened?**

A Cloud tenant can display the self-hosted email/password form after an
instance session check returns no session. This gives Cloud users the
wrong login method.

**Expected behavior**

An active Cloud session renews tenant access automatically. A signed-out
user signs in through Cloud. Staging and production use their own
configured Cloud origins. Self-hosted instances keep their instance
login form.

**Steps to reproduce**

1. Open a Cloud tenant task or an `/auth?next=...` link.
2. Keep the Cloud session active but make the instance session check
return 401.
3. Observe the instance password form instead of Cloud session recovery.

**Deployment mode**

Cloud-managed authenticated instances. No database or server API
changes.

Searched related authentication PRs. Native self-hosted OIDC support in
#10411 is a separate feature; this change uses the existing Cloud entry
contract.

## What Changed

- Wait for deployment metadata before showing an instance login form.
- Use the health response's Cloud origin and stack slug for session
recovery.
- Preserve the tenant path, query, and fragment. Reject external and
recursive login return targets.
- Limit automatic recovery per tab. Show a manual Cloud retry after
failed recovery. Show service failures as errors.
- Keep self-hosted login and local trusted access. Add focused tests,
browser regressions, deployment documentation, and an unavailable-state
design example.

## Verification

- `pnpm -r typecheck` and `pnpm build` passed.
- UI typecheck and `pnpm check:token-gates` passed after the final UI
edits.
- All UI tests passed: 639 files, 6,777 tests.
- 63 focused Vitest tests passed across Auth, CloudAccessGate, Cloud
links, and recovery coordination.
- `pnpm exec playwright test --config tests/e2e/playwright.config.ts
tests/e2e/cloud-auth.spec.ts`: 5 passed. The tests use the real tenant
UI and database with a simulated Cloud HTTP endpoint. They cover both
Cloud origins, direct auth/task links, no password-form flash, preserved
URLs, reload, and self-hosted login.
- Hands-on browser test used the real Cloud gateway and a fresh tenant
build with disposable local data. Active Cloud session plus a forced
missing instance session returned to the task. An expired tenant cookie
also renewed automatically and returned to the task. Removing both
sessions reached the real Cloud email/social login UI. Persistent
failure stopped at the retry screen; retry succeeded after removing the
injected fault. The fixture used a loopback transport adapter and a
simulated signed-out OIDC issuer. No production session or deployment
was changed.
- The default local browser startup hit the host's embedded PostgreSQL
resource limit. The passing run used a separate disposable database on
the test PostgreSQL process.
- The full local `pnpm test:run` sweep was stopped after about 31
minutes once CI completed the full suite. It had reported 33 failures in
the unchanged runner API unit/integration files; both files pass in
isolation (1,749 + 28 tests). The local sweep did not reach the later
workspace/serialized groups. CI completed all of those groups
successfully.
- Greptile reviewed commit `d603fd4e39455de44da9dae81b72197096c0e1e8` at
5/5 with its only thread resolved. All CI gates are green on this
commit, including all general/serialized server groups, workspace tests,
Runner checks, typecheck, build, and all eight browser shards
([run](https://github.com/paperclipai/paperclip/actions/runs/36446230697)).

## Risks

- Recovery depends on valid Cloud origin and stack metadata. Incomplete
metadata shows an unavailable message instead of a password form.
- Browsers with session storage disabled use the manual Cloud link,
since automatic retries cannot be bounded across documents.
- The external identity provider's email/social login was not completed
in this local test. Existing Cloud authentication owns that flow.
- No migration, credential format, membership rule, or production
deployment changes.

## Model Used

OpenAI GPT-6 through Codex. The exact served variant and context-window
limit are not exposed in this session. Used reasoning, repository tools,
code execution, and browser testing.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 13:45:27 -05:00
DottaandPaperclip 270afd2fb8 feat(ui): show running commit in staging account menu (#14410)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The account menu shows the signed-in user's identity.
> - Staging users need to know which server commit is running after a
deploy.
> - The health endpoint already returns that commit, but the menu does
not show it.
> - This pull request adds the short commit below the email on staging
hosts.
> - Users can open the menu to check a deploy without opening deployment
tools.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

The account menu on staging instances.

**Current behavior**

The menu shows the user's name and email. It does not show the running
server commit.

**Proposed behavior**

On `*.staging.paperclip.app`, show `SHA 8751e2d` below the email. Use
the current `/api/health` commit. Show the full SHA on hover and link to
the commit on GitHub. Refresh the health query when the menu opens. Hide
the label on other hosts and when commit metadata is unavailable.

**Reason and benefit**

A user can confirm which commit a staging instance runs after an
automatic deploy.

**Breaking changes**

None. The server already returns the commit field.

Refs #14060 for related account-menu work. This change adds deployment
information only.

## What Changed

- Add the existing health response commit field to the UI type.
- Share the staging host check and a separate health query across both
account-menu variants. A menu refresh failure leaves the access gate
health state unchanged.
- Show a short SHA below the email. Link to the full commit on GitHub
and include the full SHA in its accessible name and hover title.
- Document the staging label and cover staging hosts, other hosts,
missing metadata, refresh on reopen, and request failure isolation.

## Verification

- `pnpm exec vitest run ui/src/components/SidebarAccountMenu.test.tsx` —
24 tests passed.
- `pnpm check:token-gates` — passed.
- `pnpm -r typecheck` — passed.
- `pnpm build` — passed. The UI build and typecheck also passed again
after review fixes.
- Full test matrix — passed on the latest commit in
[CI](https://github.com/paperclipai/paperclip/actions/runs/36447810742).
The local `pnpm test:run` was stopped before completion after CI
finished the same suites. The focused local tests, full local typecheck,
and full local build passed.
- Browser shard 6 passed on one rerun. Its first attempt lost part of
the draft text in the existing attachment-receipt reload test. No code
changed for the rerun.
- Rendered the real account menu in a local browser fixture with a
staging hostname condition and mocked health response. Confirmed the SHA
fits below the email in the dark menu.
- Greptile review — 5/5, all review threads resolved.
- Manual check after deployment: open the menu on a staging host and
compare the SHA with `/api/health`. Open the menu again after a deploy
to refresh it. Confirm the label is absent on production and localhost.

## Risks

- Low risk. Each menu opening on staging can make one additional health
request.
- The host check applies to `*.staging.paperclip.app`. Other staging
domains will need an explicit update.
- The label identifies the running server commit. It can briefly show
cached data while the request completes.

## Model Used

OpenAI Codex, GPT-6. The exact model variant and context window are not
exposed in this session. Used code execution, repository inspection, and
browser inspection.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#123` / `Refs #123` OR (b) described the issue in-PR following the
relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 13:43:30 -05:00
DottaandPaperclip 14795136f5 fix(runtime): finalize and recover sandbox workspace exports safely (#14402)
Serialize native workspace finalization, validate streamed archives within bounded limits, and quietly recover unsafe exports from saved results. Preserve exact allocations for exhausted transient failures and provide export-only retry without rerunning the provider.

Consolidates #14314, #14315, #14329, and #14334 while preserving the already-merged finalization label changes.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-28 10:57:32 -05:00
DottaandPaperclip 8751e2de46 fix(ui): distinguish finalization recovery from live observation (#14326)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The task board shows which recovery actions are active.
> - A native run can stop while a person must repair its workspace.
> - The board previously called that state “Recovery in progress.”
> - The label implied that work would continue without operator action.
> - This pull request derives the label from the recovery owner and live
continuation.
> - Operators can distinguish scheduled recovery from a repair that
needs attention.

## Linked Issues or Issue Description

**What happened?**
A blocked task showed “Recovery in progress” after native finalization
stopped and no automatic continuation remained.

**Expected behavior**
Show “Recovery needed” for an idle board repair or when no live recovery
path exists. Show “Recovery in progress” while the recorded continuation
can run, including an explicitly admitted export whose exact callback is
executing even if the old recovery action remains board-owned.

**Steps to reproduce**
1. Complete a native run whose workspace export cannot be recovered
automatically.
2. Inspect the task recovery action and badge.
3. Compare the board-owned action with the old “Recovery in progress”
label.

Related work: #14314 serializes native workspace finalization and fences
stale recovery outcomes. This change reports the recovery action that
currently owns the task.

## What Changed

- Show “Recovery needed” for board-owned active-run recovery unless the
exact native export callback is positively verified as executing.
- Require the native continuation run and a live or future continuation
before showing progress.
- Project native activity from the exact company, source issue, and run.
Include active workspace export while the original heartbeat remains
failed.
- Require an executing callback before using a running export row as
evidence. Preserve activity for long exports and clear it when the
callback joins.
- Give native resume its own card explanation. Preserve ordinary
watchdog observation behavior.
- Remove the redundant ownership sentence from all six recovery-card
explanations that used it.
- Document the labels and add regression cases for stopped, scheduled,
and active recovery.

## Verification

- Copy-only follow-up (`c68aef04c`): all 167 focused recovery UI tests
and token gates pass. No UI occurrence of the removed sentence remains.
`pnpm -r typecheck`, `pnpm build`, and current-head CI pass (54
successful checks, two optional Storybook checks skipped). Greptile is
5/5 with no open review threads. The duplicate local `pnpm test:run` was
stopped after the full CI suite passed; it did not complete locally.
- Original regressions: seven failures before the change, then 52
focused cases pass.
- Review regressions: seven UI failures and nine database failures
before the follow-up. All 75 database/API recovery tests, 167 UI tests,
and 18 workspace lifecycle/finalizer tests pass. A further three RED
cases cover explicit board retry activity; one RED case rejects orphaned
running export rows after controller loss. Wrong company, issue, run,
service, phase, and completed-operation cases remain inactive.
- Final recursive typecheck, production build, token gates, and complete
local suite coverage pass. Embedded PostgreSQL startup/socket failures
passed in isolated retries with the canonical test environment; no
expected behavior was weakened. The recovery regression added during the
earlier full run passed in its final complete 75-case file.
- Before the copy-only follow-up, all 56 CI checks passed on
`c2f84cd89c905cda85c53aaf5bb83b7250900fe6`; Greptile is 5/5 with no
unresolved review threads. The final native-activity staging repeat
passed on integrated source `2bedd0f23bf4698b1f8b818f6796900647030427`.
- Verified on a separate staging instance: a real failed Daytona
workspace export retains its board-owned repair action and displays
“Recovery needed” in the task list. The repair card remains actionable
without starting another provider turn.

- Real staged export-only repair: the actual task list showed “Recovery
in progress” while the original run had a positively identified running
export operation, then Done after exact copyback of all 20,000
nonce-bound files. The accepted result and full provider
session/turn/terminal envelopes remained unchanged. All 17 independent
final checks passed; the browser downloaded the exact 19-byte result.
The separate fixture was cleaned up with independent provider-absence
verification. A control transport process restarted during repair; it
did not submit another provider turn.

- Final deployed-source repeat on
`d884e1ab046cc76004e35e6091e9e6e2c918c9eb`: explicit per-turn ephemeral
Daytona allocation, actual browser export repair, and a saved full
activity projection referencing the exact executing export operation
with no scheduled retry. The task list showed recovery in progress, then
Done; all 21 final checks passed, including 20,000 exact host files,
unchanged provider provenance, and provider deletion only after
committed copyback. The downloaded 19-byte result matched independently.
This integrates #14334; no source change was required here.

## Risks

- The label depends on the persisted recovery action. A separate runtime
defect can still stop work; this change makes that condition visible.
- Future recovery kinds must supply a valid continuation path before
they can display progress.

## Model Used

OpenAI Codex, based on GPT-6, with code execution, browser testing, and
subagent tool use. The runtime does not expose an exact serving model
variant or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 09:33:39 -05:00
DottaandPaperclip 890d11137f fix(ui): register artifact tabs without opening the panel (#14193)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Tasks keep agent outputs in the Artifacts tab.
> - An output can arrive while the user writes a message or reads a
document.
> - Opening the side panel on arrival interrupts that work, especially
on mobile.
> - This pull request adds the Artifacts tab without opening the panel
or changing the selected tab.
> - Users can open their outputs when they choose.

## Linked Issues or Issue Description

**What happened?**
New agent outputs opened the task side panel or mobile drawer. An
arrival could also replace the selected document or workspace file.
Existing outputs did not always register an Artifacts tab.

**Expected behavior**
Register one Artifacts tab for existing and new outputs. Keep a closed
panel closed. Preserve composer focus, the selected tab, and document or
file links.

**Steps to reproduce**
Open a task from the inbox. Close its side panel. Enter a message draft.
Create an agent output in that task. The panel must stay closed and the
draft must keep focus. Open the panel to see the Artifacts tab. Repeat
on a mobile viewport.

**Paperclip version or commit**
Base commit: 0f14d2612.

Related work: #11226 and #11551.

## What Changed

- Register existing outputs and later arrivals without opening the panel
or selecting Artifacts.
- Keep open documents, workspace-file links, and the tab launcher
unchanged.
- Handle each document deep-link request once so query refreshes
preserve later manual selection.
- Deduplicate attachment and work-product arrivals. Preserve dismissed
tabs across repeated refreshes.
- Add desktop and mobile browser regression tests. Update artifact
presentation documentation.

## Verification

- The closed-panel regression failed before the fix in unit and
real-browser tests.
- All 158 focused UI tests pass, including the original deep-link cases.
- UI typecheck and token gates pass on this branch. Full typecheck and
production build passed on the passive-arrival candidate before the
existing PR integration.
- The two local desktop/mobile browser cases pass against real
API-created artifacts.
- Desktop and mobile staging checks pass on the combined staging
candidate. Artifact arrival preserved a closed pane, draft text, and
composer focus. Explicitly opening the pane showed the Artifacts tab.
- All 56 current-head check contexts are successful or intentionally
skipped at `17ad904455b9378552f07a6f6e51402c6d164688`, including full
typecheck, test shards, build, and browser suites. Greptile is 5/5 on
that commit with no unresolved review threads.
- The interrupted local broad validation was resumed; the remaining
serialized 64 files and 1,035 tests pass. Local database startup
failures passed after stale test resources were released.

## Risks

- Outputs no longer reveal the panel automatically. Users open the panel
to view them.
- The document request guard must still allow a new explicit deep link.
The regression tests cover this case.
- There are no API or database changes.

## Model Used

- OpenAI Codex, GPT-6, with code editing, shell tools, GitHub tools, and
browser verification. The exact serving model ID and context-window size
are not exposed in this session. Earlier implementation model metadata
is not available.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 08:38:00 -05:00
DottaandPaperclip 7e0f051d3b Align task property icons with assignee avatars (#13319)
## Thinking Path

> - Paperclip is the open source app that people use to manage AI agents
for work.
> - The task properties panel helps operators inspect and update task
data.
> - The status, assignee, and project rows used different leading icon
sizes.
> - The mixed sizes made the rows look uneven.
> - The assignee avatar already provides the correct 24px visual size.
> - This pull request sets the status and project visuals to that same
size.
> - The benefit is a tidy and consistent properties panel.

## Linked Issues or Issue Description

**What happened?**

The task properties panel showed a 12px status icon, a 24px assignee
avatar, and a 16px project tile.

**Expected behavior**

The status, assignee, and project rows must use the same 24px leading
visual size.

**Steps to reproduce**

1. Open a task.
2. Open the Properties panel.
3. Compare the Status, Assignee, and Project rows.

## What Changed

- Keep the Status glyph at 16px and use spacing to give it the 24px
avatar footprint.
- Set the Project row tile to the 24px small tile size.
- Add a regression check for the Status row size.

## Verification

- `pnpm exec vitest run ui/src/components/IssueProperties.test.tsx` (72
tests passed)
- `pnpm check:token-gates` (all gates clean)
- `pnpm --filter @paperclipai/ui typecheck` (passed)
- Full repository typecheck and build reached the Rust runner step, but
this environment does not contain `cargo`.
- The full test command was stopped after unrelated chat integration
tests did not finish. The focused UI suite passed.

## Risks

- Low risk. This change only changes visual sizes in three property
rows.
- The wider status footprint can use slightly more horizontal space in a
narrow panel.

> This fix does not overlap with planned core work in `ROADMAP.md`.

## Model Used

- OpenAI Codex, GPT-5.6, tool use and code execution enabled.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-27 11:35:39 +00:00
DottaandPaperclip d423fc4411 fix(ui): use large mobile selector modals (#14250)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The new task dialog lets an operator assign work on a phone.
> - The mobile picker used fixed coordinates inside the Radix wrapper.
> - iOS Safari already makes fixed coordinates relative to the visual
viewport.
> - The old code added the visual viewport offset a second time.
> - This pull request removes the second offset and gives the mobile
wrapper full viewport geometry.
> - The benefit is a visible picker and search field when the iOS
keyboard opens.

## Linked Issues or Issue Description

**What happened?**

The assignee and project pickers could move outside the visible screen
in the new task dialog on iOS.

**What did you expect to happen?**

The picker and its search field must stay visible above the bottom edge
and the software keyboard.

**Steps to reproduce**

Open the new task dialog on an iPhone. Select Assignee or Project. The
picker can move outside the visual viewport when Safari pans the page.

**Paperclip version or commit**

The problem exists after the change in PR #13343.

**Deployment mode**

The problem affects the board UI in local and authenticated modes.

Refs #13343

## What Changed

- Use visual viewport local coordinates for the new task dialog.
- Give the mobile Radix popper wrapper full viewport geometry.
- Render both shared selector components as large, titled mobile modals.
- Keep desktop selectors as anchored popovers.
- Cover the new-task assignee, new-task project, composer assignee, and
generic searchable selector in Storybook.
- Add iPhone-sized Storybook states for the assignee and project
pickers.
- Update viewport tests for the corrected coordinate model.

## Verification

- `pnpm --filter @paperclipai/ui exec vitest run
src/components/NewIssueDialog.test.tsx
src/components/InlineEntitySelector.test.tsx`
- `pnpm check:token-gates`
- Playwright used the iPhone 14 device profile. All mobile modals were
inside the 390 by 664 CSS viewport.
- Each mobile modal was `x=16, y=16, width=358, height=632`.
- A desktop check kept the anchored popover at `width=320, height=211`.

## Risks

- Low risk. The CSS only changes narrow mobile viewports.
- The full-screen popper wrapper does not receive pointer events. The
picker still receives pointer events.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. See `CONTRIBUTING.md`.

## Model Used

- OpenAI GPT-5.6, Codex agent with tool use and code execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal task
ID
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant Storybook documentation to reflect my
changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-27 11:33:24 +00:00
DottaandPaperclip 74508357fa fix(ui): keep mobile task pickers in view (#14249)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The new task dialog lets an operator assign work on a phone.
> - The mobile picker used fixed coordinates inside the Radix wrapper.
> - iOS Safari already makes fixed coordinates relative to the visual
viewport.
> - The old code added the visual viewport offset a second time.
> - This pull request removes the second offset and gives the mobile
wrapper full viewport geometry.
> - The benefit is a visible picker and search field when the iOS
keyboard opens.

## Linked Issues or Issue Description

**What happened?**

The assignee and project pickers could move outside the visible screen
in the new task dialog on iOS.

**What did you expect to happen?**

The picker and its search field must stay visible above the bottom edge
and the software keyboard.

**Steps to reproduce**

Open the new task dialog on an iPhone. Select Assignee or Project. The
picker can move outside the visual viewport when Safari pans the page.

**Paperclip version or commit**

The problem exists after the change in PR #13343.

**Deployment mode**

The problem affects the board UI in local and authenticated modes.

Refs #13343

## What Changed

- Use visual viewport local coordinates for the new task dialog.
- Give the mobile Radix popper wrapper full viewport geometry.
- Keep the mobile picker at the visible bottom edge.
- Add iPhone-sized Storybook states for the assignee and project
pickers.
- Update viewport tests for the corrected coordinate model.

## Verification

- `pnpm --filter @paperclipai/ui exec vitest run
src/components/NewIssueDialog.test.tsx
src/components/InlineEntitySelector.test.tsx`
- `pnpm check:token-gates`
- Playwright used the iPhone 14 device profile. Both picker boxes were
inside the 390 by 664 CSS viewport.
- The assignee picker box was `x=16, y=366, width=358, height=282`.
- The project picker box was `x=16, y=366, width=358, height=282`.

## Risks

- Low risk. The CSS only changes narrow mobile viewports.
- The full-screen popper wrapper does not receive pointer events. The
picker still receives pointer events.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. See `CONTRIBUTING.md`.

## Model Used

- OpenAI GPT-5.6, Codex agent with tool use and code execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal task
ID
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant Storybook documentation to reflect my
changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-27 11:02:15 +00:00
DottaandCodie 01d9a12185 fix: make keyboard shortcut enablement a personal preference (#14141)
Store keyboard shortcut enablement per user and expose it in Profile settings.

Co-Authored-By: Codie <Codie@users.noreply.github.com>
2026-09-26 11:50:04 -05:00
Devin FoleyandPaperclip 7f3c06dac4 refactor(ui): remove the legacy Cloud organization switcher (#14061)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Its sidebar provides navigation between organizations.
> - The generic organization switcher slot lets installed plugins own
that navigation.
> - Both Core UI shells still contain the old Cloud portfolio menu.
> - Cloud now uses its private Account plugin for this menu.
> - This pull request removes the duplicate Cloud menu and keeps the
built-in company menu.
> - This reduces Cloud-specific code without changing the plugin host
contract.

## Linked Issues or Issue Description

Refs #13832 and #13854. Related #14060 changes the account popup, not
this organization switcher.

**What existing behavior does this improve?**

Organization navigation in both sidebar shells.

**Current behavior**

Core retains Cloud portfolio fetching, stack rows and stack-entry links
behind the plugin replacement slot.

**Proposed behavior**

The installed switcher plugin owns Cloud navigation. Core lists local
companies when no usable replacement exists. The Members-page Cloud
invitation action keeps its existing portfolio API; it is a live caller,
not an old-image fallback.

## What Changed

- Remove Cloud portfolio queries, stack rendering and Cloud
creation/entry branches from both built-in menus.
- Remove the unused stack-entry URL helper.
- Keep company selection, ordering, invitations, logout and plugin error
handling. Hide local company creation on managed hosts, where the server
forbids it.
- Update switcher tests and the navigation contract.

## Verification

- `pnpm -r typecheck` passed, including Rust checks with the installed
Cargo toolchain on PATH.
- `pnpm exec vitest run --project @paperclipai/ui`: 637 files and 6,727
tests passed.
- Focused switcher, plugin host and Cloud link tests: 26 passed after
the managed-host creation guard.
- `pnpm check:token-gates` passed.
- Full `pnpm test:run` was attempted, then stopped after failures in
unchanged server tests. Targeted reproduction found an ancestor
skills-directory collision for Slack and macOS EACCES errors renaming
the company skills cache. Other local failures appeared in email
connector skill setup and a process-turn test. This is not a local
full-suite pass; clean Linux CI covers the complete suite.
- Full `pnpm build`, UI production build and Storybook build passed. All
[latest-head CI
checks](https://github.com/paperclipai/paperclip/actions/runs/36201814444)
passed, including the full Linux test matrix, browser tests, typecheck,
build and release checks. Greptile is 5/5 with no unresolved comments.

## Risks

- A managed host without a usable switcher plugin now gets the ordinary
company menu. It no longer gets the old Cloud portfolio menu, and local
company creation remains unavailable there.
- The Cloud portfolio endpoint remains required by the Members-page
invitation action. This PR does not remove that live endpoint or change
its authorization.
- No database, authentication, plugin protocol or deployment changes.

## Model Used

OpenAI Codex, GPT-6, with reasoning, repository inspection and code
execution. The exact deployment identifier and context-window size are
not exposed by this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-25 22:58:52 -07:00
DottaandPaperclip 96bf004a79 fix: use persisted state for lifecycle continuation and retry budgets (#13888)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Its control plane decides when a task can continue, wait, stop, or
complete.
> - Legacy continuation could change when an agent changed its wording
without changing task state.
> - Shared attempt counts also let repair and infrastructure retries
affect each other's limits.
> - This pull request uses persisted state and separate, bounded
allowances for these decisions.
> - If automatic repair stops, the task explains what happened and
offers a guarded retry.
> - Paired tests and real-provider evaluations verify that Stop,
approvals, ownership, and spending limits remain authoritative.

## Linked Issues or Issue Description

Related work: Refs #13761, Refs #11126, Refs #13610. These cover
obsolete continuation dispatch and retry storms. Open and closed issues
and PRs were searched for related lifecycle, continuation, and retry
work.

**What happened?**
Legacy continuation depended on English wording and progress heuristics.
Repair, failure retry, and productive continuation could consume shared
counts. When bounded repair stopped, the task showed a technical
recovery message without a clear next action.

**Expected behavior**
Persisted disposition and owned execution paths determine the next
action. Missing disposition prompts bounded agent repair. Explicit work
mode determines planning mode. Narrative changes and raw activity counts
cannot replenish allowances. An exhausted repair shows a readable
notice. An explicit retry checks current controls and preserves the
assigned agent.

**Steps to reproduce**
Run `pnpm test:lifecycle-baseline`. The paired probes keep structured
state constant while varying completion, planning, blocker, and progress
prose. Run the explicit `lifecycle-baseline` and
`continuation-accounting` Product E2E suites for real-provider coverage.
In Storybook, open **Design previews / Recovery notice** to inspect the
production component's normal, pending, acknowledged, unavailable,
failure, and mobile states.

## What Changed

- Hide the image attachment button, icon, and drop/paste hint in answer
composers. Image paste and drop support remains available.
- Merge current master and retain both browser regression sets. Use a
production-stamped service worker in the offline recovery browser
fixture.
- Share one state-based legacy continuation decision across immediate,
delayed, and recovered dispatch. Bind bounded repairs to their source
run and episode.
- Remove title and description wording from work-mode authority. Agents
can still write requested plans in execution mode.
- Persist separate failure-retry and productive-continuation counters.
Disposition repair and resource waits cannot consume or reset those
allowances.
- Validate delayed repair identity, then recheck current gates before
provider dispatch. Fence native startup cancellation.
- Show **Agent needs attention**, a plain-language explanation, **Retry
agent**, and expandable details in both task interfaces. Report request
progress, acknowledgement, and errors inline.
- Store typed recovery notice metadata. Recognize older active notices
only through exact stored action and run IDs. Notice text never grants
retry authority.
- Use the existing recovery-action endpoint for retry. Recheck current
action, status, owner, agent availability, dependencies, active runs,
pending questions and confirmations, approvals, pause controls, and
budget. Duplicate requests do not wake twice.
- Add component, page, route, database, contract, and Storybook
coverage. Keep the scenario inventory and executable evals here.
Historical reports and snapshots live in the [commit-pinned
paperclip-evals
archive](https://github.com/paperclipai/paperclip-evals/blob/ce3e5afcd4a1184650f586a2b5b8be5874c66c8b/experiments/2026-09-lifecycle-authority/README.md).
- Preserve unsaved project fields while the same project URL changes to
its canonical alias. Do not reuse data across projects or companies.
This separate fix addresses the repeated repository-editor browser
failure without changing the browser test.
- Keep the development service worker from intercepting Vite module
reloads. Update the connection-intent browser fixture to record progress
and completion through the agent API.

## Verification

Merge preparation on September 25, commit
`c1e8e4b7ddd9fbc4913ed55ce21b8e12906c2f97`:

- Merged master `bd2030932` and resolved the browser test-list conflict
by keeping both sets of regressions.
- Deterministic lifecycle baseline: 1,090/1,090 assertions passed; no
failures, skips, or missing selected evidence. Unit 423, runner 184,
database integration 397, grading 86.
- Browser support: 17/17 passed. The offline recovery test first failed
with an unstamped development worker, then passed with the production
stamp. Its assertions are unchanged.
- Focused interaction UI and offline fallback tests: 19/19 passed.
Verified the custom-answer composer in Storybook: no attachment controls
or hint; entering an answer enables Next.
- Recursive typecheck, production build, token gates, and diff checks
passed. The worktree is clean. No new real-provider campaign was run.
- Current CI and review: [Current PR CI
passed](https://github.com/paperclipai/paperclip/actions/runs/36166011243):
55 successful checks and two optional Storybook skips. Greptile scored
this exact commit 5/5. Hiding the question attachment controls is an
intentional UI change; paste/drop remains available.

Earlier recovery UI verification, commit
`21be0fec0e90e86b6d662b8ee4831847cd041cdb`:

- Recursive typecheck, production build, token gates, and diff checks
passed.
- Focused UI coverage: 338 tests passed across six suites (336 before
the interaction guard, with the two affected suites rerun at 149 passed
after it). Covers both task interfaces, the real page mutation,
pending/error acknowledgement, stale state, and unavailable controls.
- Recovery database integration: 352 tests passed before the interaction
guard. The complete recovery-action and mutation-route suites passed 181
tests after it. The two new pending question/confirmation regressions
failed before the fix and passed afterward, including
resolved-interaction controls. Shared validator suite: 31 passed. E2E
catalog suites: 34 passed.
- Browser inspection passed for light/dark themes, mobile layout,
expandable details, pending retry, acknowledgement, failure, and
disabled retry. Storybook renders the production component; its request
is simulated.
- The broad local run hit two chat callback-order wait failures and was
stopped after all CI unit/database/runner shards passed. Both local
failures passed when rerun without the competing full-suite process.
- CI exposed a repeated project-repository draft-loss race during
canonical redirects. A new unit regression failed before the fix; all
nine project-page tests now pass, including controls for other projects
and companies. Both unchanged repository browser tests passed against a
fresh local server. UI typecheck, production UI build, and token gates
passed after this fix.
- [Earlier PR CI
passed](https://github.com/paperclipai/paperclip/actions/runs/36072486798)
on `21be0fec0e90e86b6d662b8ee4831847cd041cdb`: 55 successful checks, two
optional Storybook skips, and no failed or pending checks. The
repository browser shard passed with the production fix. Greptile is 5/5
on this exact commit with no unresolved review threads. The PR is
mergeable.

Historical, source-qualified lifecycle evidence:

- Lifecycle baseline: 1,074 assertions. Native session coverage: 447
tests. Product E2E support: 515 tests. Browser support: 11 tests. Full
earlier verification is retained in the archive.
- [Real-provider campaign: 8/8 passed, zero
retries](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35881382080-1/index.html),
source `e88d210417280140b44a36449027290adcb1aeaa`. Evidence and cleanup
checks passed. This includes deliberately exhausted repair cases that
correctly remain blocked; it does not mean every task finished Done.
This campaign predates the recovery UI change.
- Archive migration verified all 16 original JSON files byte-for-byte
and all 24 checksum entries. App tests do not need private archive
access. [Archive PR
#27](https://github.com/paperclipai/paperclip-evals/pull/27) is merged.

## Risks

- Agents that omit durable disposition receive at most two repair
attempts by default. Prose-only completion exposes missing state rather
than silently changing scheduling.
- A retry is an explicit board action. The server rechecks current
controls. A successful response confirms the task returned to To do; it
does not claim that the provider has already started.
- Existing notice metadata remains valid. Only older active notices with
matching structured evidence receive the new UI. Historical notices
without that evidence keep their existing rendering. No schema migration
is required.
- Old run records require conservative retry accounting. Tests cover old
counters, alternating retry lanes, restarts, and exhausted repairs.
- Historical snapshots require private `paperclip-evals` access. The app
index retains public campaign links. Live campaigns qualify specific
sources and scenarios; no new real-provider campaign has run for the
recovery UI commit.

> This fixes existing lifecycle and recovery behavior and does not
duplicate planned core work.

## Model Used

OpenAI GPT-6 through Codex assisted implementation, reasoning, code
execution, and review. The exact serving model ID and context window are
not exposed in this task. Historical real-provider evaluations used
Codex model `gpt-5.6-sol`, separately from the implementation assistant.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-25 15:28:11 -07:00
Devin FoleyandPaperclip bd6caf51bb fix: preserve restore failure results and stop unsafe retries (#14035)
Preserve agent output and earlier execution errors when workspace restore fails. Report the restore phase and confirmed saved-plan links. Require verified repair before retrying unsafe archives, while preserving approval states and the retry budget.

Verified with full CI, 506 focused regression tests, and Greptile 5/5 with all review threads resolved.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-25 11:34:31 -07:00
DottaandPaperclip bd20309323 fix(ui): keep mobile task pickers above the keyboard (#14022)
## Thinking Path

> - Paperclip helps people manage AI agents and their tasks.
> - The New Task dialog lets a person select an assignee and a project.
> - A mobile browser reduces the visible viewport when the software
keyboard opens.
> - The dialog accepted invalid viewport data, and the pickers stayed
inside a transformed container.
> - This behavior could collapse the dialog or move a picker search
field above the visible area.
> - This pull request validates viewport data and puts mobile pickers in
the visible viewport.
> - The benefit is that a mobile user can see and use each picker while
the keyboard is open.

## Linked Issues or Issue Description

**What happened?**

On a mobile device, the Assignee and Project pickers in the New Task
dialog could move above the visible viewport. A short invalid viewport
value could also collapse the dialog to a line.

**Expected behavior**

The dialog and each open picker must stay in the visible viewport while
the software keyboard is open.

**Steps to reproduce**

1. Open the New Task dialog in a mobile browser.
2. Open the Assignee picker or the Project picker.
3. Focus the picker search field so that the software keyboard opens.
4. Observe that the picker can move above the visible viewport.

**Paperclip version or commit**

`efce9356b5`

**Deployment mode**

Local dev and hosted browser UI.

## What Changed

- Ignore zero, negative, and non-finite Visual Viewport measurements.
- Keep the last safe dialog geometry until the browser gives a valid
measurement.
- Put entity pickers outside the transformed dialog container.
- Size and position the mobile picker from the dialog Visual Viewport
values.
- Add unit tests for invalid viewport recovery.
- Add Chromium tests for the Assignee and Project picker states.

## Verification

- `pnpm exec vitest run ui/src/components/NewIssueDialog.test.tsx`
passed with 34 tests.
- The Chromium viewport test passed 35 of 35 runs with five repeats and
no retries.
- `pnpm --filter @paperclipai/ui typecheck` passed.
- `pnpm check:token-gates` passed.
- `pnpm build-storybook` passed.
- `pnpm -r typecheck` passed.
- `pnpm build` passed.
- A full local Vitest run reached restricted workspace-runtime tests
that require sibling worktree and runtime writes. Remote CI will run the
supported test environment.

## Risks

- Risk is low because the new picker layout applies only to mobile
widths.
- The layout depends on Visual Viewport data when the browser supplies
valid values.
- Unit and browser tests cover invalid data, mobile pickers, tablet
layout, and desktop layout.

> This change fixes a focused UI bug. It does not duplicate a core
feature in `ROADMAP.md`.

## Model Used

- OpenAI Codex with GPT-5. The agent used reasoning, repository tools,
code execution, and browser automation. The runtime did not expose the
context window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-25 08:37:26 -07:00
DottaandPaperclip efce9356b5 fix(ui): offer recovery when the app fails before React starts (#13970)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The browser must load its JavaScript before React can render a task.
> - A failed import can stop that process before the React error
boundary exists.
> - The HTML entry then leaves an empty page with no recovery action.
> - This PR adds a small recovery screen that works without React.
> - The user can retry the same page and return to saved task content.

## Linked Issues or Issue Description

Refs #13824 and #13895. This is a follow-up to their browser startup
investigation.

**What happened?**

Interrupting the app bundle or a required import leaves an empty React
root. A startup exception has the same effect. A React error boundary
cannot handle these failures because React has not started.

**Expected behavior**

The page must explain the startup failure and offer a manual retry. A
late successful load must dismiss the recovery message without a reload.

**Steps to reproduce**

1. Open a saved task in the browser.
2. Abort the application bundle request, or make a required module
return HTTP 503.
3. Observe the empty page before this change. With this change, use
Reload page after the fault clears and verify the saved task and
comment.

**Paperclip version or commit**

The failing regression baseline used master at `8781f06a8`.

**Deployment mode**

Local source build and compiled UI. Tests cover both initial navigation
and a page controlled by the production service worker.

The exact cause of the older intermittent Vite stall remains
unconfirmed. Forty app loads and thirty replays of retained responses
did not reproduce it. This PR fixes the missing recovery path; it does
not claim to remove that historical cause. A normal HTTP 304 response is
not a failure.

## What Changed

- Add an inline startup guard and recovery screen in the HTML entry. It
does not depend on the app module graph.
- Show a manual reload action after a startup error or after 30 seconds
without rendered root content.
- Remove the notice, timer, observer, and error listeners when the app
starts. Never reload automatically.
- Keep the recovery screen outside the React root so it cannot satisfy
app-readiness checks.
- Add browser tests for interrupted imports, a stalled import, an
evaluation error, service-worker-controlled retry, repeated offline
retry, and cleanup after successful startup.
- Return a static, uncached HTML retry screen when a
service-worker-controlled navigation fails offline. It contains no task
content.
- Add a full-app test that retries an interrupted compiled bundle and
checks the saved task, comment, composer, route, and absence of agent
runs.
- Document the coverage and the limits of the historical diagnosis.

## Verification

- Red baseline: four recovery cases failed; the normal-startup case
passed. After the change, all five recovery cases passed. The review
found an offline retry gap; that additional case failed before the
worker fix and passed afterward.
- Full provider-free browser-support suite: 16 passed.
- Compiled-app browser tests: four passed, including saved-task reload,
interrupted-bundle recovery, slow-CPU service-worker reload, and sidebar
navigation.
- Expanded service-worker, offline response, PWA, and worker build-ID
unit tests: 37 passed. The two old plain-text offline expectations were
reproduced as failures and updated for the HTML retry contract.
- UI production build, full local repository typecheck (`pnpm -r
typecheck`), runner-E2E typecheck, and design token checks passed.
- Manual browser check: a temporary server failed the compiled bundle
once. The recovery screen appeared. Clicking Reload page restored the
same saved task, comment, and composer.
- Full local `pnpm build` passed.
- Full local `pnpm test:run` was attempted with a bounded deadline and
stopped after it timed out. Workspace runtime/cleanup tests reported
timeouts on this host. The monolithic local run is not a pass. The
focused tests above and the complete Linux CI run provide the successful
verification.
- Final-head [CI
run](https://github.com/paperclipai/paperclip/actions/runs/36072201966)
passed. All 53 check runs succeeded; the two Storybook jobs were
intentionally skipped. The legacy security status also passed.
- Greptile reviewed `f83e0f51fb760541d83353f2c1df4e182f3948f9`: 5/5.
Both review findings are fixed and resolved.

## Risks

- The guard only handles startup before React renders root content.
Existing React boundaries handle later rendering errors.
- A slow startup can show the message after 30 seconds. A later
successful render removes it; the page does not reload by itself.
- The fallback uses native HTML when the app stylesheet is unavailable.
- The worker changes only its offline navigation response. It returns
static HTML with a reload button and `Cache-Control: no-store`. Its
cache allowlist, private-response protections, task state, provider
prompts, and grading rules stay unchanged.
- This does not establish or fix the unknown cause of the historical
intermittent Vite stall.

## Model Used

OpenAI GPT-6 through Codex. The session exposes the GPT-6 family but not
an exact served model ID or context window size. Used reasoning, code
editing, shell tools, and browser testing. No subagents were used.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-24 19:36:45 -05:00
DottaandPaperclip 8781f06a87 feat(connections): enable MCP aggregators by default (#13964)
## Thinking Path

> - Paperclip helps people manage AI agents for work.
> - Connections let those agents use external services with explicit
access rules.
> - Zapier, Arcade, Composio Connect, and Executor already have setup
and runtime support.
> - Their experimental switch still blocks discovery and setup by
default.
> - This pull request removes those gates and the Settings toggle.
> - Users can connect these providers without enabling an experiment.

## Linked Issues or Issue Description

Refs #13755. Refs #13941.

**What existing behavior does this improve?**
Apps browsing, inline setup, and agent connection search for the four
MCP aggregators.

**Current behavior**
An instance must enable the MCP aggregators experiment before users or
agents can start setup.

**Proposed behavior**
All four providers are available by default on local and managed
instances. Old stored and managed values still parse but cannot disable
them.

## What Changed

- Remove the aggregator gates from Apps, inline setup, server setup, and
agent search.
- Remove the Settings toggle and its UI hook.
- Retain the old setting key only for upgrade compatibility. Normalize
it to true and ignore managed overrides, as Apps already does.
- Replace opt-in fixtures with default-on coverage. Test old false
values, all four setup flows, provider choice, and the removed toggle.
- Update current connector guidance and remove the opt-in from the
runner acceptance fixture.

## Verification

- 306 focused tests passed across eight files: shared remote MCP
contracts; server remote MCP lifecycle, aggregator fallback, settings
normalization, and managed overlay; UI Apps browsing, setup, and
experimental settings.
- Server and UI TypeScript checks passed.
- UI token gates and `git diff --check` passed.
- The full local suite was not run, per the maintainer's instruction.
All 54 CI checks passed; two checks were skipped. One unrelated
workspace-preview readiness timeout passed on one failed-shard retry.
- The setup fixtures use simulated MCP responses. This change does not
claim new live provider acceptance.

## Risks

- Existing instances now show all four providers, even if the old flag
was false. This is intentional.
- External provider choice, credentials, company isolation, agent
grants, and tool policies still apply. Showing a connector does not
authorize an external account.
- No data migration is required. The compatibility key keeps old managed
configuration documents valid.
- Historical Zapier live acceptance remains incomplete in the existing
evidence report. The maintainer explicitly requested the default-on
rollout for all four existing providers; the report records that scoped
exception.

## Model Used

OpenAI GPT-6 (`gpt-6-astra`) through Codex, with reasoning, repository
tools, and test execution. The context window size is not exposed in
this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-24 17:31:32 -05:00
DottaandPaperclip aa8fc86331 feat(connections): prefer native apps and ask users to choose external providers (#13941)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Connections give agents governed access to external services.
> - Native connections should remain the first choice for a supported
app.
> - Other apps may be available through an external MCP provider.
> - The user must know which external provider handles the connection
and choose it before setup.
> - This pull request adds ranked alternatives and server-authored
instructions to connection search.
> - Agents can follow the returned instructions while Paperclip
validates saved choices and access.

## Linked Issues or Issue Description

Related: #13879, which fixed inline MCP provider setup. This PR adds
discovery and provider selection on top of that work.

**Subsystem affected**

Cross-cutting: shared connection contracts, server search and intent
services, native runtime, CLI, inline setup UI, and evals.

**Problem or motivation**

An agent cannot offer a clear external-provider choice when Paperclip
has no native connection for an app. Adding provider-specific branches
to the core prompt would make those instructions harder to maintain.

**Proposed solution**

Prefer a native connection. Otherwise return verified alternatives in
Composio, Arcade, Executor, Zapier order. Include an external-service
disclosure, a question with None, and the next instruction in the search
result. Validate the saved human choice before creating a selected
fallback setup card. Reuse existing provider accounts and verify
underlying app access separately.

**Alternatives considered**

Do not silently choose a provider. Do not claim that broad execution
tools prove support for every app. Reuse existing questions and
connection intents rather than add another connection model.

**Roadmap alignment**

Extends the existing MCP Tool Gateway & Apps and Agent evals & feedback
capabilities. The MCP aggregators experiment remains the gate. No
duplicate provider-routing PR was found in the public search.

## What Changed

- Add a dated support index and authorized cached-tool evidence for
external routes.
- Return provider questions and next-step instructions from
`connections_search`.
- Preserve pending choices and declines across continuation. Validate
company, task, agent, human, app, and current route eligibility.
- Carry the selected app into new setup and account reuse, validate
explicit provider requests against persisted human messages, and
distinguish provider readiness from app authorization.
- Sync native, MCP, REST, and CLI contracts. Keep core agent
instructions provider-neutral.
- Add production-component Storybooks, focused database tests, and three
real-agent browser eval cases.
- Record the plan, observed failures, fixes, passing evidence, and
acceptance limits.

## Verification

- Latest head `586f0e6cd`: 54 checks passed, 2 skipped; Greptile 5/5 and
all review threads resolved.

- After rebasing on master `18dac1e1e`: 64 focused shared, validator,
route-contract, and database tests passed; server typecheck passed.
- Embedded-browser test drive on the rebased head: native Jira card,
HubSpot external-provider question, Arcade account reuse, one actual MCP
read against a local synthetic fixture, reload persistence, and None
preventing further calls. A real OpenAI-backed agent performed discovery
and continuation.
- UX observation: the agent initially combined mutually exclusive
request fields; the server rejected it and the agent recovered without
changing access. This extra retry remains visible in the transcript.
- After rebase: 23 focused eval grader/catalog tests, affected
TypeScript checks, token gates, production UI build, and Storybook build
passed. Full local tests are intentionally excluded at the maintainer's
request.
- Before rebase: four browser/real-agent attempts passed: native Jira,
None, and reuse of the second provider on two Codex profiles.
- Browser evals used an isolated deterministic MCP fixture through the
real Paperclip gateway. They do not prove production compatibility with
all four providers.
- Review `Apps / Connections / Provider choice` in Storybook. Choose
Arcade, continue through Access, and verify the app name,
external-service disclosure, and URL configuration.
- The detailed verification report is
`doc/connections/2026-09-23-aggregator-routing-verification.md`.

## Risks

- The public support index is finite and can age. Account capability and
app authorization still require verification after selection.
- Existing installed-tool permissions remain in effect. Provider choice
is not a new execution permission boundary. An early Mini attempt
skipped search; clearer provider-neutral instructions made the targeted
rerun pass. This is not a measured reliability rate.
- Explicit requests skip provider confirmation only when a clear
persisted human message or saved provider choice supports them. Other
phrasing falls back to confirmation; the agent query alone is not
consent. These routes do not add tool permissions.
- No database migration or legacy Composio broker is added.
Real-provider acceptance remains separate from fixture proof.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, code
execution, and browser tools. The exact deployment variant and
context-window size are not exposed in this session. Product evals
separately used the repository's primary Codex and Codex Mini profiles;
those agents supplied test behavior, not independent provider
compatibility proof.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-24 16:36:49 -05:00
DottaandPaperclip 18dac1e1ef feat(connections): add experimental memory providers and remote MCP access (#13942)
## Thinking Path

> - Paperclip manages AI agents and their work.
> - Connections give agents governed access to external tools.
> - Agents need durable memory across tasks and execution environments.
> - Mem0, Zep, Supermemory, Cognee, and Honcho provide hosted memory
tools.
> - This pull request adds their setup flows behind an experimental
toggle.
> - It also delivers assigned MCP tools through the native remote Codex
runner.
> - Operators can connect a provider once and use the same governed
tools locally or in Daytona.

## Linked Issues or Issue Description

**Problem or motivation**

The Apps catalog lacks a complete set of memory providers. Remote native
Codex agents also need access to assigned managed MCP tools without
receiving provider credentials.

**Proposed solution**

Add five memory connectors behind the disabled-by-default Experimental
memory connectors setting. Use the existing connection setup and
permissions UI. Default their tools to Allowed. Preserve company
boundaries, operator permission changes, provider scopes, and audit
attribution.

**Alternatives considered**

Direct provider credentials in each sandbox would duplicate setup and
bypass the managed gateway. The remote runner instead uses its existing
protocol channel to call the gateway on the server.

**Roadmap alignment**

This maintainer-requested experiment supports the Memory / Knowledge and
Connected Apps roadmap areas. It adds provider connections without
introducing a separate memory UI. Related connector authoring
documentation is tracked in #13692; no duplicate memory-provider
implementation was found.

## What Changed

- Add provider definitions, official branding, and the experimental
setting for all five providers.
- Use OAuth for Zep and Supermemory, API credentials for Mem0 and
Honcho, and a bundled Cloud API bridge for Cognee with no runtime
downloads or subprocesses.
- Default memory tools to Allowed and classify destructive actions
explicitly.
- Fix personal remote credential resolution and propagate provider tool
errors.
- Relay assigned managed MCP tools to remote native Codex through the
runner protocol. Recheck current authority for each call and rotate
stale tool contracts.
- Add 21 Storybook states and complete OAuth walkthrough fixtures.
- Document provider research, sanitized tool inventory, and live local
and Daytona proof.

## Verification

- Passed workspace typecheck: `pnpm -r typecheck`.
- Passed production build: `pnpm build`.
- Full local `pnpm test:run`: 13,273 passed, with failures from
process/readiness timeouts under parallel load. Reran all 16 affected
suites with one worker: 456 passed, leaving two macOS `/var` versus
`/private/var` path assertions. Both passed with `TMPDIR=/private/tmp`.
No test failures remain unverified. Latest-head remote CI passes all 54
checks (two optional Storybook jobs skipped). Greptile is 5/5 with no
unresolved findings.
- Latest Cognee gateway regression: 71 passed, including public
deployment without a runtime host and immediate recovery after a
provider error. Bundled bridge tests: 25 passed.
- Browser setup and real agent tasks exercised all five providers. Mem0,
Cognee, Zep, and Honcho have successful store/retrieve proof.
- Supermemory now has scoped read/write consent. Local storage and real
Daytona write, document read, and semantic recall passed; indexing
completion was verified before claiming success.
- All five providers were exercised through a real Daytona sandbox and
its native runner MCP relay. Provider credentials remained on the
server. The final bundled Cognee bridge also passed a fresh Daytona
store/recall run. All disposable sandboxes were removed and verified
absent after testing.
- Storybook is rebuilt and contains the experimental toggle, catalog,
setup, permissions, and error states. The Zep and Supermemory access
steps advance correctly.

## Risks

- Provider OAuth scopes and plan limits remain independent of Paperclip
tool permissions. An Allowed tool can still be rejected by the provider.
- Providers can queue memory indexing; save acceptance does not prove
that semantic recall is ready.
- Remote tool contracts must stay synchronized with current connection
authority. Regression tests cover revocation and stale contracts.
- This change has no database migration. Existing connections remain
usable when the experimental catalog toggle is disabled.

## Model Used

OpenAI Codex, GPT-6 Astra (`gpt-6-astra`), with reasoning, repository
edits, code execution, and browser automation. The exact context window
size is not exposed in this session. Live acceptance agents used
`gpt-5.6-sol`.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-24 15:54:38 -05:00
DottaandPaperclip b0155a681a feat(slack): connect Paperclip conversations and scheduled messages (#13920)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Slack conversations use the same tasks and agents as the Paperclip
board.
> - A board reply must reach that Slack conversation and let the agent
continue the work.
> - An assigned agent also needs its Slack tools during normal tasks and
scheduled routines.
> - Both paths must keep the linked user's authority, delivery rules,
and conversation history.
> - This pull request adds those paths and reduces setup friction for
Slack bots.

## Linked Issues or Issue Description

**Subsystem affected**

Server orchestration, Slack connector tools, shared contracts, and chat
setup UI.

**Problem or motivation**

Replies entered in Paperclip did not provide a complete round trip to
the linked Slack thread. Slack and board wakeups could select different
model sessions for the same task. Agents also lacked their assigned
Slack tools outside Slack-origin work, which prevented a routine from
sending its responsible user a briefing. Inviting a bot could leave the
new channel disabled.

**Proposed solution**

Mirror human board messages with author attribution and route the agent
result to the same thread. Use the same session key across both entry
points. Supply Slack tools to the connection's assigned agent in normal
tasks and routines, using the current responsible user's verified link.
Enable newly invited channels while preserving explicit disabled
choices. Add a browser-agent setup prompt to the Slack wizard.

**Alternatives considered**

A separate Slack scheduler or task dispatcher would duplicate existing
Paperclip workflows. Reusing the connection owner's identity would grant
the wrong authority. Replaying old channel history could start
unintended work. This change uses ordinary task wakeups, routine
dispatch, and Slack's original invitation mention event instead.

**Roadmap alignment**

Extends the shipped Scheduled Routines and governed Apps capabilities.
It does not add a separate task lifecycle. Related work: #13828 and
#13809. Related test stabilization: #13877. The existing plugin
Slack-control proposals are separate from this built-in connector
change.

## What Changed

- Queue human Paperclip messages for the original Slack thread with
display-name attribution and stable delivery identities. Require the
author’s current linked Slack identity and recheck access before
delivering messages or agent replies.
- Apply pause, dependency, cancellation, and closed-workspace guards
before explicit Board sends request work and again when the durable
outbox dispatches it.
- Route agent results back to Slack and preserve model-session
continuity, including replies that reopen completed tasks.
- Resolve assigned Slack connections for normal agent tasks and
routines. Recheck the responsible user's link, membership, and
permissions at execution.
- Add `slack_open_dm` for the responsible user's bot DM and request the
`im:write` scope.
- Enable newly discovered invited channels. Keep explicit OFF choices
and normal admission and deduplication rules.
- Add a copyable Slack setup prompt for a computer-use agent, with
Storybook coverage. Share the prompt-button component with GitHub.
- Update Slack tool documentation and runtime instructions.
- Stabilize the mobile project browser test by waiting for the final
canonical route before editing, preserving all persistence assertions.

## Verification

- Live staging: invited the bot after the first mention. The channel
became enabled and the bot answered that original mention.
- Live staging: a normal Paperclip reply appeared in Slack with author
attribution. The agent completed the calculation and replied once in the
original thread and in Paperclip.
- Live staging: a codeword entered in Slack was recalled from Paperclip.
A following Slack calculation used the result from the Paperclip turn.
Run metadata confirmed the same model session for both entry points.
- Live staging: a scheduled routine used `slack_open_dm` and
`slack_post_message` to deliver one DM. The existing app was reinstalled
with `im:write`. The test routine was paused after verification.
- Before the master merge: 397 focused feature tests passed. The
continuity fix passed all 76 issue comment/update route tests and six
focused route/integration cases. Typecheck, build, and token gates
passed.
- Review fixes: 47 focused integration cases passed, covering link
revocation/replacement, private membership removal, guarded outbox
dispatch, concurrent workers, lost scheduler responses, a real
one-connection pool, exact reply provenance, and attachment retries. All
99 issue-comment route tests and the Slack catalog browser test passed.
- Full local typecheck, production build, token gates, and
module-boundary checks passed. The full local test command passed 25,915
tests before a 15-second timeout in
`issue-thread-interaction-routes.test.ts`; that entire suite passed on
isolated rerun (81 tests). Remaining serialized coverage is provided by
the current-head CI shards.
- An unchanged Cursor adapter test hit its 10-second limit in CI; all
five tests in that file passed on a local rerun in 3.11 seconds, and the
failed CI shard passed on its single retry.
- The preview-server readiness test passed a local rerun (28 tests). The
mobile-project readiness fix passed three repetitions of both browser
tests (6/6).
- Final commit `64ac0d9897f4353375996f1b1b38e5040bdeb0a0`: all CI gates
passed, including all eight browser shards, all server/chat suites,
typecheck, build, runner checks, and security checks. Greptile reviewed
this exact commit at 5/5; all review threads are resolved.

## Risks

- Human messages on a Slack-linked task now publish to its Slack thread.
The task banner states this behavior. Incoming Slack messages and
internal agent bookkeeping must not echo back.
- Normal tasks and routines can now use the assigned bot. Authority
remains bound to the current responsible user's link; it does not fall
back to the connection owner. Revocation, private-context limits, and
queued-write checks still apply.
- Existing Slack apps need `im:write` and a reinstall to open DMs. Other
existing capabilities remain available without that scope.
- New invited channels default to enabled. Explicit disabled choices
remain disabled. Channels created by bot tools still require a person to
enable responses.
- No database migration or new provider credentials are required.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, code execution, GitHub CLI,
and browser tools. The runtime does not expose a more specific model
build identifier or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-24 10:51:45 -05:00
Devin FoleyandPaperclip c341588bdb fix(ui): add Cloud invitations to the Members page (#13922)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - People manage collaborators from the Members page.
> - Cloud manages invitations outside the tenant's local invitation
system.
> - The local Invites tab is hidden on Cloud, so this page has no way to
invite a person.
> - This pull request adds an Invite people action for the current Cloud
stack's owner or admin.
> - The action opens the existing Cloud People settings for that stack.

## Linked Issues or Issue Description

**What happened?**

A Cloud owner opens Organization Settings → Members and finds no
invitation action. The tenant-local Invites tab is hidden, and the page
does not link to Cloud's invitation flow.

**Expected behavior**

Cloud owners and admins can start an invitation from Members.

**Steps to reproduce**

1. Sign in to a Cloud-managed instance as the current stack's owner or
admin.
2. Open Organization Settings → Members with `company.invites` hidden.
3. Look for an invitation action beside the page heading.

**Paperclip version or commit**

`7b7c4d4172d6aac14919e2682b702ae87bc17653`.

**Deployment mode**

Cloud-managed, authenticated.

Related search: #2388 proposes broader member-management UI. This change
only connects the existing Members page to Cloud invitations. No
duplicate Cloud invitation action PR was found.

## What Changed

- Add **Invite people** beside the Members heading for the current Cloud
stack's owner/admin.
- Read the role from the authenticated Cloud portfolio. Ownership of
another stack does not enable the action.
- Navigate to the current stack's People settings on the configured
Cloud origin. Keep the local Invites tab hidden when configured.
- Cover allowed roles, denied roles, loading, failed refresh, missing
configuration, current-stack selection, and self-hosted behavior.
Document the navigation contract.

## Verification

- Focused Members and Cloud link tests: 23 passed.
- Full UI suite: 6,669 passed across 634 files.
- UI typecheck and `pnpm check:token-gates`: passed.
- `pnpm build`: passed.
- `pnpm -r typecheck`: passed.
- All PR CI checks passed, including the full general, serialized,
browser, and runner test jobs. The duplicate local repository-wide `pnpm
test:run` was stopped after CI completed; it is not reported as a local
pass.
- Manual acceptance after tenant rollout: an owner/admin opens Members,
selects **Invite people**, and reaches the same stack's Cloud People
settings. A member does not see the action.

## Risks

- The action needs a tenant app update before it appears on an existing
stack.
- The portfolio request must identify the current stack and its role.
The action stays hidden when that information is unavailable or the
request fails.
- Cloud rechecks invitation authorization at the destination. No schema,
API, or invitation-acceptance behavior changes.

## Model Used

- OpenAI Codex, GPT-6, with reasoning, code execution, and repository
tools. The session does not expose a more specific model identifier or
context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-24 08:48:04 -07:00
Devin Foley 7b7c4d4172 fix(ui): recover an archived Cloud stack from its own health probe (#13913)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - On Paperclip Cloud, a stack whose organization is archived is served
through the Cloud harness router.
> - An archived stack answers every tenant request — including the SPA's
own health probe — with a 423 archived status page.
> - If the SPA is already loaded (a stale tab, or a load that raced the
stack's archival), that 423 surfaces as a dead "Failed to load health"
screen with no way out.
> - This pull request recovers it the same way an expired tenant session
already does: a top-level reload re-enters the Cloud harness, which
redirects a fresh browser navigation on an archived host to the
portfolio.
> - The benefit is that an archived stack never traps the user on a
broken health screen.

## Linked Issues or Issue Description

No public issue exists. Description follows the bug template:

**What happened?**

A Cloud stack that becomes archived while its SPA is open (or is opened
in a stale tab) shows a "Failed to load health" screen. The SPA's health
probe receives the harness's 423 archived status page and has no
recovery path, so the browser is stuck on a dead screen.

**Expected behavior**

The browser should leave the archived stack for the Cloud portfolio, as
it already does for an expired tenant session — no dead "Failed to load
health" screen.

**Steps to reproduce**

1. On a Cloud-managed instance, open an organization's SPA in the
browser.
2. Archive that organization (or open a stale tab for one that was
archived).
3. The SPA's health probe returns 423 (archived) and the page shows
"Failed to load health" with no way out.

**Paperclip version or commit**

`master` (Paperclip Cloud managed instances).

**Deployment mode**

Cloud-managed (`authenticated`). Self-hosted instances are unaffected:
the 423 archived status page only originates from the Cloud harness
router that fronts managed stacks; a self-hosted server serves its own
health.

**Subsystem affected**

UI: tenant document recovery (`ui/src/lib/tenant-session-recovery.ts`),
which the health/client/heartbeats/audit API surfaces already route
error responses through.

## What Changed

- `isArchivedStackRecoveryError(status, body)`: true for a 423 whose
body carries a `statusPage.code === "archived"`.
- `isTenantDocumentRecoveryError` unifies the existing 401
tenant-session codes with the new archived case; the recovery
coordinator now fires on either.
- The recovery action is unchanged — a single top-level reload — so an
archived stack re-enters the harness and is redirected to the portfolio.
The existing single-reload guard prevents any loop, and only the exact
`archived` code triggers it (suspended/deleted/other 423s pass through
as normal errors).

## Verification

- `pnpm exec vitest run src/lib/tenant-session-recovery.test.ts
src/api/auth.test.ts src/api/audit.test.ts` — pass.
- New tests: the predicate accepts a 423 archived body and rejects
suspended/deleted/error-shaped/503/null; the coordinator triggers the
same single top-level reload for an archived body and stays inert for
unrelated responses.
- `pnpm exec tsc --noEmit` in `ui/` is clean.
- Pairs with the harness-side redirect
(paperclipai/paperclip-cloud#545): the harness redirects fresh
navigations on an archived host to `/orgs`; this makes an already-loaded
SPA trigger that navigation on its own health 423.

## Risks

- Low. The recovery path and its single-reload guard are unchanged; only
the set of conditions that trigger it grows by one exact code.
Non-archived 423s and all other statuses behave as before.

## Model Used

- Claude Fable 5 (`claude-fable-5`), via Claude Code CLI, extended
thinking and tool use enabled.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-23 21:09:28 -07:00
Devin FoleyandPaperclip 94a0aa7726 fix(ui): hide workspace isolation controls for managed hosts (#13907)
## Thinking Path

> - Paperclip manages AI agents and their work.
> - Workspace isolation keeps task checkouts separate.
> - Managed hosts can enable isolation and hide its experimental
toggles.
> - Project, task, and routine forms still expose choices that override
that policy.
> - This pull request adds an operator visibility key for those
controls.
> - Workspace access stays available, and execution keeps its existing
policy.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

Operator control over workspace isolation settings in the UI. This
follows the settings list cleanup in #13905.

**Current behavior**

Hiding the experimental isolation toggles leaves project policy editors,
task selectors, routine and pipeline overrides, recovery actions, and
workspace configuration visible.

**Proposed behavior**

Set `PAPERCLIP_HIDDEN_SETTINGS=workspaces.isolation` to hide these
controls. Keep workspace navigation, status, files, and runtime access.
Hide experimental toggles separately. Instances that do not set this key
keep their controls.

**Reason and benefit**

Users on managed hosts should use the host's isolation default. A hidden
form must not submit a stale draft that overrides it.

## What Changed

- Add the UI-only `workspaces.isolation` key to the shared visibility
registry.
- Hide project workspace policy, task and subtask selectors, routine and
pipeline overrides, and isolated re-issue actions.
- Hide the workspace Configuration tab and redirect direct links to
workspace issues.
- Omit hidden new-task and routine overrides. Keep explicit task/subtask
workspace launch context, saved policies, and automatic branch values
for workspace routine runs.
- Wait for the health visibility policy before showing controls. Keep
workspace access and all execution APIs available.
- Document the key and test visibility, form payloads, deep links, and
unchanged workspace access.

## Verification

- All 52 CI checks pass on `50cc770d33` (two expected skips). The branch
is mergeable. Greptile is 5/5 with no unresolved comments.
- `pnpm -r typecheck` passed.
- `pnpm build` passed.
- `pnpm check:token-gates` passed.
- Targeted UI checks passed: 346 tests across 14 suites, including
hidden project/task controls, stale task drafts, routine branch
defaults, recovery actions, configuration deep links, and workspace
access.
- Shared settings-visibility tests passed: 11 tests.
- `pnpm test:run` was run and stopped after reproducing five failures in
unchanged server tests: two `chat-channels.integration` cases
(linked-request provenance and direct external-chat finals) and three
`company-skills-service` cases (runtime refresh, concurrent download,
and explicit update). The earlier local run for #13905 showed the same
failures. The full local suite is not claimed green. Targeted UI/shared
checks pass. All PR CI shards, including the affected chat and skills
suites, pass.
- Reviewed the diff for secrets, private links, and run artifacts.

## Risks

This key changes UI visibility only. It does not reject API calls or
change feature values. Operators must enable isolation and its default
through their existing policy mechanism. Older app versions ignore the
new key until upgraded. Removing the key restores the controls. No
schema changes.

## Model Used

OpenAI GPT-6 (Codex), with reasoning, repository tools, and test
execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either linked existing issues or described the issue in-PR
following the relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links
- [x] My branch name describes the change and contains no internal
ticket id
- [x] I have run tests locally; targeted checks pass (full-suite
limitation documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-23 19:38:07 -07:00