Commit Graph
250 Commits
Author SHA1 Message Date
DottaandPaperclip d6df12cef6 fix(ui): add icons to chat connection choices (#15574)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Connectors let people chat with agents and let agents use external
tools.
> - Slack setup asks the user to choose between those two uses.
> - Both choices currently use text alone.
> - This pull request adds an agent avatar and a hammer to the shared
chooser.
> - The icons help users identify each choice at a glance.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

The chat-or-tools choice in connector setup. Related setup work: Refs
#15413.

**Subsystem affected**

ui/ — React board UI and Storybook.

**Current behavior**

Both choice cards show a title and description with no identity icon.

**Proposed behavior**

Show the existing agent avatar beside chat. Show a hammer beside tools.
Align both icons and keep the existing navigation.

**Reason and benefit**

Users can tell the two choices apart with less reading.

**Breaking changes**

None. The shared chooser applies the icons wherever that choice appears.
Only Slack currently offers both paths in the catalog. Other chat-only
providers keep their direct setup path.

## What Changed

- Add the shared agent avatar to the chat choice.
- Add a decorative hammer icon to the tools choice.
- Add desktop and mobile Storybook stories that render the production
page.

## Verification

- Token gates pass.
- Focused catalog, routing, and UI contract tests pass, including the
new chooser accessibility and navigation test.
- Checked the production chooser in Storybook at desktop and mobile
widths.
- Local repository typecheck, build, and Storybook build pass.
- All 56 checks pass on commit
`0f0673f1bb5cd340084f3a5229983da049f1c8db`, including the full CI test
suite and browser shards. Greptile reports 5/5 with no findings.
- Started the full local test command. Stopped that duplicate run after
the full CI suite passed. The focused local tests completed
successfully.
- In Storybook, open Connections → Slack → Automatic setup → 00 · Choose
connection. The mobile variant is in the same group.

## Risks

Low risk. This changes card layout and decorative images. Existing
button text, click handlers, credentials, and permissions stay the same.
The avatar is generic because the user has not selected an agent yet.

## Model Used

OpenAI GPT-6 through Codex. Used code inspection, edits, shell tools,
and browser tools. The session does not expose the exact runtime model
ID or context window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-08 09:41:46 -05:00
DottaandPaperclip d66acb7ac1 feat: automate Slack bot app setup and installation (#15413)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Chat connectors give each agent a customer-owned bot and task-backed
conversations.
> - Manual Slack setup requires app creation and copying durable
credentials.
> - Operators need a shorter setup that an assisting agent can use
safely.
> - This pull request creates the app through Slack's Manifest API and
installs it through OAuth.
> - Durable registration state supports recovery without creating
another app.
> - A four-screen wizard, automatic avatar upload, and OAuth account
linking reduce setup work.
> - Connector settings and per-turn tool guidance support daily use
after installation.

## Linked Issues or Issue Description

**Subsystem affected**

Native Slack bot setup, company secret storage, chat connector
management, and agent tool guidance.

**Problem or motivation**

New Slack bots require manual app creation and copying a signing secret
and bot token. Interrupted setup can create duplicate apps. The setup
and management screens contain unnecessary controls. Agents also need
guidance for native questions, files, thread replies, and governed Slack
actions.

**Proposed solution**

Use a temporary app-configuration access token to create a
customer-owned app. Save durable secrets in the vault. Bind OAuth to the
initiating actor, company, endpoint, registration revision, scopes, and
configured origins. Preserve manual and existing-app recovery. Link the
installing user's account, send a welcome DM, and advance from saved
server evidence. Keep request-URL recovery instructions available if
automatic connection detection waits.

**Alternatives considered**

The Slack CLI adds installation requirements. Socket Mode changes
transport. A shared Paperclip-owned app changes app ownership. These
alternatives are outside this change.

**Roadmap alignment**

This extends existing chat connectors and secrets capabilities. Related
public work: #14037 and #13954 cover Slack MCP prerequisites and user
OAuth. No duplicate bot-registration PR was found.

## What Changed

- Share one reviewed manifest builder between automatic registration and
manual setup.
- Add replay-safe migration 0318 and company-bound registration state
with vault references and uncertain-creation recovery.
- Add registration, installation, callback, and resume APIs with
short-lived, single-use OAuth state.
- Save installation credentials before downstream checks and preserve
bot identity constraints.
- Reduce automatic setup to four screens. Keep advanced app details,
manual recovery, and existing-app setup.
- Upload the agent avatar with the Paperclip dark background. Link the
OAuth installer's account and send setup DMs.
- Show agent and connector-owner avatars. Simplify settings, access, and
conversation screens.
- Discover joined Slack channels and enable them by default. Start a
task from a bare mention and admit same-thread follow-ups.
- Refresh Slack tool guidance each turn. Add native-form, file,
approval, and delivery regressions plus manual model probe definitions
and sanitized acceptance records.
- Update deployment/database docs, OpenAPI, redaction, removal cleanup,
production Storybook stories, and provider browser tests.
- Merge current master and move the registration migration after its
latest migration without rewriting published commits.

The completed Slack success view intentionally has a single centered
**Done** action and no **Save & exit**, as explicitly requested by the
product owner. `DESIGN.md` records this exception; unfinished setup
steps retain the aligned wizard footer.

## Verification

- Passed after the master merge: repository typecheck, full build,
Storybook build, design-token gates, module-boundary gates, and
migration generation.
- Passed: all 352 focused Slack deterministic tests and all 14 affected
provider browser tests. Browser tests use controlled provider fixtures
and a separate throwaway instance.
- Passed on current head `c5d01e0e2`: the complete GitHub test matrix
(general server, chat, all workspaces, serialized server, and Runner),
all eight browser shards, typecheck/release registry, build, canary dry
run, security checks, and policy gates. There are 52 passing checks and
no pending or failing checks.
- Greptile completed on the exact current head with 5/5 and no
actionable findings or open review threads.
- Local repair verification passed 93 focused tests, including same-app
reinstall after revocation and rejection of consent started before
revocation, the AgentMail browser journey, and repository typecheck.
Local build and Storybook build also passed. The redundant local
full-suite rerun was stopped after the complete current-head CI matrix
passed.
- Real Slack setup and agent replies were exercised in the authorized
isolated test drive during the setup iteration.
- The ten additional model probes were attempted with legacy
`codex_local`, `gpt-5.6-sol`: five passed, two failed, and three were
partly verified. Native runtime is not qualified. See
`server/src/services/connectors/slack/evals/2026-10-08-acceptance.md`
for evidence and limits.
- Passing model probes cover native forms, downloaded file bytes, bare
mentions with thread replies, explicit posts/reactions, and saved
approval denial.
- The controlled uncertain-write probe found wrong delivery-check IDs.
The canvas fallback attempt used an invented tool name. Search
pagination/native search, a private-source denied-tool receipt, and
distinct board/webhook origins remain unqualified.

Reviewer path: enable Chat connectors, start Slack chat setup, select an
agent, enter an app-configuration access token, and approve Slack
installation. Send a message to the bot and confirm that setup advances
to success. Inspect settings and allowed channels. See
`doc/connections/SLACK-AUTOMATIC-SETUP.md` for deployment and recovery.

## Risks

- Slack app creation has no provider idempotency guarantee. A timeout
after dispatch stays uncertain until the operator checks Slack.
- OAuth needs a stable public HTTPS board origin. Webhook ingress may
use a separate configured HTTPS origin. Workspace policy can delay
installation.
- Migration 0318 can replay safely on instances that applied the earlier
development migration.
- OAuth installation now links the installer to the initiating Paperclip
user. Identity checks and company access rules still apply.
- Joined channels now enable bot responses by default. Linked-user
authorization and per-action approval rules still apply.
- Model behavior has the documented delivery-check and canvas fallback
failures. A passing CI run does not establish that every model probe
passed.
- Removing the connection does not delete the customer's Slack app. No
new first-party telemetry is added.

## Model Used

OpenAI Codex, GPT-6 family, with reasoning, repository tools, code
execution, and browser verification. The runtime does not expose a more
specific authoring model ID or context-window size. The live bot probes
used OpenAI `gpt-5.6-sol` through `codex_local` in legacy mode.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-08 06:33:31 -05:00
DottaandPaperclip fc6304dfe5 feat(runner): add experimental OpenAI Dot provider over MCP Events (#15402)
## Thinking Path

> - Paperclip manages AI agents, tasks, permissions, and execution
budgets.
> - Paperclip Runner gives each provider the same admitted task and tool
authority.
> - OpenAI Dot runs outside the local process tree and needs
asynchronous work delivery.
> - The merged MCP gateway supplies OAuth consent and signed event
delivery.
> - A personal assistant grant cannot safely stand in for an assigned
agent.
> - This pull request adds a separate Dot agent connection and a durable
Rust Runner bridge.
> - The operator can assign work to Dot and inspect its accepted work,
tool receipts, and result.

## Linked Issues or Issue Description

**Agent or provider**

OpenAI Dot, as an experimental provider of the existing Paperclip Runner
adapter.

**Why this adapter is useful**

An operator can assign normal Paperclip tasks to an existing Dot. Dot
can read its mailbox, request work on an assigned task, use admitted
task tools, and submit a result. Paperclip keeps company scope,
checkout, approvals, known budget limits, and activity attribution.

**How the agent is invoked**

A dedicated `/mcp/runner` OAuth resource pairs one Dot grant with one
agent. A signed MCP mailbox event wakes Dot. Dot explicitly accepts the
assignment. The Rust Runner owns the durable turn and operation
receipts. The first release supports self-hosted instances with a local
Runner controller.

**Additional context**

This extends the merged public MCP gateway from #14846 and the assistant
invitation and device-consent work from #14933. This also integrates the
merged assistant tool and configuration expansion in #15380. Dot retains
its dedicated agent resource and cannot receive personal configuration
permission. The public assistant connection remains a personal
connection.

## What Changed

- Add a durable Rust Dot provider and its TypeScript Runner driver.
- Add closed PRP v3 external-provider operations and native execution
input v6.
- Add company-scoped pairing, mailbox, assignment, and operation
records.
- Reuse merged browser/device consent, client metadata verification,
webhook admissions, refresh, secret rotation, and warm-standby gates.
- Keep Dot scopes, issuer, grants, event workers, and tool access
separate from personal assistant access.
- Add Dot configuration, pairing, readiness, and consent UI. Keep agent
grants out of the personal Connections entry.
- Regenerate the Dot-only migration after master. Preserve published
gateway migrations. Make the new migration safe to reapply.
- Document setup, recovery, accounting limits, evidence, and remaining
account qualification.
- Reverify reconnect callbacks and wake outstanding work with a fresh
mailbox reference; preserve the existing assignment and operation
receipts.
- Clean up Dot bindings and waiting runs on OAuth revoke and
refresh-token replay. Old grants cannot revoke replacement bindings.
- Restore the pairing reference when an unsaved agent form is reopened;
document board-only pairing routes in OpenAPI.
- Accept a clean Rust exit after the acknowledged shutdown receipt.
Unexpected exits still require recovery.
- Clear the cached binding after a successful revoke so a failed
connection refresh cannot restore it.
- Add production-component Storybook states and screenshots for pairing
and connection review. All preview account data is synthetic.
- Persist normalized completion, serialize Dot turns and durable work
admission, and poll subscription readiness.
- Serialize mailbox writes and cursor reads; retain paused fence
acknowledgement without task authority.
- Authorize admitted review runs without changing the worker assignee.
Include the fenced assignment ID in production stop notices.

## Verification

- This PR integrates master `4a8178e9c`. Dot migration
`0317_messy_famine.sql` follows the published history and is safe to
reapply. The merge preserves the reserved migration connection,
batch-commit handling, private task checks, task monitors, and native
accounting.
- Local workspace typecheck, full build, and UI token gates pass. The
server typecheck passes after the review fixes. Database and native
executor regressions pass.
- All twelve real Rust/PostgreSQL Dot integration tests and twelve Dot
driver tests pass. The tests cover native document writing and
finalization, durable replay, queue admission, mailbox ordering,
admitted reviews, stale authority, production stop references, and
paused acknowledgements.
- Current head `d0e7e0626` passes all 57 checks: 53 pass and four are
intentionally skipped. This includes full typecheck, build, tests, Rust
Runner verification, browser E2E, release verification, and Canary Dry
Run. Greptile rates this exact head 5/5. All review threads are
resolved.
- The full local root test run is slower than the sharded CI run and has
not completed. The full CI test gates pass on the current commit.
Focused local regressions pass.
- Real-account pairing and event delivery on this base commit remain
unqualified. Live account and setup proof are recorded in the follow-up
#15414.

The following screenshots use synthetic preview data. They show the
production pairing component and do not qualify a real account or the
full agent setup journey.

![Synthetic pairing
preview](https://github.com/paperclipai/paperclip/blob/codex/dot-events-prototype/doc/screenshots/openai-dot-runner/pairing.jpg?raw=true)

![Synthetic connected
preview](https://github.com/paperclipai/paperclip/blob/codex/dot-events-prototype/doc/screenshots/openai-dot-runner/connected.jpg?raw=true)

## Risks

- This base adapter uses `PAPERCLIP_ENABLE_OPENAI_DOT=1` plus Public MCP
and Paperclip Runner. The separate experimental-settings follow-up in
#15414 replaces this environment flag with saved operator settings.
- Dot does not expose provider token usage or cost. The operator must
acknowledge external billing. Known Paperclip budget gates still apply.
- Cancellation fences Paperclip authority. It does not confirm that Dot
stopped all external activity.
- Assigned skill files and third-party MCP bindings are unsupported and
reject admission. There is no mounted workspace, model selector, or
provider thread identifier.
- Hosted agent-broker and remote controller deployments are not
qualified.
- The new migration follows the merged master history. Existing
prototype databases still need the normal master migration history
before this Dot-only migration.

## Model Used

OpenAI Codex, based on GPT-6. The exact deployment ID and context window
size are not exposed in this session. Capabilities used: reasoning,
repository editing, code execution, and test inspection.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #123` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub references)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-07 19:11:50 -05:00
DottaandPaperclip a7a244ab33 feat: add company decision models with permission and cost controls (#15473)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Optional product features need small, typed model decisions.
> - Each company needs to choose which shared API connection pays for
those decisions.
> - Calls must retain user and task permissions, budget limits, and cost
attribution.
> - This pull request adds a managed decision service, setup UI, and
request history.
> - Features can check availability cheaply and keep their existing
behavior when decisions are unavailable.

## Linked Issues or Issue Description

**Subsystem affected**

Server services, shared contracts, database accounting, Company
Settings, and Costs.

**Problem or motivation**

Paperclip has no common decision-model service. Adding provider calls
within each feature would duplicate credential access, permission
checks, and billing rules.

**Proposed solution**

Let a connection manager configure one company decision model. Support
OpenAI Decisions and Jev through OpenRouter. Provide a fixed setup test
and metadata-only history. Default company-sponsored background
decisions to on during setup, and preserve a saved off setting.

**Alternatives considered**

Per-feature credentials would duplicate existing connection management.
Personal overrides and provider fallback chains add permission and
billing complexity; they remain deferred.

**Roadmap alignment**

Reviewed ROADMAP.md and searched open PRs. This extends existing
connection access and budget accounting. Product features that call the
service remain outside this change. No matching decision-model service
PR was found.

## What Changed

- Add company settings, an internal `decisionModelService`, local
availability checks, and trusted human, agent/run, and system contexts.
- Pin Vercel AI SDK provider dependencies and adapt boolean, choice, and
ordered-score decisions for both providers. Bound requests and time;
disable paid retries.
- Add durable invocation metadata and agentless decision ledger charges.
Preserve fractional cents, pricing evidence, dispatch identity, and
unresolved billing holds.
- Share the company accounting lock and apply company, agent, and
project budgets. Settle charges once, retain unknown holds, and recover
interrupted calls without resubmission.
- Reuse connection setup and management UI. Add a Decisions view under
Costs, production-component Storybook coverage, database migration, and
service documentation.

## Verification

- Passed 166 current-code tests covering the decision service/provider,
setup component, Costs, OpenAPI, and every failure from the earlier
broad run. Coverage includes native SDK wire formats, refusals, billed
malformed responses, permission and secret-rotation races, identity
changes, concurrent budget admission, unresolved holds, agent/task
deletion, and stale setup feedback.
- Passed 170 existing connection, cost, budget, heartbeat-accounting,
and profile regression tests.
- Passed repository typecheck, production build, and design token gates
after integrating master. Verified the generated migration on a fresh
test database and upgraded the populated preview database from the
branch's earlier migration without losing settings or usage.
- Ran the required full `pnpm test:run`: its general phase completed
with 16,354 passed and 10 failures across five files while this branch
was still being updated. Every reported failure passes in the
current-code rerun; the serialized phase did not run after that failure.
The full GitHub CI suite passed on `b1b856a88`: general and serialized
tests, browser shards, runner checks, typecheck, production build,
packaging/canary, and policy gates. [CI
evidence](https://github.com/paperclipai/paperclip/actions/runs/37673106360).
- Passed the full-shell Storybook setup-to-history interaction test
again after integrating master. Greptile rates the final revision 5/5
with zero unresolved threads.
- Walked through the running app: empty setup, add each provider, save,
reload, run all three sample questions, inspect fractional charges in
history, switch provider while sponsorship is off, disable, and
reconnect. Checked mobile settings. These tests used the actual UI,
vault, server, SDKs, and database with simulated upstream responses.
- Live paid setup tests remain unverified: this environment has no
authorized OpenAI/OpenRouter credentials available. No mocked test is
presented as live provider evidence.

Reviewer journey: Company Settings → General → Decision model.
Add/select a shared API connection, save, run the billed sample, open
View usage, then disable decisions and verify Run test is disabled after
reload.

## Risks

- The SDK decision interface is experimental. Pinned versions and
wire-format tests limit upgrade drift.
- The migration allows agentless service charges and reservations.
Existing agent cost-reporting APIs still require an agent, and decision
receipts stay separate from run reconciliation.
- Timeouts can have unknown provider charges. Holds remain until an
audited accounting correction resolves them.
- OpenAI prices use a versioned Decisions rate snapshot; OpenRouter
costs use provider receipts. Unknown pricing is retained as unknown.
- A configured company authorizes background spending by default. Setup
explains this, and managers can turn it off.
- Live provider account/model availability still needs the two
credentialed acceptance checks.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository inspection, code
execution, and browser tools. The exact serving revision and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-07 15:53:13 -05:00
DottaandPaperclip ae6f95ed7a feat: add personal primary agents (#15470)
Add a personal primary agent per company and user. Initialize it from the first human-created agent, expose profile-only switching with confirmation, and use it after recent choices for task and Chat defaults.

Persist authenticated preferences, preserve lifecycle and membership rules, keep selections out of shared audit events, and document the API contract. Include the reviewed Storybook surfaces and regression coverage for concurrent choices, onboarding, cross-device updates, and browser journeys.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-07 13:53:46 -05:00
DottaandPaperclip 0830737030 feat(ui): add CSV table previews to task file tabs (#15445)
## Thinking Path

> - Paperclip helps people manage AI agents and inspect their work.
> - Task file tabs show the files that agents produce.
> - Markdown has a readable view, but CSV files show only source text.
> - People need to scan CSV reports without a separate app.
> - This pull request adds a table view to attachment and workspace file
tabs.
> - View buttons let people inspect the original source beside download.

## Linked Issues or Issue Description

**Subsystem affected**

ui/ — task file tabs.

**Problem or motivation**

CSV exports are difficult to scan as source text. Related work:
https://github.com/paperclipai/paperclip/pull/14297 added text
attachment tabs.

**Proposed solution**

Render CSV as a table by default. Add row numbers, sticky headers, and
record counts. Keep rendered and raw icons beside download in task tabs.

**Roadmap alignment**

This is a small improvement to existing task file previews. It does not
add a new core subsystem.

## What Changed

- Add a shared CSV parser and table preview.
- Handle quoted commas, escaped quotes, UTF-8 BOM, CRLF, and multiline
cells.
- Show at most 500 data rows and 100 columns. Show a notice for larger
files. Keep the complete download.
- Add CSV attachment and workspace previews. Put workspace tab view
controls in the header.
- Add parser and interaction tests, Storybook examples, and product
documentation.

## Verification

- 28 focused tests pass across the parser, attachment tabs, and file
viewer.
- UI typecheck, UI build, and token gates pass.
- Playwright checked desktop and mobile rendered → raw → rendered flows
using the real attachment preview component in Storybook. Screenshots
are recorded on the driving task.
- Repository-wide typecheck and build were attempted. Both require
cargo, which is absent in this environment.
- The earlier repository-wide local test run stopped after a server test
race failure. The prior commit passed all GitHub CI jobs.
- Review fixes add source-line navigation coverage and CSV boundary
tests. All 22 file-viewer tests pass. UI typecheck and token gates pass.
- All 54 active GitHub checks pass on the final commit. Two Storybook
checks are skipped. Greptile gives 5/5 with no actionable findings.
- The final version preserves HTML previews from master and uses the
shared toolbar control. All 35 focused tests, UI typecheck, UI build,
and token gates pass.

## Risks

- The first CSV record supplies column headers. Files without headers
use that same rule.
- The table preview is bounded. Download and raw view retain the
complete file.
- No API, database, or dependency changes.

## Model Used

- OpenAI Codex, GPT-6. The runtime does not expose a more precise model
ID or context window. Used code editing, shell tools, and browser
checks.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-07 09:17:18 -05:00
DottaandPaperclip 1640c5b6ab feat(ui): render HTML artifacts in a secure sandbox (#15447)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents deliver reports as task attachments and workspace files.
> - The board already renders Markdown, but it shows HTML reports as
source or rejects their preview.
> - Reports can need inline scripts to build charts and tables.
> - Artifact scripts must not read board cookies, storage, or the parent
page.
> - This pull request adds an opaque-origin HTML preview and view
controls beside Download.
> - Users can explore a report, inspect its source, and download the
original file.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

The task attachment panel and workspace file viewers.

**Subsystem affected**

ui/ and the server workspace file preview service.

**Current behavior**

HTML attachments show source text. Workspace HTML files cannot be
previewed.

**Proposed behavior**

Render self-contained HTML reports in a sandboxed iframe. Place eye and
code controls beside Download. Preserve the original file for raw view
and download.

**Reason and benefit**

Users can read and filter agent reports in Paperclip without giving
artifact scripts access to the board session.

**Breaking changes**

HTML previews now open rendered. Scripts and styles must be embedded in
the report. Remote resources, API requests, forms, popups, and host
navigation are blocked. Workspace HTML content uses the existing bounded
UTF-8 JSON response.

**Additional context**

Related closed proposals: #3293 and #4857. This change uses the current
attachment and workspace viewers and adds no report-serving endpoint.
The Artifacts and Work Products roadmap item is complete; this improves
its existing preview behavior.

## What Changed

- Add a shared HTML iframe renderer with `sandbox="allow-scripts"` and
no `allow-same-origin`.
- Install a restrictive CSP before artifact markup. Keep inline report
scripts and styles.
- Add shared rendered/raw icon controls beside Download in the
attachment panel, workspace panel, and file sheet.
- Return workspace HTML as bounded text inside JSON. Keep download and
path-access protections.
- Add Storybooks for reports, security probes, workspace viewers, a task
journey, and mobile layouts.
- Document the security boundary and browser test command.

## Verification

- `pnpm -r typecheck` and `pnpm build` pass.
- Token gates and the static Storybook build pass.
- Four browser tests pass. They cover report filters, raw mode, task
entry, workspace viewers, cookies, storage, host DOM, resource requests,
forms, popups, and host navigation.
- The focused UI suite passes 26 tests. The file-resource server suite
passes 36 tests.
- Local full-stack acceptance passes with an actual 52 KB HTML report.
Its chart and filter render. Raw mode preserves the source. Download
bytes match the original file.
- The complete local UI suite passes: 698 files and 7,754 tests.
- `pnpm test:run` was attempted. Its server group recorded two unrelated
timeouts and eight connector failures. All failed cases pass in isolated
reruns, including the complete 388-test connector suite. The remaining
local run was stopped after the full CI test coverage passed, to release
test database resources. The full local command did not complete
cleanly.
- The separate local CLI group passes 511 tests. Three worktree database
cases fail to start embedded PostgreSQL because this Mac has exhausted
its shared-memory allocation. One isolated rerun fails at the same
database startup step. No CLI code was changed. The corresponding CI
test jobs pass.
- All 56 PR checks are clean on
`c771255f1a2e86bd8825f22eafb5d14a5e342cce`. Greptile gives 5/5 with no
actionable findings. The security scan passes. The branch has no merge
conflict.
- Review the **HTML artifacts** section in Storybook. Run `pnpm exec
playwright test --config tests/html-preview/playwright.config.ts` for
the browser checks.

## Risks

- Never add `allow-same-origin` to this iframe. Its opaque origin is the
cookie, storage, and host-DOM boundary.
- Reports that depend on CDN assets need to embed their dependencies.
- The frame can navigate itself. Sandbox restrictions remain after
navigation. This is not full network isolation.
- Inline scripts can consume browser resources. This change does not
isolate CPU or memory use.
- No database migrations or API shape changes are required.

## Model Used

OpenAI Codex, based on GPT-6, with reasoning, repository editing, code
execution, and browser tool use. The runtime does not expose a more
specific model ID or context window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass — feature and UI checks
pass; the full local run has the infrastructure limits recorded above.
All CI test jobs pass.
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-07 09:01:05 -05:00
DottaandPaperclip b315580640 feat: add private task creation and sharing controls (#14718)
Add private task creation and sharing controls, effective access explanations, private project management, and locked task references.

Constrain the mobile composer plus menu to the viewport and explain named inherited privacy on hover, focus, or tap. Include 125 production-component Storybook stories and thirteen interactive user journeys.

Final head passes all CI checks and Greptile Apex 5/5 with no findings.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-07 07:27:52 -05:00
DottaandPaperclip b67db12d90 feat: add durable storage for private tasks (#14717)
Add private task and project ownership, downward access grants, and immutable run/workspace provenance. Apply migration 0313 with bounded batch commits and concurrent indexes. Keep the enforcement and sharing changes in their dependent PRs.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-07 06:53:20 -05:00
Devin FoleyandPaperclip 799e4d556f fix: make accounting durable and synchronize cost reporting (#14997)
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-06 22:26:48 -07:00
DottaandPaperclip 99a9de9940 fix(mcp): personalize assistant connections and hide revoked grants (#15411)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Assistant connections let people use their organization from another
assistant.
> - Setup begins with an invitation and returns to a list of connected
assistants.
> - Client-only labels obscure the authorizing person, while revoked
rows clutter that list.
> - This pull request shows the person’s avatar, names connections by
owner and client, and hides revoked rows.
> - It also shortens the copied invitation while keeping the approval
instructions.

## Linked Issues or Issue Description

Related: #14933 and #15380.

**What existing behavior does this improve?**

Invitation copy and assistant connection management, in the
organization’s Connections screen and the account-wide management page.

**Subsystem affected**

Cross-cutting: shared MCP connection types, server profile projection,
UI and Storybook.

**Current behavior**

Connections are labeled only with a client name such as “Codex.” Revoked
connections remain visible. The invitation includes an extra sentence
about agent identity.

**Proposed behavior**

Show the authorizing person’s avatar and use names such as “Dotta’s
Codex connection.” Hide revoked rows after successful revocation and
when loading retained revoked grants. Failed revocation leaves the
connection visible. Remove the extra identity sentence from invitation
copy.

**Reason and benefit**

Make connection identity clear and keep the list focused on usable
connections.

**Breaking changes**

The connection response adds optional `user` metadata with name and
image. Older servers remain usable. Names and revoked-row visibility
change in the UI; OAuth client identity, authorization and audit
retention remain unchanged.

## What Changed

- Shorten the shared invitation text.
- Project the authorizing person’s name and avatar through the
user-scoped connection endpoint, without returning email or credentials.
- Reuse the existing Identity component and owner naming conventions
across the connection page, catalog card and account-wide list.
- Hide revoked grants and remove a successfully revoked row from the
shared cache, even if the subsequent refresh fails.
- Update documentation, regression tests and production-page Storybook
fixtures and revocation journeys.

## Verification

- `pnpm -r typecheck`, `pnpm build`, `pnpm build-storybook` and `pnpm
check:token-gates` pass. Final UI type checks also pass.
- Focused consent and connection UI tests: 31 pass, including company
filtering, legacy metadata, custom client names, user-initial fallbacks,
failed revocation, retained revoked rows and refresh failure after
successful revocation.
- Existing MCP regression suite: 6 pass; 71 database checks are skipped
locally because embedded PostgreSQL cannot start on this machine. The
added database check verifies user-profile isolation and retained
revocation history; CI runs these checks.
- Browser verification with Storybook fixtures: owner avatar and name
render; revocation removes the selected row in both production pages,
leaves other connections visible, and restores the empty state after the
last revocation.
- All 54 current-head CI checks pass, with two optional Storybook jobs
skipped. CI includes database, browser, runner, typecheck, build and
clean-install canary coverage.
- Greptile reviewed commit `1368d79e1066b418712224378d89d64c2b11cb86`:
5/5, no actionable findings or unresolved threads.
- The full local `pnpm test:run` was stopped after complete CI passed.
Local database coverage remains unavailable because embedded PostgreSQL
cannot start; no full local-suite pass is claimed.

## Risks

Low risk. The additive profile field is optional for compatibility.
Revoked grants are filtered only from management UI and retained for
audit. Revocation failure does not hide an active connection. No
authorization scopes, token handling or schema changes.

## Model Used

OpenAI GPT-6 through Codex, with code editing, command execution and
browser verification. The exact deployment ID and context-window size
are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 22:09:41 -05:00
DottaandPaperclip 2d0c138122 Expand direct assistant MCP tools for work and configuration (#15380)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - People also use assistants in Codex, Claude, and other MCP clients.
> - The existing assistant connection can read work and create tasks or
comments.
> - It cannot edit tasks, exchange files, or manage normal agent and
project settings.
> - These operations must retain the person's permissions and
Paperclip's execution rules.
> - This pull request adds an explicit operation registry and separately
consented configuration access.
> - Assistants can manage work without receiving credentials or runner
authority.

## Linked Issues or Issue Description

**Subsystem affected**

Cross-cutting: assistant MCP, domain routes, consent UI, storage, and
Product E2E.

**Problem or motivation**

A connected assistant cannot update tasks, maintain documents, attach
files, or configure existing agents, projects, and skills. Users must
leave the assistant for these routine actions.

**Proposed solution**

Add named tools and a restricted API registry to direct connections.
Require separate configuration consent. Reuse domain routes and retry
receipts. File uploads save the attachment when the byte transfer
succeeds.

**Alternatives considered**

Arbitrary REST forwarding would expose administration and credential
operations. Runner impersonation would bypass execution ownership.
Separate upload completion calls add unnecessary client state.

**Roadmap alignment**

Checked ROADMAP.md and related MCP pull requests. This extends the
human-authorized connection from #14933. It does not replace the runner
or introduce agent impersonation.

Companion Cloud routing and directory isolation:
https://github.com/paperclipai/paperclip-cloud/pull/678.

## What Changed

- Add task editing, finish/block, documents/revisions, deliverables,
agent settings/instructions, projects/repositories, and skills/files.
- Add an allowlisted API search/call registry with identical field
restrictions, scopes, and retry identities.
- Add unchecked configuration consent. Existing write grants retain
their current authority.
- Add hashed, expiring file transfer tickets and atomic upload receipts.
No completion call is required.
- Preserve company boundaries, human attribution, active native
execution ownership, and execution review gates.
- Add protocol/domain tests, consent stories, and eight paid Product E2E
workflows.
- Repair two CI fixture races: await cold route setup before assertions,
and wait for asynchronously loaded connection copy. Both fixture suites
pass (24 + 48 tests).

## Verification

- Consent revision: one write-access checkbox controls requested work
and configuration permissions in browser and device flows. All 16
consent tests, UI typecheck/build and token gates pass. Updated
interactive stories cover default approval, opt-out and viewer
restrictions. The paid browser helper uses the new exact label. Real
GPT-5.4 Mini Product E2E passes 2/2 at
`64f96373118eb190f8cba1c2ab17cb979555f3ad` (configuration + permission
denial), campaign `local-2026-10-07T00-51-14-337Z`, no automatic
retries, cleanup passed; $0.04149375 estimated assistant cost plus
unpriced worker usage. Raw results, usage and source fingerprints are
retained in the worktree. UI and Product E2E typechecks pass.

- Prior head `2f246d4b74f1f98c75ebcb37ae6753a748237fac`: all 52 checks
pass; two optional Storybook checks skip. Greptile 5/5 on that head, no
unresolved review threads. Final consent head
`64f96373118eb190f8cba1c2ab17cb979555f3ad` also has all checks passing
and Greptile 5/5 with no unresolved threads. The unchanged Cursor
sandbox test had one 10-second timeout, passed in local isolation, and
passed its single CI rerun; the failed attempt remains in [the CI
run](https://github.com/paperclipai/paperclip/actions/runs/37554106934).
The existing chat retry-denial browser test had one visibility failure;
its single rerun passes, and the failed attempt remains in [the CI
run](https://github.com/paperclipai/paperclip/actions/runs/37542735691).

- Full workspace `pnpm -r typecheck` and `pnpm build` pass at final
runtime source `b2196fae1`. UI token gates pass.
- 139 MCP/OAuth/transfer/privacy tests and 76 grader calibration tests
pass, including one-connection PostgreSQL OAuth and concurrent upload
retries.
- Paid Product E2E: all eight expanded cases qualified across Mini,
Haiku and Sonnet. A merged-source repeat passed 23/24; one Haiku cell
timed out before application startup. Final affected-case qualification
passes 9/9 on all three models with grader v16, including the failed
cell. Automatic retries disabled; failures, costs, source hashes and
independent durable-state/file assertions are retained in [the
verification
record](doc/plans/2026-10-06-expanded-assistant-mcp-verification.md).
- Actual Codex CLI, Claude Code and OpenCode clients completed local
reads/mutations. Codex wrote a report, Claude updated it in a later
conversation, and OpenCode uploaded/downloaded a file with matching
SHA-256 and registered the attachment. Revoking the CLI grant rejects
subsequent bridge initialization.
- Butter staging is verified on final runtime `b2196fae1`
([deployment](https://github.com/paperclipai/paperclip-cloud/actions/runs/37538432138)).
A fresh OpenCode workspace fetched the copied invitation, configured
remote MCP, started OAuth and reached real consent with configuration
unchecked. Invalid transfer tickets return 403 through Cloud. Human
approval for the new persistent staging grant is pending; hosted
task/file success is not yet claimed. The final transaction fix is
deployed.
- Full local `pnpm test:run` passed 15,614 general-server tests but
stopped on two macOS timeouts. The heartbeat test passed in isolation;
the existing 40,000-file Git stress fixture timed out again. Its Linux
CI lane passes. Later local full-suite phases did not run after the
timeout; this is not an all-green local full-suite claim.
- Instructions and security limits are in `doc/public-mcp.md`; the saved
plan is `doc/plans/2026-10-06-expanded-assistant-mcp-tools.md`.

## Risks

- This expands the experimental direct MCP surface. Explicit schemas and
domain permissions must stay synchronized.
- Migration 0311 adds transfer tickets and upload receipts. Expired
orphan cleanup must not remove committed attachments.
- Configuration requires a new consent request containing that scope;
the single write-access choice controls it alongside work mutations.
Refreshing an old grant does not add it.
- The public directory keeps its original ten tools through the
companion Cloud change.
- Hosted consent/work proof remains the final delivery gate. The PR
stays draft while approval of the new staging grant is pending; code
checks and review are green. Merging is a separate action.

## Model Used

OpenAI Codex (GPT-6, tool use and code execution). The exact serving
model ID and context window are not exposed in this session. Paid
evaluation models: gpt-5.4-mini, claude-haiku-4-5-20251001;
claude-sonnet-4-6.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused suites; full-suite
macOS limitation disclosed above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 20:29:58 -05:00
Devin FoleyandPaperclip 228f0e2807 fix(ui): make pull request review waits visible (#15375)
## Thinking Path

> - Paperclip lets people supervise agent work and review its outputs.
> - A task can wait for a human to review a pull request while an agent
schedules checks.
> - The PR work product and detected external links use separate display
paths.
> - A private PR can disappear from the useful sidebar view when its
provider lookup fails.
> - The monitor countdown says when the next check runs but omits the
saved review request.
> - This change shows the saved PR and requested review in the task's
existing surfaces.

## Linked Issues or Issue Description

**What happened?**

A task kept checking whether a private GitHub PR had merged. Its work
product requested board review, but the Properties sidebar used detected
external objects instead. The PR URL was stored only in metadata, which
the PR refresh path did not read. The composer showed a generic monitor
countdown with no PR link or review action.

**Expected behavior**

Show saved PR links even if provider access fails. During a GitHub
monitor wait, make outstanding review requests visible with the next
check time and a status-check action. Stop requesting review when a PR
merges or closes.

**Steps to reproduce**

1. Register a PR work product with `reviewState: needs_board_review` and
its URL in `metadata.url`.
2. Schedule an external-service monitor for GitHub.
3. Open the task with external-object lookup unavailable or unable to
access the private repository.
4. Inspect Properties and the composer wait strip.

Related work: #14469 added rich artifact cards. #8759 addresses
attention on monitored task blockers. This change uses the existing work
products and monitor action; it adds no blocker or approval mechanism.

## What Changed

- Show saved PRs in Properties, including metadata-only links,
independent of external-object availability.
- Deduplicate equivalent GitHub PR URLs while preserving the saved
navigation link, and put explicit review requests first.
- Keep provider status and freshness on the combined PR row; keep
private PR links usable when lookup fails.
- Show the saved PR review request above the composer and in the
existing monitor banner during a GitHub monitor wait.
- Label the existing monitor action `Check status` and show check
failures inline.
- Suppress review prompts for merged, closed, or archived PRs even if
their review flag is stale.
- Refresh PR metadata using `metadata.url` and the existing `repository`
alias.
- Share saved work-product reads across the thread, Properties, and
Artifacts. Refresh GitHub in a separate query and enrich only matching
PR versions.
- Start a fresh saved-row request on live invalidation so late provider
responses cannot hide new artifacts or changed review requests.
- Cover stalled GitHub lookups in both panels so refresh latency cannot
hide saved work.
- Scope monitor-check mutation state to the task so failures and late
responses do not leak across navigation.
- Refresh GitHub status when the displayed run finishes, including when
saved PR rows have not changed.
- Clarify that PR review and external release handoffs need a saved
human-input interaction with an agent assigned for continuation; a flag,
monitor, or handoff comment alone does not create that card.

## Verification

- 386 tests pass across eleven affected UI and server suites on
`184c7d8746`. Regressions cover saved links, provider status, cold-cache
loading, panel reopening, monitor errors during navigation, and live
updates during provider refresh.
- Three live-update regressions and two run-completion regressions fail
before their fixes and pass afterward. These tests use the real API
client's GET coalescing and abort handling.
- Full repository typecheck and build passed on the merged parent
`1a49112dd2`. The latest UI changes pass UI typecheck and token gates;
CI also verifies the latest build. Capability contract/inventory checks
pass.
- Storybook build and browser checks passed before the query race fix.
Browser checks cover desktop and 390px mobile review waits, plus video,
mixed-file, and empty artifact galleries after background refresh.
- The earlier full local `pnpm test:run` was stopped under disk
pressure. It reported failures outside the changed suites; an isolated
skill-cache run reproduced three existing macOS permission failures. The
full local suite was not rerun for this follow-up.
- Latest-head CI passes on `184c7d8746`, including the browser shards
and canary dry run. Apex is 5/5 with no actionable findings and no
unresolved threads. All eight reported findings are addressed.

## Risks

- This uses saved PR review state. When GitHub access fails, the saved
state can remain stale until the agent updates it. The link remains
visible and the status check remains available.
- `Check status` wakes the existing monitor owner. It does not merge a
PR, accept an approval, or mark the task done.
- No schema, permissions, or scheduler behavior changes.

## Model Used

OpenAI GPT-6 via Codex, with reasoning, repository analysis, code
editing, and test execution. The exact deployment model ID and context
window were not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass — 277 affected tests
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 16:03:13 -07:00
DottaandPaperclip e38d6d16b6 feat(connections): add advanced provider setup and live browser qualification (#15341)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Users connect accounts and choose an agent harness and model.
> - The runtime change in #14970 supports custom providers on those
connections.
> - Normal setup must stay simple while advanced users can choose a
compatible gateway.
> - Shared connector rows and access controls keep these choices
consistent.
> - This pull request refines the agent setup UI and adds review stories
and repeatable browser qualification.
> - The qualification checks real tools and downloaded outputs, not only
a successful run status.

## Linked Issues or Issue Description

Refs #14970, #37, #13083, #14104, #14565, #12692.

The core implementation in #14970 is merged. This branch incorporates
its squash commit and targets `master`. Both PRs contain our
implementation. #14016 is a reference only and is not a dependency. This
PR has 96 changed files.

## What Changed

- Complete model-provider connector presentation beside other
connectors. Each row uses the existing Connect action and connection
list. Tags are stored without category UI. The base PR includes the
provider forms and routes.
- Show persistent Subscription, API Key, and Advanced choices. Label
Advanced as Custom Gateway. Reuse provider logos, connection lists, and
permissions controls. Default access to the organization and all agents
when permitted; keep narrowing controls under Advanced.
- Keep Configure reachable before subscription sign-in, so users can
select a supported environment when the default cannot sign in. Testing
and saving still require a connection. Show the execution environment in
Configure. Preserve the confirmed Connect choice. Editing a method,
credential, saved account, or advanced choice requires that current
choice to connect before testing or saving. Use matching model and
thinking-effort dropdowns and retain connection icons in selected
values.
- Preserve the new harness model default when switching an existing
OpenCode agent to Codex or Claude, and resolve user-selected model names
with the effective harness.
- Load popular OpenRouter models through the shared connection-model
discovery path. Keep explicit model lists and manual model entry
available.
- Group onboarding, connection setup, agent runtime, management,
recovery, and production-component stories under AI Connections /
Provider routing.
- Add an explicit-only provider-connections browser suite for managed
local or existing local/staging targets. Use private browser profiles
and credential handoffs. Support human-assisted subscription sign-in
without sharing passwords or tokens in reports.
- Verify persisted connection identity, runtime probes, tool execution,
exact artifact bytes, completion, and context-dependent follow-up.
Retain source/model provenance, cost bounds, closed error diagnostics,
original failures, and cleanup evidence.
- Add Gemini startup-model and skill-root fixes, Grok private-history
detection, ACP filesystem regression fixtures, selected-workspace
handling for local Hermes, and artifact-helper workspace fallback.
- Keep managed Grok runtime homes disposable. Remove host-side
transcript retention/restoration because private file modes do not
isolate same-user agent processes. Ignore earlier development archives
and use a fresh task handoff when history is unavailable. Verify the
absence of restored transcripts with a separate same-user process.
- Capture stopped-run diagnostics before deleting an attached-company
fixture agent. Track creation and owned sign-in receipts; revoke only
this attempt's accounts and never adopt a concurrent campaign's newly
created account. Preserve failure signals and final status through
cleanup.
- Require the requested environment in the saved agent and every run,
including follow-ups. Reject a forced incompatible target. Keep one
cancellation state through startup, every cell, reporting, and teardown
for SIGINT, SIGTERM, and SIGHUP. Stop further paid cells after
interruption. Document qualification limits.

## Verification

- Current head `b3bb3e94d577d43d9965a6b9daba039f599b2e49` includes
master `d9f600043`. The security fix in `a758fde31` passes full
workspace typecheck, production build, and 119 connection/Grok
regressions. The unchanged UI passes all 126
configuration/model-discovery tests and token gates. The final
published-guide correction passes Grok adapter typecheck. Earlier head
`eebd8225c` passed the complete deterministic runner suite (1,404 Vitest
tests and 128 Node tests) and all CI jobs. Current-head CI run
`37520147514` passed all 47 jobs, including the full sharded Vitest and
browser matrix, production build, and canary dry run. All 55 checks
completed: 53 successes and two expected skips. The current-head
security scan passed, Greptile is 5/5, and no review threads remain
open.
- A separate same-user process reproduced reading a restored Grok
transcript before the security fix. The regression now finds no
transcript. Existing fresh-session fallback and ordinary session
metadata behavior pass.
- The final account-choice and cleanup fixes pass 85 setup tests and 26
qualification-harness tests. Regressions verify that editing a
connection invalidates confirmation, Configure remains reachable before
sign-in, diagnostics are captured before fixture deletion, and
concurrent campaigns cannot adopt or revoke each other's accounts. UI
and E2E typechecks pass.
- The Storybook build and actual Chromium production-component stories
passed during this change. Review the neighboring AI Connections /
Provider routing stories, regular connector rows, three connection
modes, model discovery, and the single execution-environment control in
Configure.
- Cancellation smoke verified authenticated cleanup before browser close
for SIGINT, SIGTERM, and SIGHUP. Regressions cover interruption during
startup and reporting, missing-file ACP resource errors, and preserved
permission denials. Both ACP runtime versions and 54 ACPX/Grok
regressions passed. The deterministic connection-intent browser suite
passed two tests.
- Historical local qualification retained 43 passing API/gateway cells
out of 46, with downloaded outputs and follow-up receipts. These
attempts span earlier builds; they do not qualify this exact commit or
staging. Subscription combinations, Gemini overloads, and the unresolved
follow-up failure remain recorded rather than counted as passing.
- Use `pnpm test:e2e:runner -- --list --suite provider-connections` to
inspect the matrix. Follow `tests/runner-e2e/PROVIDER-CONNECTIONS.md`
for credentials, target URL, sign-in assistance, budget, evidence, and
cleanup. Paid live tests remain opt-in.

## Risks

- The core implementation in #14970 is merged. This PR adds no database
migration of its own.
- Subscription login needs an interactive provider session. Dedicated
accounts and staging qualification remain follow-up work; this PR does
not certify every login combination for production.
- Managed Grok transcript resume is deferred until provider history has
an OS isolation or authorized broker solution. Follow-ups start fresh
with Paperclip task context; earlier live Grok results do not qualify
this behavior.
- Gemini CLI 0.58.0 has an upstream ACP new-file error conversion
defect. Live overloads and one unresolved follow-up timeout remain
recorded. The stock CLI is unchanged, and those cases are not marked as
passing.
- Real-provider tests spend credits and use private credential/evidence
directories. The launcher requires explicit selection and checks target
ownership. It must not attach to a developer's database by accident.
- OpenClaw Gateway, Hermes Gateway, Claude Managed, AWS AgentCore,
Process, HTTP, and legacy ACPX local remain outside custom provider
setup.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, code
execution, and browser testing. The exact deployment model ID and
context window size were not exposed in this session.

## Checklist


- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 15:21:22 -05:00
88ff98b83d feat(apps): add read-only Enterpret MCP with OAuth and org tokens (#13906)
## Thinking Path

> - Paperclip agents need governed access to Enterpret's hosted MCP
tools for customer feedback.
> - The official Enterpret MCP is read-only. Enterpret Agent’s beta
write MCP is a separate service and is outside this PR.
> - This PR supports both organization auth tokens and personal browser
OAuth for the official MCP.
> - Enterpret committed to fixing OAuth scope behavior and reducing its
token-revocation cache window. The provider fixes will follow after
connector deployment and do not gate this release. Paperclip functional
validation remains required. New tools stay quarantined on refresh.
> - This PR combines the original connector and Vivek's token-path work
on current master, with live self-hosted QA recorded in the validation
guide.

## Linked Issues or Issue Description

No public issue covers this connector. This PR incorporates the work
from #14737. The independent OAuth scope-provenance fix is #14059; it is
not required for the organization-token path. #13890 is a merged
connector precedent. #13692 provides the connector-authoring skill.

**Problem or motivation**

The generic MCP setup can reach Enterpret, but it has no reviewed
Enterpret entry, method-specific setup, artwork, or policy defaults. The
original connector branch had also fallen behind master.

**Proposed solution**

Publish the Enterpret catalog entry with an organization auth token as
the selectable method. Support browser OAuth through instance-local DCR
on the same read-only endpoint. Default `run_graph_query` to Ask first
and quarantine tools newly found on refresh.

## What Changed

- Added the Enterpret app definition, generated registry, official icon,
Browse card, setup copy, tests, and Storybook states.
- Combined #14737 and resolved the current-master conflicts, including
the connector permission audit.
- Preserved catalog quarantine on token reconnect and updated expiry
guidance to follow the token's dashboard expiry. The QA token displayed
three months, contradicting the former six-month copy.
- Enabled personal OAuth for the official read-only MCP. Removed the
write-capability inference and OAuth draft-only setup restriction. The
beta Agent endpoint is excluded.
- Preserved tool quarantine during token reconnect, with a regression
test for existing quarantined and newly discovered tools.
- Added OAuth scope and revocation-cache disclosures and an interactive
Storybook retry check.
- Merged master `d9f600043` into the existing PR branch without
rewriting history. Resolved four conflicts, regenerated the combined
registry, and retained the model-provider assertions with a catalog
count of 66.
- Recorded live QA and its limits in
[doc/connections/ENTERPRET.md](https://github.com/paperclipai/paperclip/blob/950bacc52/doc/connections/ENTERPRET.md).

## Verification

Current head: `6c73d3084202ccbb25080e7a4e9c3467bc1341dd`. All 513
focused connector, discovery, OAuth, policy, shared-definition, and
UI-handoff tests passed across 11 suites. Full repository typecheck,
build, Storybook build, token gates, module boundaries, Node policy,
no-git-push policy, and typecheck-build-gaps passed. Release-registry
tests passed 129/129 with modern Bash. The four failures with macOS Bash
3.2 also reproduced on untouched master.

The full test suite passed in CI on this exact head, including every
general-test and serialized-server shard and all eight E2E shards. The
duplicate local `pnpm test:run` was stopped after these remote test
lanes passed; it did not complete locally. The combined registry was
regenerated. Regeneration also reproduces existing Linear and AgentMail
definition drift on untouched master; those unrelated changes were
excluded.

On an isolated self-hosted Paperclip instance, an organization token
connected and discovered eight actions. `get_organization_details`
returned the intended organization before and after reconnect. Catalog
refresh, invalid-token failure and valid-token recovery, activity
logging, local disable, and local removal were exercised. No feedback
quotes or graph query were requested. Paperclip revoked both QA
connection secrets and archived the QA connections.

The token-path agent access preview correctly showed an action Off, but
the board Test endpoint still invoked it as the board user. This shared
Test-path mismatch is documented; it is not evidence that a real agent
can bypass policy. An actual token-path agent-session denial, a real
agent process, expiry recovery, server/VPS, and Cloud remain unverified.

Enterpret's dashboard confirmed the QA token was revoked, but the same
token immediately reconnected and completed a safe read. Vivek explained
this as a 24-hour validity cache and committed to reducing the window.
The new maximum delay and deployment remain unverified. The earlier
OAuth check reported `mcp:write` and `email` for an `mcp:read` request.
This is a scope mismatch; it did not prove access to write tools. On
2026-10-01 the account holder relayed Vivek’s clarification that the
official MCP is read-only and the beta Agent write MCP is separate.

## Risks

The organization token can read the organization's customer feedback,
and provider-side revocation did not immediately stop access. Operators
should check each token's displayed expiry and remove Paperclip's local
connection when retiring access. The account holder confirmed the
provider fixes are fast follows after deployment. Release can proceed
with the current scope reporting and up-to-24-hour revocation delay
documented, once fresh OAuth authorization/refresh, real-agent
allow/deny, CI, and review pass. Retest the reduced revocation window
after Enterpret deploys it. Scope labels alone must not be used to claim
write access.

Greptile is 5/5 on head `6c73d3084` with no actionable findings. All
current-head CI checks pass, including the canary dry run. Live
agent-session denial and fresh OAuth refresh remain explicit deployment
QA items. The board Test panel does not prove agent authorization.
Provider scope-reporting and revocation-cache fixes are accepted fast
follows after deployment.

## Model Used

Original connector and validation record: Claude Opus 5 through Claude
Code. Integration, current-master reconciliation, and token-path QA:
OpenAI Codex, GPT-6 family (exact host variant not exposed), using
repository tools and an isolated runtime. Contributor commits retain
their authors.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either linked existing issues or described the issue in-PR
- [x] I have not referenced internal or instance-local Paperclip issues
or links
- [x] My branch name describes the change and contains no internal issue
ID
- [x] I have run relevant local tests and checks
- [x] I have updated relevant documentation
- [x] I have considered and documented the risks above
- [x] All current-head CI gates are green
- [x] Greptile is 5/5 with no actionable follow-ups

If squash-merging, preserve contributor credit in the squash message:

Co-Authored-By: Vivek Kaushal <kaushalvivek@users.noreply.github.com>

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Vivek Kaushal <vivek@enterpret.com>
2026-10-06 14:35:51 -05:00
DottaandPaperclip 22a3ea3414 Invite assistants from Connections with scoped browser and device consent (#14933)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Public MCP lets people use their organization from an external
assistant.
> - Operators need a visible control for this experimental access.
> - Hosted users should select an organization once and then approve its
permissions.
> - This pull request adds the setting, invitation-first setup, and
browser or device consent.
> - Connections provides a copyable invitation with public instructions
that grant no access.
> - Users reach browser consent from their assistant and return to
inspect or revoke access.

## Linked Issues or Issue Description

Builds on merged foundation #14846. This PR now targets master. Related
settings convention: #13905.

**Current behavior**

The preview uses an environment variable to enable MCP. Hosted consent
repeats organization selection. Assistant access has no entry in
Connections, so users must already know the endpoint and how to reach
consent.

**Proposed behavior**

An administrator enables Settings → Experimental → Assistant connections
(MCP). A hosted connection shows the selected organization and its icon,
then asks for permissions. Requested write access starts checked when
the user’s role permits it; the user can opt out before connecting.
Direct instance connections show an organization picker with the first
available organization selected. The selection stays fixed across
refetches and still requires an explicit Connect action. Connections
includes Assistant Connection (MCP). Its setup page explains the
canonical endpoint, client configuration, browser authentication, and
connected access. It connects as the current person and does not select
or impersonate an agent.

**Reason and benefit**

Operators manage access with the other experiments. Users select one
organization, and both the UI and server enforce that choice.

**Breaking changes**

The old enable variable has no effect. Preview operators must enable the
setting once. Apply the additive consent-request migration before
deploying the tenant, then deploy the compatible Cloud broker. Existing
direct requests and grants keep their behavior.

## What Changed

- Simplify OAuth and device consent: show the Paperclip logo beside
“Connect {client} to Paperclip”, fall back to “your assistant”, and show
the identifying origin plus its favicon below, with the callback URL
also visible when different. Remove the hosted-organization creation
action. Default to the first available organization without silently
changing it on refetch; preserve company restrictions and write
opt-outs. Keep the button row contained on narrow screens.
- Make Copy invitation the primary action, using the shared animated
AgentSetupPrompt and a collapsed manual setup section with icon-labeled
line tabs. Remove redundant link actions, copy-status text, the extra
first-prompt well and revocation explanation from the setup page. Serve
shared version-aware HTML and Markdown instructions without private
organization data.
- Support guarded Client ID Metadata Documents alongside dynamic
registration, and include authorization response issuer identification.
- Add RFC 8628 device authorization with separately hashed codes,
expiry, shared request quotas, persistent polling backoff and atomic
redemption. Reuse human consent, role checks, scoped grants, audit and
revocation.
- Add CLI device login and a local stdio bridge. Store credentials
separately with private permissions and serialize rotating refreshes.
- Add device consent stories and five cold-start paid Product E2E cases
with independent grant, configuration and durable-work assertions.

- Add `enablePublicMcp` to the settings validator, normalizer, feature
catalog, and toggle UI. Check it live for OAuth, tools, subscriptions,
and event delivery. Keep connection management and revocation available
while disabled.
- Default the MCP origin to the existing auth public URL, with strict
validation and an explicit override.
- Persist the optional OAuth `company_id` restriction. Describe only
that company and reject approval for any other company, even if the
person belongs to both. Keep active-membership and role checks.
- Show the Paperclip icon and a large organization icon during consent.
Return the saved company logo through the company-scoped request
response and reuse the standard fallback icon. Use the requested concise
permission labels: “Read all of your Paperclip data” and “Allow write
access and creating tasks as me”. Use concise permission copy, retain a
compact client and callback-origin disclosure, and remove the footer
link.
- Default requested write access on for eligible roles. Preserve opt-out
across organization changes and refetch, reset defaults for a new
request, and submit read-only access when the request or role does not
allow writes. Align the shared checkbox with its label.
- Use organization wording in consent, management, settings, and
walkthroughs. Keep the organization fixed for hosted requests and retain
direct-instance choice.
- Add an Assistant Connection (MCP) card to the Connectors catalog, a
setup page in the app shell, and a return link from Experimental
settings. Include Codex, Claude Code, OpenCode, and generic remote MCP
instructions.
- Read the live gate and canonical server URL through authenticated
setup metadata. Show only the current person’s grants for the selected
organization, refresh after consent, and support revocation. Surface
catalog status failures with an explicit retry action; do not present
them as an empty connection list. Opening setup grants no authority.
- Start the eight guided chapters in Connections. Keep presenter notes
and chapter controls around real product pages in the app shell. Explain
the terminal, consent, delegation, retrieval, and revocation handoffs.
Mark conversation examples as illustrative. Cover first use, client
setup, connected, loading, and error states. Keep the existing consent
and management stories.
- Keep the paid-eval setup and browser helper aligned with the setting
and consent button.

## Verification

- Warm-standby integration fix `ec64ea05e`: public MCP ingress now
follows the Cloud claim guard; MCP and discovery paths return 503
instead of SPA HTML while unclaimed. Event polling checks the in-memory
claim before reading the persisted experimental setting. All 97 focused
OAuth/Cloud tests and server typecheck pass, including new request and
timer regressions for idle-before-claim and resume-after-claim behavior.
Fresh review is 5/5 with no unresolved threads, and all security scans
pass on this final head. All browser shards, typecheck, build, canary
installation and other test groups passed on the first attempt. The
unchanged Cursor sandbox default-command test timed out at 10 seconds;
the exact test passed locally without edits in 587 ms. The single
failed-job retry passed, with the original failure retained in workflow
37500711895. All 54 final-head checks pass on
`ec64ea05e9a03e2179d4e2f84c2de03761f7ce26` (two optional Storybook jobs
are intentionally skipped).

- Final master integration `8457828fc`: merged foundation #14846 and
current master, preserving the invitation changes and all 33 files from
the two newer upstream changes. No migration renumbering was required.
All 95 focused OAuth/Cloud integration tests, full recursive typecheck
and token gates pass. All CI gates passed on that integration head;
review identified the warm-standby issue fixed above.

- Security-review fix `8c1d0b696`: commit shared global/per-source
admission before outbound CIMD work, preserve failed-attempt receipts,
and validate resource/scope before fetching. Added migration
`0305_chubby_vin_gonzales.sql` and six concurrent/adversarial regression
cases. All 69 OAuth/metadata tests, 26 migration checks, full recursive
typecheck and production build pass. The security scanner passed that
commit. Follow-up `87f9658e7` limits only actual cache-miss fetches; 18
authorization requests sharing one proxy across two service instances
use just two fetches. All 70 OAuth/metadata tests and server typecheck
pass after that refinement. Final follow-up `8ebeae84c` reports
admission-storage failures as retryable HTTP 503 instead of invalid
client metadata. Its regression proves no outbound request before
admission and successful retry after storage recovers. All 71
OAuth/metadata tests and server typecheck pass. Final-head security
scanning passes; Greptile is 5/5 with no unresolved findings. CI passed
all browser shards, typecheck, build, token gates and canary
installation. One unchanged adapter-utils bridge test raced a
response-file write (expected a JSON error, received the safe
file-changed error). The exact test passed locally without edits. The
single failed-job retry passed; the original failure is retained in
workflow 37490609192. All 54 checks now pass on final head
`8ebeae84ca77c0cf7ac12c2006f0f8743fe50e0b`, with security scan and fresh
Greptile 5/5 and no unresolved threads. Foundation #14846 subsequently
merged as `e34abee670069cca84afb2efb86041bce7dccbec`; the final
integration above now targets master.

- Integration with current master: preserved the new Connections source
filters and pagination, kept all eval suites, and regenerated the
consent/device snapshots as migrations 0303/0304. All four MCP migration
SQL hashes are unchanged from the staging versions. Full recursive
typecheck and production build, 132 focused UI tests (including catalog
filtering), 89 server authorization/settings tests, 26 migration tests,
120 eval calibration tests and token gates pass. Review follow-up
`4820ce74c` also keeps active assistant grants in Installed, with
pending/error recovery and revocation/company-isolation coverage. All 76
setup/catalog tests, UI typecheck and token gates pass after that fix.
The unchanged signoff browser test timed out waiting for a heartbeat in
CI at `4820ce74c`; the exact test passed locally without code changes,
and the preceding CI head passed that shard. That same unchanged test
failed at the reviewer stage in the next CI run. All five signoff tests
passed three times locally (15/15), without test changes. All eight
browser shards pass at final head `8ebeae84c`; no browser-test edits or
failed-browser-job retries were needed.

- Setup-page refinement at `9ab009178`: all 17 focused setup/consent
tests pass, along with UI typecheck, production build, Storybook build
and token gates. Browser exercised the shared prompt preview and client
tab switching, and the updated InvitationCopied Storybook interaction
checks its clipboard fixture. All final-head CI checks pass at
`9ab009178`, with no unresolved review findings. Deployed successfully
to Butter in
https://github.com/paperclipai/paperclip-cloud/actions/runs/37475189524.
Verified the actual page, tab switching and line styling, removed
actions/copy, and successful native copy/paste of the complete Butter
invitation into a local-only test field. The existing Claude grant was
left intact.
- Consent follow-up at `dc8e9fd11`: all 10 consent tests and token gates
pass. UI typecheck and production build passed again at `4e4d5e4d9`;
Storybook build and eval-helper typecheck passed for `28101cf91`.
Follow-ups let the primary button wrap on narrow screens, preserve a
distinct callback URL, and use only bundled icons to avoid pre-consent
requests to client-selected sites. Browser-verified the real consent
component in desktop and 320px mobile stories, including default
selection, write access and preserved opt-out. Updated E2E
heading/default-selection helpers. All CI checks passed at `dc8e9fd11`,
with review 5/5 and no unresolved threads. The Butter preview
publication needed a retry because npm initially accepted the DB package
before making it visible; the retry succeeded and `dc8e9fd11` deployed.
Verified a fresh, unapproved native Codex CIMD request on Butter:
default organization/write selection, known-client heading and icon,
distinct callback origin, and removed creation action. No grant was
approved for this UI check. Prior paid runs below retain their exact
source provenance; this UI-only follow-up did not rerun paid
qualification.
- Source-pinned paid matrix at
`2992ef2710f47230e7f484c709c6ba02524f884c`: **15/15 passed**, five cases
each on GPT-5.4 Mini, Claude Haiku and Sonnet. Campaign
`local-2026-10-06T02-41-14-462Z`. Covers cold start, existing config,
unavailable host, denied consent and reconnect/later retrieval, with
independent configuration/grant/task/run/document assertions. Original
failures, transcripts, source fingerprints and billing remain retained.
- Final instruction follow-up `cda8178af`: **3/3 cold starts passed** on
Mini, Haiku and Sonnet. Campaign `local-2026-10-06T02-58-30-041Z`.
Latest `0637b9f1c` shares that same guidance across HTML, Markdown and
manual UI after review; generated Markdown is verified byte-identical to
the paid-evaluated version. Shared build, server/UI typechecks, token
gates and 63 auth/metadata tests passed again. Every CI gate passed at
prior HEAD `0637b9f1c`, with review 5/5 and no unresolved threads.
- Other focused checks: 11 CLI credential/refresh-lock tests, 120 eval
calibration tests, server/UI/eval typechecks, token gates and Storybook
build passed. Full recursive typecheck and production build passed
during implementation; CI also passed them at `2992ef271`.
- Local full-suite limitations: a large-file Git streaming test times
out on this Mac, and broader CLI/route runs hit DB hook timeouts. Fresh
MCP reruns passed, and the corresponding CI groups passed. No claim that
the local full suite is green.
- Actual clients: Codex 0.153.4 and Claude Code 2.1.245 reach CIMD
consent; device CLI reaches verification/consent. New grants await human
approval. Existing local OpenCode retrieved a saved result in a fresh
conversation through its previously approved grant.
- Fresh OpenCode 1.18.17 on Butter: started with no MCP config, received
the exact copied invitation, read public setup, configured its server
and started PKCE consent. Its shell command timed out; background retry
reached the client's own callback deadline while approval remained
pending. Latest instructions cover that handoff. **No completed Butter
read/delegation/result retrieval is claimed.**
- Cloud companion
https://github.com/paperclipai/paperclip-cloud/pull/672 passes
checks/review and deployed. Anonymous setup and device-protocol routing
verified. Core `2992ef271` deployed successfully and the actual Claude
web flow now reaches consent. Its extra JWT-bearer metadata is filtered
to implemented grants; unsupported token grants remain rejected. Final
`0637b9f1c` deployed successfully to Butter in
https://github.com/paperclipai/paperclip-cloud/actions/runs/37409195300;
live HTML and Markdown both contain the final guidance. The superseded
instruction-only build was canceled before deployment. This is a
core-only staging preview; private Cloud plugins are omitted. ChatGPT
web is signed out, so browser connector use is unverified.
- Screenshot gallery begins at Butter's dashboard and distinguishes real
setup/pending consent from local reuse and fixtures. It records the
timeout finding. New persistent access needs human confirmation before
the remaining actual-client acceptance work.
- Manual path: Connectors → Assistant Connection (MCP) → Copy invitation
→ paste into assistant → configure and start authorization → sign in and
approve → verify `paperclip_connection` → delegate → retrieve the saved
report later.
- Plan and instructions: `doc/plans/2026-10-05-assistant-invitations.md`
and `doc/public-mcp.md`.

## Risks

- Apply additive, replay-safe migration `0304_curvy_shadow_king.sql`
before using device authorization. The public setup link carries no
credential. Device codes and tokens stay private; neither sharing
instructions nor installing a plugin authorizes access.
- Apply additive migration `0305_chubby_vin_gonzales.sql` before
deploying the shared metadata admission gate. It retains at most 60
short-lived, hashed-source receipts per instance and rejects excess
attempts with 429.
- CIMD metadata fetching is a new external-input boundary. It requires
HTTPS, exact client ID and redirect validation, bounded responses and
guarded DNS/network access. Client names remain self-reported.
- Device support is per-instance. The central Cloud broker retains its
existing grant support. Host installation and tool reload capabilities
vary by client; instructions describe manual settings and restart
requirements.

- Consent names the registered client in its heading and displays its
identifying origin below. Known-origin icons are bundled; all other
origins show a neutral site icon without contacting client-selected
sites. Client names are self-reported; the callback origin is the
recipient check. The Cloud chooser also displays the original client and
receiving origin before tenant handoff.

- A user who accepts the preselected write permission can create tasks
and comments. Task creation and comments can start or wake agents and
use execution budget; the consent label uses the concise wording
explicitly requested by the maintainer. Scope requests, role checks, and
the final Connect action still apply.
- Migration `0303_supreme_garia.sql` adds one nullable UUID column with
`IF NOT EXISTS`. Requests without a company restriction keep the
direct-instance picker. The binding stays recorded if its company is
deleted; consent then fails closed.
- Deploy tenant support before the Cloud broker sends `company_id`.
Unknown or inaccessible organizations must never fall back to a
different company.
- The setting defaults off. Disabling access does not cancel work
already delegated. Existing tokens and unexpired subscriptions can
resume when enabled again; revocation remains separate.
- The catalog entry is visible for discovery while the feature is off.
Setup instructions, OAuth, and tool execution remain gated. No access is
granted by viewing the entry.
- Assistant sign-in starts in the external client so it owns PKCE and
callback state. Client command syntax can change and links to official
setup documentation are included.
- An authenticated instance and valid public URL are required. Hosting,
paid execution, and store publication remain separate rollout steps.

## Model Used

OpenAI GPT-6 in Codex, with tool use and code execution. The exact
serving model version and context window are not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused checks pass;
unrelated local full-suite timeouts are explicitly recorded above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 12:35:16 -05:00
DottaandPaperclip f2715e02bb fix(native): surface model capacity errors and retry automatically (#15347)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Native runs preserve provider events and results, then decide task
state.
> - Codex can fail a turn because the selected model is temporarily at
capacity.
> - Paperclip displayed a generic native-session failure and left the
task in recovery.
> - A known capacity failure needs a clear message and a durable delayed
retry.
> - This pull request uses committed terminal evidence, the
status-effect ledger, and existing dispatch gates.
> - The task can continue automatically without an unbounded retry loop
or a model switch.

## Linked Issues or Issue Description

Related: #13993 adds launcher capacity deferrals before provider
startup. This change handles a committed native Codex `serverOverloaded`
terminal after provider work has started.

**What happened?**
A native Codex turn ended with `codexErrorInfo: serverOverloaded` and
“Selected model is at capacity. Please try a different model.” Paperclip
saved the result, showed a generic failure, and required recovery
instead of waiting and retrying.

**Expected behavior**
Show the capacity error directly. Schedule bounded retries after a short
delay. Preserve saved work and honor current execution gates.

**Steps to reproduce**
1. Run a native Codex task.
2. Commit a `turn.failed` event with `serverOverloaded`, then accept a
failed result for the same turn.
3. Inspect the task error and its recovery state. The regression test
reproduces this without a paid provider call.

**Paperclip version or commit**
Reproduced against master `16b7db35ffa0f9a95913c8cbdeea3d595435691f`.

## What Changed

- Classify capacity failures from committed runner events and the pinned
execution identity. Preserve accepted results and display the specific
capacity error.
- Atomically persist one scheduled successor with the status decision.
Retry after one minute, then two minutes. Share the existing failure
budget and stop after two automatic retries.
- Reuse dispatch gates for ownership, task holds, dependencies, budget,
and locks. Wait for predecessor execution, finalization, and cleanup
before claiming a retry.
- Preserve pending reviewer authority and suppress retries after
reassignment or a successor claim.
- Consume the failed run's resume receipt and delivered wake input.
Rebuild ordinary continuation from the failed run so explicitly resumed
tasks can retry without borrowing one-run authorization.
- Label scheduled retries “Model at capacity.” Add a Storybook example
and avoid duplicate punctuation in the existing retry card.
- Add regression coverage and document the runtime contract.

## Verification

- Targeted recovery and UI suites: 125 tests passed. Additional final
cleanup and lock regressions passed.
- Latest-head continuation and authorization regressions: 57 tests
passed, including all four capacity integration tests against an
isolated PostgreSQL database.
- Rechecked all four capacity integration tests with the full runner's
isolated `PAPERCLIP_HOME`, config, temporary directory, and serial fork
settings: passed.
- `pnpm -r typecheck`: passed. Final server typecheck passed.
- `pnpm build`: passed.
- `pnpm check:token-gates`: passed.
- Browser: checked the real Storybook card. It shows the capacity
message, automatic retry time, and existing Retry now action.
- `pnpm test:run`: attempted and restarted after an interruption. The
resumed run started before the final review correction and was stopped
after recorded workspace/native test failures and the five-minute Git
streaming timeout already documented on master. Exit 130; no complete
local full-suite pass is claimed. Current-head isolated capacity tests
and the complete CI suite pass.
- A combined local recovery-suite attempt also hit PostgreSQL
initialization failures at this macOS host's global shared-memory limit
(32 slots). The focused recovery suites passed separately; final-head
capacity tests also pass with the full runner's isolated environment
settings.
- Latest-head GitHub checks: all 56 checks green, including general and
serialized test shards, browser E2E, typecheck, build, Runner
verification, and canary dry run. No merge conflicts.
- Greptile: 5/5 on `813a470b8c18c05aeb7e31e63e573c3a3a2f5cac`, with zero
unresolved findings after fixing the resumed-task continuation issue.

## Risks

- Capacity retries can repeat a task turn after partial work. They start
a fresh provider session with task history and wait for predecessor
cleanup. They do not resume the failed turn.
- Retries retain the configured model unless an operator changes
configuration. Persistent overload consumes the existing failure budget
and then needs an explicit retry or model change.
- Only run-bound native Codex `serverOverloaded` failures qualify.
Usage-limit exhaustion, model/account incompatibility, unbound text, and
unknown provider failures retain their current recovery behavior.
- No schema migration is required.

## Model Used

- OpenAI GPT-6 through Codex. The exact model ID and context-window size
are not exposed to this session. Used reasoning, repository tools, code
execution, and browser verification. No subagents.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 10:59:18 -05:00
DottaandPaperclip 0e0b63e5a5 feat(connections): add experimental task-pinned AI routing (#14967)
## Thinking Path

> - Paperclip manages AI agents and their work.
> - AI Connections separate account access from models and harnesses.
> - A pool must act as one connection while retaining each task’s
account.
> - Core must enforce member access and preserve session and recovery
rules.
> - A plugin supplies rotation policy without receiving credentials.
> - This change adds durable routing and native connector setup and
management.

## Linked Issues or Issue Description

**Subsystem affected**

AI Connections, Connectors, plugins, run dispatch, and session
compatibility.

**Problem or motivation**

Operators need to rotate new tasks across saved accounts while each task
keeps its account and session. Pool setup must fit the existing
connector catalog and account workflow.

**Proposed solution**

Add an experimental router binding, a capability-gated plugin hook, and
transactional task pins. Plugins declare native pooled connectors
through `aiConnectionRouter`. Core hosts the existing-account picker,
ordering step, and account settings. Related usage contract: #14936.
Companion private plugin:
https://github.com/paperclipai/paperclip-cloud/pull/643.

**Roadmap alignment**

This extends Apps and AI Connections. Core supplies generic enforcement
and native connector UI; the private plugin owns rotation and quota
policy. The prior duplicate search found no matching router
implementation.

## What Changed

- Add a router binding without changing existing concrete bindings. Keep
the instance flag and new pools disabled by default. Require manual
operator configuration. Show no routing toggle in Experimental settings
on either open-source or Cloud installs, even after routing is enabled.
- Persist company-scoped pools, one shared cursor per pool, and pins
keyed by company, pool, agent, and task. Commit pins and cursor advances
together with revision checks and bounded retries. Persist run-ID
affinity before allocation.
- Pass only authorized metadata and normalized usage to plugins. Core
retains credential handling, member access checks, runtime
qualification, and recovery evidence. Probe outside locks with a shared
15-second budget and freshness cache.
- Resolve routing before credential preparation and backend selection.
Preserve pins through turns, session resets, removed members, and quota
waits. Retain admitted recovery after disable or uninstall.
- Separate credential session epochs from token generations. Verified
refresh preserves the epoch; reconnect and manual replacement change it.
Include the credential slot ID in session and usage-cache identity, so
reconnecting an indexed legacy account invalidates its old session even
when both epochs are zero.
- Validate pool member installations before accepting saved-agent
bindings and recheck compatibility when the harness changes. Install
only authorized members in the new-agent transaction and record their
IDs in local activity. Pool membership cannot install a restricted
shared connection.
- Preserve pool bindings when agents hire teammates through either
creation API or native caller runtime inheritance. Block stale manager
credential references; retain explicit child authentication precedence
and reject incompatible inherited pools.
- Add native connector registration through plugin metadata. Reuse the
Connectors catalog, setup header, account header, sidebar, dialogs, and
usage display. Setup selects and orders saved connections. Advanced
settings hold usage rules and member runtime defaults. New-account setup
opens in another tab.
- Use revision-checked pool archival from the Connectors catalog and
account page. Keep task pins, cursors, recovery evidence, and underlying
connections. Reject ordinary connection updates or removals that bypass
pool revisions.
- Add pool selectors, composer models, override notes, quota status, run
details, activity records, and local run-log records. Keep
session-adoption copy minimal.
- Show **Used by** below the pool connections. List current company
agents with shared avatars and profile links. Include paused agents;
exclude terminated agents and agents using another pool.
- Add Core stories for the generic connector workflow and runtime
surfaces. Cloud stories reuse these production routes and tokens through
a preview-only alias.

## Verification

- Final head `73cb953bca30ed83e4505dd820edd9b5edffd28b`: full workspace
`pnpm -r typecheck`, `pnpm build`, and `pnpm check:token-gates` pass
locally.
- All 422 focused connector/settings/shared-contract/migration tests and
all 156 database-backed AI connection, hiring, reconnect, and
durable-routing cases pass (69 hiring cases rerun after the final
auth-precedence fix). The merged shared contract retains connection
instructions and pool metadata. The pool migration is generated at
sequence 0299 after the latest upstream migrations; this PR makes no
lockfile changes.
- All four full-app Playwright tests pass on the final head after a cold
restart and migration, against the installed private plugin and isolated
database, with no route or pool-API mocks. They cover hidden routing
controls after manual opt-in, native pool creation, ordering, membership
edits, rename, paused defaults, enabling/save/refresh persistence, stale
edits, cancellation/removal, preserved underlying accounts, unavailable
routers, and Used by avatars and profile links. Exact command:
`PAPERCLIP_CONNECTION_POOL_E2E=1
AI_CONNECTIONS_TEST_COMPANY_ID=a37b9625-5ecf-4e29-8081-04df3d6e7d6f
AI_CONNECTIONS_TEST_URL=http://127.0.0.1:3108 pnpm exec playwright test
--config tests/ai-connections-app/playwright.config.ts
connection-pools.spec.ts`.
- [Native setup, ordering, and management
screenshots](https://github.com/paperclipai/paperclip/pull/14967#issuecomment-6006976278)
address the review follow-up. [Earlier selector, quota, and run-detail
screenshots](https://github.com/paperclipai/paperclip/pull/14967#issuecomment-5971537316)
show the runtime surfaces. Core previews: `pnpm --filter @paperclipai/ui
storybook`, then **Connectors / Pool host** or **AI Connections /
Connection pools**. Cloud owns its host-backed plugin stories; both
repositories’ Operator Setup Required story assertions pass.
- Live acceptance used OpenAI/Codex and Anthropic/Claude ACPX, resumed
both exact sessions after restart, preserved pinned accounts through
explicit reset and controlled quota deferral/recovery, and committed
only two allocations across fourteen runs. A later UI-created task test
again rotated OpenAI then Anthropic and resumed OpenAI through
follow-up/restart/quota recovery. That later Anthropic execution was
blocked by its saved OAuth token expiring (provider 401). No live usage
probes ran.
- The full local `pnpm test:run` was attempted earlier and did not
complete because of macOS embedded PostgreSQL bootstrap/shared-memory
failures and the 40,000-file Git fixture timeout. The focused database
suites above now pass; full-suite verification is provided by the split
CI lanes. The preceding CI run had one runtime readiness timeout; it
passes locally both alone and inside the larger runtime suite. That
larger local suite also encountered an embedded PostgreSQL setup failure
and two macOS temporary-path alias assertions; those two assertions pass
with canonical TMPDIR=/private/tmp. All final-head CI checks are
terminal green, including full general/serialized server suites, Runner
checks, browser E2E shards, canary verification, build, and typecheck.
Greptile is 5/5 on that exact head with no unresolved threads.

## Risks

- The migration adds routing tables and a credential epoch column.
Install the private plugin only with the compatible Core contract.
- Routing and each pool require opt-in. Production distribution and
fleet defaults remain unchanged.
- Unknown usage stays eligible. Known pinned exhaustion waits; revoked
access requires operator repair.
- Legacy adapters require compatible members. Runner model and effort
overrides remain limited by qualified backend support.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository editing, code
execution, and browser testing. The exact deployment model ID and
context window are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes:` / `Closes:`
/ `Refs:` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (targeted suites;
full-suite limitations are reported above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-05 20:44:37 -05:00
DottaandPaperclip 4857799a88 feat(connections): deliver saved instructions to authorized agent turns (#15216)
Persist optional connection instructions and deliver authorized snapshots to agent execution prompts. Keep provider templates with each app definition, preserve edits and opt-outs, and replace sessions when guidance or access changes.

Use shared production settings across setup and Permissions, with source visibility in agent Instructions. Add the initial memory-provider defaults and managed Honcho workspace configuration. Include migration 0298 and regression coverage for generic providers, runtime delivery, authorization, and catalog regeneration.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-05 18:16:00 -05:00
DottaandPaperclip 2c43b39167 refactor(ui): share task composer and project/worktree controls (#15091)
Use the shared composer for new tasks, including project and worktree selection, rich mentions and slash commands, and remembered task settings.

Save project preferences after successful task updates. Show known agent models by name and use Default for unknown runtime defaults.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-05 18:06:40 -05:00
DottaandPaperclip ad0c4f0767 fix(ui): keep mobile task output above the composer (#15282)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - People read agent output in the task conversation.
> - Mobile tasks scroll the document and keep the composer above the
footer.
> - The document body keeps a fixed height while the thread grows.
> - The old observer misses late thread growth, and a layout scroll
event can cancel follow mode.
> - This pull request observes the thread and preserves follow mode
until the reader scrolls up.
> - The benefit is visible new output without repeated manual scrolling.

## Linked Issues or Issue Description

**What happened?**

New task output can move behind the mobile composer while the reader is
already at the bottom. Late content layout does not resize the
fixed-height body. A scroll event after layout growth can also clear
follow mode before the resize observer runs.

**Expected behavior**

Keep new output visible above the composer while the reader follows the
bottom. Hold the reading position after an upward scroll. Resume
following when the reader returns to the bottom.

**Steps to reproduce**

1. Open a long task conversation on mobile and scroll to the bottom.
2. Let a streamed row or delayed content grow after the parent renders.
3. Check whether the latest output stays above the composer.
4. Scroll up to read earlier output, then return to the bottom.

**Paperclip version or commit**

Reproduced by regression tests against master at 6c36c07a4.

**Deployment mode**

Built from source. This is a mobile UI bug and does not depend on an
adapter or database.

Related work: #15228, #13229, and #13095. The search found no duplicate
fix for mobile document auto-follow.

## What Changed

- Observe the conversation box as well as the body for late layout
growth.
- Preserve bottom-follow intent when a layout scroll event arrives
without an upward scroll.
- Reconcile document scrolling when the viewport resizes.
- Add regressions for late growth, event ordering, reading position,
return to bottom, and viewport changes.
- Add a bounded Storybook streaming fixture with the production thread,
composer, and mobile footer.

## Verification

- 244 focused tests pass across `useWindowAutoFollow`,
`TaskMessageScroller`, and `TaskChatThread`. Three new regressions fail
before the fix.
- All 7,439 UI tests pass across 683 files (`pnpm exec vitest run
--project @paperclipai/ui`).
- `pnpm check:token-gates`, `pnpm -r typecheck`, and `pnpm build` pass.
- The full repository test shards and browser E2E checks pass in CI on
commit 43354eb38. The duplicate local `pnpm test:run` was stopped after
37 minutes once those CI test shards passed; that local run did not
complete.
- Browser check at 402 by 874: controlled streamed output stayed at
document bottom, with the newest paragraph above the composer. Scrolling
up held position as output grew. Returning to bottom resumed following.
- Open Storybook's **Task chat / Mobile streaming follow / Streaming**
story to repeat the check. The fixture makes no agent or provider calls.
- Greptile is 5/5 on commit 43354eb38. All review threads are resolved.
- All CI checks pass on the final commit, including the canary packaging
gate.

## Risks

- Scroll event order differs between browsers. The regression tests
cover a scroll event before resize reconciliation and deliberate upward
scrolling during growth.
- The change affects the mobile document scroll owner. Desktop scrolling
stays covered by its existing tests.
- Existing same-task targets and browser history restoration still pass.

## Model Used

OpenAI GPT-6 through Codex. The exact serving variant and context window
are not exposed in this session. Capabilities used: reasoning, code
editing, tool execution, and browser verification.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-05 17:33:29 -05:00
DottaandPaperclip b43073d11f feat(connections): sync and group accounts managed by aggregators (#15254)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Connections give agents governed access to external tools.
> - Aggregator gateways can expose accounts that users already connected
upstream.
> - The Apps catalog did not show those accounts or their current
provider status.
> - Separate cards and setup tasks also made account ownership unclear.
> - This pull request discovers upstream accounts and groups them under
one app card.
> - Users can find connected apps while each provider keeps control of
its accounts.

## Linked Issues or Issue Description

**Subsystem affected**

Connections across the database, shared contracts, server, and board UI.

**Problem or motivation**

Users cannot see which apps are connected through a saved aggregator
gateway. Native and upstream accounts need one app card. Discovery must
preserve company, user, gateway, and credential boundaries.

**Proposed solution**

Sync account metadata from Composio, Arcade, and supported Executor
gateways. Keep upstream account management in each provider. Use source
chips and search to browse the catalog. Preserve native setup and the
gateway's existing access policy.

**Alternatives considered**

Creating a local executable connection for each upstream account would
duplicate authorization state. Using an agent task for routine Composio
setup would add an unnecessary step. The board now calls the saved
gateway directly for that setup.

**Roadmap alignment**

This extends the shipped Connected Apps and MCP Tool Gateway features in
ROADMAP.md. The duplicate search found no open PR for managed account
discovery.

Related work: Refs #13755, Refs #13941, Refs #14725, Refs #13855. Open
PR #12906 covers adjacent toolkit routing work.

## What Changed

- Add provider-neutral discovery, sync, and refresh APIs. Preserve the
Composio API paths.
- Cache observations by company, saved gateway, viewing user, and
credential version. Retain stale observations after failed or incomplete
scans.
- Add optional Arcade account sync credentials in the vault. Discover
Executor accounts through its supported inventory interface.
- Group native and upstream accounts in one app card. Imported account
menus open their provider. Gateway menus own refresh and sync setup.
- Add Paperclip, Composio, Arcade, Installed, and All chips. Show 50
catalog entries per page. Keep connected accounts above discovery. Keep
explicit provider searches scoped.
- Simplify Composio app setup and refresh its connected app list on the
gateway Permissions page.
- Add a compact agent access card and task creation defaults for
connection setup. Preserve explicit blocks and approval policies.
- Add two replay-safe migrations, service and UI tests, Storybook
journeys, and acceptance stories.

## Verification

- Passed the repository typecheck, full build, token gates, and
migration ordering check.
- Passed the focused provider adapter, connection interaction, and
catalog tests after rebasing onto master.
- Passed all nine database sync and migration replay tests using a
disposable database on the test-drive PostgreSQL cluster. Removed that
database after the run.
- Verified Arcade cursor pagination against its official Go SDK and
passed all eight adapter tests, including short and incomplete pages.
- Passed all 45 interaction tests after making the exact requested tools
and their Allowed/Ask first permissions visible before granting access.
Verified the compact card in Storybook.
- Passed the complete UI suite on the final code: 683 files and 7,432
tests, including the corrected Composio destination assertions. Passed
130 focused tests for the UUID, management-link, and health-status
corrections.
- Passed 22 Composio setup/sync tests, 23 connection-intent service
tests, and the connection migration test in separate disposable
databases. Database startup alone was substituted; the suites exercised
their real SQL and services.
- Passed all 10 OpenAPI route checks and the full-stack
connection-intent browser test, including scoped consent, agent
continuation, and task completion.
- The local full runner encountered embedded PostgreSQL startup failures
on this loaded macOS host. The earlier in-flight run also held the
pre-fix Arcade transform; a fresh run of the final provider suite
passes. The final-head CI is queued during GitHub’s active Actions
incident: https://www.githubstatus.com/. The previous run also lost
several runners simultaneously; its real catalog assertion failures are
fixed and the fresh complete UI suite passes.
- Tested the real test-drive server in the embedded browser with a live
Composio gateway. Detected Airtable and Circleback. Verified refresh
progress, account rows, source chips, search scope, and 50-entry
pagination.
- Arcade and Executor coverage uses provider fixtures. Live credentials
were unavailable.
- Storybook builds successfully and includes grouped native/provider
accounts, stale and unavailable discovery, optional Arcade setup, and
mobile states. The acceptance document records the simulated and live
coverage separately.

- Greptile reviewed final commit
`217b024c27b5933e773ce9419c4e92b1032042c6` at 5/5. All six review
threads are resolved, security scans pass, and the PR has no merge
conflicts. The outstanding remote checks are `ci / Select trusted
runner` and `review`, queued by GitHub. They need to complete before
merge.

## Risks

- Provider response changes can break inventory discovery. Failed scans
retain observations and show stale status.
- Composio scans only the supported catalog and can take time. Large
inventories run in the background with progress and a bounded lease.
- Arcade requires a project API key and user ID when the gateway cannot
supply them. This key is used only for discovery.
- Executor discovery depends on the server's exposed inventory tools.
Unsupported servers report unavailable discovery.
- Cached account rows do not grant access or create executable
connections. Gateway policies still govern tool use. Account deletion
and per-app authorization remain upstream.
- The migrations add tables and one nullable column. Replay preserves
existing rows and company-scoped foreign keys.

## Model Used

OpenAI Codex, based on GPT-6. The session does not expose a more
specific serving model ID or context limit. Used reasoning, repository
tools, code execution, and browser verification.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-05 14:57:04 -05:00
DottaandPaperclip 7498705642 fix(ui): stabilize mobile task reading and document navigation (#15228)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - People read tasks and send instructions from phones as well as
desktop browsers.
> - The mobile footer let page text show through its labels, and small
text fields made Safari zoom on focus.
> - Small scroll changes made the footer switch direction, and its
changing page padding moved the conversation.
> - Task pages also showed comments before question cards and run
history arrived, so the composer and reading position moved again.
> - Document links also used native navigation, which reset the reading
position or reloaded a task through its UUID URL. Desktop tabs were
crowded in the mobile drawer.
> - This pull request keeps navigation steady, opens documents in the
mounted task, and gives mobile readers a full-height panel with a
vertical tab selector.
> - The benefit is a stable task view and smoother scrolling on mobile.

## Linked Issues or Issue Description

**What happened?**

The mobile footer was translucent. Safari zoomed when a person focused a
small text field. The footer switched abruptly while scrolling. A large
task could show saved comments, then move the page again when a question
card or run history arrived. In a local test with delayed responses, a
late question moved the mobile composer by about 374 pixels. Opening a
plan from the feed could reset the view or reload the task through a
UUID link. The mobile document drawer left part of the feed exposed
above small desktop tab controls.

**Expected behavior**

The footer has an opaque surface and moves smoothly after deliberate
scrolling. Text fields do not cause automatic focus zoom. A task shows
its initial conversation and composer together at the final scroll
position. Background refreshes keep the existing conversation visible.
Document links open in the mounted task with its existing cache and
reading position. Mobile documents fill the viewport, show a clear close
button, and offer a vertical list of open tabs.

**Steps to reproduce**

1. Open a task with many long comments and a pending question in iOS
Safari.
2. Delay its interactions, activity, and runs responses by different
amounts.
3. Reload the page and watch the conversation and composer move as each
response arrives.
4. Scroll down and back up, including small direction changes and edge
bounce.
5. Focus the task composer, search field, and new-task title and
description.
6. Open a plan or another task document from the feed, including a link
that uses the task UUID. Close the panel and check the reading position.
7. Open several documents on a phone. Switch tabs and close both active
and inactive tabs.

**Paperclip version or commit**

Developed from `1c07b5903` and rebased onto `59015846a`.

**Deployment mode**

Built from source. Tested in an isolated local test drive with iOS 26.5
Simulator Safari and Chrome. A temporary local proxy delayed independent
responses for the layout test.

Related work found in the duplicate search:

- Refs #14727. It made saved replies appear before supporting history.
This PR keeps its parallel requests and narrows the tradeoff in favor of
a stable first layout.
- Refs #14667. This open PR takes a different approach with per-run
placeholders and retries. This PR fixes the observed question/composer
movement and mobile navigation behavior.
- Refs #13095 and #13597. These earlier fixes added task scroll anchors
and skipped transcript waits for scheduled retries.
- Refs #6550. Earlier mobile board polish.
- Refs #9467. This related open PR changes list and generic tab reflow.
The task-pane selector uses a separate component.

## What Changed

- Give the mobile footer an opaque semantic surface.
- Set a base-size floor for editable text on touch devices to prevent
Safari focus zoom. Preserve larger title text.
- Share mobile scroll tracking between both layouts. Accumulate scroll
distance, ignore edge bounce and changed document bounds, and update
once per frame.
- Use shared motion tokens for the footer and composer. Keep page
padding stable and honor reduced motion.
- Wait for the initial question cards, attachments, work products,
activity, runtime selection, plan, and relevant transcript history
before the first reveal. Skip scheduled retries and older runs outside
the initial comment window.
- Bound the first reveal to 15 seconds. A stalled supporting request
leaves saved conversation and the composer accessible with an explicit
loading notice.
- Keep concealed mobile history from stretching the document. Keep the
composer mounted but concealed until the same reveal. Keep both visible
during later refreshes.
- Route first and repeated same-task document clicks in place. Recognize
UUID and identifier links. Preserve the thread history entry and feed
position. Keep modifier clicks, downloads, external links, and classic
document behavior.
- Give the mobile task panel the full viewport and safe-area padding.
Use a visible X and 44-pixel touch controls. Replace the horizontal tab
strip with a vertical selector that wraps titles and supports keyboard
focus.
- Add four interactive Storybook states for a few tabs, long names, many
tabs, and the last tab. Reuse the production selector and tab
controller.
- Add navigation and tab regressions, update first-reveal regressions,
and document the behavior in `DESIGN.md`.

## Verification

- 392 tests passed across the seven focused task-loading, scroll,
mobile-navigation, layout, and composer suites. After review fixes, all
339 tests across the four affected suites passed, including
stalled-loading fallback on mobile and desktop and the motion-token
catalog.
- All 442 focused document, tab, task-thread, and scroll tests pass. The
final click-propagation cleanup also passes all 136 task-detail tests.
UI typecheck, UI production build, and `pnpm check:token-gates` passed.
- All four cases in `artifact-tab-arrival.spec.ts` and
`text-attachment-tabs.spec.ts` pass locally, covering desktop and mobile
selection, composer focus, document rendering, and downloads of the
original bytes.
- `pnpm --filter @paperclipai/ui build-storybook` passed. Open the
mobile tab stories under `Prototypes/Task detail/Mobile tabs`.
- A local diagnostic proxy measured cached plan content at about 250 ms
after the first click. The HTML load count and task request count did
not change. Feed scroll stayed at the same position. First, repeated,
and UUID document links were tested at phone and desktop widths.
Task-reference links close their preview before the document reader
opens.
- In Chrome at desktop and phone widths, the delayed-response task
showed one complete reveal. The late question no longer moved an already
visible composer.
- In iOS Simulator Safari, verified the large-task reload, opaque
footer, navigation hide/reveal, and search/new-task/composer focus
without automatic zoom.
- Full workspace typecheck and build passed. The updated UI also passes
typecheck, production build, and token gates.
- All 54 checks pass on the final commit
`bf5c9914e61833e7cc8794a69d72cf8c7057b952` (two additional checks are
intentionally skipped). Greptile reviewed that commit at 5/5, and all
review threads are resolved.
- Full local `pnpm test:run` was attempted but stopped after
server-fixture failures. Embedded PostgreSQL startup failure reproduced
in an isolated native-interaction fixture after five startup attempts.
The broad run also reported a rapid Slack callback ordering test
failure. These server paths are unchanged by this PR, and their CI
shards pass on the latest head. The full local suite is not claimed as
passing.
- A localhost proxy stalled the activity response for 30 seconds. Chrome
revealed the available conversation after the 15-second deadline at both
desktop and phone widths, kept the composer accessible, and cleared the
loading notice when the response arrived.

## Risks

Slow initial history requests can delay the first conversation reveal by
up to 15 seconds. If that deadline expires, late data can change the
available conversation while a loading notice remains visible. The
reveal waits only for runs in the initial comment window, and later
refreshes do not conceal an existing conversation. The larger editable
text can change line wrapping on phones. Mobile navigation and composer
motion use shared tokens and respect reduced-motion settings. Mobile tab
selection changes the control layout. Document links retain URL history
while sharing the task reading position; other tasks and external links
keep their normal navigation behavior.

I checked `ROADMAP.md`. This is a fix for existing UI behavior.

## Model Used

OpenAI GPT-6 in Codex. The runtime does not expose a more specific model
ID or context-window size. The agent used reasoning, code editing,
terminal tools, and Chrome and iOS Simulator testing.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-05 12:33:34 -05:00
DottaandPaperclip 59015846ae fix(chat): keep dismissed task questions in the feed (#15229)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents ask users questions in task chats and Agent Chat.
> - Agent Chat keeps unanswered questions as compact entries in the
feed.
> - Regular task chats still show a pending composer badge after
dismissal.
> - A page reload can also open a dismissed question form again.
> - This pull request applies the same feed behavior to both chat types
and saves dismissal per person and task.
> - Users can continue the chat and return to the original question
later.

## Linked Issues or Issue Description

**What happened?**

Dismissing a question in a regular task chat leaves a pending composer
badge. The question can return after a page reload. The earlier feed
behavior only applied to Agent Chat.

**Expected behavior**

A dismissed question stays in the feed. Its form and pending badge leave
the composer. Reload preserves dismissal. Opening the feed entry
restores the original question and draft answer.

**Steps to reproduce**

1. Open a regular task chat with a pending question.
2. Select an option, then dismiss the form.
3. Reload the page. Check that the composer stays clear.
4. Open the question in the feed. Check that the draft is restored.
5. Submit the answer. Check that the answered receipt appears.

**Paperclip version or commit**

Reproduced on master at `a386a599983519eb1d399f8b770bfccdb2a74762`.

**Deployment mode**

Browser UI in local and authenticated instances. This change does not
depend on the agent adapter.

Related work: Refs #14613, which added the Agent Chat feed behavior.
Refs #9141, which validates real answers on the server. Refs #11434,
which tracks comment-driven changes to interaction state. This PR
changes question presentation only.

## What Changed

- Show compact unanswered question entries in regular task chats.
- Exclude durable questions from composer pending counts and navigation.
- Save dismissal in local storage per person and task. Merge the latest
saved IDs so dismissals from another tab survive reload. A new question
can still open its form.
- Keep the original question pending and answerable. Keep approval and
permission controls.
- Run the question-history regressions in both chat modes. Add reload,
new-question, user/task scope, and stale-tab coverage.
- Share the real-component Storybook fixture. Add a regular task test
drive and an interactive dismissal/answer scenario.
- Update the planning-mode browser test to check a dismissed question in
the feed after reload and on mobile.
- Update the design rules, implementation spec, and preview
instructions.

## Verification

- 329 focused thread, composer, and interaction-card tests pass.
- `pnpm -r typecheck` passes. The UI typecheck also passes after the
review fix.
- `pnpm build` passes for the full repository.
- UI build, Storybook build, and token gates pass. The UI build and
token gates were rerun after the review fix.
- Greptile gives final commit `98f71f221` a 5/5 score. The current-head
check passes, and there are no unresolved review threads.
- All 56 current-head checks are terminal: 54 pass and two conditional
Storybook jobs skip. The CI run includes the full test shards, build,
typecheck, browser tests, aggregate verification gate, and package
canary.
- The stale-tab regression fails in both chat modes before the review
fix and passes after it.
- `pnpm exec playwright test --config tests/e2e/playwright.config.ts
tests/e2e/planning-mode-visual-verification.spec.ts` passes against a
throwaway local server. It checks dismissal, reload, task navigation,
and desktop/mobile planning controls.
- The duplicate local `pnpm test:run` attempt was stopped after the full
CI test suites passed. It did not complete locally.
- Manual browser test: select Green, dismiss, reload, reopen, submit the
saved answer, and inspect the answered receipt. Also send a new message
while the unanswered question remains in the feed. The test drive uses
real UI components with fixture response callbacks.
- To repeat the browser test, run `pnpm storybook`. Open **Chat &
Comments → Task Chat Unanswered Questions → Test Drive**.

## Risks

- Dismissal is a browser-local preference. It does not sync to another
browser or device. Clearing local storage removes it.
- When local storage is unavailable, dismissal lasts for the mounted
thread only.
- Unanswered questions can accumulate in the feed. They stay pending
until answered or resolved through the existing API.
- No database migration, API change, or change to approval permissions.

## Model Used

OpenAI Codex, `gpt-6.1-sol`, with xhigh reasoning, repository editing,
code execution, and browser control. The context-window size is not
exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-05 10:55:38 -05:00
DottaandPaperclip db72ad4c73 fix(ui): explain branch artifacts without remote links (#15051)
## Thinking Path

> - Paperclip helps people manage agent work and inspect its results.
> - The issue artifact list presents those results as work-product
cards.
> - A branch can have no remote URL or structured metadata.
> - The current card hides its summary and has no action in that case.
> - This pull request shows the summary and provides an accessible
details action.
> - The card labels the absent link without making a false link.

## Linked Issues or Issue Description

**What happened?**
A branch work product without a URL showed a short title and an icon. It
did not show the saved summary or provide an action.

**Expected behavior**
The card must show what the branch contains. People must be able to open
the saved details. It must not imply that a remote URL exists.

**Steps to reproduce**
1. Open the artifact list of an issue with a branch work product that
has `url: null` and `metadata: null`.
2. Find the branch card. Its summary and details action are missing
before this change.

Related work: #12717 introduced rich work-product cards.

## What Changed

- Keep linkless branch cards concise: show the title and no-link state.
Keep the full summary behind an optional Saved description toggle.
- Label a branch with no URL as having no remote link. Let a person open
its details with a keyboard-accessible button.
- Add a focused interaction test and two Storybook states for the
no-link branch.

## Verification

- `pnpm --filter @paperclipai/ui typecheck` passed.
- Focused Vitest tests passed: 45 tests in two files.
- `pnpm --filter @paperclipai/ui build` passed.
- `pnpm --filter @paperclipai/ui build-storybook` passed.
- `pnpm check:token-gates` passed.
- The Storybook collapsed and expanded states rendered in Chromium. Both
screenshots were captured.
- Full workspace typecheck could not finish: the runner Rust package
requires `cargo`, which is not installed in this environment.

## Risks

- Low risk. The change affects card text and the action for records
without a URL. It does not change the saved data or routes.
- Long saved summaries on other linkless work products expand the card
height after a person selects Details.

## Model Used

- OpenAI Codex coding agent. The execution environment does not expose
the exact model ID or context window to this agent. It used code
execution and a browser to validate the change.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
(the managed execution workspace fixes this branch name)
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
(Storybook examples)
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-03 21:05:48 -05:00
DottaandPaperclip 1c3abf5075 fix(ui): use available space for composer labels (#15050)
## Thinking Path

> - Paperclip helps people manage AI agents and their work.
> - The task composer shows the next assignee and model before a message
starts work.
> - Fixed width limits shortened both labels when the composer had
unused space.
> - People could not read the selected agent and model even when the
full text could fit.
> - This pull request makes the capsule use the available composer
width.
> - It keeps the ellipsis when Plan mode or a narrow layout causes real
space pressure.
> - The benefit is clearer run settings without damage to the compact
composer layout.

## Linked Issues or Issue Description

**What happened?**

The task composer truncated the assignee name at 6rem and the complete
assignee and model capsule at 16rem. It did this even when the composer
had more available space.

**Expected behavior**

The composer must show the complete assignee and model labels when they
fit. It must use an ellipsis only when another control or a narrow
viewport limits the available width.

**Steps to reproduce**

1. Open a task composer with a long assignee name and a long model name.
2. Use a wide desktop layout.
3. Observe that the old capsule shortened both labels while unused space
remained.

**Paperclip version or commit**

Current `master` before this change.

**Deployment mode**

Local dev (`pnpm dev`).

**Installation method**

Built from source.

**Agent adapter(s) involved**

Not adapter-specific. This is a core UI bug.

**Database mode**

Not database-related.

**Access context**

Board.

**Additional context**

The Storybook cases cover a wide composer and a narrow composer with
Plan mode.

## What Changed

- Removed the fixed maximum width from the assignee and model capsule.
- Removed the fixed maximum width from the assignee label.
- Kept overflow ellipsis behavior when the parent row has insufficient
space.
- Added stable label selectors and focused component coverage.
- Added Storybook cases for complete labels and Plan-mode truncation.

## Verification

- `pnpm --filter @paperclipai/plugin-sdk build`
- `pnpm --filter @paperclipai/ui typecheck`
- `vitest run
ui/src/components/task-chat/ComposerRunSettingsPicker.test.tsx`
- `node scripts/check-token-gates.mjs`
- Storybook production build under Node.js 24.20.0
- Captured and inspected the two new Storybook cases.
- The complete repository typecheck and build reach the Rust Runner
step. This local environment does not have `cargo`. Hosted CI supplies
the Rust toolchain.
- The repository test suite reaches workspace-runtime tests. This local
runner does not allow their required temporary home directories. Hosted
CI supplies a writable test home.

## Risks

- Low risk. The change only removes fixed width limits from one flex
item.
- A very narrow composer can still shorten both labels. This is the
intended fallback.
- The Storybook constrained case verifies that Plan mode and Send remain
usable.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- OpenAI GPT-5 through the Paperclip Codex runner. The run used
reasoning, tool use, code execution, browser automation, and image
inspection.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-03 12:39:18 -05:00
DottaandPaperclip 7d59de6113 feat(connections): probe provider usage limits on demand (#14936)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Connections store the AI accounts used by legacy and native runners.
> - Subscription accounts can reach session, weekly, model, or paid
usage limits.
> - Operators need to read these limits for a specific stored account
before making a routing decision.
> - This pull request adds an on-demand usage probe to the connection
service and account detail.
> - The result preserves provider limits, reset times, paid usage, and
unknown values for later consumers.

## Linked Issues or Issue Description

**Subsystem affected**

Shared contracts, the connection service and API, and the account detail
UI.

**Problem or motivation**

Managed AI accounts lack a common operation to read their current usage
limits. A local harness probe can read a different login from the
account selected for an agent.

**Proposed solution**

Add `aiConnectionService.probeUsage()` and a board-only connection usage
endpoint. Probe the selected credential grant on request. Support Codex,
Claude, and Grok subscriptions, plus OpenRouter API key limits.

**Alternatives considered**

Harness-specific automatic polling would couple the read to execution
and can read ambient credentials. This change uses the managed
connection credential and leaves scheduling and admission decisions to
later work.

**Roadmap alignment**

This extends the existing Personal & Shared AI Accounts capability. It
adds no routing or quota enforcement. Related: Refs #14459 for managed
OpenAI quota reads; Refs #14781 and Refs #13379 for downstream pacing
and budget work. This operation reads one requested account across all
three subscription providers.

## What Changed

- Add typed usage snapshots and a probe capability flag to managed AI
connections.
- Normalize Codex, Claude, Grok, and OpenRouter responses. Keep model
scopes, provider admission, reset periods, and paid allowances separate.
Preserve unknown values.
- Enforce company membership, credential audience, grant identity, and
connection lifecycle before reading the stored secret.
- Add a board-only `GET
/api/companies/:companyId/ai-connections/:connectionId/usage` endpoint
with `no-store` responses.
- Add manual **Check usage** and **Refresh** actions to account details.
Show compact usage bars, resets, admission and overage status; remove
repeated descriptions and account-default copy. Clear previous results
during a new request or error.
- Add Storybook previews using the production account components for all
four providers, initial checks, loading, and permission errors.
- Add provider, authorization, runner selection, API, and UI coverage.
Document provider sources and live qualification.

## Verification

- Initial provider, authorization, selection, API, and UI validation
passed (96 focused tests): `pnpm exec vitest run
server/src/services/ai-connection-usage.test.ts
server/src/__tests__/ai-connections.test.ts
ui/src/components/ai-connections/AiConnectionUsagePanel.test.tsx
server/src/__tests__/openapi-routes.test.ts`.
- `pnpm -r typecheck` passes for the initial implementation. After
simplifying the UI, 9 usage-panel and date-helper tests, UI typecheck,
token gates, and Storybook build pass. The initial feature module
boundary check also passed.
- Real Codex, Claude, and Grok credentials were saved to encrypted
disposable connections. The actual usage HTTP route returned 200 with
`status: ok`. Legacy and native runner selection checks passed. The
tests started no model turn and exchanged no refresh token. The
disposable databases and vaults were removed.
- Live Claude responses added structured scoped limits. Live Grok
responses omitted included-plan usage. Tests now cover both shapes and
preserve the Grok omission as unknown.
- The full workspace build passes. A full local test run hit a heartbeat
feedback timeout. That case passes in isolation. The duplicate local run
was stopped after all remote checks passed. The Slack ordering and
OpenCode transport CI flakes also pass in isolation and on the CI rerun.


- Current head: `ff3d479029a1c4248190323e221b2803cfb0d79d`. All 54
active checks pass. Two Storybook checks are intentionally skipped by
the workflow. Greptile is 5/5 with no unresolved review findings; the
branch is mergeable.

## Risks

- Subscription usage endpoints can change. Credentials can lack
usage-read permission. The probe returns explicit errors without fresh
limits in these cases.
- A successful probe can contain partial data. Missing utilization or
admission remains unknown. An enabled paid-usage switch does not prove a
funded balance.
- This change adds no migration. It does not change runner admission or
automatic provider selection. Provider requests use fixed endpoints,
disabled redirects, bounded response sizes, and a 15-second deadline.

## Model Used

OpenAI Codex, GPT-6, with reasoning, file editing, shell execution, and
HTTP tools. The session does not expose the exact runtime model variant
or context window size. Real provider credentials were used only for the
authorized live checks.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-02 11:47:14 -05:00
DottaandPaperclip 6c1a75da49 feat(connections): make AgentMail a default connection with inline setup (#14772)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Connections give agents access to external services.
> - AgentMail needs both a saved key and an inbox assigned to the agent.
> - Chat requests offered a setup link instead of an inline card and
could treat a saved key as complete.
> - Inbox setup also hid address conflicts behind a generic server error
and a separate review step.
> - This pull request makes AgentMail a default connection, adds the
inline card, reduces setup to two steps, and shows conflicts beside the
address.
> - Shared native dropdown styles also give every caret a consistent
inset.

## Linked Issues or Issue Description

**What happened?**

AgentMail requests in chat did not show a usable inline connection card.
Manual setup required extra screens, ignored saved account keys, and
could trap new-address setup in a locked inbox dropdown. Agent selectors
omitted the avatar from the selected value. A taken address could
produce an HTTP 403 from AgentMail and appear as an internal server
error. Native dropdown arrows also touched the right edge of their
fields.

**Expected behavior**

Make AgentMail available as a default connection. Ask for the API key
inline, with a direct link to its provider page. Default human access to
the company and agent access to the requesting agent. Resume the agent
only after an assigned inbox is active. Manual setup should ask for an
agent and email address, then finish. Address checks should run as the
user types. Taken addresses should show clickable alternatives. A domain
dropdown beside the name should prefer a verified custom domain. Setup
should suggest authorized saved AgentMail keys and show agent avatars in
the picker and selected value.

**Steps to reproduce**

1. Ask an agent to connect AgentMail when it has no assigned inbox.
2. Check that an inline API-key card appears and links to the provider's
API-key page.
3. Open AgentMail setup, choose an agent, and request an address that is
already taken.
4. Correct the inline error, refresh, and finish setup with the same
request ID.
5. Inspect native dropdown carets in light, dark, disabled, and
right-to-left states.

Uses the bounded provider-error parser merged in #14768. Related work:
#13256 introduced AgentMail; #14725 expanded connection search.

## What Changed

- Stop recurring email queries for tasks that have no email thread.
Share the query between the thread provider and activity view. Keep
email-task updates and invalidation-based discovery.
- Make AgentMail available without the experimental chat setting. Keep
the catalog, setup and management routes, agent Channels tab, task email
feed, receiving worker, and agent tools available by default. Other
experimental chat providers stay gated.

- Make the email address and copy icon a single clickable action with
the shared Copied! confirmation. Add View inbox linking directly to the
matching AgentMail console inbox, with the address encoded as one URL
path segment.

- Reorganize inbox Settings around the copyable email address, usage
instructions, and receiving status. Move reconnect credentials into a
disclosure and separate the Disconnect action. Add production Settings
stories for active, paused, unassigned-address, revoked, webhook,
long-address, mobile, and reconnect states. Show repair controls when
the inbox has an error. Keep usage instructions tied to an active inbox
with an address.

- Add AgentMail channel intents and an inline key field with the direct
API-key URL.
- Keep setup and retry state tied to the interaction. Require an active
inbox for completion. Preserve company and agent access checks.
- Reduce manual setup to agent selection and email selection. Put the
domain dropdown beside the address and default to a verified custom
domain. Preserve explicit choices across reloads. Keep receiving
settings under Advanced options.
- Check the initial address and edits after a 350 ms pause. Abort
superseded requests and ignore stale responses. Show clickable
suggestions and retain known creation conflicts across reloads.
- Add a company-scoped, manager-only address check using the saved
credential. Search the visible inbox list instead of fetching an
uncreated inbox: live AgentMail retains negative lookups that can break
subsequent access-key creation. Unlisted addresses remain unknown;
creation is authoritative.
- Suggest labeled saved AgentMail keys in both manual setup and the
inline card. Filter by company, provider, active credential, and
current-user grants on the server. Prefer an account key and preserve
the selected key or an explicit new-key choice across refresh. Use
verified scope metadata and bounded concurrent checks for legacy keys.
Never return secret values.
- Catch an inbox-only key before the email step. Allow its existing
inbox only after an explicit choice. Recover old locked drafts at the
key picker. Save the replacement key before retiring an empty draft,
then use a new setup URL so refresh preserves the switched account; stop
if cleanup fails. Preserve already allocated addresses and their
original accounts.
- Use the shared AgentSelect in email setup. Show the canonical agent
avatar in each option and the selected value, including other consumers
of the shared component. Add regression coverage for legacy and current
Lucide agent-mention icon formats.
- Start each catalog Add connection with a fresh setup identity. Honor
Finish setup's exact draft/account/address instead of resuming an
unrelated browser draft. Return Cancel and Done to Connectors and Email
settings to the inbox. Group the task/thread explanation in a How it
Works card.
- Route AgentMail catalog removal through the email inbox control API,
including unfinished drafts. Refresh both the catalog and inbox views.
- Render each inbox management tab separately. Access uses the saved
account grants and agent controls; Conversations and Activity use the
shared persisted email feed. Activity lifecycle actions use the email
API. Reconnect returns to inbox Settings. Conversation failures show a
retry instead of a false empty state. Email delivery recovery stays in
the task.
- Map documented provider address conflicts to a field error. Preserve
actionable messages for other failures.
- Preserve non-secret draft fields across refresh, scoped to the
requested agent. Never save API keys in browser storage. Resume partial
inbox creation with the original agent, address, and request ID.
- Show an already-created address with explicit retry and new-address
recovery instead of locked inputs. Preserve the original inbox and
resumable draft when choosing another address. Distinguish runtime-key
404 errors and log safe provider status/operation/code.
- Apply final agent access once within email setup authorization for a
new account whose original installs are unchanged. Preserve later
permission edits and reused account installs. Support in-place retry of
progress loading.
- Let a failed inline setup change keys after retiring an empty draft.
Persist its replacement setup identity without storing secrets. Recover
a server-saved account when refresh interrupts the save response, while
preserving intentional account changes.
- Render the production setup in Storybook and add error, recovery, and
mobile states.
- Inset native select carets in shared CSS. Preserve custom icons,
listboxes, keyboard behavior, and forced-color controls.
- Add browser regression coverage and an AgentMail Product E2E case with
persisted-state and rendered-card evidence.

## Verification

- Full `pnpm -r typecheck`, `pnpm build`, `pnpm check:token-gates`, and
`git diff --check` passed after the default-availability change.
- All 485 focused tests passed. These cover setup, management, catalog
and route gates, connection intents, email authorization, Cursor
execution, and the OpenAPI contract. All 39 email integration tests run
with the experimental chat setting off.
- The shared polling change passed four behavioral tests, UI typecheck
and build, and token gates.
- `tests/e2e/agentmail.spec.ts` passed with the actual server setting
off. This full-stack browser test uses simulated provider responses. It
covers catalog entry, saved keys, editable address and domain controls,
creation, conflicts, retry, all management tabs, clipboard feedback, the
provider link, and task email rendering.
- In the live local browser, Add connection reached the editable email
step with the saved account key. The verified custom domain was selected
by default. Both domain choices worked. The existing inbox Settings page
remained available. Both active inboxes completed new mail checks with
the setting off. No new provider inbox or email message was created for
this pass.
- Earlier live provider acceptance covered creation on a verified custom
domain, Finish connecting on the reported draft, successful mail checks
after refresh, and catalog removal of disposable draft and active
connections. Clicking the email address copied the exact address and
showed Copied!. View inbox opened the same inbox in AgentMail’s console.
No email messages were sent.
- Production setup and Settings Storybook builds and interactions
passed. Settings states include active, paused, unassigned, revoked,
webhook, long-address, mobile, and reconnect. Receiving and
revoked-access stories had zero accessibility violations.
- Full local `pnpm test:run` on an earlier revision completed with
14,709 passing, 87 skipped, and four transient failures. All four failed
cases passed in focused reruns without product changes. That serial full
local command was not repeated after each follow-up. The latest-head
full CI suite is the final test gate.
- CI found an obsolete browser assertion that hid every channel when the
flag was off. Updated it to keep AgentMail and the Channels surface
visible while preserving the GitHub chat route gates. All 11 provider
browser tests passed locally after scoping the Channels selector to the
agent sidebar. Two initial local attempts stopped at temporary Postgres
initialization. The passing run used a separate disposable database on
the existing local Postgres server; it was removed after the test.
- Updated the remaining sidebar and aggregator discovery assertions for
default AgentMail availability. Ordinary task fixtures now return no
email thread. All 128 sidebar/task-page tests and all 42 aggregator
tests passed locally.
- Latest head `b42bb4cd5cc9f2d01a99ab8026832d5a956ea85f`: full CI
passed, with 54 successful checks including Snyk and two intentional
Storybook skips. The CI run is
https://github.com/paperclipai/paperclip/actions/runs/37020833647. A
fresh Greptile review scored 5/5 with no unresolved threads. Live model
evaluations and inbound/outbound email delivery were not run.

## Risks

- AgentMail no longer needs experimental opt-in. Setup still requires a
human to connect an account and assign an inbox. Inline setup creates an
inbox after a human submits a new or saved key. Company access, agent
access, inbox assignment, and completion checks remain enforced.
- AgentMail read APIs cannot prove global address availability. The
visible-list check is bounded to 100 entries and cannot see inboxes
outside the key’s scope. The UI reports this limitation, suggests
alternatives without claiming they are free, and keeps final creation
conflicts inline. Lookup outages show an error without preventing the
authoritative creation attempt.
- Native select CSS affects the whole app. Custom-icon selects and
multi-row lists are excluded. Forced-color mode keeps the browser caret.
- Saved-key discovery uses stored verified scope metadata and checks
authorized legacy credentials concurrently within a shared three-second
deadline. Provider outages mark legacy choices unavailable; users can
still enter another key. Final use rechecks authorization and provider
access.
- No database migration or transport default change. Live connection
remains the default.

## Model Used

OpenAI Codex, GPT-6, with reasoning, tool use, and code execution. The
exact served model ID and context-window size are not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused suites; full-suite
limitation documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green (latest head
`b42bb4cd5cc9f2d01a99ab8026832d5a956ea85f`)
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
(latest head `b42bb4cd5cc9f2d01a99ab8026832d5a956ea85f`)
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-02 10:01:15 -05:00
DottaandPaperclip 6d654f63d1 feat(apps): make MCP action test results readable (#14859)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Connected Apps let an operator control which MCP actions an agent
can use.
> - The Permissions page lets the operator run a real action as an
agent.
> - The Test dialog displayed the nested MCP response as escaped JSON.
> - A useful result was hard to read, even when the action worked.
> - This pull request renders known MCP content as a readable preview
and keeps the raw response available.
> - The benefit is faster validation without losing the data needed to
diagnose a failure.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

The per-action Test dialog on a connection's Permissions page.

**Subsystem affected**

ui/ — React board UI.

**Current behavior**

The dialog shows the gateway response as an escaped JSON blob. Text
content that contains JSON stays inside a string. The obsolete
connection Test page also keeps a separate set of stories.

**Proposed behavior**

The dialog uses structured MCP content when present. It parses JSON text
blocks when possible. It shows compact tables, cards, fields, or plain
text. It keeps the full raw response behind a control and opens that
view for errors or unknown block shapes. Stories exercise the
Permissions page dialog, and the obsolete Test page and stories are
removed.

**Reason and benefit**

An operator can inspect a successful action result at a glance and still
inspect the exact gateway response when a call fails or looks wrong.

**Breaking changes**

No API or stored data changes. The Test dialog presentation changes. The
raw response stays available.

**Additional context**

I tested a read-only Notion search through the real Permissions page.
The dialog showed three result cards and the raw response control
worked. Storybook uses invented example data.

No directly matching public issue or open PR was found in the GitHub
search.

## What Changed

- Render structured MCP output and JSON text content in the action Test
dialog.
- Show wide rows as cards, keep short rows as tables, and retain the raw
response for diagnosis.
- Remove the obsolete connection Test page and its stories.
- Add focused dialog tests and Permissions page Storybook cases for
success, errors, mixed blocks, and malformed blocks.
- Document the Test dialog result behavior in the connection playbook.
- Keep agent mention icons visible when the Lucide icon node is
unavailable in server rendering, which repaired a repeatable CI failure.

## Verification

- `pnpm -r typecheck` — passed.
- `pnpm exec vitest run --project @paperclipai/ui` — passed (7,111
tests).
- `pnpm exec vitest run
ui/src/pages/apps/app-detail/ActionTestDialog.test.tsx` — passed (11
tests).
- `pnpm exec vitest run --project @paperclipai/ui
ui/src/components/MarkdownBody.test.tsx` — passed (53 tests).
- `pnpm test:run` — started, then stopped after the review fixes changed
the head; the full sharded suite passed in CI.
- `pnpm build` — passed.
- `pnpm check:token-gates` — passed.
- Use a connected MCP app. Open Permissions, select a read action, and
run Test. Inspect the preview and the raw response control.

## Risks

- MCP tools can return provider-specific block shapes. Unknown blocks
open the raw response so the operator can inspect the exact result.
- Row and field previews limit visible data. The raw response preserves
the complete result.

> This is a targeted improvement to the existing Connected Apps item in
`ROADMAP.md`.

## Model Used

OpenAI Codex, GPT-6. The session used tool access, code execution, and
browser validation. The exact deployment ID and context window were not
exposed to the session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-01 11:52:26 -05:00
DottaandPaperclip 527e146980 feat(ui): share animated agent setup prompts (#14862)
Use the shared animated prompt-copy control across setup, invitations, webhooks, and task handoffs. Preserve first-click copying, clipboard recovery, and logo continuity. Add Storybook coverage and restore mention icon masks for the current Lucide data shape.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-01 11:46:14 -05:00
467125fafb feat(connections): one-screen connector setup with stated defaults (#14811)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Agents use Connections (the Apps catalog) to act in services like
Notion, GitHub, Google Workspace and Railway
> - Each connector asked the user to answer setup questions before it
went to the provider. Most of the questions already had the correct
answer selected
> - ROADMAP.md lists "simpler setup" for Apps and Connections as ongoing
work. This change continues that work
> - This pull request removes the questions that Paperclip can answer
itself. It states the defaults in one line and moves the choices behind
"Change" and onto the Permissions tab
> - The benefit is that most connectors take one click in Paperclip and
then the provider's own consent screen

## Linked Issues or Issue Description

No public issue exists. This is the description, from the enhancement
template.

**What existing behavior does this improve?**
The setup flow for tool connectors in the Apps catalog.

**Subsystem affected**
Apps and Connections: `ui/src/features/connections`,
`ui/src/pages/apps`, the `packages/shared` app definitions, and the
OAuth routes in `server/src/routes/tool-access.ts`.

**Current behavior**
Every connector opened with an Access step. The step asked who can use
the connection and which agents get it, and both answers were already
selected. 18 connectors also asked "How do you want to connect?" when
Paperclip could rank the methods. The Google apps and Postman also asked
"What should Paperclip be able to do?" before sign-in. The four gateway
connectors (Zapier, Arcade, Composio, Executor) used a separate two-step
wizard. Asana was pinned to a customer-owned OAuth app, so the user had
to register an app in Asana's developer console. The "Set all" control
on the Permissions tab changed only one action. After the user approved
access, Railway's consent page showed "you can close this window" and
did not return to Paperclip.

**Proposed behavior**
One screen per connector, with one primary button. The screen states the
defaults in one sentence, for example "Connects for everyone in your
organization, available to all agents". A "Change" link opens one
Advanced panel. When the provider's metadata allows dynamic client
registration, Paperclip registers a client itself. Connecting lands on
the Permissions tab. On that tab, "Set all" changes every action in the
group.

**Reason and benefit**
The user makes fewer decisions before the connection exists. Most
choices are easier to make after the connection, on the Permissions tab,
where a change has an immediate effect.

**Breaking changes**
None. No schema or API change. Existing connections keep their settings.

## What Changed

- **No Access step.** `ConnectionSetupFlow` no longer has the Access
step. The flow shows the resolved default above the primary button and
on the completion screen. The access controls moved into one Advanced
panel. The panel opens automatically only when a setting in it is
required.
- **A default method for every app.** The flow always picks the ranked
default method. Alternate methods are in the Advanced panel. The Google
and Postman capability choice is not asked before sign-in. The
write-capable method is the default.
- **Gateway connectors.** `RemoteMcpProductionSetup` (Zapier, Arcade,
Composio, Executor) no longer has its own Access step. Its commit path
and the main commit path use one helper, `askFirstCatalogEntryIdsFor`,
for server-suggested defaults.
- **Dynamic registration from live metadata.**
`canRegisterOAuthClientDynamically` now allows registration when the
provider advertises a registration endpoint, even if the catalog entry
lists only customer-owned clients. The Asana and Linear definitions and
catalog text match live probes. Asana issues clients for loopback
callbacks only, so a hosted deployment still needs an Asana app.
- **Connection setup states.** New
`packages/shared/src/connection-setup-state.ts` sorts each method into
`instant`, `authorize`, `paste` or `register`. The gallery card verb
("Connect" or "Add key") comes from this resolver and the instance's
ownership availability.
- **Generic MCP.** The generic path no longer asks "Does it need a key?"
first. A credential challenge from the server shows the key field.
- **Permissions tab.** Each action row shows its risk level. Each group
has a "Set all" control. The control sends one change for the whole
group. Before, each row's save started from the same render, so the
saves overwrote each other. The Zapier/Arcade/Composio/Executor setup
screen had the same defect.
- **OAuth callback interstitial.** A cross-site browser navigation to
`/api/tools/oauth/callback` gets a small same-origin "Finishing your
connection…" page. That page repeats the request, and the repeat does
the code exchange. Railway's consent page replaces itself after about
two seconds, and the code exchange plus tool discovery takes longer than
that. The interstitial uses only a meta refresh, because the OAuth code
is single-use. Requests without `Sec-Fetch-Site: cross-site` take the
old path.
- **Linear registers through its MCP server.** Linear pins the console
endpoints at `linear.app`. Pinned endpoints now replace discovery only
when the method cannot register, or when the connection has an
operator-entered client. So a Linear connection now finds the
registration endpoint at `mcp.linear.app`.
- **Own-OAuth-app recovery stays on the one-click screen.** When the
method also accepts a customer-owned client, the client fields are in
the Advanced panel. The panel opens after a failed sign-in. "Try again"
resumes the draft with the operator's client.
- **E2E specs** follow the one-screen flow. The Access-step clicks are
removed, the specs open **Change** before they pick agents, and they
expect GitHub's **Add key** verb.
- **Default permissions do not change.** New connections still allow
every action. The user can set actions to Ask first or Off on the
Permissions tab.

## Verification

- `cd ui && npx vitest run src/pages/apps src/features/connections
--no-file-parallelism`
- `cd packages/shared && npx vitest run src/app-definitions.test.ts
src/connection-setup-state.test.ts`
- `cd server && npx vitest run src/__tests__/tool-access-service.test.ts
src/__tests__/remote-mcp-connectors.test.ts`
- `pnpm check:token-gates`
- New tests:
- `PermissionsPanel.group.test.tsx` checks that "Set all" sends one
change for the whole group. It fails on the old code.
  - `action-permissions.test.ts` checks the group update.
  - `connection-setup-state.test.ts` checks the four setup states.
- A server test checks that a cross-site callback gets the interstitial
and does not use the OAuth state, and that the same-origin repeat
completes the connection.
- Manual check on a hosted staging deployment. GitHub, Google Drive,
Composio, Notion, PostHog and Railway each connected from one screen and
returned to the Permissions tab. On Railway, "Set all" changed all 65
write actions, and the change remained after a reload.
- Visual changes: snapshot baselines are intentionally not updated. See
the `doc/design/DECISION-SHEET.md` entry "Per-change snapshot
verification demoted to dormant (Jul 13 2026)".

## Risks

- **Fewer confirmation clicks.** Organization-wide access is the
default, and the user does not confirm it on a separate step. This was
already the preselected answer. The flow shows the default before the
user clicks and again after the connection.
- **Google write scope.** Google apps now request the write-capable
scope by default. A narrower scope needs a new sign-in.
- **Dynamic registration from live metadata.** A provider can advertise
registration and then reject a redirect URI. Asana rejects hosted
callbacks, for example. In that case registration fails, and the
customer-owned client path remains available for recovery.
- **Callback interstitial.** The OAuth callback adds one same-origin
step for cross-site browser navigations. Browsers without `Sec-Fetch-*`
headers use the old direct path.
- Chat and bot connectors (Discord, Telegram, Microsoft Teams, iMessage)
do not change.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- Claude Opus 5.5 (Anthropic), model ID `claude-opus-5-5`, used through
Claude Code with tool use (shell, file editing, browser automation) and
extended thinking. It wrote the code, the tests and this description. A
human product owner directed the work and tested it by hand.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: scotttong <squadbot000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 23:13:47 -07:00
DottaandPaperclip 33f2b3a159 fix: separate GitHub tools and code review bot connections (#14750)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The Connectors catalog lets people give agents tools or connect
agents to conversations.
> - GitHub put these two uses behind one card and an extra choice.
> - People should choose the connection they need from the catalog.
> - This pull request keeps GitHub for tools and adds GitHub Code Review
Bot as a separate card.
> - Each card opens its setup directly. Both use the existing connection
code.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

GitHub connector discovery and setup.

**Current behavior**

With chat connectors enabled, GitHub opens a menu that asks whether to
use tools or create a bot. Saved tools and bots share the same catalog
entry.

**Proposed behavior**

GitHub opens tool account access. GitHub Code Review Bot opens agent
selection. Saved bots and drafts appear under the bot card.
Chat-disabled instances show only GitHub tools.

**Reason and benefit**

The catalog names the two uses and removes an extra setup choice. The
bot keeps the existing GitHub provider, credentials, endpoint IDs, setup
steps, and runtime.

**Additional context**

Related work: https://github.com/paperclipai/paperclip/pull/12843 and
https://github.com/paperclipai/paperclip/pull/14594 established GitHub
account identity. This change preserves that tool flow. No duplicate
catalog split was found.

## What Changed

- Split the generated app definitions into GitHub tools and GitHub Code
Review Bot. Reuse the existing GitHub logo and channel method.
- Open bot setup directly, including old resume and reconnect links.
- Put existing bot endpoints and drafts under the bot card. Hide
duplicate internal chat applications.
- Keep pasted GitHub URLs mapped to the tool connection.
- Add seven Storybook states for the catalog, saved connections,
disabled chat, both setup paths, mobile, and light mode.
- Fix narrow-screen bot rows so the label cannot overlap status and
setup actions.
- Update catalog, route, browser, and API tests, plus the GitHub
connector guide.

## Verification

- [Hosted
Storybook](https://d1p6rlowie26tp.cloudfront.net/storybook/branches/codex~2Fgithub-review-connection/?path=/story/connections-github-and-code-review-bot--catalog):
seven states built from this branch. The deployment passed its
public-file verification.
- All GitHub checks pass on `d13a2cd53561645bb2a15c6f8e75a61a936d6459`.
Two optional Storybook jobs skip under their normal trigger rules; the
manual Storybook deployment passes. The branch has no merge conflicts.
- Greptile: 5/5 on the current head, with no review comments or
unresolved threads.
- `pnpm -r typecheck`, `pnpm build`, `pnpm check:token-gates`, and `pnpm
build-storybook` passed. The final Storybook fixture also passed UI
typecheck and the hosted build.
- Targeted catalog, URL matching, routing, grouping, brand, and chat UI
contract tests passed.
- GitHub provider browser tests: 2 passed. These cover direct tool setup
and the bot setup and management lifecycle with provider responses
mocked.
- Embedded-browser test on an isolated local instance: opened both
cards, selected an agent, saved a bot draft, and resumed the same
endpoint under the bot card after a reload.
- Storybook Tool Setup and Bot Setup assertions pass in the published
preview. Chat Disabled assertions pass locally. Inspected mobile and
light mode, including the draft-row layout and official GitHub marks.
- Local full-suite limitation: `pnpm test:run` was not clean. A
cross-company route assertion failed in the aggregate run and passed in
isolation; a workspace-runtime test reached its 30-second hook timeout.
Some isolated database reruns skipped when the embedded-PostgreSQL
availability probe failed. The local aggregate was stopped after CI
completed. The corresponding full CI suites pass all 360 tool-access
tests and all 162 workspace-runtime tests.
- No live GitHub authorization or installation was performed. The
isolated instance correctly stopped at the cloud enrollment or public
HTTPS prerequisites.

## Risks

- Low scope: catalog presentation and routing change. There is no
database migration or provider credential change.
- Existing GitHub bot URLs now open bot setup directly. The tool route
remains `/apps/connect?source=github`.
- The bot remains behind the existing chat-connectors feature flag.
Existing endpoints retain `provider: github`.
- Channel applications are represented by endpoint rows. Regression
tests cover legacy bot applications, tools, active bots, and drafts
together.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, code execution, and
embedded-browser tools. The exact deployed model ID, context window
size, and reasoning setting are not exposed to this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 17:15:02 -05:00
DottaandPaperclip d432dc7fa3 Add GitHub-synced skill sources (#14713)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Company skills supply instructions and files to those agents.
> - GitHub imports already exist, but users cannot manage repositories
as skill sources.
> - Repository refresh also needs caller-authorized access and complete
local packages.
> - This pull request adds Sources inside Skills and reuses GitHub
connections from Apps.
> - Installed snapshots let agents use skills without fetching GitHub
during a run.
> - Manual refresh preserves skill identity and leaves failed imports on
their last good version.

## Linked Issues or Issue Description

**Subsystem affected**

Cross-cutting: skills UI, server, database, shared contracts, and
runtime materialization.

**Problem or motivation**

Users keep skills in GitHub repositories. They need a clear way to
select, import, and refresh those skills. Existing imports do not expose
repository management or consistently preserve supporting files.

**Proposed solution**

Add company-scoped skill sources. Browse repositories from all
accessible GitHub connections, or paste a public repository or branch
URL. Select whole skill packages, inspect included files and reference
warnings, and install complete, immutable snapshots. Refresh each source
manually.

**Alternatives considered**

Project repository settings hide the workflow from Skills. A second
GitHub connector would duplicate credentials and grants. Upstream
editing and PR creation are separate work.

**Roadmap alignment**

This implements the Skills Manager direction in ROADMAP.md. The
maintainer requested this scope and reviewed the component and full-app
journey stories before implementation.

Related reports: Refs #10285, Refs #10949, Refs #13464. Related work:
#14356, #13656, #9268.

## What Changed

- Add source and entry records, an idempotent migration, company-scoped
APIs, and legacy GitHub import adoption.
- Reuse current caller grants and credential refresh. Combine and
deduplicate repository inventories across accessible connections. Pasted
public URLs also prefer the active user’s authorized connections. Tokens
stay in the Git child environment, never argv or disk.
- Fetch a shallow Git snapshot at one immutable commit. Scan the full
local tree, including hidden and nested folders. Read Git objects
without checkout or archive transformations and enforce nested package
boundaries.
- Bound Git downloads to 128 MiB and three minutes. Cancel active
process groups and remove incomplete downloads. Preserve cancellation
and deadlines while progress drains; close stalled HTTP progress streams
after 30 seconds. Reuse caller-scoped temporary snapshots for
preview/import after reauthorization.
- Index repository package boundaries once and cap expanded work at
1,000 packages, 10,000 files, and 100 MiB, including repeated copies of
shared blobs. Bound path depth and the shared path index. Discovery
keeps audited manifests without retaining all package bodies.
- Resolve moving branches before fetching so unchanged discovery reuses
caller-scoped snapshots. Limit active scans, scan frequency, and new
downloads per caller and company; quotas apply before metadata reads and
across connections, and cached scans do not consume the download quota.
- Store complete versions with script content, binary bytes, and
executable modes. Preserve these through copies, runtime caches, and
runner packaging.
- Stage downloads before publication. Use source leases, revision
checks, and transactional activity records. Keep installed versions
after failures, upstream deletion, deselection, and disconnect.
- Add the approved import flow, Sources page, selection tree,
provenance, read-only Studio behavior, and saved return from GitHub
setup.
- Add package manifests, commit-pinned file previews, and separate
runtime requirements and reference warnings. Supporting files are
included together; nested skills remain independently selectable.
Preview requests reauthorize the caller and re-audit package content.
- Show installed skills as compact links beneath each source. Repository
titles open GitHub. Keep Refresh, Select skills, and Disconnect source
in a three-dot menu. Source rows omit the branch, imported count, and
refresh timestamp; action alignment and repository titles work at narrow
widths.
- Stream discovery metadata over an opt-in NDJSON response. Show
measured Git download progress and real package/file counts, animate
newly checked skills, support cancellation, and require a complete scan
before selection. Keep the existing JSON API.
- Retain component stories and add a separate full-app journey story
group. Include fixed progress states and interactive scan,
large-repository, interruption, and saving stories.
- Update Skills documentation and product contracts. Suppress private
GitHub skill references in telemetry. Privacy review requested for the
telemetry changes.

## Verification

- Local repository typecheck, full build, token gates, and Storybook
build passed during this work. Focused transport, authorization,
scanner, persistence, route, and UI tests pass. The final UI refinement
passes all eight focused UI tests, UI typecheck/build, and token gates.
The scanner resource and repeated-discovery fixes pass 132 focused
scanner, transport, authorization, source-service, route, and rate-limit
tests, plus server typecheck/build. Full-suite verification comes from
CI; the older full local Vitest run was stopped after unrelated chat
failures and a font-test failure, all of which passed in fresh focused
runs. At commit `1098d5996`, all 54 active checks pass; two optional
Storybook jobs are skipped. CI covers repository typecheck, build, the
full test suites, browser shards, and the canary dry run. Greptile is
5/5 with no open findings; the security scan also passes.
- Adversarial scanner tests verify repeated-blob byte accounting with
and without declared sizes, package/file/path caps, one-time repository
indexing, metadata-only discovery audits, and nested package boundaries.
Additional tests cover branch movement, snapshot reuse, caller/company
quotas, isolation across connections, active-lease cleanup, quota
recovery, and rejection before any metadata API call.
- Real Git tests verify hidden paths, exact binary bytes, executable
modes, export-ignore preservation, symlink/submodule reporting, pinned
commits, caller-scoped cache reuse, cancellation, cleanup, and
credential isolation. Regression tests hold both download slots with
permanently blocked progress callbacks, verify timeout/cancellation
cleanup and retry, and exercise HTTP backpressure cancellation. Access
tests cover automatic public-URL connection selection and revoked
grants. Database tests verify company and grant audiences.
- Live isolated browser test: the public `anthropics/skills` scan now
completes and discovers all 20 skills without connecting an account.
Imported canvas-design with all 83 files, opened it from Sources, and
verified the installed binary-font preview/download control. Package
previews also expose the complete file inventory before import.
Cancelled an active Git download and retried successfully to all 20
discovered skills; the browser displayed measured download progress. The
current audits reject four other packages; eligible selections remain
importable.
- Browser checks verify the simplified source rows at desktop and narrow
widths, keyboard navigation into the actions menu, Refresh from the
menu, selection, and fixture disconnect with installed skills retained.
Storybook includes a menu-open checkpoint and a 320px layout.
- Storybook includes receiving/preparing download checkpoints and a
timed full-app import journey, plus cancellation, retry,
large-repository, and saving states. Streaming tests cover split UTF-8
frames, incomplete streams, late responses, cross-company requests, HTTP
errors, and JSON compatibility.
- Earlier live acceptance on this PR imported `stitch-skill` with
`DESIGN.md`, assigned it to an agent, disconnected its source, and ran a
successful Studio test that read both installed files. An editable copy
changed independently. Both Skills variants, mobile selection, and
return from GitHub setup were exercised.
- Private access, revoked credentials, OAuth success return,
binary/script preservation, concurrent refresh, transaction rollback,
version pins, and legacy adoption have automated coverage. A real
private-repository OAuth grant was not created during this test.

## Risks

- The migration groups recognizable legacy imports without provider
calls. Their first successful refresh completes the local package
snapshot.
- Reference checks are advisory. They cover Markdown links and explicit
relative resource paths, not arbitrary runtime dependency graphs.
Preview text is capped at 64 KiB; imported bytes remain complete.
- Git must be installed on the server. Shallow fetches still download
the branch snapshot, including files outside selected packages.
Downloads have size/time/concurrency limits. Temporary caches are
bounded and caller-scoped. GitHub API quota still applies to repository
metadata and the connection picker; content no longer uses per-file API
requests. Failed scans retain installed content.
- Sources depend on the current caller's GitHub access. A saved
connection does not grant access to another person's token.
- GitHub script support and immediate manual refresh are explicit
maintainer-approved requirements. The operator trusts the selected
repository and accepts upstream script and executable-mode changes on
refresh. Static audits are not a sandbox or a guarantee of safe code;
agents may later invoke installed helpers under their runtime
permissions. Import and refresh do not execute scripts, hooks, package
installation, or builds. Raw URL and skills.sh imports keep their prior
script restrictions.
- Source originals remain read-only. Refresh affects subsequent unpinned
runs; explicit pins and active runs retain their versions.
- The telemetry change removes source-managed GitHub identifiers from
skill-reference events. It introduces no event or field. Please review
the privacy boundary.

## Model Used

OpenAI Codex, based on GPT-6, with reasoning, code execution, and
browser tools. The exact serving model ID and context-window size are
not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 13:32:58 -05:00
DottaandPaperclip 3c561642b4 fix(chat): resolve approvals and preserve unanswered questions (#14613)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents ask for decisions and optional details through cards in chat.
> - A clear approval in a message can leave the matching card pending.
> - An unanswered question can also block an unrelated later reply.
> - Decisions need a saved source message, while optional questions need
to remain answerable in history.
> - This pull request records conversational decisions and lets users
move on from questions and answer them later.

## Linked Issues or Issue Description

**What happened?**

Native Claude and Codex could act on approval in chat while the original
approval card stayed pending. Pending question forms stayed above the
composer, were absent from history, and could suppress later chat
replies. A late native question answer could wait for a finished run to
reconnect.

**Expected behavior**

The active agent records a clear approval or refusal against the exact
card and user message. Ambiguous replies do not grant consent. Users can
send another message without answering a question. The question remains
pending in history and can be reopened and answered later. The saved
answer reaches the agent.

**Steps to reproduce**

1. Ask an agent to propose work with a confirmation card, then approve
it in chat.
2. Check that the original card records that approval before work
starts.
3. Ask an interactive question, send an unrelated message, and reload.
4. Open the unanswered question from history and submit an answer.

Related work: #14408 added completion delivery. #14607 tests completion
reporting turns. Neither records conversational answers on approval
cards.

## What Changed

- Add a confirmation endpoint backed by a user comment, with schema
validation, OpenAPI discovery, and native Plan-mode access. Ask mode
remains read-only.
- Check company, active run, actor, current session, message provenance,
revision, and resolver policy. Save the decision and audit in one
transaction. Retries do not repeat effects. Emit resolution telemetry
after commit.
- Give fresh and resumed chat turns the actual pending confirmation
identities. Teach agents to save clear conversational decisions before
acting and to clarify ambiguity.
- Keep unanswered Agent Chat questions as compact history entries. A
newer user message closes the old form. Question cards never contribute
to composer pending counts or navigation, including after dismissing a
fresh form. The history card is the sole reminder; clicking it restores
that exact form and draft.
- Preserve Agent Chat questions when later messages or questions arrive.
Historical ordinary inputs no longer gate later chat replies.
Current-run requests, task execution, and governed approvals keep their
gates. Remove the special acknowledgement-publication proof helpers that
this rule replaces.
- Route answers to finished native runs through durable fresh-wake
delivery, with existing idempotency and source-question context. Settle
late replies against contiguous completed conversation turns and freeze
their history replay; failed, unhandled, and newly arriving messages
remain actionable.
- Add real-component Storybook scenarios, database and UI regressions,
and a three-turn native Claude/Codex E2E case. Capture distinct,
UI-ready screenshots and report the individual assertions.

## Verification

- Focused decision/publication/UI regressions after merging master: 288
passed; subsequent UI draft, failed-send, and conversation checks: 199
passed.
- Native question and durable delivery regressions: 106 passed,
including all four terminal run states and exactly-once late delivery.
Seven targeted regressions fail against the original implementation and
pass with the fix.
- Latest conversation/decision/native-delivery regressions after the
master merge: 121 passed. Covers completed progress, missing or failed
intervening turns, new messages during a late reply, stale sessions, and
frozen retry/replay boundaries. Four new assertions fail before the
ordering fix.
- E2E support suite after the master merge: 792 passed. Negative
controls reject expired cards, wrong questions/answers, stale or missing
replies, unrelated clarification forms, and unexpected tasks.
- The embedded-browser walkthrough caught one additional defect:
dismissing a fresh question still showed a composer badge. Both Cancel
and close-button regressions failed before the fix. The fix at
`65f2ade12` passes 170 chat-thread tests and 792 E2E support tests.
After merging master, 232 chat-thread/confirmation tests, server/UI
typechecks, and token gates pass. The preview and two-provider live E2E
pass at `e5512a206`; Greptile is 5/5 with zero unresolved threads at
that commit. All 55 checks are now successful at `e5512a206` (four
conditional checks skipped), including the aggregate verification gate
and clean-install canary test. The first attempt was interrupted by
simultaneous CI worker shutdowns; one failed-job rerun passed without
code changes.
- [Published
Storybook](https://d1p6rlowie26tp.cloudfront.net/storybook/branches/codex~2Fchat-approval-resolution/?path=/story/chat-comments-agent-chat-unanswered-questions--moved-on):
nine real-component scenarios. Manually exercised move on, reopen,
preserve draft, answer later, answer one of multiple questions, and a
custom mobile answer in the embedded browser. Retested fresh Cancel and
close-button dismissal in the updated build, then reopened and submitted
the preserved Green selection and inspected its answered receipt. Static
preview has no live model/backend; its callbacks are fixture responses.
- [First live
campaign](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36714504406-1/)
reproduced the late-answer completion-state defect on both providers
despite correct saved answers and acknowledgements. It also exposed a
valid imperative clarification rejected by the old oracle. Both issues
are fixed with regression controls; this failing run is retained as
evidence.
- [Four-cell
qualification](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36717804064-1/)
passed 4/4 at `2bf8a1009`: unanswered-question return and ambiguous
confirmation, each on native Claude and Codex. Inspected saved state,
source-message decisions, visible cards, and agent replies. Both
late-answer chats settled to waiting; no unrequested tasks were created.
[Final branch
rerun](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36719666238-1/)
passed 2/2 at `142630720`: the same unanswered-question journey after
merging master, plus an additional screenshot and browser assertion for
the actual late-answer acknowledgement.
- [Composer-reminder
E2E](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36727006818-1/)
passed 2/2 at `5b62c52d9`: native Claude and Codex, three turns each,
with explicit no-badge assertions before and after reload. Inspected
saved pending/answered state, both screenshots with a clear composer,
and actual Blue acknowledgements; all five behavioral matchers passed
per provider and neither created tasks. Cost coverage is partial; this
is bounded workflow qualification.
- [Fresh-dismissal
E2E](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36742773318-1/)
passed 2/2 at `e5512a206`: native Claude and Codex, including fresh
Cancel, clear composer, reopen, unrelated message, reload, late Blue
answer, and actual agent acknowledgement. All five behavioral matchers
pass per provider. Inspected the fresh-dismissal screenshots and saved
pending/answered identity; neither created tasks. Cost coverage is
partial (4/6 runs).
- Prior evidence remains available in [the earlier
campaign](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36642252725-1/).
Its early loading screenshot and overwritten final capture prompted the
UI-ready, distinct screenshot fixes.

## Risks

- The model interprets intent. The server verifies permission and
provenance; it does not infer consent from text. Ambiguous and unrelated
replies are not approvals.
- Historical questions can accumulate. They remain visible, pending, and
answerable; no automatic answer or expiry is invented.
- The change to completion gates is scoped to Agent Chat and ordinary
historical inputs. Current-turn and governed approvals retain their
existing controls.
- Live qualification is limited to the selected stories. Broader native
onboarding finalization remains separate work.
- No database migration. Telemetry adds no fields or values; the
contract and README document the commit boundary. Privacy review was
requested on the PR.

## Model Used

OpenAI Codex, GPT-6 family, with reasoning, repository tools, code
execution, and browser-test orchestration. The exact model ID and
context-window size are not exposed to this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 11:46:46 -05:00
DottaandPaperclip 1b48e73e0b feat(ui): add secondary navigation for agent chat (#14706)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agent Chat already provides a persistent conversation with each
agent.
> - Its shortcuts share the primary navigation and do not give chats a
dedicated place.
> - People need to find agents, start a chat, and switch conversations
without moving the page layout.
> - This pull request adds a secondary chat sidebar and a landing page
around the existing chat surface.
> - The same conversation, composer, history, and context panel remain
in use.

## Linked Issues or Issue Description

Refs #13283 and #13420. This extends the existing experimental Agent
Chat navigation after review of the component and page stories. It
supports the CEO Chat roadmap item through the existing task-backed
conversation model.

**Subsystem affected**

The board UI and the company-scoped conversation list API.

**Current behavior**

Chat shortcuts sit inside the primary navigation. There is no dedicated
landing page with a searchable conversation list. A separate landing
header also moves the sidebar when an agent is selected.

**Proposed behavior**

Show a Chat entry in primary navigation. Keep a searchable agent sidebar
beside the chat content. The plus button starts or reopens the current
user's single conversation with that agent. Keep the header and sidebar
in the same positions before and after selection.

**Reason and benefit**

People can find agents and return to persistent conversations without
leaving the chat area or creating duplicate chats.

**Breaking changes**

The experimental chat navigation changes. Explicitly adding a chat now
resolves its conversation immediately. Direct visits to unused agent
chat URLs remain read-only. The existing per-agent routes and message
contracts remain compatible. No database migration is required.

## What Changed

- Add an account- and company-scoped conversation list endpoint with the
existing access checks, feature gate, and OpenAPI entry.
- Add the live secondary sidebar, landing page, avatars, search, loading
states, errors, and retry controls.
- Make the agent picker wait for chat creation and display failures.
Existing agents reopen the same conversation. A dismissed selection
cannot close a reopened picker or navigate over a newer choice.
- Preserve recent-activity ordering and terminated agents’ chat history.
Scope live list refreshes to the current user’s conversation events. A
failed historical-agent lookup leaves healthy chats usable and offers a
focused retry.
- Keep the sidebar and header stable across chat routes. Keep mobile
selection in the navigation drawer.
- Use the production components in Storybook. Prepare the theme and
mobile viewport before mounting the page to avoid the startup flash.
- Update product documentation, the design guide, and navigation tests.
Replace old browser expectations for stars and recent shortcuts with
persistent conversation and layout coverage.

## Verification

- `pnpm -r typecheck`, `pnpm build`, and `pnpm check:token-gates` pass
after rebasing onto master.
- Focused UI tests pass, including 105 sidebar, picker, and live-update
checks after review fixes. The 33 conversation service and route tests
and 10 OpenAPI checks pass, including ownership, feature gating, and
concurrent creation.
- The full local test run passed 14,163 server tests before three
environment or timeout failures. The embedded Postgres startup,
connector socket, and native runner failures all passed direct reruns.
- Browser test-drive verification covers a real provider reply, add and
reopen, persisted history after reload, no-match search recovery, mobile
drawer dismissal, and top-aligned context panels.
- Browser measurements confirm that the sidebar has the same position
and dimensions on the landing page and an agent conversation.
- Storybook builds and its add-and-reopen interaction passes.
- The revised browser regression passes locally against a freshly built
throwaway instance. It covers stable sidebar geometry, add/reopen
uniqueness, drafts, search, history, and terminated-agent history after
reload. The full CI browser suite also passes.
- Latest commit `b323577d9523180104df4000eaceedea2772608c`: all 54
completed checks pass, including the complete server/workspace/browser
suites, aggregate verification, build/typecheck, security scans, and
canary packaging. The two Storybook jobs are skipped by their workflow
conditions. [CI
run](https://github.com/paperclipai/paperclip/actions/runs/36714052050).
- Greptile reviewed this same commit at 5/5 with no remaining actionable
findings; all review threads are resolved.
- Reviewer path: enable Agent Chat, click Chat, use plus to choose an
agent, send a message, switch away, and reopen that agent. One
conversation must remain, with its history intact.

## Risks

- The new sidebar lists persistent conversations instead of starred and
recent shortcuts.
- Chat creation is asynchronous. Errors stay visible in the picker, and
delayed responses cannot navigate into a previous company or account.
- The shell adjustment is limited to chat routes and preserves the
existing conversation implementation.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, code execution, and browser
tools. The session does not expose the exact API model ID or
context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#123` / `Refs #123` OR (b) described the issue in-PR following the
relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 07:35:32 -05:00
DottaandPaperclip d72389bee2 feat: add Browser Use Cloud connector and live task browsers (#14627)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The Apps gateway gives agents governed access to external tools.
> - Browser Use Cloud can run browser work, but a tool result alone does
not let a person watch or take over.
> - A task needs a durable browser session, a visible viewer, and
recorded costs.
> - This pull request adds a Browser Use Cloud v4 connection and
interactive browser tabs on tasks.
> - People can follow the work, interact with the page, and retain the
browser after the agent finishes.

## Linked Issues or Issue Description

**Problem or motivation**

Agents need governed access to Browser Use Cloud. People need to see and
interact with the same browser from the task. A browser must remain
available after a run finishes and appear at the correct point in the
task feed.

**Proposed solution**

Add a native REST connection for the v4 API. Bind each session to its
company, task, agent, and credential grant. Open its interactive viewer
in the task side panel. Record provider costs as financial events. Use
`browser-use-cloud` as the app and connector key. Keep its skill with
the connector and deliver it only with authorized connection tools.

**Alternatives considered**

A v3 MCP connection would expose tools without the v4 lifecycle
integration. An external viewer link would leave the task. A fixed
viewer size would prevent pages from responding to changes in the task
pane.

**Roadmap alignment**

This extends the governed Apps gateway and Connected Apps roadmap. It
uses the existing task, grant, secret, approval, and financial records.
The work was requested by the maintainer. A search found no duplicate
Browser Use connector PR or issue.

## What Changed

- Add the Browser Use Cloud app, brand asset, API-key connection, and
profile settings under the `browser-use-cloud` key.
- Bundle the `browser-use-cloud` skill with the connector. Keep it out
of global `skills/` discovery. Deliver it only with authorized task/run
connection tools. Remove retired connector skill keys from runtime
overlays and preserve unrelated browser skills.
- Expose seven v4 tools through the governed gateway and deliver them to
native and CLI agents.
- Persist sessions, browsers, runs, event and recovery cursors, shutdown
leases, and cumulative cost accounting. Recover uncertain paid starts
without replaying them.
- Enforce task ownership, credential grants, approvals, revoked access,
and budget limits.
- Add interactive task browser tabs and compact chronological feed
entries. Retain the viewer across tab switches and keep visible idle
browsers open.
- Add debounced automatic viewport fitting, standard size presets, and a
viewer ownership lease.
- Add lifecycle, authorization, accounting, viewport, UI, and Storybook
coverage.
- Add an idempotent database migration after the current master
migration. Preserve deployed migration hashes. Migrate pre-release Cloud
connection and financial keys without replacing grants, credentials, or
browser history.
- Document provider behavior, live acceptance results, and the lack of
documented passkey forwarding.

## Verification

- Full workspace typecheck and production build pass on the updated
branch.
- Token gates, brand asset validation, module boundaries, and migration
ordering pass.
- Cloud tests verify global skill exclusion, authorized task/run
delivery, unassigned agents, disabled connections, revocation, adapter
isolation, and secret exclusion. The existing AgentMail connector
assignment test also passes.
- Migration replay runs twice against existing browser work and
financial records. It preserves the records and avoids duplicate costs.
- The focused provider, app catalog, OpenAPI, connection gateway, and
migration regression suites pass. Recovery coverage includes lost
replies, process crashes, provider rejection, and browser arrival
acknowledgement.
- All 54 checks pass on `2974b5f03641ad0cea3c941d8c02579316fa8c92`,
including the full test matrix, browser E2E shards, build, typecheck,
security, and release canary. Two optional Storybook jobs are skipped.
- Greptile is 5/5 on the same commit, with zero unresolved review
threads. The corrected review uses the actual master-to-head diff.
- The local `pnpm test:run` started and was stopped after the full CI
matrix passed. It did not complete locally; the full-suite result above
comes from CI.
- Earlier live acceptance used an isolated company with a capped
provider credential. The agent opened paperclip.ing, the embedded viewer
accepted navigation, and the same browser stayed available after
completion and tab switches.
- The local Storybook build passes. Stories cover the panel, footer,
feed entries, settings, lifecycle failures, and viewport modes with an
offline viewer fixture.

## Risks

- Browser Use charges for hosted work. Provider caps and local budget
checks reduce exposure; reported costs can arrive after work completes.
- Viewer and CDP URLs grant access to the browser. The server validates
and restricts them. They are excluded from agent results and durable
event data.
- Runtime resizing of v4 agent browsers uses a provider option confirmed
by live testing but absent from its published agent schema. Resizing
during a click may invalidate coordinates. Fixed presets remain
available.
- Viewport ownership is process-local and resets on restart. The
lifecycle and accounting records remain in the database.
- The original intermittent embedded-viewer stall has not been fully
diagnosed. A bounded reconnect and active-session recovery cover the
observed failure paths.
- Live tests did not cover every revocation, approval, rate-limit, or
restart case. Deterministic integration tests cover those paths. Passkey
forwarding is not claimed.
- Unknown create outcomes keep the credential available for cleanup.
Run-list absence cannot prove a paid POST was rejected, so recovery
stays pending until it can identify provider work.

## Model Used

OpenAI Codex, GPT-6. Used reasoning, repository search, code execution,
browser interaction, and test tools. The exact serving model ID and
context-window size were not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 07:13:16 -05:00
Devin Foley f38b5693f6 fix: always enable keyboard shortcuts (#14643)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The web UI has keyboard shortcuts for the inbox, task lists, cases,
and task detail, plus global shortcuts such as `c`, `/`, `?`, `[`, and
`]`
> - Shortcut enablement was an instance-wide General setting until
#14141 moved it to a per-user preference that defaults to off
> - The move did not carry the old instance value over, so every
existing user lost shortcuts on upgrade and had to find a new toggle
under Profile settings
> - A toggle that only turns off a standard, input-safe feature costs a
setting, a database column, two API routes, and a React context for
little benefit
> - This pull request removes both the instance setting and the personal
preference and enables keyboard shortcuts for every signed-in user
> - The benefit is one less thing to configure, no silent loss of
shortcuts on upgrade, and less code to maintain

## Linked Issues or Issue Description

Refs #14141 (the change that introduced the personal preference).

**What existing behavior does this improve?**

Keyboard shortcuts in the web UI stay off unless each user turns them on
in Profile settings.

**Subsystem affected**

Web UI shortcuts, Profile settings, instance general settings, the
`/api/auth/preferences` routes, and the `user` table.

**Current behavior**

Shortcuts default to off per user. #14141 moved the toggle from Instance
settings → General to Profile settings and did not carry the old
instance value over. Users who had shortcuts on lost them after the
upgrade and had to find the new toggle.

**Proposed behavior**

Keyboard shortcuts are always enabled for every signed-in user. There is
no instance setting and no personal preference. Shortcuts already ignore
key presses inside text inputs and modal dialogs, so an opt-out is not
needed.

**Reason and benefit**

Fewer settings, no silent loss of shortcuts on upgrade, and removal of a
database column, two API routes, a query hook, and a React context that
existed only to gate this feature.

**Breaking changes**

`GET` and `PATCH /api/auth/preferences` are removed. `PATCH
/api/instance/settings/general` no longer accepts `keyboardShortcuts`;
that schema is strict, so the key now returns 400.
`instance.general.keyboardShortcuts` is no longer a valid
`PAPERCLIP_HIDDEN_SETTINGS` key; the parser ignores unknown keys with a
warning.

## What Changed

- Removed the Keyboard shortcuts section from Profile settings, the
`useUserPreferences` hook, `queryKeys.auth.preferences`, and
`authApi.getPreferences` / `authApi.updatePreferences`.
- Removed `GeneralSettingsContext`. The inbox, legacy inbox, task list,
legacy task list, cases, and task detail pages no longer gate their key
handlers.
- Removed the `enabled` option from `useKeyboardShortcuts`. The app
shell always registers the global shortcuts.
- Removed `GET` and `PATCH /api/auth/preferences`, their OpenAPI
entries, and the `currentUserPreferencesSchema` /
`updateCurrentUserPreferencesSchema` validators.
- Removed `keyboardShortcuts` from `InstanceGeneralSettings`, the
general settings zod schema, the settings service defaults, and
`HIDEABLE_GENERAL_SECTIONS`.
- Added migration `0289_drop_user_keyboard_shortcuts`, which drops
`user.keyboard_shortcuts`.
- Updated `AGENTS.md`, `doc/SPEC.md`, `doc/SPEC-implementation.md`, and
`docs/deploy/environment-variables.md`.
- Parsed the stored general settings row with
`instanceGeneralSettingsSchema.strip()` in the feedback vote path, so a
retired key left in the row cannot reset the sharing preference to
`prompt` and overwrite the stored choice.
- Kept every bare global shortcut (`c`, `?`, `[`, `]`, `/`) out of open
modal dialogs in `useKeyboardShortcuts`; only `/` had that guard before.
- Updated the affected tests and added a Profile settings test that
asserts the toggle is gone, a hook test for the modal dialog guard, and
a feedback service regression test for the retired-key case.

## Verification

- Typecheck passes for `@paperclipai/shared`, `@paperclipai/db`
(including the migration numbering and safety checks),
`@paperclipai/server`, and `ui`.
- `pnpm exec vitest run
server/src/__tests__/instance-settings-routes.test.ts
server/src/__tests__/openapi-routes.test.ts
server/src/__tests__/auth-routes.test.ts
server/src/__tests__/sentry.test.ts` → 119 passed.
- `pnpm exec vitest run ui/src/components/Layout.test.tsx
ui/src/pages/ProfileSettings.test.tsx ui/src/pages/IssueDetail.test.tsx
ui/src/pages/Inbox.test.tsx ui/src/pages/Cases.test.tsx
ui/src/hooks/useKeyboardShortcuts.test.tsx ui/src/pages/Agents.test.tsx
ui/src/pages/InstanceGeneralSettings.test.tsx` → 286 passed.
- `pnpm exec vitest run packages/shared/src/settings-visibility.test.ts`
→ 16 passed.
- `pnpm exec vitest run ui/src/hooks/useKeyboardShortcuts.test.tsx` → 7
passed.
- `pnpm exec vitest run server/src/__tests__/feedback-service.test.ts`
(embedded Postgres) → the new retired-key test passes with the fix and
fails without it.
- Manual: sign in with no settings changed, open the inbox, press `j`
and `k` to move the selection, press `?` to open the cheatsheet. Open
Settings → Profile and confirm there is no Keyboard shortcuts section.

## Risks

- The migration drops a column. It uses `DROP COLUMN IF EXISTS`, and the
column has no readers after this change. If you roll back to a build
from before this PR after the migration has run, re-add the column
first: `ALTER TABLE "user" ADD COLUMN "keyboard_shortcuts" boolean
DEFAULT false NOT NULL;`. The older build's ORM selects that column when
it loads users.
- Any external client that still sends `keyboardShortcuts` to `PATCH
/api/instance/settings/general` receives a 400. No in-repo client does.
- Stored `instance_settings.general.keyboardShortcuts` values are
stripped on read and ignored.
- Users who never turned the toggle on now get shortcuts. The handlers
skip text inputs, contenteditable regions, and modal dialogs, so typing
is unaffected.

## Model Used

Claude Fable 5.1 (`claude-fable-5-1`) in Claude Code, with extended
thinking and tool use.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-29 21:28:06 -07:00
DottaandPaperclip e912f0df53 fix(ui): open text attachments in task tabs (#14297)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agent work often ends with a Markdown or plain-text file.
> - Task attachments currently open outside the task panel.
> - Users need to inspect those files while keeping the task
conversation in view.
> - This pull request opens text attachments in task tabs and adds
rendered, raw, and download controls.
> - The same controls work in the mobile task drawer.

## Linked Issues or Issue Description

**What happened?**
Opening a text attachment did not put its content in a task tab.
Markdown files had no in-task rendered/raw toggle.

**Expected behavior**
Open Markdown and text attachments in one reusable task tab. Show
Markdown as rendered content or raw text. Download the original file.

**Steps to reproduce**
1. Upload a Markdown file and a plain-text file to a task comment.
2. Open each attachment from the task conversation or artifact list.
3. Switch Markdown between Rendered and Raw. Download both files.
4. Repeat at a mobile viewport width.

Related work: #14193 controls artifact tab arrival. This change adds
text attachment content tabs.

## What Changed

- Route text attachment opens from conversation and artifact cards into
task tabs.
- Add a text attachment panel with accessible Rendered, Raw, and
Download controls.
- Preserve ordinary links for other file types.
- Support the selected attachment in the mobile drawer.
- Keep text-tab actions on the current rich artifact cards, including
CSV previews.
- Render attachment image references and diagram source without loading
media URLs.
- Add browser regression tests, component tests, Storybook examples, and
usage documentation.

## Verification

- Full workspace typecheck, production build, Storybook build, and UI
token gates pass locally.
- All 6,960 UI tests pass. The additional media regression passes
against the real Markdown renderer and fails before the fix. CSV
coverage verifies direct downloads and text tabs after preview.
- Both desktop and mobile browser cases pass locally. They check
rendered/raw Markdown, literal plain text, reusable tabs, review
controls, and exact original download bytes. The local fixture used a
separate database port because an existing socket occupied the default
range.
- The full CI test matrix passes on
`82be5efbef26927b237a031725bb3d7fa79f637f`. The duplicate local `pnpm
test:run` was stopped after this CI result; it did not complete locally.
- Greptile is 5/5 on the final commit. All review threads are resolved,
and the security scan passes.
- All 54 final-head checks pass, including the canary dry run. The two
optional Storybook jobs are skipped.

## Risks

- Text attachment links now open in the task panel. Other content types
keep their existing link behavior.
- File display still depends on the existing authenticated attachment
route. There are no API or database changes.
- Raw text is displayed as text, including strings that look like HTML.
Rendered Markdown keeps media references inert.

## Model Used

OpenAI Codex, based on GPT-6, with code execution, browser testing, and
subagent tool use. The runtime does not expose an exact serving model
variant or context-window size. Recovered earlier implementation changes
were reviewed and tested; their exact model metadata is unavailable.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-29 16:40:10 -05:00
DottaandPaperclip 29c8fb0b66 fix(composer): match effort options to adapter execution (#14576)
## Thinking Path

> - Paperclip is an open source control plane for AI agents.
> - The issue composer lets operators choose a model and effort for a
task message.
> - The picker must show only effort levels that the selected adapter
can execute.
> - The Codex Runner fix in #14568 exposed similar gaps in other
adapters.
> - Grok hid a supported control. Claude used one range for all models.
Kimi ACP showed a control that it ignored.
> - This pull request aligns the picker with each adapter and adds
matching stories.
> - Operators can see and save supported effort settings without false
controls.

## Linked Issues or Issue Description

**What happened?**

The composer hid Grok effort. It showed an incomplete effort range for
some Claude models and a slider for Claude Haiku. It showed Kimi effort
on the default ACP engine even though that engine drops the value.

**Expected behavior**

The composer should show only effort levels supported by the selected
model and execution engine. Grok effort should reach its adapter as
`reasoningEffort`.

**Steps to reproduce**

1. Select a Grok agent and `grok-4.7`. Observe that the picker has no
effort slider.
2. Select Claude Sonnet 5 or Haiku 4.5. Observe the generic low to high
slider.
3. Select a Kimi agent on ACP. Observe the slider even though ACP
ignores effort.

Related fix: #14568.

## What Changed

- Use Claude and Grok adapter capability helpers in the production
picker and Storybook preview.
- Map Grok composer effort to `reasoningEffort` in per-message
overrides.
- Show Kimi effort only when its agent uses the CLI engine, including
when it runs its default model.
- Show Grok effort when it runs its default model.
- Add design and production stories for Claude, Grok, and Kimi ACP and
CLI cases.
- Document the picker rules and add focused tests.

## Verification

- `pnpm --filter @paperclipai/ui exec vitest run
src/components/task-chat/composer-run-settings.test.ts`
- `pnpm --filter @paperclipai/ui typecheck`
- `pnpm check:token-gates`
- `pnpm -r typecheck`, `pnpm test:run`, and `pnpm build` passed on the
first commit.
- The focused test, UI typecheck, token check, and Storybook build
passed again after the review fixes.
- Browser smoke: production stories show effort for default Grok and CLI
Kimi, and hide it for ACP Kimi and Claude Haiku.

## Risks

- Existing Kimi ACP agents lose a slider that could not change
execution. Their stored effort override remains unchanged until the next
edit.
- Claude and Grok ranges follow their adapter catalogs. A provider can
change model support before the catalog is updated.

## Model Used

OpenAI Codex based on GPT-6. The exact runtime model ID and context
window are not exposed in this session. Reasoning, tool use, and code
execution were used.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-29 10:34:29 -05:00
da887ea3e9 fix(runner): honor Codex effort selected in composer (#14568)
## Thinking Path

> - Paperclip manages AI agents that work on assigned tasks.
> - The task composer lets a person choose an assignee, model, and
effort for the next run.
> - A Paperclip Runner agent can use Codex as its provider.
> - The composer hid Codex effort for that agent because it checked only
the older Codex adapter.
> - The native Runner input also did not carry an effort choice to
Codex.
> - This pull request carries the chosen effort from the composer to
each Codex turn.
> - People can now select a supported effort and get the effort they
selected.

## Linked Issues or Issue Description

Refs #14322

**What happened?**

The composer showed a model but no effort slider when the assignee used
Paperclip Runner with the Codex provider. A task-level model override
also did not reach the native Runner input.

**Expected behavior**

The composer shows effort choices for a known Codex model. The next
native Codex turn uses the selected model and effort.

**Steps to reproduce**

1. Open a task composer.
2. Select an agent that uses Paperclip Runner with the Codex provider.
3. Select a known Codex model such as `gpt-6-astra`.
4. Open the assignee and model picker. The effort slider is missing
before this change.

## What Changed

- Show known Codex effort levels for Paperclip Runner Codex assignees.
- Save the task effort override in the native run input and send it to
Codex on each turn.
- Apply the task's merged model and effort overrides when the native run
starts.
- Apply a task model override for OpenCode Runner without changing the
agent's provider.
- Add Runner effort tests and desktop and mobile Storybook cases.

## Verification

- `pnpm -r typecheck` passed.
- `pnpm build` passed.
- `pnpm build-storybook` passed.
- `pnpm check:token-gates` passed.
- Focused UI, server, Runner contract, and Codex driver tests passed.
- The full CI test matrix, build, typecheck, and canary dry run passed
on the latest head.

## Risks

- Native Runner inputs add an optional Codex effort field to the current
v5 input. Older inputs keep their previous behavior.
- A known model rejects an effort that its catalog does not support.
Unknown models do not show a slider.

> This fixes an existing composer bug. I checked `ROADMAP.md`; it does
not describe this bug as planned work.

## Model Used

OpenAI Codex, GPT-6. The exact deployment ID and context window are not
exposed in this session. The model used reasoning, code execution, and
repository tools.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: OpenAI GPT-6 Astra <noreply@openai.com>
2026-09-29 09:41:05 -05:00
DottaandPaperclip 24beb00575 feat(runner): add rich ACP transport and durable interaction foundation (#14430)
Add shared rich ACP transport, durable questions and permissions, verified provider packaging, and bounded activity and plan presentation. Keep Cursor, Copilot, and Pi pending their separate provider qualification.

Persist interaction settlement before publication, fence failed writes until fresh recovery, and preserve owned-process cleanup. Incorporate reviewed mainline integration with extended harness coverage.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-29 08:56:21 -05:00
DottaandPaperclip 61b3fd57a6 fix(ui): recover gracefully during server restarts (#14560)
Show a clear reconnecting state during server restarts and retry safe access
checks every five seconds. Preserve open drafts, wait for initial startup
readiness, and keep authentication failures separate from temporary outages.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-29 08:31:05 -05:00
DottaandFry 3ca196b0a6 feat(agents): persist agent files across tasks without revision history (#14420)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - An agent needs personal files across tasks and sessions.
> - AGENTS.md is one file in that directory. Supporting files need the
same persistence.
> - The Instructions Editor and agent runs must share one current
directory.
> - Concurrent runs should apply only the files they change. The last
sync of the same file wins.
> - This pull request uses existing file transport and removes temporary
copies after sync.
> - Old instruction-only sessions keep their restore contract. New saves
do not create revision history.

## Linked Issues or Issue Description

Refs #14325. This replaces its revision-oriented design with persistent
agent files. Keep #14325 unmerged.

Transport prerequisite #14416 merged first at
`d172197117a14b80a1eb2d2835a0e7cce2679656`. This PR now targets master
and remains below 100 changed files.

Related work: #4513 and #8798 cover instruction tooling. This change
handles run synchronization, cross-task personal files, browser editing,
and old-session restoration.

## What Changed

- Keep one current directory per company and agent. Point AGENT_HOME at
a temporary working copy for each active run. Keep task files and
provider HOME separate.
- Restore text, binary files, and nested folders through workspace
transport. Exclude remote agent files from task Git snapshots with a
self-ignoring file inside the reserved runtime directory; never write
through repository-controlled Git metadata.
- Collect after the provider and child processes have stopped. Keep
resumable conversation state.
- Apply changed and deleted files under the agent lock. The last sync
wins for the same file. Unrelated concurrent changes survive.
- Remove temporary copies after successful sync, rejected sync, and
staging failure. Register ownership before copying so restart recovery
can remove interrupted preparation. Retry transient synchronization up
to three times. Preserve the original remote lease reference until
deletion succeeds; restart cleanup never acquires a replacement sandbox.
Do not create captured directories or a conflict-review queue for new
runs.
- Keep browser editing, stale-draft protection, and streaming binary
downloads. Keep the instruction entry and text editor limited to 1 MiB.
- Keep historical agent-folder sync failures on their affected runs
instead of repeating them above current saved instructions. Preserve
legacy candidate review and current browser-save errors. Avoid duplicate
quota warnings while retaining separate sync failures when they describe
a different problem.
- Require target-scoped caller grants for peer instruction access, while
preserving self edits, responsible-user checks, and protected-change
consent.
- Treat full storage as a nonblocking run warning, never an agent pause
or run-admission failure. Restore already-over-quota saved folders so
ordinary agent cleanup can recover; warn on each run until cleanup. The
run detail view shows the warning.
- Allow 256 MiB per file, 2 GiB per directory, and 100,000 entries. Hash
large files as streams. Check editor-save quotas with metadata instead
of hashing unrelated files.
- Preserve old native inputs, instruction-only copies, paths, digests,
and pending legacy candidates. Adopt old revision heads once. New writes
do not append history rows.
- Add idempotent migration 0287 and verify upgrades from the preview
tables and receipts.
- Add nine interactive stories under **Agents / Persistent files**,
including automatic incoming edits, stale browser drafts, and
storage-limit diagnostics.

## Verification

- Merge candidate: `4f5390107ec6ffd80a76d1d2e85530e66f21d079`, after
merging current master and the landed transport prerequisite.
Integration required no manual conflict resolution; the feature remains
99 changed files. Full workspace typecheck, production build, token
gates, and 715 focused tests passed on this merge candidate. Fresh
Greptile review is 5/5 with no unresolved findings. All 55 checks
passed, with four conditional skips, including the build, typecheck,
browser E2E, and canary dry run. A single retry recovered four jobs
interrupted by runner shutdowns; no source changes were required.
- Historical-warning UI fix: all 6,834 UI tests across 640 files passed,
including regression coverage for three old failures, legacy preserved
edits, and warnings scoped to the affected run. Full workspace
typecheck, production build, Storybook build, and token gates passed.
Browser-verified Storybook playtests passed for Historical Failures
After Successful Save, Storage Limit, and Full Storage Run Warning.
- Review follow-ups at `4e20c9fb2`: all 18 focused tests passed,
including external Git directories, linked worktrees, symlinks,
hardlinks, and distinct I/O failures alongside storage warnings. Server
and UI typechecks, token gates, and the production build passed.
- Storage warning regressions at `0724f3012`: all 33 directory tests and
all five heartbeat-list tests passed, with no skips in their successful
runs. They cover repeated runs while full, an already-over-quota saved
folder, cleanup, warnings retained after unrelated save failures, and
bounded warnings in large result JSON. Server typecheck passed after the
final warning fixes.
- Full workspace typecheck, production build, and token gates passed
during this follow-up. Product E2E harness: 631 tests passed across 52
files; harness typecheck passed. Earlier native session/context and
directory/legacy collection suites passed 537 tests; Runner
unit/transport suites passed 329 tests.
- **Real E2E at `0724f3012` (before this follow-up):** legacy local
Codex and native Daytona Codex each passed six tasks, one server
restart, seven independent assertions, and cleanup verification. Both
prove browser-to-agent edits, agent-to-browser edits, nested/binary
restoration, per-file last-sync-wins, a successful run after an
oversized save rejection, and cleanup clearing the warning.
- Native local Codex also passed the six-task quota flow before the
final warning-retention fixes. That pass began at `918d1ed02` while the
bounded-result warning fix was being edited, so it is not claimed as
exact-final-head evidence. Its final-head rerun failed during embedded
PostgreSQL bootstrap before any provider run: the macOS host had 87,365
of 87,381 SysV semaphores occupied. No unrelated services or kernel
limits were changed.
- The final-source report intentionally records **2/3 cells passed**,
preserving the blocked native-local attempt:
`tests/runner-e2e/results/agent-files-quota-final-20260928-report/`.
Earlier failed attempts and provenance notes remain under
`tests/runner-e2e/results/agent-files-quota-final-20260928-input/` and
the original campaign directories.
- Daytona used immutable image
`ghcr.io/paperclipai/paperclip-daytona-runner@sha256:5643f0d801417cae3581833a1a3bc6715b325e028602738d2652c44cac5dc6bf`
and its exact Linux runner binary. Controller source is `0724f3012`;
image source is recorded separately.
- Legacy-session compatibility and all three ACP Stop/resume browser
regressions passed on the prior validated feature head
`169fab46d5af21caa2269b4c1b29b69c933a6951`. They assert the same
provider session is retained and interrupted writes are not replayed.
Migration upgrade tests also passed earlier.
- Nine interactive stories are under **Agents / Persistent files**,
including **Full Storage Run Warning**. Its playtest and visual browser
inspection passed; the warning states that runs continue and the editor
remains available.
- Prior-head checks on `4e20c9fb2`: 55 passed, two conditional jobs
skipped, no failures or pending checks. All eight browser E2E shards and
their aggregate passed. Fresh Greptile review is 5/5 with no findings;
all review threads are resolved, the security scan passed, and GitHub
reports no merge conflicts.
- The broad local follow-up test run was interrupted after host
semaphore exhaustion affected isolated PostgreSQL instances. It also
encountered the existing macOS long-path fixture failure and two timeout
failures. This is not a claim that the full local suite passed. Logs are
retained; focused storage/warning tests passed.

## Risks

- A later sync can overwrite an earlier edit to the same file, including
a saved browser edit. There is no text merge or retained version. This
is the intended last-sync-wins policy.
- A save that exceeds a storage limit is rejected and its temporary copy
is discarded. The run itself continues normally, and later runs restore
the last saved files with a warning until cleanup. Transient sync
failures get bounded retries. An I/O failure partway through a sync can
leave some files updated; a failed receipt does not claim whole-folder
success.
- Larger folders increase copy time, network traffic, and temporary disk
usage. Active runs still need working copies. Terminal runs do not
accumulate archives. Operators must provision disk for agents and
configured concurrency; these limits are not company-wide quotas.
- A restored old native session remains instruction-only until a fresh
session starts. Its original conflict fence and existing pending
candidates remain compatible.
- Provider processes close at the collection boundary. Conversation
resume remains available, but warm process reuse is lost.
- Backups must include the instance filesystem and database. External
bundles keep their existing behavior until explicitly moved to managed
storage.

## Model Used

OpenAI Codex, GPT-6 family. The session does not expose a more specific
model ID or context-window size. Reasoning, code execution, and browser
tools assisted this change. Real provider E2E uses `gpt-5.6-sol`.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Fry (Paperclip) <noreply@paperclip.ing>
2026-09-29 08:25:56 -05:00
0be2afcca6 feat(ui): improve task composer controls and pending input (#14322)
## Thinking Path

> - Paperclip lets operators assign tasks to AI agents and review their
work.
> - The task composer controls the next message and its assigned agent.
> - Operators needed a way to choose that agent's model and effort
without leaving the composer.
> - The old mode selector, upload button, and input cards made the
mobile composer crowded and hid normal messaging during a pending
decision.
> - Harnesses publish different model and effort capabilities, so the
picker must follow the selected agent.
> - This pull request adds one responsive composer flow, keeps pending
cards visible above it, and protects Codex ACP authentication in the
local test path.
> - Operators can choose run settings, send a message, and answer a
pending card as separate actions.

## Linked Issues or Issue Description

**Subsystem affected**

Task composer UI, issue thread interactions, Codex ACP credential
handling, and Storybook.

**Problem or motivation**

The composer did not expose model or effort for the selected agent.
Mobile actions wrapped poorly. Pending questions and confirmations
replaced the composer. A local Codex ACP test could also reuse host
authentication after the managed key was removed.

**Proposed solution**

Put assignee search, model search, exact model IDs, effort, and fast
mode in one picker. Use a mobile dialog. Replace the direct-upload plus
action and separate mode selector with an Add menu and removable Plan or
Ask chips. Place pending interaction cards above the usable composer.
Keep these cards pending after an ordinary message unless their creator
asks for comment superseding. Replace managed ACP auth files atomically
and isolate the test key from host credentials.

**Roadmap alignment**

ROADMAP.md does not list an overlapping composer milestone. This change
improves the existing task and review flows.

## What Changed

- Added the combined assignee, model, and effort picker to both task
composers. Search matches agent name, role, and harness. The server uses
a curated Codex list by default and honors instance-declared models.
Manual IDs remain available.
- Added an effort slider for known model capabilities, a conditional
Codex fast control, and reset. The picker opens in a modal on mobile.
- Added the Add menu for files, supported goals, Plan mode, and Ask
mode. Plan and Ask are exclusive removable chips. Keyboard mode cycling
remains available.
- Adjusted mobile spacing, avatars, wrapping, and Send placement.
Removed the composer divider.
- Moved pending question, confirmation, review, and related cards above
the composer. Ordinary comments now leave question and confirmation
cards pending by default. The onboarding prompt retains explicit comment
superseding.
- Updated the Storybook composer group with responsive states and the
production picker. Added UI, service, route, and browser regression
coverage.
- Isolated Codex ACP API-key authentication, skipped subscription auth
merge and shared-home copy-back for remote API-key runs, and replaced
the managed auth file atomically.

## Verification

- `pnpm -r typecheck` — passed on the final local head.
- `pnpm check:token-gates` — passed on the final local head.
- `pnpm exec vitest run server/src/__tests__/adapter-models.test.ts
ui/src/components/task-chat/ComposerRunSettingsPicker.test.tsx` — 31
tests passed, including role and harness search, declared Codex models,
and filtering general OpenAI models.
- `pnpm exec vitest run
server/src/__tests__/issue-thread-interactions-service.test.ts` — 74
tests passed.
- `pnpm exec vitest run
packages/adapters/codex-local/src/server/acp.test.ts` — 42 tests passed,
including remote API-key copy-back isolation.
- `pnpm test:run` — attempted locally; the embedded PostgreSQL test
database could not initialize on macOS. The isolated
`heartbeat-run-event-sequencing` suite reproduced that environment
failure. GitHub CI runs the full test matrix for this head.
- `pnpm build` — passed on the final head. `pnpm build-storybook` passed
after the last UI change; only server code, tests, and docs changed
afterward.
- Live local test drive — Codex ACP ran a task with a managed API key.
The test agent was restored to its default ACP configuration afterward.
- Review the interactive stories under the top-level Composer group with
`pnpm storybook`. Check a narrow desktop width and mobile Plan, Ask,
picker, and pending-question states.

## Risks

- A pending card stays open when an ordinary comment changes the
discussion. Its creator can set `supersedeOnUserComment: true` when a
new comment should replace it.
- Model and effort overrides persist on the task until reset or changed.
An unlisted manual model ID may fail when the provider runs it.
- Some harness catalogs do not report effort support. The picker hides
effort for those models.
- No database migration is required.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

OpenAI GPT-6 via Codex. This runtime does not expose the exact model ID
or context window to the task. The model used code execution and browser
tools.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: OpenAI Codex <codex@openai.com>
2026-09-28 23:02:39 +00:00
DottaandPaperclip 24c58e479a Improve task artifacts with rich cards and editable stories (#14469)
Render eight artifact card types from real task records and share them with editable Storybook stories. Preserve document review and media/file actions, and load bounded CSV previews on request.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-28 15:38:06 -05:00
DottaandPaperclip d423fc4411 fix(ui): use large mobile selector modals (#14250)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The new task dialog lets an operator assign work on a phone.
> - The mobile picker used fixed coordinates inside the Radix wrapper.
> - iOS Safari already makes fixed coordinates relative to the visual
viewport.
> - The old code added the visual viewport offset a second time.
> - This pull request removes the second offset and gives the mobile
wrapper full viewport geometry.
> - The benefit is a visible picker and search field when the iOS
keyboard opens.

## Linked Issues or Issue Description

**What happened?**

The assignee and project pickers could move outside the visible screen
in the new task dialog on iOS.

**What did you expect to happen?**

The picker and its search field must stay visible above the bottom edge
and the software keyboard.

**Steps to reproduce**

Open the new task dialog on an iPhone. Select Assignee or Project. The
picker can move outside the visual viewport when Safari pans the page.

**Paperclip version or commit**

The problem exists after the change in PR #13343.

**Deployment mode**

The problem affects the board UI in local and authenticated modes.

Refs #13343

## What Changed

- Use visual viewport local coordinates for the new task dialog.
- Give the mobile Radix popper wrapper full viewport geometry.
- Render both shared selector components as large, titled mobile modals.
- Keep desktop selectors as anchored popovers.
- Cover the new-task assignee, new-task project, composer assignee, and
generic searchable selector in Storybook.
- Add iPhone-sized Storybook states for the assignee and project
pickers.
- Update viewport tests for the corrected coordinate model.

## Verification

- `pnpm --filter @paperclipai/ui exec vitest run
src/components/NewIssueDialog.test.tsx
src/components/InlineEntitySelector.test.tsx`
- `pnpm check:token-gates`
- Playwright used the iPhone 14 device profile. All mobile modals were
inside the 390 by 664 CSS viewport.
- Each mobile modal was `x=16, y=16, width=358, height=632`.
- A desktop check kept the anchored popover at `width=320, height=211`.

## Risks

- Low risk. The CSS only changes narrow mobile viewports.
- The full-screen popper wrapper does not receive pointer events. The
picker still receives pointer events.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. See `CONTRIBUTING.md`.

## Model Used

- OpenAI GPT-5.6, Codex agent with tool use and code execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal task
ID
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant Storybook documentation to reflect my
changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-27 11:33:24 +00:00
DottaandPaperclip 74508357fa fix(ui): keep mobile task pickers in view (#14249)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The new task dialog lets an operator assign work on a phone.
> - The mobile picker used fixed coordinates inside the Radix wrapper.
> - iOS Safari already makes fixed coordinates relative to the visual
viewport.
> - The old code added the visual viewport offset a second time.
> - This pull request removes the second offset and gives the mobile
wrapper full viewport geometry.
> - The benefit is a visible picker and search field when the iOS
keyboard opens.

## Linked Issues or Issue Description

**What happened?**

The assignee and project pickers could move outside the visible screen
in the new task dialog on iOS.

**What did you expect to happen?**

The picker and its search field must stay visible above the bottom edge
and the software keyboard.

**Steps to reproduce**

Open the new task dialog on an iPhone. Select Assignee or Project. The
picker can move outside the visual viewport when Safari pans the page.

**Paperclip version or commit**

The problem exists after the change in PR #13343.

**Deployment mode**

The problem affects the board UI in local and authenticated modes.

Refs #13343

## What Changed

- Use visual viewport local coordinates for the new task dialog.
- Give the mobile Radix popper wrapper full viewport geometry.
- Keep the mobile picker at the visible bottom edge.
- Add iPhone-sized Storybook states for the assignee and project
pickers.
- Update viewport tests for the corrected coordinate model.

## Verification

- `pnpm --filter @paperclipai/ui exec vitest run
src/components/NewIssueDialog.test.tsx
src/components/InlineEntitySelector.test.tsx`
- `pnpm check:token-gates`
- Playwright used the iPhone 14 device profile. Both picker boxes were
inside the 390 by 664 CSS viewport.
- The assignee picker box was `x=16, y=366, width=358, height=282`.
- The project picker box was `x=16, y=366, width=358, height=282`.

## Risks

- Low risk. The CSS only changes narrow mobile viewports.
- The full-screen popper wrapper does not receive pointer events. The
picker still receives pointer events.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. See `CONTRIBUTING.md`.

## Model Used

- OpenAI GPT-5.6, Codex agent with tool use and code execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal task
ID
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant Storybook documentation to reflect my
changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-27 11:02:15 +00:00
DottaandCodie 01d9a12185 fix: make keyboard shortcut enablement a personal preference (#14141)
Store keyboard shortcut enablement per user and expose it in Profile settings.

Co-Authored-By: Codie <Codie@users.noreply.github.com>
2026-09-26 11:50:04 -05:00
DottaandPaperclip 96bf004a79 fix: use persisted state for lifecycle continuation and retry budgets (#13888)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Its control plane decides when a task can continue, wait, stop, or
complete.
> - Legacy continuation could change when an agent changed its wording
without changing task state.
> - Shared attempt counts also let repair and infrastructure retries
affect each other's limits.
> - This pull request uses persisted state and separate, bounded
allowances for these decisions.
> - If automatic repair stops, the task explains what happened and
offers a guarded retry.
> - Paired tests and real-provider evaluations verify that Stop,
approvals, ownership, and spending limits remain authoritative.

## Linked Issues or Issue Description

Related work: Refs #13761, Refs #11126, Refs #13610. These cover
obsolete continuation dispatch and retry storms. Open and closed issues
and PRs were searched for related lifecycle, continuation, and retry
work.

**What happened?**
Legacy continuation depended on English wording and progress heuristics.
Repair, failure retry, and productive continuation could consume shared
counts. When bounded repair stopped, the task showed a technical
recovery message without a clear next action.

**Expected behavior**
Persisted disposition and owned execution paths determine the next
action. Missing disposition prompts bounded agent repair. Explicit work
mode determines planning mode. Narrative changes and raw activity counts
cannot replenish allowances. An exhausted repair shows a readable
notice. An explicit retry checks current controls and preserves the
assigned agent.

**Steps to reproduce**
Run `pnpm test:lifecycle-baseline`. The paired probes keep structured
state constant while varying completion, planning, blocker, and progress
prose. Run the explicit `lifecycle-baseline` and
`continuation-accounting` Product E2E suites for real-provider coverage.
In Storybook, open **Design previews / Recovery notice** to inspect the
production component's normal, pending, acknowledged, unavailable,
failure, and mobile states.

## What Changed

- Hide the image attachment button, icon, and drop/paste hint in answer
composers. Image paste and drop support remains available.
- Merge current master and retain both browser regression sets. Use a
production-stamped service worker in the offline recovery browser
fixture.
- Share one state-based legacy continuation decision across immediate,
delayed, and recovered dispatch. Bind bounded repairs to their source
run and episode.
- Remove title and description wording from work-mode authority. Agents
can still write requested plans in execution mode.
- Persist separate failure-retry and productive-continuation counters.
Disposition repair and resource waits cannot consume or reset those
allowances.
- Validate delayed repair identity, then recheck current gates before
provider dispatch. Fence native startup cancellation.
- Show **Agent needs attention**, a plain-language explanation, **Retry
agent**, and expandable details in both task interfaces. Report request
progress, acknowledgement, and errors inline.
- Store typed recovery notice metadata. Recognize older active notices
only through exact stored action and run IDs. Notice text never grants
retry authority.
- Use the existing recovery-action endpoint for retry. Recheck current
action, status, owner, agent availability, dependencies, active runs,
pending questions and confirmations, approvals, pause controls, and
budget. Duplicate requests do not wake twice.
- Add component, page, route, database, contract, and Storybook
coverage. Keep the scenario inventory and executable evals here.
Historical reports and snapshots live in the [commit-pinned
paperclip-evals
archive](https://github.com/paperclipai/paperclip-evals/blob/ce3e5afcd4a1184650f586a2b5b8be5874c66c8b/experiments/2026-09-lifecycle-authority/README.md).
- Preserve unsaved project fields while the same project URL changes to
its canonical alias. Do not reuse data across projects or companies.
This separate fix addresses the repeated repository-editor browser
failure without changing the browser test.
- Keep the development service worker from intercepting Vite module
reloads. Update the connection-intent browser fixture to record progress
and completion through the agent API.

## Verification

Merge preparation on September 25, commit
`c1e8e4b7ddd9fbc4913ed55ce21b8e12906c2f97`:

- Merged master `bd2030932` and resolved the browser test-list conflict
by keeping both sets of regressions.
- Deterministic lifecycle baseline: 1,090/1,090 assertions passed; no
failures, skips, or missing selected evidence. Unit 423, runner 184,
database integration 397, grading 86.
- Browser support: 17/17 passed. The offline recovery test first failed
with an unstamped development worker, then passed with the production
stamp. Its assertions are unchanged.
- Focused interaction UI and offline fallback tests: 19/19 passed.
Verified the custom-answer composer in Storybook: no attachment controls
or hint; entering an answer enables Next.
- Recursive typecheck, production build, token gates, and diff checks
passed. The worktree is clean. No new real-provider campaign was run.
- Current CI and review: [Current PR CI
passed](https://github.com/paperclipai/paperclip/actions/runs/36166011243):
55 successful checks and two optional Storybook skips. Greptile scored
this exact commit 5/5. Hiding the question attachment controls is an
intentional UI change; paste/drop remains available.

Earlier recovery UI verification, commit
`21be0fec0e90e86b6d662b8ee4831847cd041cdb`:

- Recursive typecheck, production build, token gates, and diff checks
passed.
- Focused UI coverage: 338 tests passed across six suites (336 before
the interaction guard, with the two affected suites rerun at 149 passed
after it). Covers both task interfaces, the real page mutation,
pending/error acknowledgement, stale state, and unavailable controls.
- Recovery database integration: 352 tests passed before the interaction
guard. The complete recovery-action and mutation-route suites passed 181
tests after it. The two new pending question/confirmation regressions
failed before the fix and passed afterward, including
resolved-interaction controls. Shared validator suite: 31 passed. E2E
catalog suites: 34 passed.
- Browser inspection passed for light/dark themes, mobile layout,
expandable details, pending retry, acknowledgement, failure, and
disabled retry. Storybook renders the production component; its request
is simulated.
- The broad local run hit two chat callback-order wait failures and was
stopped after all CI unit/database/runner shards passed. Both local
failures passed when rerun without the competing full-suite process.
- CI exposed a repeated project-repository draft-loss race during
canonical redirects. A new unit regression failed before the fix; all
nine project-page tests now pass, including controls for other projects
and companies. Both unchanged repository browser tests passed against a
fresh local server. UI typecheck, production UI build, and token gates
passed after this fix.
- [Earlier PR CI
passed](https://github.com/paperclipai/paperclip/actions/runs/36072486798)
on `21be0fec0e90e86b6d662b8ee4831847cd041cdb`: 55 successful checks, two
optional Storybook skips, and no failed or pending checks. The
repository browser shard passed with the production fix. Greptile is 5/5
on this exact commit with no unresolved review threads. The PR is
mergeable.

Historical, source-qualified lifecycle evidence:

- Lifecycle baseline: 1,074 assertions. Native session coverage: 447
tests. Product E2E support: 515 tests. Browser support: 11 tests. Full
earlier verification is retained in the archive.
- [Real-provider campaign: 8/8 passed, zero
retries](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35881382080-1/index.html),
source `e88d210417280140b44a36449027290adcb1aeaa`. Evidence and cleanup
checks passed. This includes deliberately exhausted repair cases that
correctly remain blocked; it does not mean every task finished Done.
This campaign predates the recovery UI change.
- Archive migration verified all 16 original JSON files byte-for-byte
and all 24 checksum entries. App tests do not need private archive
access. [Archive PR
#27](https://github.com/paperclipai/paperclip-evals/pull/27) is merged.

## Risks

- Agents that omit durable disposition receive at most two repair
attempts by default. Prose-only completion exposes missing state rather
than silently changing scheduling.
- A retry is an explicit board action. The server rechecks current
controls. A successful response confirms the task returned to To do; it
does not claim that the provider has already started.
- Existing notice metadata remains valid. Only older active notices with
matching structured evidence receive the new UI. Historical notices
without that evidence keep their existing rendering. No schema migration
is required.
- Old run records require conservative retry accounting. Tests cover old
counters, alternating retry lanes, restarts, and exhausted repairs.
- Historical snapshots require private `paperclip-evals` access. The app
index retains public campaign links. Live campaigns qualify specific
sources and scenarios; no new real-provider campaign has run for the
recovery UI commit.

> This fixes existing lifecycle and recovery behavior and does not
duplicate planned core work.

## Model Used

OpenAI GPT-6 through Codex assisted implementation, reasoning, code
execution, and review. The exact serving model ID and context window are
not exposed in this task. Historical real-provider evaluations used
Codex model `gpt-5.6-sol`, separately from the implementation assistant.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-25 15:28:11 -07:00
DottaandPaperclip aa8fc86331 feat(connections): prefer native apps and ask users to choose external providers (#13941)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Connections give agents governed access to external services.
> - Native connections should remain the first choice for a supported
app.
> - Other apps may be available through an external MCP provider.
> - The user must know which external provider handles the connection
and choose it before setup.
> - This pull request adds ranked alternatives and server-authored
instructions to connection search.
> - Agents can follow the returned instructions while Paperclip
validates saved choices and access.

## Linked Issues or Issue Description

Related: #13879, which fixed inline MCP provider setup. This PR adds
discovery and provider selection on top of that work.

**Subsystem affected**

Cross-cutting: shared connection contracts, server search and intent
services, native runtime, CLI, inline setup UI, and evals.

**Problem or motivation**

An agent cannot offer a clear external-provider choice when Paperclip
has no native connection for an app. Adding provider-specific branches
to the core prompt would make those instructions harder to maintain.

**Proposed solution**

Prefer a native connection. Otherwise return verified alternatives in
Composio, Arcade, Executor, Zapier order. Include an external-service
disclosure, a question with None, and the next instruction in the search
result. Validate the saved human choice before creating a selected
fallback setup card. Reuse existing provider accounts and verify
underlying app access separately.

**Alternatives considered**

Do not silently choose a provider. Do not claim that broad execution
tools prove support for every app. Reuse existing questions and
connection intents rather than add another connection model.

**Roadmap alignment**

Extends the existing MCP Tool Gateway & Apps and Agent evals & feedback
capabilities. The MCP aggregators experiment remains the gate. No
duplicate provider-routing PR was found in the public search.

## What Changed

- Add a dated support index and authorized cached-tool evidence for
external routes.
- Return provider questions and next-step instructions from
`connections_search`.
- Preserve pending choices and declines across continuation. Validate
company, task, agent, human, app, and current route eligibility.
- Carry the selected app into new setup and account reuse, validate
explicit provider requests against persisted human messages, and
distinguish provider readiness from app authorization.
- Sync native, MCP, REST, and CLI contracts. Keep core agent
instructions provider-neutral.
- Add production-component Storybooks, focused database tests, and three
real-agent browser eval cases.
- Record the plan, observed failures, fixes, passing evidence, and
acceptance limits.

## Verification

- Latest head `586f0e6cd`: 54 checks passed, 2 skipped; Greptile 5/5 and
all review threads resolved.

- After rebasing on master `18dac1e1e`: 64 focused shared, validator,
route-contract, and database tests passed; server typecheck passed.
- Embedded-browser test drive on the rebased head: native Jira card,
HubSpot external-provider question, Arcade account reuse, one actual MCP
read against a local synthetic fixture, reload persistence, and None
preventing further calls. A real OpenAI-backed agent performed discovery
and continuation.
- UX observation: the agent initially combined mutually exclusive
request fields; the server rejected it and the agent recovered without
changing access. This extra retry remains visible in the transcript.
- After rebase: 23 focused eval grader/catalog tests, affected
TypeScript checks, token gates, production UI build, and Storybook build
passed. Full local tests are intentionally excluded at the maintainer's
request.
- Before rebase: four browser/real-agent attempts passed: native Jira,
None, and reuse of the second provider on two Codex profiles.
- Browser evals used an isolated deterministic MCP fixture through the
real Paperclip gateway. They do not prove production compatibility with
all four providers.
- Review `Apps / Connections / Provider choice` in Storybook. Choose
Arcade, continue through Access, and verify the app name,
external-service disclosure, and URL configuration.
- The detailed verification report is
`doc/connections/2026-09-23-aggregator-routing-verification.md`.

## Risks

- The public support index is finite and can age. Account capability and
app authorization still require verification after selection.
- Existing installed-tool permissions remain in effect. Provider choice
is not a new execution permission boundary. An early Mini attempt
skipped search; clearer provider-neutral instructions made the targeted
rerun pass. This is not a measured reliability rate.
- Explicit requests skip provider confirmation only when a clear
persisted human message or saved provider choice supports them. Other
phrasing falls back to confirmation; the agent query alone is not
consent. These routes do not add tool permissions.
- No database migration or legacy Composio broker is added.
Real-provider acceptance remains separate from fixture proof.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, code
execution, and browser tools. The exact deployment variant and
context-window size are not exposed in this session. Product evals
separately used the repository's primary Codex and Codex Mini profiles;
those agents supplied test behavior, not independent provider
compatibility proof.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-24 16:36:49 -05:00