Commit Graph
353 Commits
Author SHA1 Message Date
DottaandPaperclip a7a244ab33 feat: add company decision models with permission and cost controls (#15473)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Optional product features need small, typed model decisions.
> - Each company needs to choose which shared API connection pays for
those decisions.
> - Calls must retain user and task permissions, budget limits, and cost
attribution.
> - This pull request adds a managed decision service, setup UI, and
request history.
> - Features can check availability cheaply and keep their existing
behavior when decisions are unavailable.

## Linked Issues or Issue Description

**Subsystem affected**

Server services, shared contracts, database accounting, Company
Settings, and Costs.

**Problem or motivation**

Paperclip has no common decision-model service. Adding provider calls
within each feature would duplicate credential access, permission
checks, and billing rules.

**Proposed solution**

Let a connection manager configure one company decision model. Support
OpenAI Decisions and Jev through OpenRouter. Provide a fixed setup test
and metadata-only history. Default company-sponsored background
decisions to on during setup, and preserve a saved off setting.

**Alternatives considered**

Per-feature credentials would duplicate existing connection management.
Personal overrides and provider fallback chains add permission and
billing complexity; they remain deferred.

**Roadmap alignment**

Reviewed ROADMAP.md and searched open PRs. This extends existing
connection access and budget accounting. Product features that call the
service remain outside this change. No matching decision-model service
PR was found.

## What Changed

- Add company settings, an internal `decisionModelService`, local
availability checks, and trusted human, agent/run, and system contexts.
- Pin Vercel AI SDK provider dependencies and adapt boolean, choice, and
ordered-score decisions for both providers. Bound requests and time;
disable paid retries.
- Add durable invocation metadata and agentless decision ledger charges.
Preserve fractional cents, pricing evidence, dispatch identity, and
unresolved billing holds.
- Share the company accounting lock and apply company, agent, and
project budgets. Settle charges once, retain unknown holds, and recover
interrupted calls without resubmission.
- Reuse connection setup and management UI. Add a Decisions view under
Costs, production-component Storybook coverage, database migration, and
service documentation.

## Verification

- Passed 166 current-code tests covering the decision service/provider,
setup component, Costs, OpenAPI, and every failure from the earlier
broad run. Coverage includes native SDK wire formats, refusals, billed
malformed responses, permission and secret-rotation races, identity
changes, concurrent budget admission, unresolved holds, agent/task
deletion, and stale setup feedback.
- Passed 170 existing connection, cost, budget, heartbeat-accounting,
and profile regression tests.
- Passed repository typecheck, production build, and design token gates
after integrating master. Verified the generated migration on a fresh
test database and upgraded the populated preview database from the
branch's earlier migration without losing settings or usage.
- Ran the required full `pnpm test:run`: its general phase completed
with 16,354 passed and 10 failures across five files while this branch
was still being updated. Every reported failure passes in the
current-code rerun; the serialized phase did not run after that failure.
The full GitHub CI suite passed on `b1b856a88`: general and serialized
tests, browser shards, runner checks, typecheck, production build,
packaging/canary, and policy gates. [CI
evidence](https://github.com/paperclipai/paperclip/actions/runs/37673106360).
- Passed the full-shell Storybook setup-to-history interaction test
again after integrating master. Greptile rates the final revision 5/5
with zero unresolved threads.
- Walked through the running app: empty setup, add each provider, save,
reload, run all three sample questions, inspect fractional charges in
history, switch provider while sponsorship is off, disable, and
reconnect. Checked mobile settings. These tests used the actual UI,
vault, server, SDKs, and database with simulated upstream responses.
- Live paid setup tests remain unverified: this environment has no
authorized OpenAI/OpenRouter credentials available. No mocked test is
presented as live provider evidence.

Reviewer journey: Company Settings → General → Decision model.
Add/select a shared API connection, save, run the billed sample, open
View usage, then disable decisions and verify Run test is disabled after
reload.

## Risks

- The SDK decision interface is experimental. Pinned versions and
wire-format tests limit upgrade drift.
- The migration allows agentless service charges and reservations.
Existing agent cost-reporting APIs still require an agent, and decision
receipts stay separate from run reconciliation.
- Timeouts can have unknown provider charges. Holds remain until an
audited accounting correction resolves them.
- OpenAI prices use a versioned Decisions rate snapshot; OpenRouter
costs use provider receipts. Unknown pricing is retained as unknown.
- A configured company authorizes background spending by default. Setup
explains this, and managers can turn it off.
- Live provider account/model availability still needs the two
credentialed acceptance checks.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository inspection, code
execution, and browser tools. The exact serving revision and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-07 15:53:13 -05:00
DottaandPaperclip ae6f95ed7a feat: add personal primary agents (#15470)
Add a personal primary agent per company and user. Initialize it from the first human-created agent, expose profile-only switching with confirmation, and use it after recent choices for task and Chat defaults.

Persist authenticated preferences, preserve lifecycle and membership rules, keep selections out of shared audit events, and document the API contract. Include the reviewed Storybook surfaces and regression coverage for concurrent choices, onboarding, cross-device updates, and browser journeys.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-07 13:53:46 -05:00
DottaandPaperclip b315580640 feat: add private task creation and sharing controls (#14718)
Add private task creation and sharing controls, effective access explanations, private project management, and locked task references.

Constrain the mobile composer plus menu to the viewport and explain named inherited privacy on hover, focus, or tap. Include 125 production-component Storybook stories and thirteen interactive user journeys.

Final head passes all CI checks and Greptile Apex 5/5 with no findings.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-07 07:27:52 -05:00
Devin FoleyandPaperclip 799e4d556f fix: make accounting durable and synchronize cost reporting (#14997)
Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-06 22:26:48 -07:00
Devin FoleyandPaperclip 228f0e2807 fix(ui): make pull request review waits visible (#15375)
## Thinking Path

> - Paperclip lets people supervise agent work and review its outputs.
> - A task can wait for a human to review a pull request while an agent
schedules checks.
> - The PR work product and detected external links use separate display
paths.
> - A private PR can disappear from the useful sidebar view when its
provider lookup fails.
> - The monitor countdown says when the next check runs but omits the
saved review request.
> - This change shows the saved PR and requested review in the task's
existing surfaces.

## Linked Issues or Issue Description

**What happened?**

A task kept checking whether a private GitHub PR had merged. Its work
product requested board review, but the Properties sidebar used detected
external objects instead. The PR URL was stored only in metadata, which
the PR refresh path did not read. The composer showed a generic monitor
countdown with no PR link or review action.

**Expected behavior**

Show saved PR links even if provider access fails. During a GitHub
monitor wait, make outstanding review requests visible with the next
check time and a status-check action. Stop requesting review when a PR
merges or closes.

**Steps to reproduce**

1. Register a PR work product with `reviewState: needs_board_review` and
its URL in `metadata.url`.
2. Schedule an external-service monitor for GitHub.
3. Open the task with external-object lookup unavailable or unable to
access the private repository.
4. Inspect Properties and the composer wait strip.

Related work: #14469 added rich artifact cards. #8759 addresses
attention on monitored task blockers. This change uses the existing work
products and monitor action; it adds no blocker or approval mechanism.

## What Changed

- Show saved PRs in Properties, including metadata-only links,
independent of external-object availability.
- Deduplicate equivalent GitHub PR URLs while preserving the saved
navigation link, and put explicit review requests first.
- Keep provider status and freshness on the combined PR row; keep
private PR links usable when lookup fails.
- Show the saved PR review request above the composer and in the
existing monitor banner during a GitHub monitor wait.
- Label the existing monitor action `Check status` and show check
failures inline.
- Suppress review prompts for merged, closed, or archived PRs even if
their review flag is stale.
- Refresh PR metadata using `metadata.url` and the existing `repository`
alias.
- Share saved work-product reads across the thread, Properties, and
Artifacts. Refresh GitHub in a separate query and enrich only matching
PR versions.
- Start a fresh saved-row request on live invalidation so late provider
responses cannot hide new artifacts or changed review requests.
- Cover stalled GitHub lookups in both panels so refresh latency cannot
hide saved work.
- Scope monitor-check mutation state to the task so failures and late
responses do not leak across navigation.
- Refresh GitHub status when the displayed run finishes, including when
saved PR rows have not changed.
- Clarify that PR review and external release handoffs need a saved
human-input interaction with an agent assigned for continuation; a flag,
monitor, or handoff comment alone does not create that card.

## Verification

- 386 tests pass across eleven affected UI and server suites on
`184c7d8746`. Regressions cover saved links, provider status, cold-cache
loading, panel reopening, monitor errors during navigation, and live
updates during provider refresh.
- Three live-update regressions and two run-completion regressions fail
before their fixes and pass afterward. These tests use the real API
client's GET coalescing and abort handling.
- Full repository typecheck and build passed on the merged parent
`1a49112dd2`. The latest UI changes pass UI typecheck and token gates;
CI also verifies the latest build. Capability contract/inventory checks
pass.
- Storybook build and browser checks passed before the query race fix.
Browser checks cover desktop and 390px mobile review waits, plus video,
mixed-file, and empty artifact galleries after background refresh.
- The earlier full local `pnpm test:run` was stopped under disk
pressure. It reported failures outside the changed suites; an isolated
skill-cache run reproduced three existing macOS permission failures. The
full local suite was not rerun for this follow-up.
- Latest-head CI passes on `184c7d8746`, including the browser shards
and canary dry run. Apex is 5/5 with no actionable findings and no
unresolved threads. All eight reported findings are addressed.

## Risks

- This uses saved PR review state. When GitHub access fails, the saved
state can remain stale until the agent updates it. The link remains
visible and the status check remains available.
- `Check status` wakes the existing monitor owner. It does not merge a
PR, accept an approval, or mark the task done.
- No schema, permissions, or scheduler behavior changes.

## Model Used

OpenAI GPT-6 via Codex, with reasoning, repository analysis, code
editing, and test execution. The exact deployment model ID and context
window were not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass — 277 affected tests
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 16:03:13 -07:00
DottaandPaperclip a590ab769d Give managed agents persistent cryptographic identities (#15352)
Give agents persistent Ed25519 identities encrypted with the existing instance master key. Create keys transactionally for new agents and lazily before supported managed runs, expose public identities in the API and agent UI, and protect private material during runtime delivery and output persistence.

Preserve identities in recovery backups while giving imported and development-cloned agents fresh keys.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-06 13:57:03 -05:00
DottaandPaperclip 9f7057e122 feat(connections): configure custom model providers across agent harnesses (#14970)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents use a harness, a model, and a credential to run tasks.
> - Connections already store credentials and control who can use them.
> - Custom providers also need an endpoint and a supported API format.
> - A per-agent endpoint would duplicate credentials and access rules.
> - This pull request stores routing on the connection and projects it
into the harness.
> - Isolated credentials and protocol checks keep the selected
connection authoritative.

## Linked Issues or Issue Description

Refs #37, #13083, #14104, #14565, #12692.

#14016 is a reference only. This PR has its own schema, vault
persistence, routing validation, runtime projection, and tests. None of
the five commits in #14016 is an ancestor of this branch. We do not
depend on or plan to merge it. #14967 addresses task-pinned account
pools. #14422 addresses another provider integration.

This is the first of two linked PRs. Merge this connection change before
#15341, which refines agent setup and adds the qualification harness.
The split keeps each review below 100 changed files. Provider catalog
entries have usable setup forms in this PR. Local browser subscription
sign-in is included.

## What Changed

- Store non-secret routing metadata on AI connections. Vault provider
API keys, including Bedrock bearer API keys. Reject general AWS access
keys.
- Enforce company, owner, human audience, agent access, connection
status, and protocol checks before resolving credentials. Keep reconnect
destinations immutable and retain connection identity during key
rotation.
- Project OpenRouter and compatible custom endpoints into Codex, Claude,
OpenCode, and local Hermes. Carry these settings through both legacy and
native runner transports. Clear conflicting host credentials and redact
keys from diagnostics.
- Preserve older OpenRouter accounts and native personal defaults. Add
Google API-key accounts and migration `0306` for the two
provider-default constraints.
- Run local Claude and Codex subscription sign-in behind the existing
browser sign-in card. Use private attempt homes and owner-bound
completion instead of a copied terminal command.
- Seed isolated Gemini authentication and preserve OpenCode workspace
permissions. Keep the selected connection authoritative. The independent
Gemini and Grok workflow fixes are in #15341.
- Keep native OpenCode custom gateway keys in a runner-owned
selected-model proxy; the harness config contains only a session-scoped
capability. Honor runtime outgoing proxy and certificate settings.
Preserve streamed responses and revoke the proxy on close or startup
failure.
- Allow ordinary members to connect native personal accounts before an
agent exists.
- Repair routed accounts from task cards using the saved provider
destination, protocol, model aliases, and connection identity.
- Add provider catalog definitions, model discovery, pinned logos, and
complete native and routed setup forms. Allow a personal routed
connection before a new agent exists. Keep endpoint authentication keys
out of Hermes terminal children.
- Recover cancelled or restarted browser sign-in with a clear restart
action. Support no-auth endpoints without a vault credential. Add
isolation and recovery regressions and runtime documentation.

## Verification

- Updated with `origin/master` at `22a3ea341`. Migration `0306` follows
the new master migration and passes migration and snapshot checks.
- The integrated connection regressions passed 152 tests and 50 native
OpenCode driver tests, including key-free child-shell configuration
reads, authenticated/no-auth forwarding, streaming, model/path
restrictions, cancellation, outgoing proxy routing, and NO_PROXY bypass.
Provider setup has 14 passing tests. The pinned real OpenCode 1.18.34
executable also completed a turn through the proxy against a local
synthetic provider; the reusable key was absent from its config. A
second real-executable smoke passed with an HTTPS CONNECT proxy and
runtime-specific synthetic certificate trust. Certificate-file and
certificate-directory regressions pass.
- Task-card repair passed 48 tests, including OpenRouter, Bedrock, and
custom gateway reconnect cases. UI typecheck and token gates passed.
- The prior core regression set passed 133 tests across new-agent setup,
provider forms, browser sign-in, routing projection, and connection
authorization. Token gates and UI typecheck passed.
- Full workspace typecheck and production build passed again after the
latest integration and credential-proxy fix. The merged deterministic
runner E2E suite passed 1,400 Vitest tests and 128 Node tests.
- Full workspace typecheck passed on the prior linked combined
implementation. Production build, Storybook build, 1,316 browser-harness
Vitest tests, and 128 Node tests passed. Head `9d964c8d9` includes the
latest master integration and regenerated migration. This exact head
passed 54 remote checks with four expected skips and Greptile 5/5; no
review threads remain open. An unchanged server fixture had a random
six-character issue-prefix collision on its first attempt. All 245 tests
passed locally and the single CI retry passed.
- A provider-free terminal check used the cited supported Hermes source
and dummy keys. Gateway and OpenRouter terminal children could not read
the selected key.
- The broad local Vitest attempt passed 15,442 tests but was not green.
It had an embedded-Postgres startup failure, an HTTP logger timeout, an
origin socket error, and a browser cancellation wait timeout. The
cancellation wait was corrected. The relevant connection tests and the
full origin test file passed separately. Latest-head CI must pass before
merge.
- Prior credential-backed acceptance exercised task creation, tool use,
artifact delivery, completion, and context-dependent follow-up. Claude
legacy and native runners passed Bedrock with `us-east-1` and
`us.anthropic.claude-sonnet-4-6`.
- Historical local qualification retained 43 passing API/gateway cells
out of 46. Those attempts span earlier builds. They do not qualify this
exact commit or staging. All subscription combinations and staging
remain unqualified.
- Verify native subscription and API-key setup. Connect a regular
provider catalog row. Verify an incompatible harness and a changed
reconnect URL are rejected. Use #15341 for the complete browser
campaign.

## Risks

- Migration `0306` changes two check constraints. It preserves rows and
is safe to reapply. It takes normal constraint-change locks.
- Credential projection touches several harnesses. CLI upgrades can
change provider configuration and session behavior.
- The native OpenCode proxy adds a loopback hop, pins requests to the
selected model, limits request bodies to 16 MiB, rejects redirects, and
expires at session close. It prevents reusable keys in the child
configuration; it is not an OS isolation boundary against a process
debugger running as the same user.
- Custom endpoints must be reachable from the agent environment. Saving
a connection does not prove connectivity. Bedrock keys require rotation
before expiry.
- Gemini CLI 0.58.0 has an upstream ACP new-file error conversion
defect. Provider overloads and an unresolved follow-up timeout also
affect live Gemini qualification. We have not patched the installed CLI
or marked those cases as passing.
- OpenClaw Gateway, Hermes Gateway, Claude Managed, AWS AgentCore,
Process, HTTP, and legacy ACPX local are excluded. Vertex, ambient AWS
identity, arbitrary auth headers, and custom routing for other harnesses
are excluded.
- These PRs do not establish production or staging qualification for
every provider and login method.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, code
execution, and browser testing. The exact deployment model ID and
context window size were not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 13:29:03 -05:00
DottaandPaperclip 22a3ea3414 Invite assistants from Connections with scoped browser and device consent (#14933)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Public MCP lets people use their organization from an external
assistant.
> - Operators need a visible control for this experimental access.
> - Hosted users should select an organization once and then approve its
permissions.
> - This pull request adds the setting, invitation-first setup, and
browser or device consent.
> - Connections provides a copyable invitation with public instructions
that grant no access.
> - Users reach browser consent from their assistant and return to
inspect or revoke access.

## Linked Issues or Issue Description

Builds on merged foundation #14846. This PR now targets master. Related
settings convention: #13905.

**Current behavior**

The preview uses an environment variable to enable MCP. Hosted consent
repeats organization selection. Assistant access has no entry in
Connections, so users must already know the endpoint and how to reach
consent.

**Proposed behavior**

An administrator enables Settings → Experimental → Assistant connections
(MCP). A hosted connection shows the selected organization and its icon,
then asks for permissions. Requested write access starts checked when
the user’s role permits it; the user can opt out before connecting.
Direct instance connections show an organization picker with the first
available organization selected. The selection stays fixed across
refetches and still requires an explicit Connect action. Connections
includes Assistant Connection (MCP). Its setup page explains the
canonical endpoint, client configuration, browser authentication, and
connected access. It connects as the current person and does not select
or impersonate an agent.

**Reason and benefit**

Operators manage access with the other experiments. Users select one
organization, and both the UI and server enforce that choice.

**Breaking changes**

The old enable variable has no effect. Preview operators must enable the
setting once. Apply the additive consent-request migration before
deploying the tenant, then deploy the compatible Cloud broker. Existing
direct requests and grants keep their behavior.

## What Changed

- Simplify OAuth and device consent: show the Paperclip logo beside
“Connect {client} to Paperclip”, fall back to “your assistant”, and show
the identifying origin plus its favicon below, with the callback URL
also visible when different. Remove the hosted-organization creation
action. Default to the first available organization without silently
changing it on refetch; preserve company restrictions and write
opt-outs. Keep the button row contained on narrow screens.
- Make Copy invitation the primary action, using the shared animated
AgentSetupPrompt and a collapsed manual setup section with icon-labeled
line tabs. Remove redundant link actions, copy-status text, the extra
first-prompt well and revocation explanation from the setup page. Serve
shared version-aware HTML and Markdown instructions without private
organization data.
- Support guarded Client ID Metadata Documents alongside dynamic
registration, and include authorization response issuer identification.
- Add RFC 8628 device authorization with separately hashed codes,
expiry, shared request quotas, persistent polling backoff and atomic
redemption. Reuse human consent, role checks, scoped grants, audit and
revocation.
- Add CLI device login and a local stdio bridge. Store credentials
separately with private permissions and serialize rotating refreshes.
- Add device consent stories and five cold-start paid Product E2E cases
with independent grant, configuration and durable-work assertions.

- Add `enablePublicMcp` to the settings validator, normalizer, feature
catalog, and toggle UI. Check it live for OAuth, tools, subscriptions,
and event delivery. Keep connection management and revocation available
while disabled.
- Default the MCP origin to the existing auth public URL, with strict
validation and an explicit override.
- Persist the optional OAuth `company_id` restriction. Describe only
that company and reject approval for any other company, even if the
person belongs to both. Keep active-membership and role checks.
- Show the Paperclip icon and a large organization icon during consent.
Return the saved company logo through the company-scoped request
response and reuse the standard fallback icon. Use the requested concise
permission labels: “Read all of your Paperclip data” and “Allow write
access and creating tasks as me”. Use concise permission copy, retain a
compact client and callback-origin disclosure, and remove the footer
link.
- Default requested write access on for eligible roles. Preserve opt-out
across organization changes and refetch, reset defaults for a new
request, and submit read-only access when the request or role does not
allow writes. Align the shared checkbox with its label.
- Use organization wording in consent, management, settings, and
walkthroughs. Keep the organization fixed for hosted requests and retain
direct-instance choice.
- Add an Assistant Connection (MCP) card to the Connectors catalog, a
setup page in the app shell, and a return link from Experimental
settings. Include Codex, Claude Code, OpenCode, and generic remote MCP
instructions.
- Read the live gate and canonical server URL through authenticated
setup metadata. Show only the current person’s grants for the selected
organization, refresh after consent, and support revocation. Surface
catalog status failures with an explicit retry action; do not present
them as an empty connection list. Opening setup grants no authority.
- Start the eight guided chapters in Connections. Keep presenter notes
and chapter controls around real product pages in the app shell. Explain
the terminal, consent, delegation, retrieval, and revocation handoffs.
Mark conversation examples as illustrative. Cover first use, client
setup, connected, loading, and error states. Keep the existing consent
and management stories.
- Keep the paid-eval setup and browser helper aligned with the setting
and consent button.

## Verification

- Warm-standby integration fix `ec64ea05e`: public MCP ingress now
follows the Cloud claim guard; MCP and discovery paths return 503
instead of SPA HTML while unclaimed. Event polling checks the in-memory
claim before reading the persisted experimental setting. All 97 focused
OAuth/Cloud tests and server typecheck pass, including new request and
timer regressions for idle-before-claim and resume-after-claim behavior.
Fresh review is 5/5 with no unresolved threads, and all security scans
pass on this final head. All browser shards, typecheck, build, canary
installation and other test groups passed on the first attempt. The
unchanged Cursor sandbox default-command test timed out at 10 seconds;
the exact test passed locally without edits in 587 ms. The single
failed-job retry passed, with the original failure retained in workflow
37500711895. All 54 final-head checks pass on
`ec64ea05e9a03e2179d4e2f84c2de03761f7ce26` (two optional Storybook jobs
are intentionally skipped).

- Final master integration `8457828fc`: merged foundation #14846 and
current master, preserving the invitation changes and all 33 files from
the two newer upstream changes. No migration renumbering was required.
All 95 focused OAuth/Cloud integration tests, full recursive typecheck
and token gates pass. All CI gates passed on that integration head;
review identified the warm-standby issue fixed above.

- Security-review fix `8c1d0b696`: commit shared global/per-source
admission before outbound CIMD work, preserve failed-attempt receipts,
and validate resource/scope before fetching. Added migration
`0305_chubby_vin_gonzales.sql` and six concurrent/adversarial regression
cases. All 69 OAuth/metadata tests, 26 migration checks, full recursive
typecheck and production build pass. The security scanner passed that
commit. Follow-up `87f9658e7` limits only actual cache-miss fetches; 18
authorization requests sharing one proxy across two service instances
use just two fetches. All 70 OAuth/metadata tests and server typecheck
pass after that refinement. Final follow-up `8ebeae84c` reports
admission-storage failures as retryable HTTP 503 instead of invalid
client metadata. Its regression proves no outbound request before
admission and successful retry after storage recovers. All 71
OAuth/metadata tests and server typecheck pass. Final-head security
scanning passes; Greptile is 5/5 with no unresolved findings. CI passed
all browser shards, typecheck, build, token gates and canary
installation. One unchanged adapter-utils bridge test raced a
response-file write (expected a JSON error, received the safe
file-changed error). The exact test passed locally without edits. The
single failed-job retry passed; the original failure is retained in
workflow 37490609192. All 54 checks now pass on final head
`8ebeae84ca77c0cf7ac12c2006f0f8743fe50e0b`, with security scan and fresh
Greptile 5/5 and no unresolved threads. Foundation #14846 subsequently
merged as `e34abee670069cca84afb2efb86041bce7dccbec`; the final
integration above now targets master.

- Integration with current master: preserved the new Connections source
filters and pagination, kept all eval suites, and regenerated the
consent/device snapshots as migrations 0303/0304. All four MCP migration
SQL hashes are unchanged from the staging versions. Full recursive
typecheck and production build, 132 focused UI tests (including catalog
filtering), 89 server authorization/settings tests, 26 migration tests,
120 eval calibration tests and token gates pass. Review follow-up
`4820ce74c` also keeps active assistant grants in Installed, with
pending/error recovery and revocation/company-isolation coverage. All 76
setup/catalog tests, UI typecheck and token gates pass after that fix.
The unchanged signoff browser test timed out waiting for a heartbeat in
CI at `4820ce74c`; the exact test passed locally without code changes,
and the preceding CI head passed that shard. That same unchanged test
failed at the reviewer stage in the next CI run. All five signoff tests
passed three times locally (15/15), without test changes. All eight
browser shards pass at final head `8ebeae84c`; no browser-test edits or
failed-browser-job retries were needed.

- Setup-page refinement at `9ab009178`: all 17 focused setup/consent
tests pass, along with UI typecheck, production build, Storybook build
and token gates. Browser exercised the shared prompt preview and client
tab switching, and the updated InvitationCopied Storybook interaction
checks its clipboard fixture. All final-head CI checks pass at
`9ab009178`, with no unresolved review findings. Deployed successfully
to Butter in
https://github.com/paperclipai/paperclip-cloud/actions/runs/37475189524.
Verified the actual page, tab switching and line styling, removed
actions/copy, and successful native copy/paste of the complete Butter
invitation into a local-only test field. The existing Claude grant was
left intact.
- Consent follow-up at `dc8e9fd11`: all 10 consent tests and token gates
pass. UI typecheck and production build passed again at `4e4d5e4d9`;
Storybook build and eval-helper typecheck passed for `28101cf91`.
Follow-ups let the primary button wrap on narrow screens, preserve a
distinct callback URL, and use only bundled icons to avoid pre-consent
requests to client-selected sites. Browser-verified the real consent
component in desktop and 320px mobile stories, including default
selection, write access and preserved opt-out. Updated E2E
heading/default-selection helpers. All CI checks passed at `dc8e9fd11`,
with review 5/5 and no unresolved threads. The Butter preview
publication needed a retry because npm initially accepted the DB package
before making it visible; the retry succeeded and `dc8e9fd11` deployed.
Verified a fresh, unapproved native Codex CIMD request on Butter:
default organization/write selection, known-client heading and icon,
distinct callback origin, and removed creation action. No grant was
approved for this UI check. Prior paid runs below retain their exact
source provenance; this UI-only follow-up did not rerun paid
qualification.
- Source-pinned paid matrix at
`2992ef2710f47230e7f484c709c6ba02524f884c`: **15/15 passed**, five cases
each on GPT-5.4 Mini, Claude Haiku and Sonnet. Campaign
`local-2026-10-06T02-41-14-462Z`. Covers cold start, existing config,
unavailable host, denied consent and reconnect/later retrieval, with
independent configuration/grant/task/run/document assertions. Original
failures, transcripts, source fingerprints and billing remain retained.
- Final instruction follow-up `cda8178af`: **3/3 cold starts passed** on
Mini, Haiku and Sonnet. Campaign `local-2026-10-06T02-58-30-041Z`.
Latest `0637b9f1c` shares that same guidance across HTML, Markdown and
manual UI after review; generated Markdown is verified byte-identical to
the paid-evaluated version. Shared build, server/UI typechecks, token
gates and 63 auth/metadata tests passed again. Every CI gate passed at
prior HEAD `0637b9f1c`, with review 5/5 and no unresolved threads.
- Other focused checks: 11 CLI credential/refresh-lock tests, 120 eval
calibration tests, server/UI/eval typechecks, token gates and Storybook
build passed. Full recursive typecheck and production build passed
during implementation; CI also passed them at `2992ef271`.
- Local full-suite limitations: a large-file Git streaming test times
out on this Mac, and broader CLI/route runs hit DB hook timeouts. Fresh
MCP reruns passed, and the corresponding CI groups passed. No claim that
the local full suite is green.
- Actual clients: Codex 0.153.4 and Claude Code 2.1.245 reach CIMD
consent; device CLI reaches verification/consent. New grants await human
approval. Existing local OpenCode retrieved a saved result in a fresh
conversation through its previously approved grant.
- Fresh OpenCode 1.18.17 on Butter: started with no MCP config, received
the exact copied invitation, read public setup, configured its server
and started PKCE consent. Its shell command timed out; background retry
reached the client's own callback deadline while approval remained
pending. Latest instructions cover that handoff. **No completed Butter
read/delegation/result retrieval is claimed.**
- Cloud companion
https://github.com/paperclipai/paperclip-cloud/pull/672 passes
checks/review and deployed. Anonymous setup and device-protocol routing
verified. Core `2992ef271` deployed successfully and the actual Claude
web flow now reaches consent. Its extra JWT-bearer metadata is filtered
to implemented grants; unsupported token grants remain rejected. Final
`0637b9f1c` deployed successfully to Butter in
https://github.com/paperclipai/paperclip-cloud/actions/runs/37409195300;
live HTML and Markdown both contain the final guidance. The superseded
instruction-only build was canceled before deployment. This is a
core-only staging preview; private Cloud plugins are omitted. ChatGPT
web is signed out, so browser connector use is unverified.
- Screenshot gallery begins at Butter's dashboard and distinguishes real
setup/pending consent from local reuse and fixtures. It records the
timeout finding. New persistent access needs human confirmation before
the remaining actual-client acceptance work.
- Manual path: Connectors → Assistant Connection (MCP) → Copy invitation
→ paste into assistant → configure and start authorization → sign in and
approve → verify `paperclip_connection` → delegate → retrieve the saved
report later.
- Plan and instructions: `doc/plans/2026-10-05-assistant-invitations.md`
and `doc/public-mcp.md`.

## Risks

- Apply additive, replay-safe migration `0304_curvy_shadow_king.sql`
before using device authorization. The public setup link carries no
credential. Device codes and tokens stay private; neither sharing
instructions nor installing a plugin authorizes access.
- Apply additive migration `0305_chubby_vin_gonzales.sql` before
deploying the shared metadata admission gate. It retains at most 60
short-lived, hashed-source receipts per instance and rejects excess
attempts with 429.
- CIMD metadata fetching is a new external-input boundary. It requires
HTTPS, exact client ID and redirect validation, bounded responses and
guarded DNS/network access. Client names remain self-reported.
- Device support is per-instance. The central Cloud broker retains its
existing grant support. Host installation and tool reload capabilities
vary by client; instructions describe manual settings and restart
requirements.

- Consent names the registered client in its heading and displays its
identifying origin below. Known-origin icons are bundled; all other
origins show a neutral site icon without contacting client-selected
sites. Client names are self-reported; the callback origin is the
recipient check. The Cloud chooser also displays the original client and
receiving origin before tenant handoff.

- A user who accepts the preselected write permission can create tasks
and comments. Task creation and comments can start or wake agents and
use execution budget; the consent label uses the concise wording
explicitly requested by the maintainer. Scope requests, role checks, and
the final Connect action still apply.
- Migration `0303_supreme_garia.sql` adds one nullable UUID column with
`IF NOT EXISTS`. Requests without a company restriction keep the
direct-instance picker. The binding stays recorded if its company is
deleted; consent then fails closed.
- Deploy tenant support before the Cloud broker sends `company_id`.
Unknown or inaccessible organizations must never fall back to a
different company.
- The setting defaults off. Disabling access does not cancel work
already delegated. Existing tokens and unexpired subscriptions can
resume when enabled again; revocation remains separate.
- The catalog entry is visible for discovery while the feature is off.
Setup instructions, OAuth, and tool execution remain gated. No access is
granted by viewing the entry.
- Assistant sign-in starts in the external client so it owns PKCE and
callback state. Client command syntax can change and links to official
setup documentation are included.
- An authenticated instance and valid public URL are required. Hosting,
paid execution, and store publication remain separate rollout steps.

## Model Used

OpenAI GPT-6 in Codex, with tool use and code execution. The exact
serving model version and context window are not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused checks pass;
unrelated local full-suite timeouts are explicitly recorded above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 12:35:16 -05:00
Devin FoleyandPaperclip 202c2d307e fix(server): leave unclaimed warm Cloud databases idle (#15314)
## Thinking Path

> - Paperclip manages work performed by AI agents.
> - Managed deployments prepare empty applications before an owner
claims them.
> - Those applications start database pollers even though no company can
have work.
> - Health probes also query SQL, so idle databases cannot remain
suspended.
> - This pull request adds an explicit standby marker for empty,
unclaimed Cloud apps.
> - The existing signed, durable claim resumes normal processing without
restarting the app.

## Linked Issues or Issue Description

**What happened?**

An empty, unclaimed warm application runs recurring chat, email, plugin,
heartbeat, cleanup, and reconciliation queries. Its health route also
opens the database. This prevents idle database compute from suspending.

**Expected behavior**

An explicitly marked unclaimed application should keep its HTTP process
and sandbox provider plugins ready while leaving the database idle. A
successful signed claim should resume normal API behavior and background
processing. Claimed and self-hosted instances should keep their current
behavior.

**Steps to reproduce**

Start an empty Cloud-managed application and leave it unclaimed. Observe
database activity while repeatedly requesting `/api/health`. Before this
change, periodic queries continue without company data.

Related: #15153 reduces allocation during chat polling. This change
suppresses polling only for explicitly marked, empty, unclaimed Cloud
apps.

## What Changed

- Add `PAPERCLIP_CLOUD_WARM_STANDBY=1`. Check company emptiness once
after restoring the persisted Cloud runtime identity. Missing Cloud
configuration, existing data, or a persisted claim leaves normal
processing active.
- Gate recurring database pollers with an in-memory predicate. Keep
startup preparation and sandbox provider plugin loading intact.
- Serve unclaimed health probes without session or database reads and
report `warmStandby: true`. Serve standby pages/assets directly from the
UI router, bypassing session, bearer, tenant, and dynamic handlers.
Refuse API requests and all WebSocket upgrades before authentication can
query SQL or seed company data.
- Exit standby after the existing signed identity assertion commits.
Normal timers resume at their next tick; a restart restores the claim
even with stale provider variables.
- Document the marker, readiness semantics, rollout checks, and
rollback.

## Verification

- `pnpm -r typecheck` passed. A final server typecheck also passed after
adding tests.
- `pnpm build` passed.
- Focused standby, signed claim, restart, health, static/Vite routing,
hostname, HMR, and live-events suites: 72 passed after the review fixes.
Includes real HTTP upgrade admission before/after claim.
- `pnpm test:run` was attempted locally; both superseded runs were
stopped after encountering checkout/platform failures. A clean-checkout
rerun eliminated ancestor skill-directory lookup failures. The
company-skills/runtime-cache families encounter macOS
read-only-directory rename failures (`EACCES`); all three company-skills
failures reproduce on unmodified base `bf14f803d5`. The initial full run
also reported one native runner API test failure; an isolated comparison
on both revisions was blocked by local embedded PostgreSQL startup
failures. The full [Linux CI
run](https://github.com/paperclipai/paperclip/actions/runs/37423263993)
passed on final commit `5016c415ea`, including all server and workspace
test shards, browser suites, typecheck, build, and release canary. This
is not a claim that the full local suite passed.
- Isolated full server with local PostgreSQL: after startup and
connection expiry, 70 health probes, 70 page requests carrying valid
synthetic tenant credentials, and 70 rejected WebSocket upgrades over 70
seconds observed zero app database connections. The signed claim
completed in 62 ms and normal polling resumed (475 database transactions
over 12 seconds). Restart with stale provider variables restored the
durable claim. The latency is local-only, not a provider wake
measurement.
- Apex review: **5/5** on `5016c415ea`, both earlier threads resolved,
no open recommendations.
- No live-provider test or production deployment was performed. An
actual database suspension/resume canary remains required before
enabling the control-plane switch.

## Risks

- Standby health reports HTTP readiness rather than current database
connectivity. The signed claim still requires a durable database write;
claimed health checks retain the SQL probe and 503 failure behavior.
- Pollers resume at their usual intervals. A suspended database may add
claim latency. Validate the real provider before enabling the marker.
- Startup preparation and sandbox plugins remain loaded. New plugins or
background loops must respect the same standby contract.
- The marker is off by default. Remove it or set it to `0` and restart
to roll back. No schema migration or claimed-workspace inactivity policy
changes.

## Model Used

OpenAI Codex, based on GPT-6. The exact serving snapshot and configured
context-window size are not exposed in this session. Assistance included
source review, TypeScript changes, command execution, and PostgreSQL
tests.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass — targeted tests pass; full
local suite limitations are documented above, and full Linux CI is green
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-06 09:11:25 -07:00
DottaandPaperclip 9b3fe260ba fix(tasks): surface Codex ChatGPT model rejection (#15299)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for work.
> - Native Runner runs report provider errors and can also save a structured failure result.
> - Codex rejects a model that the user's ChatGPT account cannot use.
> - Master now diagnoses this rejection, but the task's compact run data omits its message. The thread can still call it a generic run failure.
> - Users need the account restriction and a clear step to repair the model selection.
> - This pull request shows the existing diagnosis as “Model unavailable” on the task.

## Linked Issues or Issue Description

**What happened?**

Codex returns HTTP 400 with `invalid_request_error` and the message `The 'gpt-6.1-sol' model is not supported when using Codex with a ChatGPT account.` Master now stores an actionable diagnosis for this rejection. The task thread still labels it “Run failed” and does not receive its error text in compact run data.

**Expected behavior**

Show the account restriction on the task and in the run error. Tell the user to choose a supported model or clear the task's model override before retrying.

**Steps to reproduce**

1. Use a Native Runner Codex agent signed in with a ChatGPT account.
2. Select a model that produces the rejection above and start a task.
3. Let the runner save its generic failed result. Inspect the task's failure marker and recovery notice.

**Additional context**

Refs: #15304. That merged PR diagnoses the provider failure and preserves worker and review recovery rules. This PR adds its task-facing message and guidance without changing that diagnosis or those rules.

Refs: #13134. That PR improves model discovery for ChatGPT accounts. This PR exposes the rejection when a configured model still fails at execution time.

## What Changed

- Return a bounded model rejection message in compact issue-run data.
- Show “Model unavailable” and model-change guidance in the task thread and recovery notice.
- Add database and UI regression tests, including both the original rejection text and master's fixed diagnosis. Document the new failure message.

## Verification

- Focused merged-branch validation: 279 tests passed across activity service, native provider failure observation and PostgreSQL integration, TaskChatThread, and ExecutionBlockerNotice.
- `pnpm -r typecheck`, `pnpm build`, and `pnpm check:token-gates` passed on the merged source tree.
- Greptile reviewed conflict-resolution commit `1db39c4e636a92e2e76ba224b71c6f1e2f556c81` at 5/5 with no findings or inline comments.
- All 55 check runs completed without failure on that commit. The two optional Storybook jobs were skipped. The legacy Snyk status passed. The branch has no merge conflicts.
- UI regression assertion: the task's failure marker says “Model unavailable” and retains the account restriction. It no longer says that this failure happened after a final response.

## Risks

- Provider recognition and recovery are owned by the existing master implementation. This PR exposes only bounded error text for failed runs with `native_provider_model_rejected`.
- Historical runs with the generic `adapter_failed` code are not reclassified. No stored run is rewritten.
- Existing retry and reconciliation gates remain in place. No schema migration or model configuration change is required.

## Model Used

OpenAI Codex, based on GPT-6, with reasoning, code execution, and browser inspection. The exact deployment model ID and context window are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context to this change
- [x] I have specified the model used (with version and capability details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before requesting merge

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-06 08:05:25 -05:00
DottaandPaperclip 0e0b63e5a5 feat(connections): add experimental task-pinned AI routing (#14967)
## Thinking Path

> - Paperclip manages AI agents and their work.
> - AI Connections separate account access from models and harnesses.
> - A pool must act as one connection while retaining each task’s
account.
> - Core must enforce member access and preserve session and recovery
rules.
> - A plugin supplies rotation policy without receiving credentials.
> - This change adds durable routing and native connector setup and
management.

## Linked Issues or Issue Description

**Subsystem affected**

AI Connections, Connectors, plugins, run dispatch, and session
compatibility.

**Problem or motivation**

Operators need to rotate new tasks across saved accounts while each task
keeps its account and session. Pool setup must fit the existing
connector catalog and account workflow.

**Proposed solution**

Add an experimental router binding, a capability-gated plugin hook, and
transactional task pins. Plugins declare native pooled connectors
through `aiConnectionRouter`. Core hosts the existing-account picker,
ordering step, and account settings. Related usage contract: #14936.
Companion private plugin:
https://github.com/paperclipai/paperclip-cloud/pull/643.

**Roadmap alignment**

This extends Apps and AI Connections. Core supplies generic enforcement
and native connector UI; the private plugin owns rotation and quota
policy. The prior duplicate search found no matching router
implementation.

## What Changed

- Add a router binding without changing existing concrete bindings. Keep
the instance flag and new pools disabled by default. Require manual
operator configuration. Show no routing toggle in Experimental settings
on either open-source or Cloud installs, even after routing is enabled.
- Persist company-scoped pools, one shared cursor per pool, and pins
keyed by company, pool, agent, and task. Commit pins and cursor advances
together with revision checks and bounded retries. Persist run-ID
affinity before allocation.
- Pass only authorized metadata and normalized usage to plugins. Core
retains credential handling, member access checks, runtime
qualification, and recovery evidence. Probe outside locks with a shared
15-second budget and freshness cache.
- Resolve routing before credential preparation and backend selection.
Preserve pins through turns, session resets, removed members, and quota
waits. Retain admitted recovery after disable or uninstall.
- Separate credential session epochs from token generations. Verified
refresh preserves the epoch; reconnect and manual replacement change it.
Include the credential slot ID in session and usage-cache identity, so
reconnecting an indexed legacy account invalidates its old session even
when both epochs are zero.
- Validate pool member installations before accepting saved-agent
bindings and recheck compatibility when the harness changes. Install
only authorized members in the new-agent transaction and record their
IDs in local activity. Pool membership cannot install a restricted
shared connection.
- Preserve pool bindings when agents hire teammates through either
creation API or native caller runtime inheritance. Block stale manager
credential references; retain explicit child authentication precedence
and reject incompatible inherited pools.
- Add native connector registration through plugin metadata. Reuse the
Connectors catalog, setup header, account header, sidebar, dialogs, and
usage display. Setup selects and orders saved connections. Advanced
settings hold usage rules and member runtime defaults. New-account setup
opens in another tab.
- Use revision-checked pool archival from the Connectors catalog and
account page. Keep task pins, cursors, recovery evidence, and underlying
connections. Reject ordinary connection updates or removals that bypass
pool revisions.
- Add pool selectors, composer models, override notes, quota status, run
details, activity records, and local run-log records. Keep
session-adoption copy minimal.
- Show **Used by** below the pool connections. List current company
agents with shared avatars and profile links. Include paused agents;
exclude terminated agents and agents using another pool.
- Add Core stories for the generic connector workflow and runtime
surfaces. Cloud stories reuse these production routes and tokens through
a preview-only alias.

## Verification

- Final head `73cb953bca30ed83e4505dd820edd9b5edffd28b`: full workspace
`pnpm -r typecheck`, `pnpm build`, and `pnpm check:token-gates` pass
locally.
- All 422 focused connector/settings/shared-contract/migration tests and
all 156 database-backed AI connection, hiring, reconnect, and
durable-routing cases pass (69 hiring cases rerun after the final
auth-precedence fix). The merged shared contract retains connection
instructions and pool metadata. The pool migration is generated at
sequence 0299 after the latest upstream migrations; this PR makes no
lockfile changes.
- All four full-app Playwright tests pass on the final head after a cold
restart and migration, against the installed private plugin and isolated
database, with no route or pool-API mocks. They cover hidden routing
controls after manual opt-in, native pool creation, ordering, membership
edits, rename, paused defaults, enabling/save/refresh persistence, stale
edits, cancellation/removal, preserved underlying accounts, unavailable
routers, and Used by avatars and profile links. Exact command:
`PAPERCLIP_CONNECTION_POOL_E2E=1
AI_CONNECTIONS_TEST_COMPANY_ID=a37b9625-5ecf-4e29-8081-04df3d6e7d6f
AI_CONNECTIONS_TEST_URL=http://127.0.0.1:3108 pnpm exec playwright test
--config tests/ai-connections-app/playwright.config.ts
connection-pools.spec.ts`.
- [Native setup, ordering, and management
screenshots](https://github.com/paperclipai/paperclip/pull/14967#issuecomment-6006976278)
address the review follow-up. [Earlier selector, quota, and run-detail
screenshots](https://github.com/paperclipai/paperclip/pull/14967#issuecomment-5971537316)
show the runtime surfaces. Core previews: `pnpm --filter @paperclipai/ui
storybook`, then **Connectors / Pool host** or **AI Connections /
Connection pools**. Cloud owns its host-backed plugin stories; both
repositories’ Operator Setup Required story assertions pass.
- Live acceptance used OpenAI/Codex and Anthropic/Claude ACPX, resumed
both exact sessions after restart, preserved pinned accounts through
explicit reset and controlled quota deferral/recovery, and committed
only two allocations across fourteen runs. A later UI-created task test
again rotated OpenAI then Anthropic and resumed OpenAI through
follow-up/restart/quota recovery. That later Anthropic execution was
blocked by its saved OAuth token expiring (provider 401). No live usage
probes ran.
- The full local `pnpm test:run` was attempted earlier and did not
complete because of macOS embedded PostgreSQL bootstrap/shared-memory
failures and the 40,000-file Git fixture timeout. The focused database
suites above now pass; full-suite verification is provided by the split
CI lanes. The preceding CI run had one runtime readiness timeout; it
passes locally both alone and inside the larger runtime suite. That
larger local suite also encountered an embedded PostgreSQL setup failure
and two macOS temporary-path alias assertions; those two assertions pass
with canonical TMPDIR=/private/tmp. All final-head CI checks are
terminal green, including full general/serialized server suites, Runner
checks, browser E2E shards, canary verification, build, and typecheck.
Greptile is 5/5 on that exact head with no unresolved threads.

## Risks

- The migration adds routing tables and a credential epoch column.
Install the private plugin only with the compatible Core contract.
- Routing and each pool require opt-in. Production distribution and
fleet defaults remain unchanged.
- Unknown usage stays eligible. Known pinned exhaustion waits; revoked
access requires operator repair.
- Legacy adapters require compatible members. Runner model and effort
overrides remain limited by qualified backend support.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository editing, code
execution, and browser testing. The exact deployment model ID and
context window are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes:` / `Closes:`
/ `Refs:` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (targeted suites;
full-suite limitations are reported above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-05 20:44:37 -05:00
DottaandPaperclip 4857799a88 feat(connections): deliver saved instructions to authorized agent turns (#15216)
Persist optional connection instructions and deliver authorized snapshots to agent execution prompts. Keep provider templates with each app definition, preserve edits and opt-outs, and replace sessions when guidance or access changes.

Use shared production settings across setup and Permissions, with source visibility in agent Instructions. Add the initial memory-provider defaults and managed Honcho workspace configuration. Include migration 0298 and regression coverage for generic providers, runtime delivery, authorization, and catalog regeneration.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-10-05 18:16:00 -05:00
DottaandPaperclip b43073d11f feat(connections): sync and group accounts managed by aggregators (#15254)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Connections give agents governed access to external tools.
> - Aggregator gateways can expose accounts that users already connected
upstream.
> - The Apps catalog did not show those accounts or their current
provider status.
> - Separate cards and setup tasks also made account ownership unclear.
> - This pull request discovers upstream accounts and groups them under
one app card.
> - Users can find connected apps while each provider keeps control of
its accounts.

## Linked Issues or Issue Description

**Subsystem affected**

Connections across the database, shared contracts, server, and board UI.

**Problem or motivation**

Users cannot see which apps are connected through a saved aggregator
gateway. Native and upstream accounts need one app card. Discovery must
preserve company, user, gateway, and credential boundaries.

**Proposed solution**

Sync account metadata from Composio, Arcade, and supported Executor
gateways. Keep upstream account management in each provider. Use source
chips and search to browse the catalog. Preserve native setup and the
gateway's existing access policy.

**Alternatives considered**

Creating a local executable connection for each upstream account would
duplicate authorization state. Using an agent task for routine Composio
setup would add an unnecessary step. The board now calls the saved
gateway directly for that setup.

**Roadmap alignment**

This extends the shipped Connected Apps and MCP Tool Gateway features in
ROADMAP.md. The duplicate search found no open PR for managed account
discovery.

Related work: Refs #13755, Refs #13941, Refs #14725, Refs #13855. Open
PR #12906 covers adjacent toolkit routing work.

## What Changed

- Add provider-neutral discovery, sync, and refresh APIs. Preserve the
Composio API paths.
- Cache observations by company, saved gateway, viewing user, and
credential version. Retain stale observations after failed or incomplete
scans.
- Add optional Arcade account sync credentials in the vault. Discover
Executor accounts through its supported inventory interface.
- Group native and upstream accounts in one app card. Imported account
menus open their provider. Gateway menus own refresh and sync setup.
- Add Paperclip, Composio, Arcade, Installed, and All chips. Show 50
catalog entries per page. Keep connected accounts above discovery. Keep
explicit provider searches scoped.
- Simplify Composio app setup and refresh its connected app list on the
gateway Permissions page.
- Add a compact agent access card and task creation defaults for
connection setup. Preserve explicit blocks and approval policies.
- Add two replay-safe migrations, service and UI tests, Storybook
journeys, and acceptance stories.

## Verification

- Passed the repository typecheck, full build, token gates, and
migration ordering check.
- Passed the focused provider adapter, connection interaction, and
catalog tests after rebasing onto master.
- Passed all nine database sync and migration replay tests using a
disposable database on the test-drive PostgreSQL cluster. Removed that
database after the run.
- Verified Arcade cursor pagination against its official Go SDK and
passed all eight adapter tests, including short and incomplete pages.
- Passed all 45 interaction tests after making the exact requested tools
and their Allowed/Ask first permissions visible before granting access.
Verified the compact card in Storybook.
- Passed the complete UI suite on the final code: 683 files and 7,432
tests, including the corrected Composio destination assertions. Passed
130 focused tests for the UUID, management-link, and health-status
corrections.
- Passed 22 Composio setup/sync tests, 23 connection-intent service
tests, and the connection migration test in separate disposable
databases. Database startup alone was substituted; the suites exercised
their real SQL and services.
- Passed all 10 OpenAPI route checks and the full-stack
connection-intent browser test, including scoped consent, agent
continuation, and task completion.
- The local full runner encountered embedded PostgreSQL startup failures
on this loaded macOS host. The earlier in-flight run also held the
pre-fix Arcade transform; a fresh run of the final provider suite
passes. The final-head CI is queued during GitHub’s active Actions
incident: https://www.githubstatus.com/. The previous run also lost
several runners simultaneously; its real catalog assertion failures are
fixed and the fresh complete UI suite passes.
- Tested the real test-drive server in the embedded browser with a live
Composio gateway. Detected Airtable and Circleback. Verified refresh
progress, account rows, source chips, search scope, and 50-entry
pagination.
- Arcade and Executor coverage uses provider fixtures. Live credentials
were unavailable.
- Storybook builds successfully and includes grouped native/provider
accounts, stale and unavailable discovery, optional Arcade setup, and
mobile states. The acceptance document records the simulated and live
coverage separately.

- Greptile reviewed final commit
`217b024c27b5933e773ce9419c4e92b1032042c6` at 5/5. All six review
threads are resolved, security scans pass, and the PR has no merge
conflicts. The outstanding remote checks are `ci / Select trusted
runner` and `review`, queued by GitHub. They need to complete before
merge.

## Risks

- Provider response changes can break inventory discovery. Failed scans
retain observations and show stale status.
- Composio scans only the supported catalog and can take time. Large
inventories run in the background with progress and a bounded lease.
- Arcade requires a project API key and user ID when the gateway cannot
supply them. This key is used only for discovery.
- Executor discovery depends on the server's exposed inventory tools.
Unsupported servers report unavailable discovery.
- Cached account rows do not grant access or create executable
connections. Gateway policies still govern tool use. Account deletion
and per-app authorization remain upstream.
- The migrations add tables and one nullable column. Replay preserves
existing rows and company-scoped foreign keys.

## Model Used

OpenAI Codex, based on GPT-6. The session does not expose a more
specific serving model ID or context limit. Used reasoning, repository
tools, code execution, and browser verification.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-05 14:57:04 -05:00
1c07b5903b feat: Chat leads the left nav, agent work beside chats, and a Combined Inbox + Task List flag (#15100)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The left nav is the main way people move between tasks, the inbox,
and Agent Chat
> - The nav has separate Inbox and Tasks rows that show overlapping
work, and Chat is one row among many
> - The side panel beside a chat shows the conversation's own artifacts,
not the work the agent did
> - People want Chat to be easy to find, and they want one place for
their task views
> - This pull request moves Chat to the top of Work, shows the agent's
tasks and artifacts beside each chat, and adds an experimental flag that
folds Inbox into Tasks
> - The benefit is a shorter nav and a chat view that shows what the
agent is working on. Both changes stay off until an operator enables
them

## Linked Issues or Issue Description

Refs #14706 (the secondary Agent Chat navigation this change builds on)
Refs #14848 (reopen the last visited agent chat)

**Subsystem affected**
UI navigation (left nav, mobile tab bar), Agent Chat side panel, task
list and inbox, and the company artifacts API.

**Problem or motivation**
Inbox and Tasks are two nav rows for overlapping work. Chat sits in the
top group with no clear home. The chat rail lists only agents you
already talked to, so you cannot see your other teammates there. The
side panel beside a chat shows only the conversation's own artifacts. It
does not show the tasks and files the agent made.

**Proposed solution**
With Agent Chat on, Chat leads the Work section and the rail lists every
eligible agent. The chat side panel opens on the agent's tasks as cards,
and the agent's artifacts are available from +. A new experimental flag,
Combined Inbox + Task List, makes Inbox a set of views inside Tasks.

**Alternatives considered**
Rebuilding the inbox inside the task list. Instead, `/issues` hosts the
existing Inbox component for inbox views and the existing task list for
status views, so all inbox behaviour stays the same.

**Roadmap alignment**
Agent Chat (ROADMAP.md, "Agent Chat (including CEO Chat)"). All changes
are behind experimental flags that are off by default.

## What Changed

- **Agent Chat nav (streamlined shell):** Chat is the first row of Work,
not a top-group row. Workspaces leaves the nav while Agent Chat is on.
The mobile tab bar is Home · Chat · + · Tasks · Agents. The legacy shell
keeps master's top-group Chat row.
- **Chat rail:** `AgentConversationsSidebar` lists every eligible agent.
The open chat is first, then conversations by recent activity, then the
rest of the roster alphabetically. Terminated agents and agents you left
are omitted unless you have history with them. The picker still marks
only real conversations as "Open chat".
- **Chat side panel:** a new default Tasks tab shows one card per task
the agent created, was assigned, commented on, or acted on, newest
first. It has the task list's filter popover and a sort control. **+ →
Artifacts** shows the agent's artifacts as cards. Cards open in a new
tab. Agent Chat off keeps the old Artifacts tab.
- **Artifacts API:** `GET /api/companies/:companyId/artifacts` accepts
`agentId`. The filter applies to documents, work products, and
attachments by the agent each result is attributed to. The shared
validator and the UI client carry the new parameter, and the OpenAPI
entry picks it up from the shared schema.
- **Combined Inbox + Task List flag (`enableCombinedInboxTasks`, off by
default):** new card in Settings > Experimental. The Inbox row goes away
and its badge moves to Tasks. A Views menu on `/issues` covers Mine,
Unread, Blocked, Recent, Everything, All, Active, Backlog, and Done.
Bare `/issues` opens the last-used view (default Mine). Links that carry
`assignee`, `workspace`, `participantAgentId`, or `q` open All so the
filter is kept. `/inbox/*` and
`/issues/{all,active,backlog,done,recent}` redirect to the matching
view. `/inbox/requests` stays its own page.
- **Task detail breadcrumb:** the view key now decides the source, so
quick-archive still works after a reload from an inbox view.
- **Docs:** `doc/PRODUCT.md` and `doc/SPEC.md` describe the chat rail,
the chat side panel, and the new flag.

## Verification

- `cd ui && npx vitest run --no-file-parallelism src/components/chat
src/components/task-side-panel/TaskSidePanel.test.tsx
src/components/AgentConversationsSidebar.test.tsx
src/components/Sidebar.test.tsx
src/components/SidebarCompanyMenu.test.tsx
src/components/Layout.test.tsx src/pages/AgentChats.test.tsx
src/pages/InstanceExperimentalSettings.test.tsx
src/lib/task-views.test.ts src/lib/issueDetailBreadcrumb.test.ts
src/pages/Inbox.test.tsx src/pages/Issues.test.tsx src/App.test.tsx
src/App.activity-routing.test.tsx
src/components/MobileBottomNav.test.tsx
src/components/CommandPalette.test.tsx`: 20 files, 356 tests pass.
- `cd server && npx vitest run
src/__tests__/company-artifacts-service.test.ts`: 13/13 pass, including
the new agent-filter test across all three artifact sources.
- The new rail test fails against the unmodified rail.
- `pnpm check:token-gates`: all gates clean.
- Manual: enable Agent Chat in Settings > Experimental. Open Chat. The
rail lists all agents. Open a chat. The side panel shows the agent's
tasks. Use **+ → Artifacts** to see the agent's artifacts. Then enable
Combined Inbox + Task List. The Inbox row goes away, and Tasks shows a
Views menu.
- Snapshot baselines are intentionally not updated. See
`doc/design/DECISION-SHEET.md`, "Per-change snapshot verification
demoted to dormant (Jul 13 2026)".

## Risks

- With both flags off, the app behaves like master. The only exception
is the API: it accepts a new optional query parameter.
- With Agent Chat on, the rail can list many agents in a large company.
It uses the agent list the app already loads, and search filters it.
- The Tasks panel reads at most 200 recently updated tasks per agent and
says so when it reaches the limit. The Artifacts panel reads at most 500
of the agent's artifacts.
- Combined Inbox + Task List changes what bare `/issues` opens for
people who enable it. Deep links with a task filter still open All.

## Model Used

- Claude (Anthropic), model ID `claude-opus-5-5`, through Claude Code
with tool use (shell, file edit, test runs). Extended thinking was
enabled.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: scotttong <squadbot000@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-04 02:13:46 -07:00
DottaandPaperclip 7d59de6113 feat(connections): probe provider usage limits on demand (#14936)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Connections store the AI accounts used by legacy and native runners.
> - Subscription accounts can reach session, weekly, model, or paid
usage limits.
> - Operators need to read these limits for a specific stored account
before making a routing decision.
> - This pull request adds an on-demand usage probe to the connection
service and account detail.
> - The result preserves provider limits, reset times, paid usage, and
unknown values for later consumers.

## Linked Issues or Issue Description

**Subsystem affected**

Shared contracts, the connection service and API, and the account detail
UI.

**Problem or motivation**

Managed AI accounts lack a common operation to read their current usage
limits. A local harness probe can read a different login from the
account selected for an agent.

**Proposed solution**

Add `aiConnectionService.probeUsage()` and a board-only connection usage
endpoint. Probe the selected credential grant on request. Support Codex,
Claude, and Grok subscriptions, plus OpenRouter API key limits.

**Alternatives considered**

Harness-specific automatic polling would couple the read to execution
and can read ambient credentials. This change uses the managed
connection credential and leaves scheduling and admission decisions to
later work.

**Roadmap alignment**

This extends the existing Personal & Shared AI Accounts capability. It
adds no routing or quota enforcement. Related: Refs #14459 for managed
OpenAI quota reads; Refs #14781 and Refs #13379 for downstream pacing
and budget work. This operation reads one requested account across all
three subscription providers.

## What Changed

- Add typed usage snapshots and a probe capability flag to managed AI
connections.
- Normalize Codex, Claude, Grok, and OpenRouter responses. Keep model
scopes, provider admission, reset periods, and paid allowances separate.
Preserve unknown values.
- Enforce company membership, credential audience, grant identity, and
connection lifecycle before reading the stored secret.
- Add a board-only `GET
/api/companies/:companyId/ai-connections/:connectionId/usage` endpoint
with `no-store` responses.
- Add manual **Check usage** and **Refresh** actions to account details.
Show compact usage bars, resets, admission and overage status; remove
repeated descriptions and account-default copy. Clear previous results
during a new request or error.
- Add Storybook previews using the production account components for all
four providers, initial checks, loading, and permission errors.
- Add provider, authorization, runner selection, API, and UI coverage.
Document provider sources and live qualification.

## Verification

- Initial provider, authorization, selection, API, and UI validation
passed (96 focused tests): `pnpm exec vitest run
server/src/services/ai-connection-usage.test.ts
server/src/__tests__/ai-connections.test.ts
ui/src/components/ai-connections/AiConnectionUsagePanel.test.tsx
server/src/__tests__/openapi-routes.test.ts`.
- `pnpm -r typecheck` passes for the initial implementation. After
simplifying the UI, 9 usage-panel and date-helper tests, UI typecheck,
token gates, and Storybook build pass. The initial feature module
boundary check also passed.
- Real Codex, Claude, and Grok credentials were saved to encrypted
disposable connections. The actual usage HTTP route returned 200 with
`status: ok`. Legacy and native runner selection checks passed. The
tests started no model turn and exchanged no refresh token. The
disposable databases and vaults were removed.
- Live Claude responses added structured scoped limits. Live Grok
responses omitted included-plan usage. Tests now cover both shapes and
preserve the Grok omission as unknown.
- The full workspace build passes. A full local test run hit a heartbeat
feedback timeout. That case passes in isolation. The duplicate local run
was stopped after all remote checks passed. The Slack ordering and
OpenCode transport CI flakes also pass in isolation and on the CI rerun.


- Current head: `ff3d479029a1c4248190323e221b2803cfb0d79d`. All 54
active checks pass. Two Storybook checks are intentionally skipped by
the workflow. Greptile is 5/5 with no unresolved review findings; the
branch is mergeable.

## Risks

- Subscription usage endpoints can change. Credentials can lack
usage-read permission. The probe returns explicit errors without fresh
limits in these cases.
- A successful probe can contain partial data. Missing utilization or
admission remains unknown. An enabled paid-usage switch does not prove a
funded balance.
- This change adds no migration. It does not change runner admission or
automatic provider selection. Provider requests use fixed endpoints,
disabled redirects, bounded response sizes, and a 15-second deadline.

## Model Used

OpenAI Codex, GPT-6, with reasoning, file editing, shell execution, and
HTTP tools. The session does not expose the exact runtime model variant
or context window size. Real provider credentials were used only for the
authorized live checks.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-02 11:47:14 -05:00
DottaandPaperclip 6c1a75da49 feat(connections): make AgentMail a default connection with inline setup (#14772)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Connections give agents access to external services.
> - AgentMail needs both a saved key and an inbox assigned to the agent.
> - Chat requests offered a setup link instead of an inline card and
could treat a saved key as complete.
> - Inbox setup also hid address conflicts behind a generic server error
and a separate review step.
> - This pull request makes AgentMail a default connection, adds the
inline card, reduces setup to two steps, and shows conflicts beside the
address.
> - Shared native dropdown styles also give every caret a consistent
inset.

## Linked Issues or Issue Description

**What happened?**

AgentMail requests in chat did not show a usable inline connection card.
Manual setup required extra screens, ignored saved account keys, and
could trap new-address setup in a locked inbox dropdown. Agent selectors
omitted the avatar from the selected value. A taken address could
produce an HTTP 403 from AgentMail and appear as an internal server
error. Native dropdown arrows also touched the right edge of their
fields.

**Expected behavior**

Make AgentMail available as a default connection. Ask for the API key
inline, with a direct link to its provider page. Default human access to
the company and agent access to the requesting agent. Resume the agent
only after an assigned inbox is active. Manual setup should ask for an
agent and email address, then finish. Address checks should run as the
user types. Taken addresses should show clickable alternatives. A domain
dropdown beside the name should prefer a verified custom domain. Setup
should suggest authorized saved AgentMail keys and show agent avatars in
the picker and selected value.

**Steps to reproduce**

1. Ask an agent to connect AgentMail when it has no assigned inbox.
2. Check that an inline API-key card appears and links to the provider's
API-key page.
3. Open AgentMail setup, choose an agent, and request an address that is
already taken.
4. Correct the inline error, refresh, and finish setup with the same
request ID.
5. Inspect native dropdown carets in light, dark, disabled, and
right-to-left states.

Uses the bounded provider-error parser merged in #14768. Related work:
#13256 introduced AgentMail; #14725 expanded connection search.

## What Changed

- Stop recurring email queries for tasks that have no email thread.
Share the query between the thread provider and activity view. Keep
email-task updates and invalidation-based discovery.
- Make AgentMail available without the experimental chat setting. Keep
the catalog, setup and management routes, agent Channels tab, task email
feed, receiving worker, and agent tools available by default. Other
experimental chat providers stay gated.

- Make the email address and copy icon a single clickable action with
the shared Copied! confirmation. Add View inbox linking directly to the
matching AgentMail console inbox, with the address encoded as one URL
path segment.

- Reorganize inbox Settings around the copyable email address, usage
instructions, and receiving status. Move reconnect credentials into a
disclosure and separate the Disconnect action. Add production Settings
stories for active, paused, unassigned-address, revoked, webhook,
long-address, mobile, and reconnect states. Show repair controls when
the inbox has an error. Keep usage instructions tied to an active inbox
with an address.

- Add AgentMail channel intents and an inline key field with the direct
API-key URL.
- Keep setup and retry state tied to the interaction. Require an active
inbox for completion. Preserve company and agent access checks.
- Reduce manual setup to agent selection and email selection. Put the
domain dropdown beside the address and default to a verified custom
domain. Preserve explicit choices across reloads. Keep receiving
settings under Advanced options.
- Check the initial address and edits after a 350 ms pause. Abort
superseded requests and ignore stale responses. Show clickable
suggestions and retain known creation conflicts across reloads.
- Add a company-scoped, manager-only address check using the saved
credential. Search the visible inbox list instead of fetching an
uncreated inbox: live AgentMail retains negative lookups that can break
subsequent access-key creation. Unlisted addresses remain unknown;
creation is authoritative.
- Suggest labeled saved AgentMail keys in both manual setup and the
inline card. Filter by company, provider, active credential, and
current-user grants on the server. Prefer an account key and preserve
the selected key or an explicit new-key choice across refresh. Use
verified scope metadata and bounded concurrent checks for legacy keys.
Never return secret values.
- Catch an inbox-only key before the email step. Allow its existing
inbox only after an explicit choice. Recover old locked drafts at the
key picker. Save the replacement key before retiring an empty draft,
then use a new setup URL so refresh preserves the switched account; stop
if cleanup fails. Preserve already allocated addresses and their
original accounts.
- Use the shared AgentSelect in email setup. Show the canonical agent
avatar in each option and the selected value, including other consumers
of the shared component. Add regression coverage for legacy and current
Lucide agent-mention icon formats.
- Start each catalog Add connection with a fresh setup identity. Honor
Finish setup's exact draft/account/address instead of resuming an
unrelated browser draft. Return Cancel and Done to Connectors and Email
settings to the inbox. Group the task/thread explanation in a How it
Works card.
- Route AgentMail catalog removal through the email inbox control API,
including unfinished drafts. Refresh both the catalog and inbox views.
- Render each inbox management tab separately. Access uses the saved
account grants and agent controls; Conversations and Activity use the
shared persisted email feed. Activity lifecycle actions use the email
API. Reconnect returns to inbox Settings. Conversation failures show a
retry instead of a false empty state. Email delivery recovery stays in
the task.
- Map documented provider address conflicts to a field error. Preserve
actionable messages for other failures.
- Preserve non-secret draft fields across refresh, scoped to the
requested agent. Never save API keys in browser storage. Resume partial
inbox creation with the original agent, address, and request ID.
- Show an already-created address with explicit retry and new-address
recovery instead of locked inputs. Preserve the original inbox and
resumable draft when choosing another address. Distinguish runtime-key
404 errors and log safe provider status/operation/code.
- Apply final agent access once within email setup authorization for a
new account whose original installs are unchanged. Preserve later
permission edits and reused account installs. Support in-place retry of
progress loading.
- Let a failed inline setup change keys after retiring an empty draft.
Persist its replacement setup identity without storing secrets. Recover
a server-saved account when refresh interrupts the save response, while
preserving intentional account changes.
- Render the production setup in Storybook and add error, recovery, and
mobile states.
- Inset native select carets in shared CSS. Preserve custom icons,
listboxes, keyboard behavior, and forced-color controls.
- Add browser regression coverage and an AgentMail Product E2E case with
persisted-state and rendered-card evidence.

## Verification

- Full `pnpm -r typecheck`, `pnpm build`, `pnpm check:token-gates`, and
`git diff --check` passed after the default-availability change.
- All 485 focused tests passed. These cover setup, management, catalog
and route gates, connection intents, email authorization, Cursor
execution, and the OpenAPI contract. All 39 email integration tests run
with the experimental chat setting off.
- The shared polling change passed four behavioral tests, UI typecheck
and build, and token gates.
- `tests/e2e/agentmail.spec.ts` passed with the actual server setting
off. This full-stack browser test uses simulated provider responses. It
covers catalog entry, saved keys, editable address and domain controls,
creation, conflicts, retry, all management tabs, clipboard feedback, the
provider link, and task email rendering.
- In the live local browser, Add connection reached the editable email
step with the saved account key. The verified custom domain was selected
by default. Both domain choices worked. The existing inbox Settings page
remained available. Both active inboxes completed new mail checks with
the setting off. No new provider inbox or email message was created for
this pass.
- Earlier live provider acceptance covered creation on a verified custom
domain, Finish connecting on the reported draft, successful mail checks
after refresh, and catalog removal of disposable draft and active
connections. Clicking the email address copied the exact address and
showed Copied!. View inbox opened the same inbox in AgentMail’s console.
No email messages were sent.
- Production setup and Settings Storybook builds and interactions
passed. Settings states include active, paused, unassigned, revoked,
webhook, long-address, mobile, and reconnect. Receiving and
revoked-access stories had zero accessibility violations.
- Full local `pnpm test:run` on an earlier revision completed with
14,709 passing, 87 skipped, and four transient failures. All four failed
cases passed in focused reruns without product changes. That serial full
local command was not repeated after each follow-up. The latest-head
full CI suite is the final test gate.
- CI found an obsolete browser assertion that hid every channel when the
flag was off. Updated it to keep AgentMail and the Channels surface
visible while preserving the GitHub chat route gates. All 11 provider
browser tests passed locally after scoping the Channels selector to the
agent sidebar. Two initial local attempts stopped at temporary Postgres
initialization. The passing run used a separate disposable database on
the existing local Postgres server; it was removed after the test.
- Updated the remaining sidebar and aggregator discovery assertions for
default AgentMail availability. Ordinary task fixtures now return no
email thread. All 128 sidebar/task-page tests and all 42 aggregator
tests passed locally.
- Latest head `b42bb4cd5cc9f2d01a99ab8026832d5a956ea85f`: full CI
passed, with 54 successful checks including Snyk and two intentional
Storybook skips. The CI run is
https://github.com/paperclipai/paperclip/actions/runs/37020833647. A
fresh Greptile review scored 5/5 with no unresolved threads. Live model
evaluations and inbound/outbound email delivery were not run.

## Risks

- AgentMail no longer needs experimental opt-in. Setup still requires a
human to connect an account and assign an inbox. Inline setup creates an
inbox after a human submits a new or saved key. Company access, agent
access, inbox assignment, and completion checks remain enforced.
- AgentMail read APIs cannot prove global address availability. The
visible-list check is bounded to 100 entries and cannot see inboxes
outside the key’s scope. The UI reports this limitation, suggests
alternatives without claiming they are free, and keeps final creation
conflicts inline. Lookup outages show an error without preventing the
authoritative creation attempt.
- Native select CSS affects the whole app. Custom-icon selects and
multi-row lists are excluded. Forced-color mode keeps the browser caret.
- Saved-key discovery uses stored verified scope metadata and checks
authorized legacy credentials concurrently within a shared three-second
deadline. Provider outages mark legacy choices unavailable; users can
still enter another key. Final use rechecks authorization and provider
access.
- No database migration or transport default change. Live connection
remains the default.

## Model Used

OpenAI Codex, GPT-6, with reasoning, tool use, and code execution. The
exact served model ID and context-window size are not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused suites; full-suite
limitation documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green (latest head
`b42bb4cd5cc9f2d01a99ab8026832d5a956ea85f`)
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
(latest head `b42bb4cd5cc9f2d01a99ab8026832d5a956ea85f`)
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-02 10:01:15 -05:00
DottaandPaperclip e00d10d5d5 fix(connections): repair stale AI defaults from agent settings (#14916)
## Thinking Path

> - Paperclip manages AI agents and controls the credentials used for
their work.
> - Managed AI connections resolve each responsible user's provider
default.
> - Agent settings created another account but kept the old default
selected.
> - A rejected provider test left the old account marked as connected.
> - Claude ACP reported a typed login failure as a generic
terminal-access error.
> - This pull request repairs the selected account or selects the new
login explicitly.
> - Agents can save and run with the repaired credential, and failed
logins request sign-in.

## Linked Issues or Issue Description

- Fixes #14831.
- Refs #13867. Environment failures remain separate from
credential-health failures.

## What Changed

- Add an agent-settings action to reconnect an unavailable personal
default in place. Keep its connection, grant, default, and agent access.
- State that a new account becomes the user's provider default. Select
its returned grant before changing the agent binding. Keep the actual
sign-in method.
- Show default-update errors and allow retry without another provider
login.
- Show the agent-access choice. Connection managers start with
company-wide access for their own tasks. Other members start with access
for the current agent.
- Use the server's connection-manager permission in the shared list
response. This includes members with a custom management grant.
- Mark credentials as needing attention after an explicit login
rejection in Test or Save. This includes API-key 401 and 403 responses.
Network, quota, and server failures keep the credential health
unchanged.
- Reuse the credential-generation check so an old failure cannot
invalidate a newer reconnect.
- Route Claude's typed provider `access` failure to the existing
login-recovery flow. Replace its generic terminal-access fallback with a
sign-in message.
- Add regression tests and update the AI Connections documentation.

## Verification

- Red: the UI tests failed on the missing reconnect action, unused
returned grant, missing access choice, and lost default-update error.
The server tests failed because rejected credentials stayed connected.
The real ACP fixture returned `acpx_turn_failed` for typed login
failures.
- Green: 156 tests passed across the AI connection, hiring, agent field,
and New Agent suites. All 37 environment-route tests passed. The Claude
ACP authentication fixtures also passed.
- `pnpm check:token-gates` passed.
- `pnpm -r typecheck` passed.
- `pnpm build` passed.
- The full local `pnpm test:run` passed 707 files and 14,503 tests, then
exited with an agent-conversation timeout and embedded PostgreSQL
startup failures in unchanged suites. The isolated conversation and
migration tests passed on rerun. Later local test groups did not run
after this failure.
- [All CI gates
passed](https://github.com/paperclipai/paperclip/actions/runs/37012669356)
on commit `38513dfe2`. This includes the full test matrix, browser
tests, typecheck, build, Runner checks, and canary dry run.
- Greptile reviewed commit `38513dfe2` and returned 5/5 with no open
findings.
- The regression tests use a real embedded database and a real ACP
fixture process. Live provider sign-in requires a valid account and was
not run.

## Risks

- Connecting a new account from agent settings changes the user's
provider default. The dialog states this before sign-in.
- The displayed access choice can allow all company agents to use the
account for its owner's tasks. Reconnect keeps the existing access.
Server permissions still control installs.
- Claude's typed `access` category maps to the provider's
`auth_required` signal. Tool and workspace request failures retain their
existing classification.
- No database migration or provider credential format changes are
required.

## Model Used

- OpenAI GPT-6 through Codex. The exact served model identifier and
context window are not exposed in this session. Capabilities used:
reasoning, repository tools, code editing, and command execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-02 08:57:52 -05:00
DottaandPaperclip efc2e6810e fix: show each task once in dashboard agent cards (#14847)
## Thinking Path

> - Paperclip helps people manage AI agents and their tasks.
> - The dashboard shows recent agent activity in compact cards.
> - Those cards use run records, so two runs for one task can create
duplicate task cards.
> - An operator needs to see each task once when scanning the dashboard.
> - This pull request selects one run per linked task before it applies
the card limit.
> - The live runs page still shows each run for run inspection.

## Linked Issues or Issue Description

**What happened?**

The dashboard showed the same task in two agent cards when that task had
both an active run and a completed run.

**Expected behavior**

The dashboard should show a linked task at most once. It should keep the
active run card when one is present.

**Steps to reproduce**

1. Start an agent run for a task that already has a completed run.
2. Open the company dashboard.
3. Observe two cards linked to the same task.

**Paperclip version or commit**

Reproduced on the pre-change master at `8b4aa0692`.

**Deployment mode**

Local dev, built from source. The bug is in the core dashboard UI and
does not depend on an agent adapter or database mode.

## What Changed

- Select distinct linked tasks from capped active and recent run samples
before applying the dashboard card limit.
- Keep separate cards for runs without a linked task.
- Preserve the dashboard's count of additional distinct cards behind the
live-runs link.
- Add UI and embedded Postgres regression tests for duplicate runs and
document the dashboard rule.
- Give the existing multi-request cross-tenant authorization test enough
time on loaded CI runners.

## Verification

- `pnpm --filter @paperclipai/ui exec vitest run
src/components/ActiveAgentsPanel.test.tsx`
- `pnpm --filter @paperclipai/ui exec vitest run
src/api/heartbeats.test.ts`
- `pnpm exec vitest run server/src/__tests__/dashboard-service.test.ts
server/src/__tests__/agent-live-run-routes.test.ts`
- `pnpm exec vitest run
server/src/__tests__/agent-cross-tenant-authz-routes.test.ts`
- `pnpm --filter @paperclipai/ui typecheck`
- `pnpm --filter @paperclipai/server typecheck`
- `pnpm --filter @paperclipai/ui build`
- `pnpm -r typecheck`
- `pnpm build`
- `pnpm check:token-gates`
- Review the dashboard with an active and a completed run on the same
task. Confirm that it shows one card. Open Live agent runs to inspect
both run records.

## Risks

- A very high volume of recent runs for one task can fill the capped
sample and leave older tasks off the dashboard. The Live runs page
remains available for full run inspection.
- The dashboard may fetch up to 50 distinct run representatives to
preserve its overflow count. The default run API response and persisted
data are unchanged.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- OpenAI Codex, GPT-6. The runtime does not expose the exact model ID or
context window size to this task. The model used reasoning, tool calls,
and code execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-01 12:18:35 -05:00
DottaandPaperclip 018993140f feat: let agents name prompt-only tasks (#14761)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Users create tasks with a title and a description.
> - A required title adds work when the prompt already explains the
request.
> - An agent can name the task once it reads that request.
> - This pull request accepts prompt-only tasks and starts them with a
short prompt slice.
> - A scoped title tool lets the assigned agent replace that slice early
without changing execution state.
> - A live browser eval checks the real agent call, saved title, audit
entry, and preservation of user titles.

## Linked Issues or Issue Description

**Subsystem affected**

Cross-cutting: task creation, shared contracts, database, server, runner
tools, and board UI.

**Problem or motivation**

Users must currently write a title before they can submit a detailed
task prompt. The agent has enough context to write a useful title
itself.

**Proposed solution**

Make the title optional when a description is present. Save the first
120 characters of the normalized prompt as a provisional title. Ask the
assigned agent to call `set_task_title` early. Use an atomic
provisional-title guard to preserve titles supplied or edited by users.
Keep explicit titles supported.

Related: #14543 and #14556 concern empty-title submission. This change
intentionally enables that submission when a prompt is present, instead
of requiring a title.

## What Changed

- Add the `titleNeedsGeneration` field with an idempotent migration.
Keep existing titles unchanged.
- Add `PUT /api/issues/:id/title` and the native and legacy
`set_task_title` tool. Enforce company access, active-run ownership,
shared, bounded retry receipts across native/HTTP calls, and
transactional audit logging. Refresh external-object links after commit,
with the same feature gate and plugin detectors as ordinary title edits.
- Add early naming guidance in Standard, Ask, and Plan task context.
Preserve the description, status, and assignment.
- Allow prompt-only root and child task creation, plus draft restoration
in the New Task dialog. Keep user titles supported.
- Add an opt-in Product E2E suite for prompt-only Standard and Ask
tasks, plus an explicit-title control. It checks actual provider calls
within the first five tools, persisted state, audit attribution, and the
reloaded UI.
- Preserve a closed vocabulary of API key maintenance phrases in
declared prose while rejecting opaque credential suffixes. Add one
bounded naming retry after wording is rejected, without treating the
rejected call as a saved title.
- Repair the native cleanup receipt check exposed during full
verification: accept matching input digests, retain legacy input checks,
and reject conflicting receipts.

## Verification

- Live Product E2E on `f43478473800e3a46b85c5ee79677efdb15108e7`: **3/3
passed** with native Codex `gpt-5.4-mini`, first attempts only,
automatic retries disabled. Standard and Ask each saved “Rotate expired
API key” on their first tool call, with matching persisted state and a
single same-run audit entry. The explicit-title control retained its
user title with zero title writes. All three verified the reloaded
browser UI.
- Campaign: `local-2026-09-30T21-30-11-021Z`. Earlier failed campaigns
are retained separately; they exposed credential-prose handling and
prompted the naming recovery fix. No failed result was regraded or
deleted.
- Reproduce with `pnpm test:e2e:runner -- --id
task-titles.runner-codex-mini.local.prompt-title-standard --id
task-titles.runner-codex-mini.local.prompt-title-ask --id
task-titles.runner-codex-mini.local.preserve-explicit-title
--max-automatic-retries 0` and an authorized provider key.
- Full `pnpm -r typecheck` and `pnpm build` passed on the latest commit.
The runner build used the configured external eval source tree.
- Product E2E unit suite: **61 files, 818 tests passed**; E2E typecheck
and UI token gates passed.
- Title API/native regressions cover prompt-only and explicit child
creation, user edits, ownership/company isolation, external reference
refresh, cross-surface retry replay, and the 64-key limit without
receipt eviction. All passed. Prompt-context coverage: **44 tests
passed**.
- Rust credential regressions: **35 tests passed**, including benign
maintenance qualifiers and opaque credential rejection in every declared
prose field. Catalog/report reconciliation: **28 tests passed**. Native
recovery: **560 tests passed**.
- Broad local `pnpm test:run`: **14,555 tests passed** in the general
server group; two suites failed to initialize embedded PostgreSQL and
the existing 40,000-file Git streaming stress test exceeded its
300-second macOS timeout. All three suites then passed in isolation (**5
tests passed**) without code or timeout changes. The original full local
command exited nonzero and is not being represented as a clean full run.
- Latest-head GitHub checks are green: **53 passed, 4 skipped, zero
failed or pending**, including all test shards and the canary packaging
dry run. Greptile reviewed the same commit at **5/5**, with zero
unresolved review threads.

## Risks

- The additive database field must reach the server and UI together. The
migration uses `IF NOT EXISTS` and defaults existing tasks to a final
title.
- Title generation depends on the assigned agent running. Tasks without
a run keep their provisional title.
- Live qualification covers the native Codex path in Standard and Ask
modes. API/legacy and Plan behavior have deterministic coverage.
- The credential-prose exception validates the entire suffix against a
closed maintenance vocabulary. Unknown suffixes, assignments, quoted
values, credential prefixes, and diagnostics retain strict checks.

## Model Used

OpenAI Codex, based on GPT-6, with reasoning, tool use, and code
execution. The exact deployment ID and context window are not exposed in
this session. The live eval uses the native Codex `gpt-5.4-mini`
profile.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 16:48:24 -05:00
DottaandPaperclip d432dc7fa3 Add GitHub-synced skill sources (#14713)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Company skills supply instructions and files to those agents.
> - GitHub imports already exist, but users cannot manage repositories
as skill sources.
> - Repository refresh also needs caller-authorized access and complete
local packages.
> - This pull request adds Sources inside Skills and reuses GitHub
connections from Apps.
> - Installed snapshots let agents use skills without fetching GitHub
during a run.
> - Manual refresh preserves skill identity and leaves failed imports on
their last good version.

## Linked Issues or Issue Description

**Subsystem affected**

Cross-cutting: skills UI, server, database, shared contracts, and
runtime materialization.

**Problem or motivation**

Users keep skills in GitHub repositories. They need a clear way to
select, import, and refresh those skills. Existing imports do not expose
repository management or consistently preserve supporting files.

**Proposed solution**

Add company-scoped skill sources. Browse repositories from all
accessible GitHub connections, or paste a public repository or branch
URL. Select whole skill packages, inspect included files and reference
warnings, and install complete, immutable snapshots. Refresh each source
manually.

**Alternatives considered**

Project repository settings hide the workflow from Skills. A second
GitHub connector would duplicate credentials and grants. Upstream
editing and PR creation are separate work.

**Roadmap alignment**

This implements the Skills Manager direction in ROADMAP.md. The
maintainer requested this scope and reviewed the component and full-app
journey stories before implementation.

Related reports: Refs #10285, Refs #10949, Refs #13464. Related work:
#14356, #13656, #9268.

## What Changed

- Add source and entry records, an idempotent migration, company-scoped
APIs, and legacy GitHub import adoption.
- Reuse current caller grants and credential refresh. Combine and
deduplicate repository inventories across accessible connections. Pasted
public URLs also prefer the active user’s authorized connections. Tokens
stay in the Git child environment, never argv or disk.
- Fetch a shallow Git snapshot at one immutable commit. Scan the full
local tree, including hidden and nested folders. Read Git objects
without checkout or archive transformations and enforce nested package
boundaries.
- Bound Git downloads to 128 MiB and three minutes. Cancel active
process groups and remove incomplete downloads. Preserve cancellation
and deadlines while progress drains; close stalled HTTP progress streams
after 30 seconds. Reuse caller-scoped temporary snapshots for
preview/import after reauthorization.
- Index repository package boundaries once and cap expanded work at
1,000 packages, 10,000 files, and 100 MiB, including repeated copies of
shared blobs. Bound path depth and the shared path index. Discovery
keeps audited manifests without retaining all package bodies.
- Resolve moving branches before fetching so unchanged discovery reuses
caller-scoped snapshots. Limit active scans, scan frequency, and new
downloads per caller and company; quotas apply before metadata reads and
across connections, and cached scans do not consume the download quota.
- Store complete versions with script content, binary bytes, and
executable modes. Preserve these through copies, runtime caches, and
runner packaging.
- Stage downloads before publication. Use source leases, revision
checks, and transactional activity records. Keep installed versions
after failures, upstream deletion, deselection, and disconnect.
- Add the approved import flow, Sources page, selection tree,
provenance, read-only Studio behavior, and saved return from GitHub
setup.
- Add package manifests, commit-pinned file previews, and separate
runtime requirements and reference warnings. Supporting files are
included together; nested skills remain independently selectable.
Preview requests reauthorize the caller and re-audit package content.
- Show installed skills as compact links beneath each source. Repository
titles open GitHub. Keep Refresh, Select skills, and Disconnect source
in a three-dot menu. Source rows omit the branch, imported count, and
refresh timestamp; action alignment and repository titles work at narrow
widths.
- Stream discovery metadata over an opt-in NDJSON response. Show
measured Git download progress and real package/file counts, animate
newly checked skills, support cancellation, and require a complete scan
before selection. Keep the existing JSON API.
- Retain component stories and add a separate full-app journey story
group. Include fixed progress states and interactive scan,
large-repository, interruption, and saving stories.
- Update Skills documentation and product contracts. Suppress private
GitHub skill references in telemetry. Privacy review requested for the
telemetry changes.

## Verification

- Local repository typecheck, full build, token gates, and Storybook
build passed during this work. Focused transport, authorization,
scanner, persistence, route, and UI tests pass. The final UI refinement
passes all eight focused UI tests, UI typecheck/build, and token gates.
The scanner resource and repeated-discovery fixes pass 132 focused
scanner, transport, authorization, source-service, route, and rate-limit
tests, plus server typecheck/build. Full-suite verification comes from
CI; the older full local Vitest run was stopped after unrelated chat
failures and a font-test failure, all of which passed in fresh focused
runs. At commit `1098d5996`, all 54 active checks pass; two optional
Storybook jobs are skipped. CI covers repository typecheck, build, the
full test suites, browser shards, and the canary dry run. Greptile is
5/5 with no open findings; the security scan also passes.
- Adversarial scanner tests verify repeated-blob byte accounting with
and without declared sizes, package/file/path caps, one-time repository
indexing, metadata-only discovery audits, and nested package boundaries.
Additional tests cover branch movement, snapshot reuse, caller/company
quotas, isolation across connections, active-lease cleanup, quota
recovery, and rejection before any metadata API call.
- Real Git tests verify hidden paths, exact binary bytes, executable
modes, export-ignore preservation, symlink/submodule reporting, pinned
commits, caller-scoped cache reuse, cancellation, cleanup, and
credential isolation. Regression tests hold both download slots with
permanently blocked progress callbacks, verify timeout/cancellation
cleanup and retry, and exercise HTTP backpressure cancellation. Access
tests cover automatic public-URL connection selection and revoked
grants. Database tests verify company and grant audiences.
- Live isolated browser test: the public `anthropics/skills` scan now
completes and discovers all 20 skills without connecting an account.
Imported canvas-design with all 83 files, opened it from Sources, and
verified the installed binary-font preview/download control. Package
previews also expose the complete file inventory before import.
Cancelled an active Git download and retried successfully to all 20
discovered skills; the browser displayed measured download progress. The
current audits reject four other packages; eligible selections remain
importable.
- Browser checks verify the simplified source rows at desktop and narrow
widths, keyboard navigation into the actions menu, Refresh from the
menu, selection, and fixture disconnect with installed skills retained.
Storybook includes a menu-open checkpoint and a 320px layout.
- Storybook includes receiving/preparing download checkpoints and a
timed full-app import journey, plus cancellation, retry,
large-repository, and saving states. Streaming tests cover split UTF-8
frames, incomplete streams, late responses, cross-company requests, HTTP
errors, and JSON compatibility.
- Earlier live acceptance on this PR imported `stitch-skill` with
`DESIGN.md`, assigned it to an agent, disconnected its source, and ran a
successful Studio test that read both installed files. An editable copy
changed independently. Both Skills variants, mobile selection, and
return from GitHub setup were exercised.
- Private access, revoked credentials, OAuth success return,
binary/script preservation, concurrent refresh, transaction rollback,
version pins, and legacy adoption have automated coverage. A real
private-repository OAuth grant was not created during this test.

## Risks

- The migration groups recognizable legacy imports without provider
calls. Their first successful refresh completes the local package
snapshot.
- Reference checks are advisory. They cover Markdown links and explicit
relative resource paths, not arbitrary runtime dependency graphs.
Preview text is capped at 64 KiB; imported bytes remain complete.
- Git must be installed on the server. Shallow fetches still download
the branch snapshot, including files outside selected packages.
Downloads have size/time/concurrency limits. Temporary caches are
bounded and caller-scoped. GitHub API quota still applies to repository
metadata and the connection picker; content no longer uses per-file API
requests. Failed scans retain installed content.
- Sources depend on the current caller's GitHub access. A saved
connection does not grant access to another person's token.
- GitHub script support and immediate manual refresh are explicit
maintainer-approved requirements. The operator trusts the selected
repository and accepts upstream script and executable-mode changes on
refresh. Static audits are not a sandbox or a guarantee of safe code;
agents may later invoke installed helpers under their runtime
permissions. Import and refresh do not execute scripts, hooks, package
installation, or builds. Raw URL and skills.sh imports keep their prior
script restrictions.
- Source originals remain read-only. Refresh affects subsequent unpinned
runs; explicit pins and active runs retain their versions.
- The telemetry change removes source-managed GitHub identifiers from
skill-reference events. It introduces no event or field. Please review
the privacy boundary.

## Model Used

OpenAI Codex, based on GPT-6, with reasoning, code execution, and
browser tools. The exact serving model ID and context-window size are
not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 13:32:58 -05:00
DottaandPaperclip 1b48e73e0b feat(ui): add secondary navigation for agent chat (#14706)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agent Chat already provides a persistent conversation with each
agent.
> - Its shortcuts share the primary navigation and do not give chats a
dedicated place.
> - People need to find agents, start a chat, and switch conversations
without moving the page layout.
> - This pull request adds a secondary chat sidebar and a landing page
around the existing chat surface.
> - The same conversation, composer, history, and context panel remain
in use.

## Linked Issues or Issue Description

Refs #13283 and #13420. This extends the existing experimental Agent
Chat navigation after review of the component and page stories. It
supports the CEO Chat roadmap item through the existing task-backed
conversation model.

**Subsystem affected**

The board UI and the company-scoped conversation list API.

**Current behavior**

Chat shortcuts sit inside the primary navigation. There is no dedicated
landing page with a searchable conversation list. A separate landing
header also moves the sidebar when an agent is selected.

**Proposed behavior**

Show a Chat entry in primary navigation. Keep a searchable agent sidebar
beside the chat content. The plus button starts or reopens the current
user's single conversation with that agent. Keep the header and sidebar
in the same positions before and after selection.

**Reason and benefit**

People can find agents and return to persistent conversations without
leaving the chat area or creating duplicate chats.

**Breaking changes**

The experimental chat navigation changes. Explicitly adding a chat now
resolves its conversation immediately. Direct visits to unused agent
chat URLs remain read-only. The existing per-agent routes and message
contracts remain compatible. No database migration is required.

## What Changed

- Add an account- and company-scoped conversation list endpoint with the
existing access checks, feature gate, and OpenAPI entry.
- Add the live secondary sidebar, landing page, avatars, search, loading
states, errors, and retry controls.
- Make the agent picker wait for chat creation and display failures.
Existing agents reopen the same conversation. A dismissed selection
cannot close a reopened picker or navigate over a newer choice.
- Preserve recent-activity ordering and terminated agents’ chat history.
Scope live list refreshes to the current user’s conversation events. A
failed historical-agent lookup leaves healthy chats usable and offers a
focused retry.
- Keep the sidebar and header stable across chat routes. Keep mobile
selection in the navigation drawer.
- Use the production components in Storybook. Prepare the theme and
mobile viewport before mounting the page to avoid the startup flash.
- Update product documentation, the design guide, and navigation tests.
Replace old browser expectations for stars and recent shortcuts with
persistent conversation and layout coverage.

## Verification

- `pnpm -r typecheck`, `pnpm build`, and `pnpm check:token-gates` pass
after rebasing onto master.
- Focused UI tests pass, including 105 sidebar, picker, and live-update
checks after review fixes. The 33 conversation service and route tests
and 10 OpenAPI checks pass, including ownership, feature gating, and
concurrent creation.
- The full local test run passed 14,163 server tests before three
environment or timeout failures. The embedded Postgres startup,
connector socket, and native runner failures all passed direct reruns.
- Browser test-drive verification covers a real provider reply, add and
reopen, persisted history after reload, no-match search recovery, mobile
drawer dismissal, and top-aligned context panels.
- Browser measurements confirm that the sidebar has the same position
and dimensions on the landing page and an agent conversation.
- Storybook builds and its add-and-reopen interaction passes.
- The revised browser regression passes locally against a freshly built
throwaway instance. It covers stable sidebar geometry, add/reopen
uniqueness, drafts, search, history, and terminated-agent history after
reload. The full CI browser suite also passes.
- Latest commit `b323577d9523180104df4000eaceedea2772608c`: all 54
completed checks pass, including the complete server/workspace/browser
suites, aggregate verification, build/typecheck, security scans, and
canary packaging. The two Storybook jobs are skipped by their workflow
conditions. [CI
run](https://github.com/paperclipai/paperclip/actions/runs/36714052050).
- Greptile reviewed this same commit at 5/5 with no remaining actionable
findings; all review threads are resolved.
- Reviewer path: enable Agent Chat, click Chat, use plus to choose an
agent, send a message, switch away, and reopen that agent. One
conversation must remain, with its history intact.

## Risks

- The new sidebar lists persistent conversations instead of starred and
recent shortcuts.
- Chat creation is asynchronous. Errors stay visible in the picker, and
delayed responses cannot navigate into a previous company or account.
- The shell adjustment is limited to chat routes and preserves the
existing conversation implementation.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, code execution, and browser
tools. The session does not expose the exact API model ID or
context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#123` / `Refs #123` OR (b) described the issue in-PR following the
relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 07:35:32 -05:00
DottaandPaperclip d72389bee2 feat: add Browser Use Cloud connector and live task browsers (#14627)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The Apps gateway gives agents governed access to external tools.
> - Browser Use Cloud can run browser work, but a tool result alone does
not let a person watch or take over.
> - A task needs a durable browser session, a visible viewer, and
recorded costs.
> - This pull request adds a Browser Use Cloud v4 connection and
interactive browser tabs on tasks.
> - People can follow the work, interact with the page, and retain the
browser after the agent finishes.

## Linked Issues or Issue Description

**Problem or motivation**

Agents need governed access to Browser Use Cloud. People need to see and
interact with the same browser from the task. A browser must remain
available after a run finishes and appear at the correct point in the
task feed.

**Proposed solution**

Add a native REST connection for the v4 API. Bind each session to its
company, task, agent, and credential grant. Open its interactive viewer
in the task side panel. Record provider costs as financial events. Use
`browser-use-cloud` as the app and connector key. Keep its skill with
the connector and deliver it only with authorized connection tools.

**Alternatives considered**

A v3 MCP connection would expose tools without the v4 lifecycle
integration. An external viewer link would leave the task. A fixed
viewer size would prevent pages from responding to changes in the task
pane.

**Roadmap alignment**

This extends the governed Apps gateway and Connected Apps roadmap. It
uses the existing task, grant, secret, approval, and financial records.
The work was requested by the maintainer. A search found no duplicate
Browser Use connector PR or issue.

## What Changed

- Add the Browser Use Cloud app, brand asset, API-key connection, and
profile settings under the `browser-use-cloud` key.
- Bundle the `browser-use-cloud` skill with the connector. Keep it out
of global `skills/` discovery. Deliver it only with authorized task/run
connection tools. Remove retired connector skill keys from runtime
overlays and preserve unrelated browser skills.
- Expose seven v4 tools through the governed gateway and deliver them to
native and CLI agents.
- Persist sessions, browsers, runs, event and recovery cursors, shutdown
leases, and cumulative cost accounting. Recover uncertain paid starts
without replaying them.
- Enforce task ownership, credential grants, approvals, revoked access,
and budget limits.
- Add interactive task browser tabs and compact chronological feed
entries. Retain the viewer across tab switches and keep visible idle
browsers open.
- Add debounced automatic viewport fitting, standard size presets, and a
viewer ownership lease.
- Add lifecycle, authorization, accounting, viewport, UI, and Storybook
coverage.
- Add an idempotent database migration after the current master
migration. Preserve deployed migration hashes. Migrate pre-release Cloud
connection and financial keys without replacing grants, credentials, or
browser history.
- Document provider behavior, live acceptance results, and the lack of
documented passkey forwarding.

## Verification

- Full workspace typecheck and production build pass on the updated
branch.
- Token gates, brand asset validation, module boundaries, and migration
ordering pass.
- Cloud tests verify global skill exclusion, authorized task/run
delivery, unassigned agents, disabled connections, revocation, adapter
isolation, and secret exclusion. The existing AgentMail connector
assignment test also passes.
- Migration replay runs twice against existing browser work and
financial records. It preserves the records and avoids duplicate costs.
- The focused provider, app catalog, OpenAPI, connection gateway, and
migration regression suites pass. Recovery coverage includes lost
replies, process crashes, provider rejection, and browser arrival
acknowledgement.
- All 54 checks pass on `2974b5f03641ad0cea3c941d8c02579316fa8c92`,
including the full test matrix, browser E2E shards, build, typecheck,
security, and release canary. Two optional Storybook jobs are skipped.
- Greptile is 5/5 on the same commit, with zero unresolved review
threads. The corrected review uses the actual master-to-head diff.
- The local `pnpm test:run` started and was stopped after the full CI
matrix passed. It did not complete locally; the full-suite result above
comes from CI.
- Earlier live acceptance used an isolated company with a capped
provider credential. The agent opened paperclip.ing, the embedded viewer
accepted navigation, and the same browser stayed available after
completion and tab switches.
- The local Storybook build passes. Stories cover the panel, footer,
feed entries, settings, lifecycle failures, and viewport modes with an
offline viewer fixture.

## Risks

- Browser Use charges for hosted work. Provider caps and local budget
checks reduce exposure; reported costs can arrive after work completes.
- Viewer and CDP URLs grant access to the browser. The server validates
and restricts them. They are excluded from agent results and durable
event data.
- Runtime resizing of v4 agent browsers uses a provider option confirmed
by live testing but absent from its published agent schema. Resizing
during a click may invalidate coordinates. Fixed presets remain
available.
- Viewport ownership is process-local and resets on restart. The
lifecycle and accounting records remain in the database.
- The original intermittent embedded-viewer stall has not been fully
diagnosed. A bounded reconnect and active-session recovery cover the
observed failure paths.
- Live tests did not cover every revocation, approval, rate-limit, or
restart case. Deterministic integration tests cover those paths. Passkey
forwarding is not claimed.
- Unknown create outcomes keep the credential available for cleanup.
Run-list absence cannot prove a paid POST was rejected, so recovery
stays pending until it can identify provider work.

## Model Used

OpenAI Codex, GPT-6. Used reasoning, repository search, code execution,
browser interaction, and test tools. The exact serving model ID and
context-window size were not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 07:13:16 -05:00
DottaandPaperclip 2f6fa3b6dc fix: recover provider authentication inside tasks (#14629)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents need a working model provider connection to run a task.
> - A provider can reject a stored credential after the task starts.
> - The failed run must ask the responsible user to repair that
connection.
> - This pull request adds that request directly to the task and reuses
Connections sign-in.
> - The user can choose an API key or subscription, then continue the
task with a fresh session.

## Linked Issues or Issue Description

**What happened?**

A run that ended with `acpx_auth_required` or another known provider
authentication error did not immediately offer an inline way to connect
the provider. A repair form could also lock the user to the failed
account's sign-in method.

**Expected behavior**

Show a provider connection card in the task as soon as the
authentication failure is saved. Allow the responsible user to connect
or repair the provider with any supported sign-in method. Keep the
connection name automatic and resume the task after successful setup.

**Steps to reproduce**

1. Run a task with a supported provider and an expired or invalid
credential.
2. Let the run fail with a provider authentication error.
3. Open the task and attempt to repair the connection.

Related work: Refs #13724 and #13726. This change adds the inline task
repair flow and method choice.

## What Changed

- Classify provider authentication failures and create one connection
request for the current task. A persisted blocked classification
suppresses automatic retries only after the repair card is created;
unsupported providers retain their existing recovery path.
- Mark only the attributed, unchanged credential as needing sign-in.
Preserve credentials that were refreshed after the failed run started.
- Reuse the provider sign-in controls inside the task. Allow API key and
subscription choices for Claude, Codex, and Grok. Keep names hidden and
generate a default from the user, provider, and method.
- Keep the existing account when reconnecting with the same method.
Create and select another account when the method changes. Validate
updates to explicit agent bindings through the normal agent save path.
- Require explicit adoption for legacy agent authentication. Validate in
the agent environment, then commit the binding, connection install,
audit, and card completion in one transaction. Keep failed setup and
account selection visible and retryable.
- Add regression tests and update the specification and Connections
documentation.

## Verification

- Fresh local verification: 199 tests passed across the inline form,
provider method selector, default naming, authentication and recovery
classifiers, run liveness, OpenAPI routes, database adoption/rollback,
and Cursor execution suites. The adoption database suite also passed
against disposable Docker PostgreSQL.
- Full repository `pnpm build` and `pnpm -r typecheck` passed on the
latest commit. Token gates are clean.
- Embedded browser: opened real task cards from seeded authentication
failures; switched Claude from API key to subscription and back;
switched Codex from subscription to API key; confirmed the name field
stays hidden. Provider sign-in was not completed with real credentials.
- The broad local `pnpm test:run` started before review fixes and was
interrupted after the working tree changed; it is not counted as a
passing full run. Fresh focused tests passed. CI supplies the full test
and browser suite results for the current commit.
- CI is green on commit `4b97a4e447045ff3d7516525a187a5d1d21e0d4c`: 54
checks passed and two Storybook checks were skipped by their path rules.
The workspace preview job passed on one rerun after a local-server
startup timeout; its rerun passed 835 tests.
- Greptile is 5/5 on the same commit with no actionable findings and no
unresolved review threads.

## Risks

- Incorrect authentication classification could prompt for a connection
unnecessarily. Tests exclude tool authorization, quota, and unrelated
runtime failures.
- A method change selects the new personal provider default, which also
applies to other agents that use that user's default. Explicit account
bindings use the existing permission and runtime validation path.
- Credential invalidation must not race with refresh or reconnect. The
code compares the saved credential generation and grant update time
under locks.
- No database migration or new credential storage format is required.

## Model Used

OpenAI GPT-6 through Codex. The exact model ID and context window size
were not exposed in this session. Capabilities used: reasoning,
repository editing, shell commands, database tests, and embedded-browser
interaction.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused tests listed
above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 06:39:37 -05:00
Devin Foley f38b5693f6 fix: always enable keyboard shortcuts (#14643)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The web UI has keyboard shortcuts for the inbox, task lists, cases,
and task detail, plus global shortcuts such as `c`, `/`, `?`, `[`, and
`]`
> - Shortcut enablement was an instance-wide General setting until
#14141 moved it to a per-user preference that defaults to off
> - The move did not carry the old instance value over, so every
existing user lost shortcuts on upgrade and had to find a new toggle
under Profile settings
> - A toggle that only turns off a standard, input-safe feature costs a
setting, a database column, two API routes, and a React context for
little benefit
> - This pull request removes both the instance setting and the personal
preference and enables keyboard shortcuts for every signed-in user
> - The benefit is one less thing to configure, no silent loss of
shortcuts on upgrade, and less code to maintain

## Linked Issues or Issue Description

Refs #14141 (the change that introduced the personal preference).

**What existing behavior does this improve?**

Keyboard shortcuts in the web UI stay off unless each user turns them on
in Profile settings.

**Subsystem affected**

Web UI shortcuts, Profile settings, instance general settings, the
`/api/auth/preferences` routes, and the `user` table.

**Current behavior**

Shortcuts default to off per user. #14141 moved the toggle from Instance
settings → General to Profile settings and did not carry the old
instance value over. Users who had shortcuts on lost them after the
upgrade and had to find the new toggle.

**Proposed behavior**

Keyboard shortcuts are always enabled for every signed-in user. There is
no instance setting and no personal preference. Shortcuts already ignore
key presses inside text inputs and modal dialogs, so an opt-out is not
needed.

**Reason and benefit**

Fewer settings, no silent loss of shortcuts on upgrade, and removal of a
database column, two API routes, a query hook, and a React context that
existed only to gate this feature.

**Breaking changes**

`GET` and `PATCH /api/auth/preferences` are removed. `PATCH
/api/instance/settings/general` no longer accepts `keyboardShortcuts`;
that schema is strict, so the key now returns 400.
`instance.general.keyboardShortcuts` is no longer a valid
`PAPERCLIP_HIDDEN_SETTINGS` key; the parser ignores unknown keys with a
warning.

## What Changed

- Removed the Keyboard shortcuts section from Profile settings, the
`useUserPreferences` hook, `queryKeys.auth.preferences`, and
`authApi.getPreferences` / `authApi.updatePreferences`.
- Removed `GeneralSettingsContext`. The inbox, legacy inbox, task list,
legacy task list, cases, and task detail pages no longer gate their key
handlers.
- Removed the `enabled` option from `useKeyboardShortcuts`. The app
shell always registers the global shortcuts.
- Removed `GET` and `PATCH /api/auth/preferences`, their OpenAPI
entries, and the `currentUserPreferencesSchema` /
`updateCurrentUserPreferencesSchema` validators.
- Removed `keyboardShortcuts` from `InstanceGeneralSettings`, the
general settings zod schema, the settings service defaults, and
`HIDEABLE_GENERAL_SECTIONS`.
- Added migration `0289_drop_user_keyboard_shortcuts`, which drops
`user.keyboard_shortcuts`.
- Updated `AGENTS.md`, `doc/SPEC.md`, `doc/SPEC-implementation.md`, and
`docs/deploy/environment-variables.md`.
- Parsed the stored general settings row with
`instanceGeneralSettingsSchema.strip()` in the feedback vote path, so a
retired key left in the row cannot reset the sharing preference to
`prompt` and overwrite the stored choice.
- Kept every bare global shortcut (`c`, `?`, `[`, `]`, `/`) out of open
modal dialogs in `useKeyboardShortcuts`; only `/` had that guard before.
- Updated the affected tests and added a Profile settings test that
asserts the toggle is gone, a hook test for the modal dialog guard, and
a feedback service regression test for the retired-key case.

## Verification

- Typecheck passes for `@paperclipai/shared`, `@paperclipai/db`
(including the migration numbering and safety checks),
`@paperclipai/server`, and `ui`.
- `pnpm exec vitest run
server/src/__tests__/instance-settings-routes.test.ts
server/src/__tests__/openapi-routes.test.ts
server/src/__tests__/auth-routes.test.ts
server/src/__tests__/sentry.test.ts` → 119 passed.
- `pnpm exec vitest run ui/src/components/Layout.test.tsx
ui/src/pages/ProfileSettings.test.tsx ui/src/pages/IssueDetail.test.tsx
ui/src/pages/Inbox.test.tsx ui/src/pages/Cases.test.tsx
ui/src/hooks/useKeyboardShortcuts.test.tsx ui/src/pages/Agents.test.tsx
ui/src/pages/InstanceGeneralSettings.test.tsx` → 286 passed.
- `pnpm exec vitest run packages/shared/src/settings-visibility.test.ts`
→ 16 passed.
- `pnpm exec vitest run ui/src/hooks/useKeyboardShortcuts.test.tsx` → 7
passed.
- `pnpm exec vitest run server/src/__tests__/feedback-service.test.ts`
(embedded Postgres) → the new retired-key test passes with the fix and
fails without it.
- Manual: sign in with no settings changed, open the inbox, press `j`
and `k` to move the selection, press `?` to open the cheatsheet. Open
Settings → Profile and confirm there is no Keyboard shortcuts section.

## Risks

- The migration drops a column. It uses `DROP COLUMN IF EXISTS`, and the
column has no readers after this change. If you roll back to a build
from before this PR after the migration has run, re-add the column
first: `ALTER TABLE "user" ADD COLUMN "keyboard_shortcuts" boolean
DEFAULT false NOT NULL;`. The older build's ORM selects that column when
it loads users.
- Any external client that still sends `keyboardShortcuts` to `PATCH
/api/instance/settings/general` receives a 400. No in-repo client does.
- Stored `instance_settings.general.keyboardShortcuts` values are
stripped on read and ignored.
- Users who never turned the toggle on now get shortcuts. The handlers
skip text inputs, contenteditable regions, and modal dialogs, so typing
is unaffected.

## Model Used

Claude Fable 5.1 (`claude-fable-5-1`) in Claude Code, with extended
thinking and tool use.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-29 21:28:06 -07:00
DottaandPaperclip 61b3fd57a6 fix(ui): recover gracefully during server restarts (#14560)
Show a clear reconnecting state during server restarts and retry safe access
checks every five seconds. Preserve open drafts, wait for initial startup
readiness, and keep authentication failures separate from temporary outages.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-29 08:31:05 -05:00
DottaandFry 3ca196b0a6 feat(agents): persist agent files across tasks without revision history (#14420)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - An agent needs personal files across tasks and sessions.
> - AGENTS.md is one file in that directory. Supporting files need the
same persistence.
> - The Instructions Editor and agent runs must share one current
directory.
> - Concurrent runs should apply only the files they change. The last
sync of the same file wins.
> - This pull request uses existing file transport and removes temporary
copies after sync.
> - Old instruction-only sessions keep their restore contract. New saves
do not create revision history.

## Linked Issues or Issue Description

Refs #14325. This replaces its revision-oriented design with persistent
agent files. Keep #14325 unmerged.

Transport prerequisite #14416 merged first at
`d172197117a14b80a1eb2d2835a0e7cce2679656`. This PR now targets master
and remains below 100 changed files.

Related work: #4513 and #8798 cover instruction tooling. This change
handles run synchronization, cross-task personal files, browser editing,
and old-session restoration.

## What Changed

- Keep one current directory per company and agent. Point AGENT_HOME at
a temporary working copy for each active run. Keep task files and
provider HOME separate.
- Restore text, binary files, and nested folders through workspace
transport. Exclude remote agent files from task Git snapshots with a
self-ignoring file inside the reserved runtime directory; never write
through repository-controlled Git metadata.
- Collect after the provider and child processes have stopped. Keep
resumable conversation state.
- Apply changed and deleted files under the agent lock. The last sync
wins for the same file. Unrelated concurrent changes survive.
- Remove temporary copies after successful sync, rejected sync, and
staging failure. Register ownership before copying so restart recovery
can remove interrupted preparation. Retry transient synchronization up
to three times. Preserve the original remote lease reference until
deletion succeeds; restart cleanup never acquires a replacement sandbox.
Do not create captured directories or a conflict-review queue for new
runs.
- Keep browser editing, stale-draft protection, and streaming binary
downloads. Keep the instruction entry and text editor limited to 1 MiB.
- Keep historical agent-folder sync failures on their affected runs
instead of repeating them above current saved instructions. Preserve
legacy candidate review and current browser-save errors. Avoid duplicate
quota warnings while retaining separate sync failures when they describe
a different problem.
- Require target-scoped caller grants for peer instruction access, while
preserving self edits, responsible-user checks, and protected-change
consent.
- Treat full storage as a nonblocking run warning, never an agent pause
or run-admission failure. Restore already-over-quota saved folders so
ordinary agent cleanup can recover; warn on each run until cleanup. The
run detail view shows the warning.
- Allow 256 MiB per file, 2 GiB per directory, and 100,000 entries. Hash
large files as streams. Check editor-save quotas with metadata instead
of hashing unrelated files.
- Preserve old native inputs, instruction-only copies, paths, digests,
and pending legacy candidates. Adopt old revision heads once. New writes
do not append history rows.
- Add idempotent migration 0287 and verify upgrades from the preview
tables and receipts.
- Add nine interactive stories under **Agents / Persistent files**,
including automatic incoming edits, stale browser drafts, and
storage-limit diagnostics.

## Verification

- Merge candidate: `4f5390107ec6ffd80a76d1d2e85530e66f21d079`, after
merging current master and the landed transport prerequisite.
Integration required no manual conflict resolution; the feature remains
99 changed files. Full workspace typecheck, production build, token
gates, and 715 focused tests passed on this merge candidate. Fresh
Greptile review is 5/5 with no unresolved findings. All 55 checks
passed, with four conditional skips, including the build, typecheck,
browser E2E, and canary dry run. A single retry recovered four jobs
interrupted by runner shutdowns; no source changes were required.
- Historical-warning UI fix: all 6,834 UI tests across 640 files passed,
including regression coverage for three old failures, legacy preserved
edits, and warnings scoped to the affected run. Full workspace
typecheck, production build, Storybook build, and token gates passed.
Browser-verified Storybook playtests passed for Historical Failures
After Successful Save, Storage Limit, and Full Storage Run Warning.
- Review follow-ups at `4e20c9fb2`: all 18 focused tests passed,
including external Git directories, linked worktrees, symlinks,
hardlinks, and distinct I/O failures alongside storage warnings. Server
and UI typechecks, token gates, and the production build passed.
- Storage warning regressions at `0724f3012`: all 33 directory tests and
all five heartbeat-list tests passed, with no skips in their successful
runs. They cover repeated runs while full, an already-over-quota saved
folder, cleanup, warnings retained after unrelated save failures, and
bounded warnings in large result JSON. Server typecheck passed after the
final warning fixes.
- Full workspace typecheck, production build, and token gates passed
during this follow-up. Product E2E harness: 631 tests passed across 52
files; harness typecheck passed. Earlier native session/context and
directory/legacy collection suites passed 537 tests; Runner
unit/transport suites passed 329 tests.
- **Real E2E at `0724f3012` (before this follow-up):** legacy local
Codex and native Daytona Codex each passed six tasks, one server
restart, seven independent assertions, and cleanup verification. Both
prove browser-to-agent edits, agent-to-browser edits, nested/binary
restoration, per-file last-sync-wins, a successful run after an
oversized save rejection, and cleanup clearing the warning.
- Native local Codex also passed the six-task quota flow before the
final warning-retention fixes. That pass began at `918d1ed02` while the
bounded-result warning fix was being edited, so it is not claimed as
exact-final-head evidence. Its final-head rerun failed during embedded
PostgreSQL bootstrap before any provider run: the macOS host had 87,365
of 87,381 SysV semaphores occupied. No unrelated services or kernel
limits were changed.
- The final-source report intentionally records **2/3 cells passed**,
preserving the blocked native-local attempt:
`tests/runner-e2e/results/agent-files-quota-final-20260928-report/`.
Earlier failed attempts and provenance notes remain under
`tests/runner-e2e/results/agent-files-quota-final-20260928-input/` and
the original campaign directories.
- Daytona used immutable image
`ghcr.io/paperclipai/paperclip-daytona-runner@sha256:5643f0d801417cae3581833a1a3bc6715b325e028602738d2652c44cac5dc6bf`
and its exact Linux runner binary. Controller source is `0724f3012`;
image source is recorded separately.
- Legacy-session compatibility and all three ACP Stop/resume browser
regressions passed on the prior validated feature head
`169fab46d5af21caa2269b4c1b29b69c933a6951`. They assert the same
provider session is retained and interrupted writes are not replayed.
Migration upgrade tests also passed earlier.
- Nine interactive stories are under **Agents / Persistent files**,
including **Full Storage Run Warning**. Its playtest and visual browser
inspection passed; the warning states that runs continue and the editor
remains available.
- Prior-head checks on `4e20c9fb2`: 55 passed, two conditional jobs
skipped, no failures or pending checks. All eight browser E2E shards and
their aggregate passed. Fresh Greptile review is 5/5 with no findings;
all review threads are resolved, the security scan passed, and GitHub
reports no merge conflicts.
- The broad local follow-up test run was interrupted after host
semaphore exhaustion affected isolated PostgreSQL instances. It also
encountered the existing macOS long-path fixture failure and two timeout
failures. This is not a claim that the full local suite passed. Logs are
retained; focused storage/warning tests passed.

## Risks

- A later sync can overwrite an earlier edit to the same file, including
a saved browser edit. There is no text merge or retained version. This
is the intended last-sync-wins policy.
- A save that exceeds a storage limit is rejected and its temporary copy
is discarded. The run itself continues normally, and later runs restore
the last saved files with a warning until cleanup. Transient sync
failures get bounded retries. An I/O failure partway through a sync can
leave some files updated; a failed receipt does not claim whole-folder
success.
- Larger folders increase copy time, network traffic, and temporary disk
usage. Active runs still need working copies. Terminal runs do not
accumulate archives. Operators must provision disk for agents and
configured concurrency; these limits are not company-wide quotas.
- A restored old native session remains instruction-only until a fresh
session starts. Its original conflict fence and existing pending
candidates remain compatible.
- Provider processes close at the collection boundary. Conversation
resume remains available, but warm process reuse is lost.
- Backups must include the instance filesystem and database. External
bundles keep their existing behavior until explicitly moved to managed
storage.

## Model Used

OpenAI Codex, GPT-6 family. The session does not expose a more specific
model ID or context-window size. Reasoning, code execution, and browser
tools assisted this change. Real provider E2E uses `gpt-5.6-sol`.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Fry (Paperclip) <noreply@paperclip.ing>
2026-09-29 08:25:56 -05:00
DottaandPaperclip 270afd2fb8 feat(ui): show running commit in staging account menu (#14410)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The account menu shows the signed-in user's identity.
> - Staging users need to know which server commit is running after a
deploy.
> - The health endpoint already returns that commit, but the menu does
not show it.
> - This pull request adds the short commit below the email on staging
hosts.
> - Users can open the menu to check a deploy without opening deployment
tools.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

The account menu on staging instances.

**Current behavior**

The menu shows the user's name and email. It does not show the running
server commit.

**Proposed behavior**

On `*.staging.paperclip.app`, show `SHA 8751e2d` below the email. Use
the current `/api/health` commit. Show the full SHA on hover and link to
the commit on GitHub. Refresh the health query when the menu opens. Hide
the label on other hosts and when commit metadata is unavailable.

**Reason and benefit**

A user can confirm which commit a staging instance runs after an
automatic deploy.

**Breaking changes**

None. The server already returns the commit field.

Refs #14060 for related account-menu work. This change adds deployment
information only.

## What Changed

- Add the existing health response commit field to the UI type.
- Share the staging host check and a separate health query across both
account-menu variants. A menu refresh failure leaves the access gate
health state unchanged.
- Show a short SHA below the email. Link to the full commit on GitHub
and include the full SHA in its accessible name and hover title.
- Document the staging label and cover staging hosts, other hosts,
missing metadata, refresh on reopen, and request failure isolation.

## Verification

- `pnpm exec vitest run ui/src/components/SidebarAccountMenu.test.tsx` —
24 tests passed.
- `pnpm check:token-gates` — passed.
- `pnpm -r typecheck` — passed.
- `pnpm build` — passed. The UI build and typecheck also passed again
after review fixes.
- Full test matrix — passed on the latest commit in
[CI](https://github.com/paperclipai/paperclip/actions/runs/36447810742).
The local `pnpm test:run` was stopped before completion after CI
finished the same suites. The focused local tests, full local typecheck,
and full local build passed.
- Browser shard 6 passed on one rerun. Its first attempt lost part of
the draft text in the existing attachment-receipt reload test. No code
changed for the rerun.
- Rendered the real account menu in a local browser fixture with a
staging hostname condition and mocked health response. Confirmed the SHA
fits below the email in the dark menu.
- Greptile review — 5/5, all review threads resolved.
- Manual check after deployment: open the menu on a staging host and
compare the SHA with `/api/health`. Open the menu again after a deploy
to refresh it. Confirm the label is absent on production and localhost.

## Risks

- Low risk. Each menu opening on staging can make one additional health
request.
- The host check applies to `*.staging.paperclip.app`. Other staging
domains will need an explicit update.
- The label identifies the running server commit. It can briefly show
cached data while the request completes.

## Model Used

OpenAI Codex, GPT-6. The exact model variant and context window are not
exposed in this session. Used code execution, repository inspection, and
browser inspection.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#123` / `Refs #123` OR (b) described the issue in-PR following the
relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 13:43:30 -05:00
DottaandPaperclip 14795136f5 fix(runtime): finalize and recover sandbox workspace exports safely (#14402)
Serialize native workspace finalization, validate streamed archives within bounded limits, and quietly recover unsafe exports from saved results. Preserve exact allocations for exhausted transient failures and provide export-only retry without rerunning the provider.

Consolidates #14314, #14315, #14329, and #14334 while preserving the already-merged finalization label changes.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-28 10:57:32 -05:00
DottaandCodie 01d9a12185 fix: make keyboard shortcut enablement a personal preference (#14141)
Store keyboard shortcut enablement per user and expose it in Profile settings.

Co-Authored-By: Codie <Codie@users.noreply.github.com>
2026-09-26 11:50:04 -05:00
Devin FoleyandPaperclip e4237c45f3 fix(ui): prevent organization title flicker during plugin loading (#13854)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The sidebar identifies the current organization.
> - An optional plugin can replace this navigation surface.
> - The built-in title appears before discovery and module loading
finish.
> - This pull request reserves the trigger until its owner is known.
> - The organization name appears once, while failures retain built-in
navigation.

## Linked Issues or Issue Description

Refs #13832. Searched related pull requests and issues; no duplicate fix
found.

**What happened?**
The organization switcher renders a provisional built-in title before an
installed replacement loads. Unrelated plugin imports can also affect
its loading state.

**Expected behavior**
Reserve the trigger with a neutral placeholder, then show the resolved
navigation surface. Keep the built-in menu on failed or absent
contributions.

**Steps to reproduce**
Install an organization-switcher contribution. Delay session, company,
contribution, and module responses. Reload the page and watch the
trigger through each stage.

## What Changed

- Reserve the trigger through account, company selection, slot
discovery, and module loading.
- Distinguish failed session lookup from pending lookup so errors retain
usable navigation.
- Load and await only contributions matching the requested slots.
Observe completion of imports started by another consumer.
- Document loading behavior and add regression coverage for loading,
failures, unrelated modules, and identity transitions.

## Verification

- `pnpm -r typecheck` passed, including Rust checks.
- `pnpm build` passed.
- All 629 UI test files passed: 6,593 tests. The 42 focused
UI/API/plugin tests also passed.
- `pnpm check:token-gates` and `git diff --check` passed.
- `pnpm test:run` was also attempted. The broad local server run was
stopped after recording skill-cache/channel fixture failures outside
this diff (for example, runtime skill source status `missing` instead of
`available`). The original cause is not established. All latest-head
Linux CI gates pass; the complete UI suite and affected local checks
pass.
- Desktop (1440px) and mobile (390px) Chromium checks passed with real
host components, dynamic module loading, and the built Account bundle.
Delayed fixture responses produced exactly two title states: empty
placeholder, then the resolved name. A slow refresh preserved the title
and trigger dimensions; absent/failed plugin fallback and Escape
dismissal passed, with zero uncaught browser errors. This is browser
component integration, not a live signed-in tenant test.

## Risks

A cold load displays a neutral placeholder until discovery completes.
Absent, ambiguous, failed, and invalid contributions still use the
built-in menu. No migrations or authorization changes. Scoped module
loading changes when an unrelated contribution is imported; each surface
loads its own matching modules.

## Model Used

OpenAI Codex, GPT-6, with reasoning, code execution, and browser
verification. The exact deployment ID and context window are not exposed
in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-22 20:33:29 -07:00
DottaandPaperclip a10702a878 feat(slack): add governed tools for Slack-origin tasks (#13828)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Chat connectors let people start and continue agent tasks from other
services.
> - A Slack conversation needs access to its surrounding discussion and
Slack collaboration tools.
> - The agent must use the linked requester's access and keep private
material within its permitted audience.
> - This pull request adds Slack tools through the existing connector
contribution and approval framework.
> - People can ask an invited bot to read a discussion, create follow-up
tasks, and collaborate in Slack.

## Linked Issues or Issue Description

**Subsystem affected**

Chat connectors, connector runtime, tool gateway, and connection
Settings/Access.

**Problem or motivation**

Slack-origin tasks can receive messages but cannot inspect the rest of a
channel or act through the originating bot. People must paste context or
configure a separate integration.

**Proposed solution**

Supply typed Slack tools and a bundled skill only to the originating
task and assigned agent. Resolve the linked requester on the server.
Check bot and requester access before reads and writes. Use existing
durable actions and approvals. Retrieved messages remain source
material.

**Alternatives considered**

Slack's user-OAuth MCP server does not replace the customer-created chat
bot. An unrestricted Web API proxy would not provide suitable permission
or publication boundaries.

**Roadmap alignment**

This extends the existing MCP Tool Gateway & Apps work with a provider
contribution. It does not add a task dispatcher or a separate Slack task
lifecycle. Related: #11144 covers generic per-user MCP grant execution;
this change binds Slack bot operations to chat-origin tasks.

## What Changed

- Add 39 typed Slack tools, a method/scope matrix, a bundled skill, and
shared native/HTTP execution.
- Bind tools to company, endpoint, task, run, assigned agent, and
admitted linked requester. Check membership and revocation on each call
and before queued writes.
- Add paginated reads, bounded history search, source links, messages,
file uploads, reactions, pins, bookmarks, topics, canvases, lists, and
approved channel operations.
- Restrict private-source publication, including automatic replies and
uploaded deliverables. Keep other people's bot DMs inaccessible.
- Reuse action receipts, idempotency, approvals, and reconciliation.
Suppress an identical explicit-send/final-reply duplicate. Return
governed results through their verified originating conversation.
- Add endpoint-bound personal search OAuth storage and lifecycle. Keep
native real-time search disabled until a runtime meets Slack's
transient-result requirements. Current runtimes use bounded history
search.
- Show capabilities, scope upgrades, and personal search authorization
in Settings/Access and Storybook. Document provider and runtime limits.

## Verification

- Current head `0eb21cba4`: CI checks pass and Greptile is 5/5 with no
unresolved findings. One unchanged rapid-callback timing test passed on
a single CI retry.
- Approval presentation regressions cover board-comment precedence and
exact Slack publication; the expanded database assertion passed in CI.
The local PostgreSQL startup probe later became unavailable, so that
final assertion was verified in CI. Slack setup and failed-run retry
browser tests also passed locally.

- Full workspace typecheck and build passed. Server typecheck/build
passed again after the approval routing fix.
- Broad local suites passed in separate groups: server 12,958 tests, UI
6,555, shared 770, skills catalog 20, and other workspace packages
2,652. CLI and serialized server checks passed after environment/timeout
retries. These are composite results, not one uninterrupted green
full-suite invocation.
- PostgreSQL authority regression covers admitted identity,
cross-company/task/agent rejection, recovery, retained-session
revocation, OAuth refresh/disconnect races, approval execution, exact
publication lineage, retries, uncertain sends, and duplicate
suppression.
- Gateway/response regressions cover separate-origin approval batches
and durable continuation. Focused provider, access, search, native
runtime, route, and AgentMail regressions pass.
- Storybook capability, missing-scope, OAuth configuration,
authorization, and disconnect states were inspected in the browser.
- Live staging: read a channel decision and full thread, create exactly
two assigned backlog tasks, add a reaction, paginate discovery to
exhaustion, and return bounded search matches with source links and
coverage.
- Live staging: create/edit/read a canvas and list, inspect the canvas
in Slack, post/edit one message, and create a channel only after
approval. New channels remain disabled for responses.
- Live staging: read a response-disabled channel from the requester's
DM; writes to that channel were denied. The test setting was restored.
- Final live retest passed: explicit file upload and exact content
read-back; approved deletion of only the disposable bot message;
continuation confirmation returned to the original Slack thread without
repeating the action.
- Optional OAuth, private multi-user boundaries, native RTS, and CLI
provider execution are not fully live-qualified. The staging agent
initially supplied malformed tool arguments; valid arguments succeeded,
and the tool/skill descriptions now emphasize UUID write keys.

## Risks

- Existing Slack apps must add scopes and reinstall for new
capabilities. Provider plans and document permissions can still restrict
operations.
- Instances need an independent `PAPERCLIP_TOOL_ACTION_SIGNING_SECRET`
for governed tool actions. The staging instance was configured with
explicit operator approval; fleet provisioning is a separate gap.
- Native RTS is not exposed on current transcript-retaining runtimes.
Bounded history scans are deliberately reported as incomplete. Inline
file reads support text/canvas content up to 256 KiB; other types return
metadata.
- Private document edits fail closed when the full audience cannot be
verified. Uncertain effects other than posts/uploads require inspection
instead of blind retries.
- Shared approval-delivery code now separates outcomes by source run to
preserve origin boundaries. No database migration is required.
- A separate completion-validator gap remains when the agent cites a
prior run's registered artifact during finalization. It asked for
registration again even though Slack delivery was confirmed. This change
does not add a connector-specific task-completion policy.

## Model Used

OpenAI GPT-6 through Codex, with repository tools, code execution, and
browser testing. The exact deployed model identifier and context-window
size were not exposed in the session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-22 15:29:07 -05:00
Devin FoleyandPaperclip df66219780 fix(ui): retry the latest failed task attempt (#13765)
## Thinking Path

> - Paperclip manages work done by AI agents.
> - The task thread lets an operator retry a failed run.
> - Legacy runs with transcript output did not get a failure marker.
> - The thread could therefore offer Try again for an older failure.
> - The server correctly reused that failure's existing retry, even when
it had already failed.
> - This change keeps the latest failure actionable and reports stopped
retry responses to the operator.

## Linked Issues or Issue Description

**What happened?**

Try again could return success without starting work. An initial setup
failure had an empty transcript. Its later retry produced output and
failed. Only the initial failure had a retry marker, so the button kept
requesting the initial failure's already-failed successor.

**Expected behavior**

Try again targets the latest failed attempt. An already-stopped retry
response shows an error and refreshes the task's run state.

**Steps to reproduce**

1. Start a legacy adapter task that fails before producing transcript
output.
2. Retry it. Let this attempt produce output and a final failure comment
before it fails.
3. Click Try again in the task thread.
4. Before this fix, the click targets the original failure and replays
the stopped successor.

**Paperclip version or commit**

Reproduced against `1483bb8bcf`; the regression is also present on the
branch base `8813a50105`.

**Deployment mode**

Authenticated server with a legacy adapter. The bug is in the shared
task UI and retry API client.

Related: #11650 adds a different recovery-notice action. This change
fixes failed-run markers and retry response handling. Searches found no
duplicate of this failure case.

## What Changed

- Render legacy failure markers even when the run has a transcript or
final comment.
- Keep later cancelled automatic retries from replacing the failed run's
retry action.
- Reject already-stopped retry responses in the API client so existing
error feedback appears.
- Refresh task run queries after both successful and failed retry
requests.
- Add regression coverage for failed and timed-out attempts, execution
gates with output, and retry response states.

## Verification

- Before the fix, the new regression tests failed: two selected the
original failure, and four accepted a stopped successor as success.
- Targeted task-thread, retry API, marker, and issue-page tests: 279
passed.
- `pnpm --filter @paperclipai/ui typecheck`: passed.
- `pnpm --filter @paperclipai/ui build`: passed.
- `pnpm check:token-gates`: passed.
- Full UI suite: 6,529 tests passed across 626 files.
- [CI run
35650385023](https://github.com/paperclipai/paperclip/actions/runs/35650385023):
all 53 checks passed, including full workspace build, typecheck,
unit/integration suites, and browser tests. The redundant local
full-workspace test run was stopped after CI passed; it is not counted
as a completed local pass.
- Greptile: 5/5 on `608ee58c99`, with no review threads or unresolved
comments. The branch is mergeable.
- `pnpm -r typecheck` and `pnpm build` were attempted. Both stop at the
Runner's Rust checks because this host has no `cargo`. The full
workspace checks passed in CI.

## Risks

Low risk. This changes UI presentation and response handling only. The
server's exact-retry idempotency, authorization, execution ownership,
and recovery gates remain in place. No schema changes or live task
mutations. Existing documentation describes this retry action; the fix
restores that behavior.

## Model Used

OpenAI GPT-6 (Codex), with reasoning, repository tools, and code
execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (targeted checks;
full-workspace limits described above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
(existing behavior restored; no documentation change needed)
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 13:31:43 -07:00
DottaandPaperclip 8813a50105 feat: run GitHub review bots through Paperclip agents (#13717)
## Thinking Path

> - Paperclip manages agent work as tasks and runs.
> - GitHub chat brings repository conversations into those tasks.
> - A review bot needs the assigned agent, its authority, and governed
provider tools.
> - The existing channel connection did not supply that review workflow
or a complete setup journey.
> - This pull request adds GitHub App setup, account access, event
prompts, task-bound review tools, and exact-commit checks.
> - Operators can inspect each review through the same task, run, and
activity systems.

## Linked Issues or Issue Description

**Subsystem affected**

GitHub chat, governed connection tools, task execution, shared/database
contracts, and connector setup UI.

**Problem or motivation**

Operators need a GitHub review bot that runs their assigned Paperclip
agent. Mentions and PR events must preserve task ownership and requester
authority. Provider publication must use the bot App identity and
enforce the configured permissions.

**Proposed solution**

Extend the existing GitHub chat connector with resumable App onboarding,
linked-member and sponsored-guest access, editable event prompts, and
governed review operations. Validate structured assessments on the
server and compute a stable Paperclip Review check for the exact head
commit.

**Alternatives considered**

A separate review scheduler would duplicate Paperclip execution and
permissions. Reusing personal GitHub credentials would change the bot
identity and credential boundary.

**Roadmap alignment**

This extends the existing Connected Apps and governed-tool
infrastructure. The project owner requested and approved this design.
Related PR #8645 imports external Codex review feedback; this change
runs an assigned Paperclip agent and publishes its results through the
existing chat connector.

## What Changed

- Include the current Paperclip instance origin in the copied setup
prompt. Storybook uses its configured Paperclip origin; callback
parameters and URL credentials are excluded.

- Add a Claude/Codex copy button in the real setup and Storybook opening
step. Its detailed prompt asks four setup questions and guides
embedded-browser setup, verification, and optional required checks.
Clipboard failure exposes selectable instructions.
- Add a tutorial that explains why App installation, review scheduling,
and required checks are separate choices.

- Add manifest registration, an existing-App path, separate installation
and repository selection, repository refresh, and explicit account
confirmation.
- Add low-trust agent guidance, effective capability verification,
member selection, and explicit restricted guests with a sponsor.
- Add configurable PR events, prompts, repository overrides, rating
thresholds, and separate formal-review permissions.
- Give the assigned agent governed App tools to read PRs, comment, begin
an assessment, submit findings, and optionally submit a formal review.
- Bind review history, root PR events, and inline replies to ordinary
tasks. Deduplicate deliveries/findings and reject stale publication.
- Link check Details to the underlying task on the current trusted
hostname, or to Reviews before task creation.
- Add schema migration 0283, API contracts, production UI, and 49
interactive Storybook states.
- Repair local lease recovery. Keep the Cloud Dockerfile identical to
master; no provider-pack layer or runtime-default environment variable
is added.
- Retry only rolled-back wake-admission transactions after transient
endpoint-lock contention. A deterministic held-lock regression proves
one accepted wake.

## Verification

- Current head: `7ba761fe007bb798400d3e62346fa964f607f0f8`, rebased on
master `d9b3a5653e41f2ee5a1345b97c86a238f7a5c8e9`. Dockerfile has zero
diff against master. Final workspace typecheck and build passed. The new
PostgreSQL migration regression passed and preserves existing relation
and constraint identities after replay.
- Greptile reviewed this exact head at 5/5. There are zero unresolved
review threads and no merge conflicts.
- All current-head checks are green: 54 passed and two conditional
Storybook jobs skipped. This includes complete server/workspace test
suites, build, typechecks, policy checks, Runner suites, browser suites,
and security status. One timing-sensitive callback-ordering test passed
in isolation and its CI shard passed one retry. The duplicate local
full-suite run was stopped after CI completed; it is not counted as a
local full-suite pass.
- Before the final Slack rebase and migration renumbering, 186 focused
GitHub tests, 14 native bootstrap cases, token gates, and Storybook
build passed. The final rebase retained the new Slack communication
guidance.
- The embedded-browser setup test copied the full detailed prompt,
including the configured Paperclip instance URL. Desktop and narrow
layouts were checked. Component tests cover successful copying and
clipboard failure with selectable text and retry.
- Live local and hosted GitHub acceptance evidence refers to application
revision `cb703ac959876a07ebf3d7a295847f9f351eb6fc`. Real agent tasks
exercised issue mentions, automatic PR reviews, inline findings,
repeated mentions, task continuation, and failing-to-passing checks
after a push. The Storybook agent generated, built, and browser-rendered
pages; missing acceptance text failed, matching text passed, and broken
JSX produced an incomplete result.
- Live cases also covered independently disabled push events, prompt
injection, duplicate signed deliveries, rapid pushes, stale-result
rejection, finding deduplication, and restart recovery. Formal reviews
were denied while disabled and published only after explicit enablement.
Check Details links pointed to the underlying task on the trusted
hostname.
- Those hosted native Claude runs used the provider-pack layer now
removed from this PR. They do not prove native Claude works on the
standard Cloud image. A replacement hosted native Codex run is not yet
verified: the disposable QA tenant has only an Anthropic AI connection.
No new staging or production deployment was made for the packaging
removal.
- Required-check merge enforcement could not be tested because the
private disposable repository's GitHub plan rejected the rules
configuration. Published success/failure/incomplete check states were
verified directly.

## Risks

- Latest master allocated migration 0282 to Slack. The GitHub migration
is regenerated as 0283 with replay-safe table/index/constraint creation;
a PostgreSQL regression verifies existing relations and constraints are
preserved. Existing preview tenants remain subject to the fleet
migration-history compatibility preflight; no bypass is introduced.

- Migration 0283 adds company-scoped configuration, registration,
review, and publication records. Existing connections retain their
behavior until reviews/tools are enabled.
- Signed webhooks and expiring registration state remain required.
Hosted installations also need the companion narrow Cloud gateway
exemptions.
- Agent assessments can be incomplete or wrong. The server enforces
coverage/result structure, current-head publication, rating policy, and
separate formal-review permission; it does not replace code-review
judgment.
- No Cloud image packaging changes are included. Remote native
ACPX/Claude and OpenCode retain their existing operator-supplied
provider-pack prerequisite. Native Codex and Codex with managed MCP
tools do not require that pack. Earlier staging deployment evidence
refers to its stated revision, not this packaging-removal head.
Production rollout and merging remain outside this change.

## Model Used

OpenAI GPT-6 through Codex, with repository, code execution, API, and
embedded-browser tools. The exact serving model ID and context-window
size were not exposed by the environment.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 14:41:19 -05:00
DottaandPaperclip d9b3a5653e feat(chat): add initial Slack communication guidance and connection menus (#13760)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Chat connectors let people use the same tasks and agent tools from
external conversations.
> - Agents need communication guidance that fits the conversation
medium.
> - That guidance belongs in the original task context, without repeated
instructions on each turn.
> - Connection owners also need clear settings and a consistent way to
remove a connection.
> - This pull request adds initial Slack guidance, optional connection
instructions, and chat connection menus.
> - The benefit is clearer Slack replies with the existing Paperclip
workflow and permissions.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

Agent replies in Slack and chat connection management in the Apps
catalog.

**Current behavior**

Slack tasks do not carry a saved communication profile. The catalog
shows a separate Manage button and does not offer removal on every chat
connection row.

**Proposed behavior**

Save Slack guidance when a new conversation creates a task. Restore that
original guidance when a model session is rebuilt. Do not append it to
ordinary follow-ups. Expose optional additional instructions in Slack
Settings. Put Manage and Remove connection in a three-dot menu for all
chat providers. Keep Finish setup visible for drafts.

**Reason and benefit**

Small answers fit in Slack. Substantial deliverables use ordinary
document or artifact tools with a useful Slack summary. Connection
settings apply to new tasks and cannot change permissions. Users can
remove both active and unfinished chat connections from the catalog.

**Breaking changes**

Two additive database columns store endpoint preferences and the initial
conversation snapshot. Existing endpoints default to empty preferences.
Existing conversations keep their original behavior. Non-Slack guidance
is unchanged.

Related public context:
https://github.com/paperclipai/paperclip/pull/13741 improves native chat
recovery. This change adds communication context to those existing
execution paths. A search found no duplicate communication-guidance PR.

## What Changed

- Add a provider-guidance registry, enabled for Slack first.
- Persist optional endpoint communication instructions and capture an
immutable snapshot when a conversation creates a task.
- Resolve guidance from the verified company-scoped connection. Restore
it for fresh native and legacy sessions without per-turn reminders,
extra model calls, or extra context queries.
- Add the Slack Settings field, validation, audit coverage, and
Storybook save/error states.
- Add Manage and Remove connection menus for all seven chat providers.
Keep the draft setup button. Require removal confirmation and allow
retry after failure.
- Add regression coverage, an active/draft menu story, and connector
documentation.

## Verification

All CI checks are green for 5f48df4e0. Greptile scored this head 5/5
with no actionable findings. No review threads remain unresolved.

- Passed `pnpm -r typecheck` and `pnpm build` on PR head 5f48df4e0.
- Passed design-token checks, UI typecheck, and all 20 catalog tests
after rebase. Tests cover all seven providers, active/draft removal,
confirmation, cache refresh, errors, and cancellation.
- Verified the active/draft menu in Storybook. The interaction test runs
without browser console errors.
- Passed focused guidance, endpoint persistence/isolation, heartbeat
trust, native context, ACPX, adapter utility, and CLI recovery tests.
Full UI and CLI groups passed (6,512 and 502 tests).
- Tested real Slack conversations on staging: concise updates with
public links, a planning question with buttons, a saved plan, a saved
report, task creation and assignment, and explicit detailed output. Old
tasks retained original preferences after an edit; a new task used the
changed preferences. Restored the staging setting afterward.
- Existing safe progress remained visible without duplicate final
replies or private reasoning.
- Broad local tests found resource/time-sensitive failures that passed
targeted reruns. One Cursor archive-download fixture failed on both this
branch and the unchanged main checkout. The full local suite is not
claimed clean. All PR-head CI test shards passed, including general,
serialized, Runner, and browser suites. The redundant local full-suite
rerun was stopped after CI completed successfully.
- Live delegation was not tested because the staging company has only
one agent. Live testing also found separate latency and runner
task-editing capability gaps; this PR does not add connector-specific
workflow behavior to hide them.

## Risks

- Prompt guidance changes the form of new Slack replies. Explicit
requests for detail still take precedence.
- The additive migration is idempotent. Conversation snapshots remain
fixed when connection settings change.
- Native and legacy recovery must preserve the initial context without
duplicates; targeted tests cover these paths.
- Removing a connection stops new work through the existing lifecycle
action. It retains Paperclip task history and does not delete the
external app or bot.

## Model Used

OpenAI GPT-6 through Codex, with repository editing, shell tools, and
browser testing. The host does not expose a more specific model ID or
context-window size. No separate model calls were added to the product.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (targeted suites; broad
local limitations are listed above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 14:14:25 -05:00
DottaandPaperclip b82661b561 refactor(connections): retire the legacy Composio broker (#13758)
## Thinking Path

> - Paperclip manages agents and their access to external tools.
> - Connectors expose these tools through a governed MCP gateway.
> - PR #13755 added a direct Composio MCP connection behind the
experimental MCP aggregators flag.
> - The old project API-key broker still created toolkit child
connections and showed a separate Services tab.
> - Keeping both paths leaves obsolete setup and session code in the
product.
> - This change removes the broker and preserves direct MCP setup,
credentials, permissions, and execution.
> - Saved legacy records fail closed and remain available for explicit
removal.

## Linked Issues or Issue Description

Related: #13755. This retirement supersedes the legacy-path fixes
proposed in #12630, #12632, #12634, and #12906. It does not close those
PRs.

**What existing behavior does this improve?**

Composio connector setup, management, and runtime dispatch.

**Current behavior**

Composio offers both direct MCP and a project API-key broker. The broker
mints sessions and creates one child connection per toolkit.

**Proposed behavior**

Offer only direct MCP. Remove the toolkit Services UI, REST routes, API
client, and session broker. Block saved legacy parent and child records
from discovery, execution, health checks, reconnect, and OAuth. Preserve
their records and credentials until the operator removes each
connection.

**Reason and benefit**

The direct MCP connector becomes the single supported Composio workflow.
Provider accounts remain managed in Composio.

## What Changed

- Remove the API-key catalog method and its generated-source definition.
- Delete Composio broker clients, session creation, account
synchronization, child lifecycle, and toolkit routes.
- Remove the Services tab, service rows, child provenance, and
cascade-removal controls. Keep Vercel provenance intact.
- Retain a shared retirement guard for stored legacy records. Show
Retired status and replacement/removal guidance in the connection list
and details; hide obsolete runtime controls.
- Preserve the experimental MCP aggregators flag and direct MCP
infrastructure.
- Replace broker fixtures with retirement tests and extend direct
Composio catalog/reconnect coverage.

## Verification

- Focused shared, server, and UI tests passed with one worker. Server
retirement tests use a name filter; no full local test suite was run, as
requested.
- Server and UI TypeScript checks passed.
- Token gates and UI build passed.
- Real browser: opened the saved Composio connection, refreshed all 11
tools, and ran the provider's read-only GitHub account-list operation
through the standard Test dialog as an agent. The provider returned
success using the existing OAuth credentials.
- See `doc/connections/COMPOSIO-BROKER-RETIREMENT.md` for scope and live
evidence.
- Storybook build passed. A fresh real agent used
`COMPOSIO_SEARCH_TOOLS` and `COMPOSIO_MULTI_EXECUTE_TOOL` to return the
actual Paperclip DeepWiki hierarchy: one success, zero errors. Gateway
audit records confirm both calls succeeded.
- Browser retirement check: a credential-free legacy fixture showed the
guidance, opened the direct MCP replacement flow, and was removed
through the standard confirmation.
- Focused regressions for the experimental settings copy and exact
OpenAPI route coverage passed. All latest-head CI checks passed (54
successful, two intentionally skipped); Greptile scored 5/5 with no
unresolved review threads. The PR has no merge conflicts.

## Risks

This intentionally breaks the old Composio project API-key and
child-connection workflow. Existing legacy records cannot run, even if
their stored status is active. Operators must create a new direct MCP
connection and choose access rules; credentials and grants are not
migrated. Remove each old record separately to delete its credentials.
No schema migration or data deletion runs automatically. Direct MCP
connections keep their existing grants and secrets.

## Model Used

OpenAI GPT-6 via Codex, with reasoning, code execution, and browser
tools. The exact runtime variant and context-window size are not exposed
in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 13:57:35 -05:00
1ef3b08714 feat(ui): integrate agent personas across the app (#13171)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - A stable agent persona is useful only when the same identity appears
across the app.
> - Lists, task messages, selectors, and activity feeds need inexpensive
static avatars.
> - Onboarding and agent headers need a larger character with
expressions and pointer tracking.
> - This pull request connects the persona foundation to those existing
views and preserves onboarding draft assignments.
> - Full-page stories and Linux checks make the placements and
performance contract reviewable.

## Linked Issues or Issue Description

**Problem or motivation**

Agents need a stable visual identity in lists, tasks, onboarding, and
configuration. External tools also need an image URL for that identity.

**Proposed solution**

Assign each agent a permanent palette from a fixed ClipLab character
library. Store the assignment on the agent. Render and cache preset PNG
URLs on demand. Use static images in dense views and one animated
character in larger placements.

**Alternatives considered**

A generated image bundle requires a separate asset build. A live
renderer in every avatar adds unnecessary work in large lists. Arbitrary
uploaded images do not provide the requested shared character system.

**Roadmap alignment**

This improves agent identity across existing control-plane views. It
preserves agent permissions, company boundaries, and status labels.
ROADMAP.md has no separate ClipLab persona milestone.

Related approaches: #2422 adds configurable image URLs and DiceBear
generation; #5578 adds optional uploaded avatars. This work uses a
fixed, versioned character library and preset URLs.

## What Changed

- Replace agent icons with static persona images across lists, the
sidebar, org charts, tasks, comments, selectors, activity, and dashboard
views.
- Put one animated character in the agent header. Let it follow the
pointer across the page, with reduced-motion and touch fallbacks.
- Add larger padded characters to agent creation. Keep the palette
stable across draft refreshes and connection retries, then reveal it
after success.
- Pass appearance through shared projections rather than fetching each
agent separately.
- Add real full-page Storybook examples for the agent list, overview,
task, dashboard, new-agent dialog, and connection page.
- Add Linux screenshot, clipping, density, and 500-avatar performance
checks.

## Verification

- `pnpm -r typecheck`, `pnpm build`, and token gates pass on the rebased
tree. Persona lifecycle tests pass.
- The rebased feature passes 38 Linux screenshot/performance checks,
including both display densities, corner pointer positions, and the
no-WebGL/no-live-download contract for 500 avatars.
- The final Linux persona suite passes all 38 visual, lifecycle,
density, and full-page checks using the standard Storybook configuration
and real on-demand avatar endpoint.
- Final local focused verification: 45 avatar/native-recovery tests
pass; UI identity/routine tests, typecheck/build, token gates, and
Storybook build pass.
- Current-head CI passes: full workspace/server tests, all serialized
server groups, typecheck/release checks, build, canary validation, and
end-to-end shards. The build passed after retrying a native-runner
concurrency-test failure; its three targeted cases also pass locally.
- Manual inspection covered stable identities in the app, header
placement, full-page mouse tracking, onboarding size, and task/dashboard
placements.


### Screenshots

Linux captures use synthetic Storybook fixtures. Full-page captures use
reduced motion. The live character, mouse tracking, and disposal are
checked separately.

<details>
<summary>Agent overview with the character in its header</summary>

<img
src="https://raw.githubusercontent.com/paperclipai/paperclip/8c68f42b268ada22b67d79d0fe1bb0a2f84ec25c/screenshots/full-page-agent-overview.png"
width="900" alt="Agent overview with the character in its header" />

</details>
<details>
<summary>Task messages and assignee identity</summary>

<img
src="https://raw.githubusercontent.com/paperclipai/paperclip/8c68f42b268ada22b67d79d0fe1bb0a2f84ec25c/screenshots/full-page-task.png"
width="900" alt="Task messages and assignee identity" />

</details>
<details>
<summary>Larger onboarding character with room for expressions</summary>

<img
src="https://raw.githubusercontent.com/paperclipai/paperclip/8c68f42b268ada22b67d79d0fe1bb0a2f84ec25c/screenshots/full-page-meet-your-next-agent.png"
width="900" alt="Larger onboarding character with room for expressions"
/>

</details>
<details>
<summary>Dashboard agent activity</summary>

<img
src="https://raw.githubusercontent.com/paperclipai/paperclip/8c68f42b268ada22b67d79d0fe1bb0a2f84ec25c/screenshots/full-page-company-dashboard.png"
width="900" alt="Dashboard agent activity" />

</details>

## Risks

- This PR depends on #13170, the persona foundation. Merge the
foundation first, then retarget this PR to master.
- Many placements change from icons to character silhouettes. Human
avatars and authoritative agent status labels retain their existing
behavior.
- Only one character can render live per view. Reduced motion,
hidden/offscreen content, touch input, and renderer failures use the
defined fallbacks.
- The full-page stories use fixture data. They do not contact a real
company or complete real provider sign-in.

## Model Used

OpenAI Codex, GPT-6 family. The exact model identifier and context
window are not exposed in this session. Used code editing, shell
execution, browser inspection, and Linux visual testing.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Tonio <tonework@gmail.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 07:57:53 -05:00
DottaandPaperclip 924f07be8c feat(chat): simplify Slack onboarding and account linking (#13638)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Chat connections let people start and continue that work from Slack.
> - Setup mixed app creation, credentials, URL verification, account
linking, and testing on the same screens.
> - People also needed a safe way to link their own Slack identity after
the first operator finished setup.
> - This pull request gives each step a clear place and keeps membership
approval separate from identity linking.
> - It also makes connection details easier to use and fixes misleading
callback health behind HTTPS proxies.

## Linked Issues or Issue Description

**Subsystem affected**
Cross-cutting: chat routes and services, shared contracts, and the Apps
board UI.

**Problem or motivation**
Slack onboarding made users find settings without enough guidance. A
second user needed operator help to link their account. Activity stopped
at 100 records, and TLS termination could mark working callbacks as
stale.

**Proposed solution**
Use six setup steps with editable app names, a generated manifest,
credential guidance, URL verification, account linking, and an optional
message test. Send each Slack user a private, expiring confirmation
link. Require company membership or an approved access request before
linking. Add cursor pagination and tolerate the internal HTTP hop in
callback diagnostics.

**Roadmap alignment**
This improves the existing connected-app surface and supports CEO Chat
without changing the task-and-comments model. The maintainer requested
and reviewed the flow during a live Slack test drive.

**Additional context**
Related work: #7, #3349, #13000, and #13620. Those cover broader chat
capabilities, older webhook paths, or plugins. This PR improves the
existing native connector's setup and account-linking flow. HTTPS
documentation was published separately in
paperclipai/paperclip-docs#128.

## What Changed

- Split Slack onboarding into six clickable sidebar steps. Keep
secondary and primary actions on one row.
- Generate the Slack creation link and read-only manifest from editable
app, bot, and command names. Add credential prefix validation and direct
instructions.
- Add live account-link status and an optional mention-based message
test.
- Add private, single-use Slack account invitations and membership
access requests. Retain cloud authentication/bootstrap checks and
enforce the chat rollout flag in all identity APIs. Default new Slack
connections to linked users only.
- Put Settings, Access, Conversations, and Activity in the sidebar.
Simplify conversation rows and remove active header badges.
- Add 25-item activity pages, stable timestamp/ID cursors, and replay
safety across pages. Preserve the legacy array API for clients without
pagination parameters.
- Fix false callback warnings when HTTPS terminates at a proxy. Keep
host, port, and path drift detection.
- Document the setup flow, pagination, callback diagnostics, and shared
wizard footer rule.

## Verification

- Passed: `pnpm -r typecheck`, `pnpm build`, and `pnpm
check:token-gates`.
- Passed: focused Slack callback and pagination integration tests; UI
clipboard, wizard, pagination, and activity tests; OpenAPI route tests.
The final access-gate fix also passes 27 focused tests covering cloud
authentication/bootstrap, nonmember invitations, token validity, and the
server-enforced rollout flag.
- Passed: all 1,002 chat integration tests, 6,356 UI tests, and all 11
provider browser scenarios (including mobile light/dark navigation).
After rebase, the identity route, sidebar, and 25 clipboard tests pass.
- The full local `pnpm test:run` was attempted. The first run found 14
Slack fixtures that needed explicit guest access; those are fixed and
the complete chat suite passes. Unrelated embedded PostgreSQL
startup/resource failures and timeouts prevented a clean full local run.
All CI checks pass on `2d858b036`, including the full chat, server,
workspace, build, typecheck, and browser suites.
- Live test drive: Slack app creation, credential setup, URL
verification, private account confirmation, mention messages, and thread
replies. Verified the callback warning clears for the existing proxied
connection.
- Review: create a Slack connection, follow the six steps, link a second
user's account, and browse older activity with Next and Previous.

## Risks

- Identity invitations carry a temporary capability. Tokens are hashed,
expire after 15 minutes, work once, and require explicit confirmation by
a company member. Access requests do not grant membership.
- New Slack connections reject unlinked people by default. Existing
connection settings remain intact.
- Activity is a live ledger. Updated action rows can move forward in
time. Older pages do not poll.
- Proxy tolerance affects health display only. Slack signature checks
and proxy authentication settings remain unchanged.
- No database migration or package-lock changes.

## Model Used

OpenAI Codex, based on GPT-6, with reasoning, repository tools, code
execution, and browser verification. The runtime does not expose an
exact model build ID or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (targeted suites; full
local-run limitations documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-18 17:23:53 -05:00
Devin FoleyandPaperclip 6fe8e30625 feat(apps): add Railway connection and governed deployment tools (#13415)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Apps gives agents governed access to external resources.
> - Operators need to inspect Railway services, read logs, deploy code,
and run container commands.
> - Railway offers hosted MCP with OAuth, but broad remote actions hide
their internal operations.
> - This PR adds a branded connection and fixed direct operations
through the existing gateway.
> - Separate SSH keys enable container commands under the same grants
and policies.
> - Operators can require approval for an action and inspect the
resulting audit record.

## Linked Issues or Issue Description

**Subsystem affected**

Apps catalog, connection setup, gateway execution, and connection
documentation.

**Problem or motivation**

Agents need Railway access through Paperclip. Operators need to grant
and revoke that access, inspect available actions, and govern deployment
and container operations without giving agents provider credentials.

**Proposed solution**

Reuse hosted MCP OAuth, vault storage, catalog discovery, grants, and
the gateway. Probe the actual credential before enabling fixed GraphQL
operations. Use a dedicated grant-owned SSH key for bounded container
commands.

**Alternatives considered**

A catalog entry alone cannot execute the missing operations. The hosted
general agent has opaque internal effects. An unrestricted CLI runtime
can bypass action policy and inherit ambient credentials.

**Roadmap alignment**

This extends the existing MCP Tool Gateway & Apps path and the Connected
Apps direction in ROADMAP.md. It does not add a plugin or parallel
connection service.

Related PRs #311, #939, and #7861 concern hosting Paperclip on Railway.
They do not add this outbound Apps connection. The separate shared
agent-picker fix is #13414 and is not included here.

## What Changed

- Add the generated Railway catalog entry, official marks, provenance,
and OAuth setup guidance.
- Add fixed service/deployment status, bounded logs, and
redeploy/restart/rollback tools. Block source deployment until the
provider can atomically bind the approved repository and commit.
- Verify API access with an explicit workspace before exposing direct
tools.
- Add grant-owned SSH key setup and a bounded runner with host
verification, target checks, isolated state, and cleanup.
- Block the opaque hosted railway-agent and accept-deploy actions.
Preserve normal Allowed defaults and Ask-first policies for other
actions.
- Quarantine new or changed Railway schemas after initial discovery,
including reconnect.
- Add provider, lifecycle, gateway, SSH, UI, and browser fixtures.
Document setup, limitations, and the release checklist.

## Verification

- Security follow-up: removed the unsafe source-deployment mutation.
Direct calls and old active catalog entries are denied before any
upstream request, including normalized aliases. Refresh marks retired
entries disabled. All 386 focused Railway, catalog and gateway tests
passed, and server TypeScript checking passed. Full [GitHub
CI](https://github.com/paperclipai/paperclip/actions/runs/35139421144)
passed on d86530ab9, including typecheck, build, all tests, runner
checks, and browser tests. Superagent passed and confirmed the P2 fix.
Greptile reviewed the same commit at 5/5 with no findings.

- CI follow-up: fixed the missing Railway SSH operation in the OpenAPI
document, including its request schema, operator-only authentication,
and error responses. The failure reproduced locally before the fix; all
403 selected API, Railway, catalog, and artwork tests passed after it.
Synced current master and resolved the catalog/artwork conflicts.

- After rebase: 440 focused provider, lifecycle, gateway, catalog, and
container-panel tests passed. AppDetail and AppsConnect passed another
196 tests.
- Full typecheck, build, token gates, and the gallery browser check
passed after rebase.
- During implementation, full build and the gallery browser check
passed. Shared generic-MCP fixtures covered OAuth callback/state/issuer
binding and failure paths.
- Local live consent and tools/list succeeded. There were 44 active
hosted actions and two blocked actions. A workspace-bound API probe and
direct project/service/environment reads succeeded. The inspected
project had no deployed services. No provider mutation ran.
- Full GitHub CI passed on commit 303340f19, including all
server/workspace test groups, typecheck, build, runtime verification,
release dry run, and browser tests. The original local full-run attempt
was incomplete; the complete automated suite is now verified in CI.

Manual review: connect Railway, review the actual actions, install for
an agent, and run a resource read through the gateway. Choose Ask first
before testing a deployment mutation. Configure a dedicated key only
when container access is needed.

**Release qualification is still open.** Live agent gateway reads/logs,
rejected and approved deployment calls, refresh/revoke, public HTTPS
consent, and SSH enrollment/commands/cleanup need an authorized
disposable service. The passing API diagnostic does not replace those
tests. See doc/connections/RAILWAY.md and RAILWAY-REVIEW.md.

## Risks

Overall risk is medium. New runtime behavior is gated to Railway
connections, but the PR changes shared catalog, credential lifecycle,
and gateway code. A regression in those paths can affect other Apps
connections. The highest-impact operations are Railway deployments and
container commands.

- Provider consent can authorize an entire workspace. Catalog labels are
not local resource allowlists. Direct tools check target membership, and
provider permissions still apply.
- Shell commands have broad internal authority. Action policy cannot
approve each internal shell step. Timeouts close the local connection
but cannot guarantee remote child-process termination.
- Log and command output may contain application secrets that pattern
redaction cannot recognize.
- Source deployment is unavailable until the provider supports atomic
repository/commit binding. Existing deployments can still be redeployed,
restarted or rolled back.
- No database migration is required. Rollback can remove promotion and
direct dispatch while preserving connection data and the generic MCP
path.
- Live Railway qualification must still pass before release acceptance.

## Model Used

OpenAI Codex, based on GPT-6, with code execution and browser testing.
An independent read-only security agent reviewed the local
implementation. The exact serving model ID and context window were not
exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-16 13:54:44 -07:00
DottaandPaperclip 728f7185f6 feat: add native in-app announcements with persistent dismissal (#13403)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Self-hosted boards need a way to show occasional product
announcements.
> - An app release should not be required to publish or withdraw a card.
> - Native card controls keep publishing consistent; the hero can use a
static image or isolated HTML/CSS animation.
> - This pull request renders a validated JSON feed with native
components.
> - It stores dismissals per account on each instance, so a closed card
stays closed across companies and browsers.
> - Named staging feeds let authors test content before production
publication.

## Linked Issues or Issue Description

**Subsystem affected**

Board application shell, announcement delivery, and user preferences.

**Problem or motivation**

Operators need a small, optional announcement card. Users need reliable
dismissal state. Authors need to test remote content without changing
the production feed.

**Proposed solution**

Add one non-modal AnnouncementWell. Fetch validated JSON and
content-addressed media through the instance server. Keep card controls
native, with optional sandboxed HTML/CSS animation in the hero. Use
stable announcement IDs for dismissal, an explicit empty manifest and
quiet 404 handling. Provide a staged publishing helper and isolated
test-drive guide.

**Alternatives considered**

Hosting the entire card as a page would move navigation and dismissal
into remote content. This change limits HTML to a scriptless, isolated
visual hero and keeps controls native. Browser-only storage would lose
dismissals across browsers, so the instance stores account preferences.

**Roadmap alignment**

ROADMAP.md has no overlapping announcement feature. A GitHub title
search found no related announcement pull requests. This work implements
a maintainer-requested feature.

## What Changed

- Add shared feed types, strict validation of every object, supported
routes, expiration and version checks.
- Add a board-only current-feed API, constrained media proxy, and
idempotent dismissal API. Store the first dismissal and its company
audit entry in one transaction.
- Cache upstream data for one hour. Use conditional requests, request
deduplication, response limits, public destination checks, and a
three-second deadline. Treat a remote 404 as an empty feed with a
fifteen-minute retry cooldown.
- Keep announcement visibility stable when focus moves to browser chrome
or another app pane; only tab visibility starts a return check.
- Add a responsive native announcement card. Respect onboarding, dialogs
and toast placement. Sync pending dismissals across tabs and retry after
reconnect or return.
- Add idempotent migrations for dismissals and validated publication
IDs, design-guide examples, static and animated Storybook examples, and
focused tests. The publication registry supports offline retries without
accepting caller-invented IDs.
- Add HTML/CSS animated heroes with static posters, automatic playback,
reduced-motion handling, strict DOMPurify validation, an empty iframe
sandbox and CSP that blocks scripts/network resources.
- Add validated staging publication, content-addressed assets, an empty
production manifest, preview fixtures, and authoring/operator
documentation.

## Verification

- The preceding implementation passed 98 targeted
shared/server/publisher/route/OpenAPI/UI tests and 127 tests including
the master rebase. The playback-control removal passes all 21
announcement UI tests, covering the rendered sandbox, fallback, reduced
motion, dismissal and slow/stale state lookups. The preceding
shared/server tests cover HTML validation and response sandbox headers.
- The playback-control removal passes UI typecheck, production UI build,
Storybook build and token gates locally. Browser verification confirms
the animated card has only its dismiss button and two links, with no
page errors. The full canonical CI matrix passed on current head
`00e416431edb610861599d50490270bbd0f3c6b6`: 32 successful checks and two
optional Storybook deployment checks skipped. This run needed no
retries. Greptile reviewed this same head at 5/5 with no outstanding
findings.
- The local canonical general-server run passed 12,063 tests before
reporting embedded-PostgreSQL startup failures in an unrelated fixture.
All 31 tests in that fixture passed across isolated retries. The UI
group passed 6,219 tests and other workspace groups passed 3,201; two
CLI database-startup failures also passed individually. Serialized
server suites were verified by the full CI matrix rather than repeating
them locally. No source changes were needed for these environment
failures.
- The real S3/CloudFront staging manifest and both media asset headers
were verified. Production remains empty/unpublished. The guide
distinguishes the preview host's disabled edge cache from production
cache requirements.
- In the isolated test-drive, the animation visibly moves without
playback controls. A 390×844 browser viewport keeps the card above
navigation. Reduced motion makes no animation request. Both themes
render correctly and browser page errors are empty. Browser fault
injection verified that scripts cannot execute and CSS cannot make
network requests; a missing animation leaves its poster and controls.
- Refresh leaves the animated card visible. Closing it persists after
reload and the API returns null. Earlier live checks verified dismissal
across browsers, company-relative CTA navigation, modal
deferral/restoration, and new-ID eligibility after restarting the same
database.
- The deployed empty feed and a real remote 404 return HTTP 200 with
null from the board API, with a usable dashboard and no announcement
popup or browser warnings.
- Authoring documentation covers staging, animated HTML constraints,
test-drive, withdrawal, ID reuse and cache-refresh steps.

## Risks

- Animation supports self-contained visual HTML/CSS and inline SVG,
without JavaScript or external resources. A static image is required.
Older builds that do not recognize the optional animation field quietly
hide that unsupported feed.
- The default feed makes an outbound request from an instance when a
board is used. Operators can disable it. Requests contain no account
IDs, company data, cookies or interaction events.
- Feed publication and withdrawal can take about 65 minutes to reach
returning users because of CDN and instance caches. Expiration also
removes visible cards locally.
- Dismissals follow an account within one instance. No-login instances
share the existing local-board identity. Separate installations do not
share state.
- Both tables are additive. A unique key prevents duplicate dismissals;
the transaction prevents duplicate first-dismissal audit entries. The
publication registry retains only validated IDs. AGENTS.md and the
implementation spec document the required exception to company scope for
these instance-level records.
- Publication was limited to separate public staging prefixes on the
existing preview host. Production remains empty/unpublished. No AWS
policies or infrastructure were changed.

## Model Used

OpenAI GPT-6 through Codex. The exact runtime model ID and
context-window size are not exposed in this session. Capabilities used:
reasoning, code editing, shell execution, tests, browser interaction,
and tool use.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-14 10:19:54 -05:00
DottaandPaperclip b2acc674be fix: enable isolated subscription login on authenticated self-hosted instances (#13344)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - AI Connections store reusable provider credentials with company and
owner boundaries.
> - Self-hosted instances can require board authentication while running
agents on the local server.
> - The login UI treated those instances as unsupported and displayed
instructions without a command.
> - Removing that restriction must not expose the server operator's
existing CLI account.
> - This pull request enables owner-scoped login attempts and reuses the
existing login UI.
> - Users can connect Codex and Claude subscriptions on an authenticated
self-hosted instance.

## Linked Issues or Issue Description

**What happened?**

On an authenticated self-hosted instance, OpenAI subscription setup
displayed terminal instructions with no command and a disabled Connect
button. Claude also could not complete the local connection flow.

**Expected behavior**

An authorized board user can prepare an isolated sign-in attempt, sign
in on the server, and save the verified account as a Connection. This
must not import another user's or the operator's ambient credentials.

**Steps to reproduce**

Run Paperclip in authenticated mode with a local environment. Open
Connections, choose OpenAI or Anthropic, and select Subscription. The
previous UI never enabled local login preparation.

Related: #13247, #13248, and #10751. This fix preserves the restriction
on remote access to the operator's ambient Claude login.

## What Changed

- Allow company-authorized users to create, check, cancel, and complete
their own isolated local login attempts.
- Keep ambient Claude credential import restricted to the local
operator.
- Support isolated Claude credential files without falling back to the
host account or mutating process-wide environment variables.
- Use Codex device authorization so sign-in does not depend on a browser
callback to the remote server's localhost.
- Gate server-host login on authenticated public deployments unless a
trusted runtime host is configured. Publish the capability through
health so setup shows supported alternatives.
- Read isolated Claude credential files through bounded,
descriptor-bound opens with ownership, permission, and symlink checks.
Try the alternate filename after malformed JSON.
- Reuse shared login instructions and lifecycle hooks in onboarding,
agent setup, and Connections. Show health-query failures explicitly.
- Document authenticated self-hosted behavior and add authorization,
isolation, lifecycle, and UI regression tests.

## Verification

- Passed 71 focused tests across connection routes, credential
isolation, legacy compatibility, the shared login hook, and agent setup.
- Passed 89 onboarding regression tests.
- Passed `pnpm -r typecheck`, `pnpm build`, Storybook build, and `pnpm
check:token-gates`.
- Completed real Codex device authorization and Claude browser
authorization on an authenticated Linux self-hosted instance. Both
accounts were detected automatically and saved as Connected. Both
completed attempt directories were removed.
- These live checks cover login, credential validation, and connection
creation. They do not establish a new model execution or long-running
refresh result.
- Review follow-up: 68 focused checks passed after rerunning one route
socket error; the full route/health rerun passed all 49 tests. The
70-test onboarding suite also passed. Final workspace typecheck,
production build, and Storybook build passed again.
- The broad local run exposed an instance-name assumption in two new
assertions. The fixture now uses an explicit non-default instance, and
all 32 connection tests passed with a different inherited instance name.
The superseded broad run was stopped; this is not a claim that the full
local suite completed. Full CI results will be recorded before merge.

## Risks

- The server must have the provider CLI installed. Users still run the
displayed command on the server that hosts Paperclip.
- Authorization checks must keep login attempts scoped to the company,
owner, provider, and reconnect target. Regression tests cover cross-user
and cross-company access.
- Existing local-trusted Claude behavior stays available. Authenticated
remote users cannot use its ambient import path.
- No database migration, dependency change, agent binding change, or
provider routing change is included.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, code execution, and browser
tools. The runtime does not expose a more specific model version or
context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-12 23:06:55 +00:00
DottaandPaperclip df984cbc2c fix: dispatch queued legacy messages with operator identity and task permissions (#13315)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Task conversations save messages that arrive during an active turn.
> - Legacy adapters deliver these messages in a later turn.
> - A run can stop before the saved queue is delivered.
> - The Interrupt button previously required an active run, so it could
not release this queue.
> - This pull request lets a board operator send the saved queue after
the run stops and retries queues missed during finalization.
> - Manual dispatch must use the clicking operator and must not require
permission to create agents.
> - The task can continue without a duplicate message or a second
execution owner.

## Linked Issues or Issue Description

**What happened?**

A legacy task retained a queued message after its run stopped. Interrupt
was disabled because the queue had no active target. Finalization and
deferred message admission can also leave a queue without a successor.

**Expected behavior**

Interrupt sends the saved messages when no runner is active. Messages
that arrive during normal completion are delivered automatically. An
uncertain previous execution still requires proof that its process or
sandbox stopped.

**Steps to reproduce**

1. Queue a user message during a legacy conversation turn.
2. Let the turn stop or simulate a server restart before queue
promotion.
3. Open the task with a deferred queue and no active run.
4. Try Interrupt. Before this change, the button is disabled.

**Paperclip version or commit**

Reproduced against `8f40b4ad4`.

**Deployment mode**

Legacy conversation adapter. The same persisted queue state is covered
with an isolated PostgreSQL fixture.

Related public work: #13275 adds active legacy interruption. #13291
addresses automatic sandbox conversation recovery. This change handles
explicit saved-queue delivery and late queue promotion.

## What Changed

- Accept a null Interrupt target while retaining queue identity,
revision, company, and assignee checks.
- Save the operator's request on the existing queue. Reuse normal
admission after verified stop, including older messages, different
authors, and queues whose original wake came from the system.
- Strip interruption authority from caller-supplied wake payloads. Only
the board queue route can persist that authority.
- Retry durable interruption requests after restart and deferred queues
after legacy cleanup.
- Let an explicit Interrupt retry cleanup for its stopped run, including
old ephemeral leases that recorded success without a provider stop
receipt. Preserve retained resources, other lease owners, and the
automatic retry limit.
- Preserve the server's waiting explanation when normalizing and
combining queue entries.
- Revalidate the consumed board queue receipt at dispatch so a different
message author does not cause setup failure.
- Use the Interrupt user's execution identity for the new run. Preserve
original message authors. Validate the receipt independently at startup
and inherit the resulting identity on retry.
- Persist authenticated board authority for ordinary manual wakes too.
Adopting someone else's queued messages cannot switch a manual run to
that author's permissions. Strip caller-supplied authority markers and
retain private conversation ownership checks.
- Keep the clicking user when a manual wake is merged into an older
deferred receipt. Update its requester and payload in the same
transaction.
- Use the same current-queue/revision API on task details and pipeline
conversations; show Interrupt after a legacy target stops.
- Keep manual wakes out of active runs, including unscoped agent wakes.
They receive their own execution identity; a matching receipt requester
is not sufficient because an exact retry can retain a different
originating identity.
- Authorize both existing-agent wake endpoints with `agent:wake`,
available to active non-viewer company members. Keep `agents:create` for
hiring. Validate the stored task and current assignee before an exact
task retry.
- Reject viewer Interrupt requests before saving intent or stopping
execution. Keep external chat retry authorization and per-action
agent/user permission checks.
- Preserve edits and discards until dispatch. Prevent another queue
promotion when the same agent already has a successor. Keep independent
reviewer recovery available.
- Suppress cancelled/failed run toasts for intentional operator
interruption. Keep ordinary runtime error notices.
- Add UI, route, admission, restart, successor ownership, and toast
regression tests. Document the behavior.
- Reuse the existing socket reservation helper for both
credential-quorum test cases after CI exposed an ambient-port collision.
This changes test preparation only; production credential staging is
still called exactly once.

## Verification

- Failing regression tests reproduced the message-author identity bug
and an operator's `agents:create` rejection before the fixes.
- All 316 focused tests pass across eight route, queue, identity,
authorization, continuation, and responsible-user suites, including the
44-test rerun of queue admission and actual startup after the final
manual-wake restriction. Regressions reproduce cross-user merging both
with and without a task, and same-requester receipt ambiguity. The
cross-company existence guard also passes both tests.
- Startup integration tests reach adapter execution under the clicking
operator and retain that identity through follow-up. Coverage includes
mixed authors, adopted queues, system-origin queues, restarts, forged or
stale receipts, viewers, suspended memberships, changed assignees,
private conversations, and caller-supplied authority markers.
- The earlier queue/cleanup/UI regression suite passed 402 tests. The
final review corrections pass another 180 tests across queue
admission/persistence, real heartbeat startup, UI API, conversation
rendering, and pipeline suites. Regression tests reproduced both review
findings before correction. The final head has a 5/5 review with no
unresolved threads. Full CI passes on `c2002979c`, including every
general and serialized server shard, all browser shards, Paperclip
Runner verification, typecheck, build, canary dry run, and the aggregate
gates.
- Full `pnpm -r typecheck`, `pnpm build`, and UI token gates pass after
the final application changes. CI identified an outdated task-page API
mock after the shared helper extraction; the fixture now exercises the
real helper, and all 131 task-page/API tests pass. The final application
build passes with the additional manual-wake restriction.
- CI exposed a pre-existing port collision in the Codex
credential-quorum fixture. It reproduced locally; both listener cases
now use the existing bounded reservation helper. All 41 credential tests
pass on rerun. One intervening local run hit a separate ambient bind
collision in the two-occupied-port case.
- The full local `pnpm test:run` attempt was stopped after host
contention caused focused-suite timeouts. The affected focused tests
passed on rerun. An expiring trace fixture and a missing
private-conversation state were corrected. The successful full CI run is
the complete-suite verification.
- Hosted Interrupt previously cleared the original queue and produced
exactly one successor with neutral interruption feedback. It exposed the
dispatch authorization defect. Retry on that earlier build was rejected
for missing `agents:create` before creating another run.
- Deployed the final application build (`38257f391`) to the scoped
hosted instance and verified readiness. The latest PR commit changes
only the credential test fixture; application code matches that
deployment. A live Retry by the same operator without `agents:create`
created one successor attributed to that operator, passing the former
dispatch permission gate. Startup then stopped at
`configuration_incomplete` because that operator has not configured
their required personal Claude Code OAuth secret; the post-deployment
run page confirms the operator identity and no provider work started,
and the My secrets UI still shows the token as not set. Provider
execution remains unverified pending that credential. No permission
grants or credentials were changed.

## Risks

Queue admission and finalization can race. The task lock, durable queue
receipt, current comment IDs, and successor guard prevent duplicate
dispatch. Process and lease stop checks, task pauses, approvals,
ownership, and budgets remain in force. The API change only allows null
on legacy Interrupt; native steering still requires an active run. No
schema migration is required. Active non-viewer board members can now
invoke existing agents without agent-creation permission. Agent
self-invocation rules, raw provider-trace admin access, task retry
scope, external chat authorization, and action-specific user/agent
permissions remain enforced.

## Model Used

OpenAI GPT-6 through Codex. The session does not expose an exact backend
model ID or context-window size. Used reasoning, repository search, code
execution, tests, and browser tools.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-12 17:11:30 -05:00
DottaandPaperclip 8d6232e7b0 feat: reuse provider sign-in across AI connection workflows (#13248)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Users connect provider accounts during onboarding and agent setup.
> - They should reuse and manage those accounts through the existing
Connectors interface.
> - A second login wizard would diverge from the established provider
workflows.
> - This pull request composes the existing sign-in components into
Connections and agent configuration.
> - Users can select accounts without changing their agent's harness or
model.

## Linked Issues or Issue Description

**Problem or motivation**
AI credentials are configured separately from Connections. Agents cannot
consistently reuse a responsible user's account or a permitted shared
account.

**Proposed solution**
Manage AI accounts with the existing Connections grants and permissions.
Keep model and harness selection independent from credential selection.
Preserve legacy authentication until validated adoption.

**Alternatives considered**
A separate credential registry would duplicate ownership and access
policy. Automatic fallback would risk using the wrong account.

**Roadmap alignment**
This extends the shipped Apps, multi-user, secrets, and agent-runtime
capabilities. The maintainer requested the feature and reviewed the UI.
Related groundwork: #11899 (connection permissions), #10910 (connection
wizard), #11692 (Claude subscription profiles), and #11854 (Codex
account rotation).

## What Changed

- Add compact AI-account management to the existing Connectors pages.
- Reuse AgentProviderConnection, AdapterLoginPanel, AdapterLoginChrome,
and authentication controllers.
- Add the shared connection picker to agent setup/settings and task
requests.
- Preserve onboarding's sequence and reuse existing accounts.
- Add local-login recovery, retry, cancellation, and React StrictMode
handling.
- Add interactive Storybook scenarios, design-guide examples, and app
acceptance checks.

This is part 2 of the AI Connections change. The runtime foundation in
#13247 is merged. This PR now targets master.

## Verification

- Updated against master `47ded8bf9`, including the landed runtime
foundation and upstream task-search changes.
- Full workspace typecheck, production build, Storybook build, and token
gates passed on the integrated branch. Final local-login changes passed
59 focused tests; new-agent and inbox regression suites passed 63 tests.
- Browser checks verified automatic local Claude account detection,
resumable Codex login commands, retry, focus restoration, and
desktop/phone layouts. Commands create their isolated directory before
invoking the CLI.
- All CI test, browser, build, packaging, and runner jobs passed on
final head `dd17d3211931dd70aaa6ea619d83a7f9966dd18e`. The fresh
Greptile review is 5/5, the security scan passed, and there are no
unresolved review threads. The final CI aggregate gates passed.
- Local general-server coverage passed 11,804 tests; three
port-collision failures passed in an isolated 25-test rerun. All 6,111
UI tests passed. CLI coverage passed 484 tests; its remaining doctor
test requires port 3199, which is occupied by an unrelated report server
on this Mac. The complete CLI suite passed in CI.
- Live browser testing verified Codex API-key reconnect inside a task
card on desktop and phone. Real provider runs resumed and completed with
unchanged connection/grant identity and agent routing.
- Tested opening, cancelling, reopening, and completing connection
creation. A regression confirms Connect another account cannot submit
the new-agent form or copy provider keys into agent settings.
- Added shared inline repair, automatic local sign-in checks, and
responsive connection dialogs. Standalone Daytona installation ignores
workspace configuration and suppresses dependency scripts. Its
standalone build also passed with CI's exact pnpm 9.15.4.
- Destructive live tests are excluded by default. Explicit opt-in, local
deployment checks, and matching disposable fixture identities are
required before any mutation.

## Risks

- Local Codex/Grok creation requires the connection-specific terminal
login command.
- Browser sign-in uses the existing supported-environment controllers.
- This update verifies live local Claude detection and Codex API-key
task repair. New subscription authorization/refresh and
independent-human/native-runner isolation were not reverified in this
update.
- No agent automatically adopts managed Connections.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, code
execution, and browser testing. The exact runtime model identifier and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-12 16:51:26 -05:00
DottaandPaperclip 4d317274ce feat(channels): add experimental iMessage Photon (#13299)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Channels connect external conversations to company tasks and agent
execution.
> - Slack, Discord, and AgentMail already provide durable delivery and
access controls.
> - People also need to reach an agent from Apple Messages and send
photos.
> - Photon provides shared Pro DMs, dedicated numbers, and authenticated
event recovery.
> - This pull request connects Photon to the existing channel services.
> - People can message an agent while Paperclip retains task ownership
and approval authority.

## Linked Issues or Issue Description

**Subsystem affected**

Cross-cutting: channel services, shared contracts, database constraints,
Apps, and agent Channels UI.

**Problem or motivation**

Paperclip has no iMessage channel. A person cannot use Apple Messages to
start a task, send a photo, or answer an agent's pending question.

**Proposed solution**

Add experimental **iMessage Photon** with Pro-compatible shared DMs or a
dedicated Photon Cloud number per agent channel. Reuse channel
admission, identity links, task generations, publication, and
interaction continuation. Keep groups disabled for shared allocation.
Dedicated lines support groups that an operator explicitly enables.
Require a fresh linked message and a published agent response before
setup completes.

**Alternatives considered**

Shared allocation has no owned phone number, so it reserves one project
and allows DMs only. Dedicated allocation reserves one stable number.
Local Mac access needs a separate deployment model. The upstream Photon
Chat SDK adapter does not persist the poll mappings and send receipts
required here. This change uses the lower-level SDK without adding
another agent runtime.

**Roadmap alignment**

This extends Connected Apps and agent communication through the existing
channel subsystem. It does not add a parallel tool connection or agent
loop. GitHub searches for Photon and iMessage found no matching provider
implementation.

**Additional context**

This ships behind the existing experimental channel gate. Dedicated-line
release qualification remains incomplete. Real Photon Pro DMs passed
task/reply, native poll, text answers, confirmation rejection, media,
restart, pause, reconnect, revocation, and removal tests. An
operator-supplied iPhone camera HEIC also passed the full round trip.
Dedicated groups remain unqualified. See [the verification
record](doc/connections/IMESSAGE-PHOTON-VERIFICATION.md) and [the
implementation plan](doc/plans/2026-09-11-imessage-photon.md).

## What Changed

- Add the provider catalog entry, shared setup contracts, and a forward
migration. A global partial index reserves the dedicated number or
shared project until its endpoint is archived.
- Add Cloud project inspection, vaulted project credentials,
selected-line token renewal, and a leased receiver. Persist checkpoint
updates under the receiver lease. Shared project replay accepts sparse
increasing sequences only after a complete recovery barrier.
- Connect DMs and enabled groups to existing task generations, sender
authorization, ordered delivery, and publication services. Keep each
iMessage conversation on its task after completion; only explicit `/new`
or `/close` releases the binding. Publish committed inbound comments
live and label their human bubbles “Sent from iMessage” in both
task-chat renderers.
- Persist immutable text/file send identities, upload receipts, poll
IDs, option IDs, per-person drafts, and canonical interaction
continuation proofs.
- Add source-bound file recovery, bounded HEIC/HEIF conversion, JPEG
previews, and related Live Photo companion video retention.
- Add the three-step setup flow and channel management surfaces with
official branding. Preserve the experimental gate and existing
pause/disconnect behavior.
- Add interactive production-component Storybooks for setup, access,
recovery, and ongoing conversations. Add provider, integration, catalog,
and browser regression coverage. Document setup, recovery, supported
boundaries, and qualification gaps.

## Verification

- Live Photon Pro, SDK 2.1.0: linked iPhone messages create a task and
receive native Codex replies in Apple Messages. Unlinked senders cannot
start work.
- Three real follow-ups each reopened the same completed task. Incoming
bubbles appeared on its open page without reload and showed “Sent from
iMessage.” The third follow-up ran after restarting the server on
`4d7222110`; the agent correctly repeated its previous reply from before
the restart.
- Native polls after restart, sequential text drafts, required-field
correction, explicit submission, approval rejection with a required
reason, and native continuation passed against Photon.
- PNG, text documents, synthetic HEIC, and a real iPhone camera HEIC
passed in both directions. The camera photo produced a 3024×4032 JPEG
preview. The native agent described it and returned the received HEIC
byte-for-byte.
- Pause/resume, reconnect, identity revocation, removal, `/status`,
`/new`, `/close`, and stale answers after close passed live. Messages
suppressed by pause did not become work on resume. Removal stopped
intake and removed credential bindings.
- All 304 focused tests passed on `4d7222110`. These cover Photon
unit/integration behavior, both task-chat renderers, live comment
hydration, completed-task continuity after restart, enabled groups,
duplicate delivery, and explicit reset/close. The selected Teams
completion-boundary regression also passed. Full workspace
typecheck/build and token gates passed for the conversation fix; the
final UI changes passed their affected typecheck/build and tests.
- All 26 new Photon Storybook Playwright cases passed in light and dark
themes, including the complete shared-DM setup journey and 390px mobile
follow-ups. UI typecheck and the Storybook build passed. These stories
use simulated Photon responses and do not replace the live evidence
above.
- The full chat-adapters browser suite previously passed all 39 cases.
Migration checks passed, and migration 0275 applied to the isolated live
instance with the earlier Photon migration already applied.
- The local full Vitest run was previously interrupted by the host's
embedded-Postgres shared-memory limit; it is not a full-suite pass. All
30 applicable CI checks passed on preceding head `7a5419cac`, with two
skipped checks and Greptile 5/5. Head `24f8e1aae` adds an explicit
required-story discovery guard to the 26 passing Storybook cases.
Greptile rates this final head 5/5 with no unresolved review threads.
All 30 applicable CI checks passed, with two optional checks skipped.
- A repeated live send key suppressed the duplicate but returned gRPC 6
/ SDK `internalError` without an original receipt. Paperclip keeps
unknown delivery unresolved. This provider behavior is covered by a
regression test.
- See [the verification
record](doc/connections/IMESSAGE-PHOTON-VERIFICATION.md) for package
versions, redacted live evidence, deterministic coverage, and remaining
qualification gaps.

## Risks

- Dedicated group qualification remains unrun; groups are disabled for
the approved Pro scope. Real iPhone camera HEIC passed transport,
preview generation, agent inspection, and return. Keep the channel
experimental; the dedicated-line release matrix remains incomplete.
- Shared recovery and attachment aliases were verified against the live
gateway. Duplicate writes currently return an error without the original
receipt; unresolved sends require operator resolution. The
implementation fails visibly on invalid replay ordering, a reset cursor,
or changed identity.
- The HEIF converter passed on macOS arm64 and in Linux CI. Windows HEIF
binaries have not been executed in this work. Linux musl has no packaged
converter. Unsupported conversion retains the original and reports the
missing preview.
- The migration adds a global reservation across companies for Photon
numbers and shared projects. Paused and revoked endpoints keep that
reservation until removal.
- Integration touches shared channel services. Existing provider browser
coverage passes; broad repository verification is recorded above.
- `pnpm-lock.yaml` is intentionally excluded under repository policy.
The repository bot owns lockfile updates. The additional Superagent
supply-chain scan is neutral/inconclusive because these new dependencies
are not yet in the committed lockfile. Its security scan passed; all
required CI checks pass.

## Model Used

OpenAI Codex, GPT-6 family, with reasoning, repository inspection, code
execution, browser testing, and tool use. The exact served model
identifier and context-window size are not exposed in this session. No
sub-agents were used.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-12 15:23:50 -05:00
DottaandPaperclip ab15aff390 feat: add experimental persistent agent chat (#13284)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Conversations must use the same tasks, controls, and execution
history.
> - Users need an ongoing chat with an agent without managing task
properties.
> - Agents should clarify and plan work, then hand execution to assigned
project tasks.
> - This pull request combines the reviewed Agent Chat stack for one
squash merge.
> - The benefit is persistent conversation with normal task governance
and shared UI.

## Linked Issues or Issue Description

**Subsystem affected**

Task lifecycle, agent runtime tools, shared task UI, and browser/paid
runner tests.

**Problem or motivation**

Users need one persistent conversation with each agent. A separate chat
store or renderer would duplicate task behavior and bypass existing
controls.

**Proposed solution**

Use a task-backed chat per company, user, and agent. Reuse the task
composer and transcript. Clarify and plan in chat, then create assigned
project tasks with the relevant plan. Keep Agent Chat behind its own
disabled-by-default experimental setting.

**Roadmap alignment**

This implements the task-backed direction in [CEO
Chat](https://github.com/paperclipai/paperclip/blob/master/ROADMAP.md#-ceo-chat).
Related proposals: #2504 and #9693. Related request: #7981. The
maintainer requested one squash merge of the complete stack.

Consolidates the reviewed runtime
[#13281](https://github.com/paperclipai/paperclip/pull/13281), backend
[#13282](https://github.com/paperclipai/paperclip/pull/13282), and UI
[#13283](https://github.com/paperclipai/paperclip/pull/13283) layers
with this PR's E2E coverage. All four layers passed CI and received
Greptile 5/5 before consolidation. This PR targets master and includes
the complete feature.

## What Changed

- Add personal canonical chat tasks with ordinary company visibility,
immutable identity, idempotent first sends, and an idle waiting state.
- Process `/new` in queue order. Preserve history, release a chat pause,
and fence old provider context and delayed writes.
- Keep chat lifecycle rules across recovery, finalization, assignment,
task lists, and rollups.
- Support research and plan revision in chat. Hand plans to ordinary
assigned project tasks before execution starts. Reject new chat
subtasks.
- Add repository-aware project creation and discovery tools, including
multiple repository IDs and GitHub URLs, authorization, idempotency, and
durable project-created cards.
- Reuse task UI components for chat, with starred/recent agent
navigation and a separate `enableAgentChat` experimental flag.
- Add deterministic browser tests and 24 paid chat cells across four
Codex/Claude profiles, with validated reports and screenshots.
- Integrate current master recovery, controller lease, queued-message,
and task UI changes. Gate chat interruption and deferred promotion on
ownership/feature policy. Guarantee lease renewal and active controls
are stopped even if teardown fails.
- Preserve master's migration 0273 and generate chat migration 0274 with
idempotent replay for development databases.

## Verification

- Prior exact heads of all four PRs passed Linux CI, including build,
typecheck, general/serialized tests, and browser E2E. Each had Greptile
5/5 and no unresolved findings.
- Integrated local verification passed: full repository typecheck and
production build, Storybook build, token gates, 340 focused UI tests,
all 20 deterministic chat browser tests, two migration replay tests, 88
focused chat/queue/native/controller tests, and provider/session
regressions including real lease expiry. These include the three
lifecycle regressions for the final admission/teardown fixes; server
typecheck also passes. Current head
`1268eda16cc2af892055917e7292f068820be135` has Greptile 5/5 with no
unresolved findings and passing security scans. All final-head CI gates
passed: build, full Runner verification, typecheck/release registry,
canary, all general/serialized test shards, and all browser E2E shards
([CI
run](https://github.com/paperclipai/paperclip/actions/runs/34696739927)).
Local PostgreSQL startup contention required serialized retries; skipped
fixtures do not count as passing coverage.
- The earlier paid campaign passed all 24 chat cells and retained 32
screenshots:
[report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-34648511170-1/index.html?report=agent-chat#suite-agent-chat).
It tested `abacbdfd2f660709ec37312cdb758284c8399d04`; it is prior
evidence, not a paid run of this integrated head.
- Manual check: enable Agent Chat in Experimental settings, open an
agent, clarify and revise a plan, then hand off to an assigned project
task. Stop a reply, send `/new`, and verify fresh context with retained
history. Disable the setting and verify agent shortcuts/new chat turns
are blocked.

## Risks

- Queue/session integration can affect retries and delayed writes. Tests
cover ownership, cancellation, reset boundaries, idle recovery, and
ordinary task behavior.
- Migration 0274 adds conversation fields and constraints. Replay is
idempotent and preserves existing development chat history.
- This combines the previously reviewed stack at the maintainer's
request. Agent Chat remains off by default and is separate from
Conference Room.

## Model Used

OpenAI Codex, GPT-6 Astra (`gpt-6-astra`), with reasoning, code
execution, browser tools, and parallel review. The exact context-window
size is not exposed in this session. Codex and Claude also ran as test
subjects in the linked paid campaign.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-12 08:56:04 -05:00
DottaandPaperclip f12b647ae8 fix: reliably interrupt and resume legacy message queues (#13275)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - A task can collect more messages while its agent works.
> - Legacy runners must stop the active process before they can receive
those messages.
> - The old Interrupt action cancelled the run but could leave the queue
idle and hidden.
> - Codex could also classify a cancelled run as successful or start a
fresh process after cancellation.
> - This pull request joins cancellation, preserves the provider
session, and dispatches the current queue after cleanup.
> - The benefit is reliable interruption with the saved message order,
edits, and deletions.

## Linked Issues or Issue Description

**What happened?**

Interrupt could strand a legacy message queue. The UI could hide pending
messages after the run stopped. A Codex signal exit could race the
cancellation write. A stale session warning could also trigger a fresh
process after an interrupted resume.

**Expected behavior**

Interrupt stops the active turn and sends the remaining messages once,
in their saved order. Deleted messages stay deleted. An interrupted
Codex turn keeps its session and does not restart itself.

**Steps to reproduce**

1. Assign a task to a legacy Codex agent that runs a long command.
2. Queue three messages. Edit one, discard another, and move the last
message first.
3. Click Interrupt in the queue.
4. Repeat the interruption while the resumed session runs another
command.

Related work: Refs #13160, which moves native queue steering into the
wake-queue module. This change fixes legacy interruption and keeps
native steering unchanged.

## What Changed

- Add a revision-checked, company-scoped endpoint for legacy queue
interruption.
- Promote only the requested queue after the provider stops and releases
its lease. Retry its persisted interrupt intent from the scheduler after
a promotion error or server restart.
- Keep pending legacy queues visible after a run stops. Use server state
for the interrupt result.
- Serialize owned process cancellation before classifying the adapter
result. Preserve late session and log metadata. Acknowledge cancellation
only when an actual process or process group was owned; scheduler
placeholders retain their normal release policy.
- Send Ctrl-C to legacy Codex. Prevent missing-session fallback once the
session has started.
- Add cancellation race, multi-actor queue order, durable retry, resume
fallback, and stale request regression tests. Document the behavior.

## Verification

- Real browser tests passed with legacy Codex CLI and ACP engines, using
Codex 0.153.4 and gpt-5.6-sol.
- All three automated ACP browser scenarios passed locally: immediate
Interrupt delivery, no replay of an unfinished write, and pause
requiring Resume. Updated the old test expectation that required a
separate “go” after Interrupt.
- Browser tests covered queued edits, deletion, reordering, deleting the
final message, and repeated interruption.
- Two consecutive CLI interrupts kept one provider session. Both stopped
processes exited. The final message arrived once.
- `pnpm -r typecheck` passed.
- `pnpm check:token-gates` passed.
- All 346 post-review scheduling, recovery, queue-route,
archived-company, worktree-suppression, and stale-queue regression tests
passed.
- All 318 process-recovery and durable-chat tests passed after the final
cancellation guard.
- Codex adapter, queue UI, issue-page, and OpenAPI contract tests
passed.
- `pnpm build` passed.
- Full local suite coverage completed with
`PAPERCLIP_IN_WORKTREE=false`, using the stable runner and its CI
shards: 618 general server suites, all 145 serialized server suites, and
all workspace groups. Every failing suite passed a targeted rerun after
the fixes, rebuilding the native test fixture, correcting macOS
temporary-path setup, or retrying setup/timing failures. Existing skips
remain.
- The original monolithic run reported failures before the final fixes;
its failed suites were rerun rather than rerunning all 618 suites again.
The final process-recovery/durable-chat regression run passed all 318
tests.
- All CI checks passed for `e30eaf787f23a5511a3cb3cdb5abbccab9ed001d`:
[run 34654820774, attempt
2](https://github.com/paperclipai/paperclip/actions/runs/34654820774/attempts/2),
including typecheck, build, all test shards, E2E, and canary. The
signoff and Cursor sandbox tests each hit a timeout in the initial
attempt; both suites passed locally, and both failed shards passed their
single CI rerun. All three corrected ACP browser scenarios passed in CI.
- Greptile reviewed `e30eaf787f23a5511a3cb3cdb5abbccab9ed001d`: 5/5, no
open review threads.

## Risks

Cancellation order affects local adapters. The tests cover signal exits,
graceful exits, adapter exceptions, termination errors, and cancellation
write errors. Embedded adapters keep their cancellation controls.
Ordinary run cancellation and task pause keep their distinct queue
policies. No database migration is required.

## Model Used

OpenAI Codex, GPT-6, with reasoning, tool use, browser testing, and code
execution. The exact serving model ID and context-window size are not
exposed in this session. The live test runner used OpenAI gpt-5.6-sol.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-11 18:26:59 -05:00
DottaandPaperclip 2083bf6f9a feat(connections): add AgentMail inboxes and email tasks (#13256)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Connections give agents controlled access to external services.
> - Experimental channels already map conversations to tasks and durable
work queues.
> - Email needs inbox ownership, recipient envelopes, delivery records,
and explicit sends.
> - This pull request adds AgentMail to that infrastructure and keeps
the provider key in the server vault.
> - Agents can receive and send email from local or sandbox execution
while the board follows each conversation in its task.

## Linked Issues or Issue Description

**Problem or motivation**

Agents need dedicated email addresses. Incoming email should become
assigned work. Internal task comments and progress must never become
outgoing email by accident.

**Proposed solution**

Add experimental AgentMail connections, an inbox assignment wizard,
durable email intake and publication, task email cards, and
authenticated API, CLI, and native runtime actions. Agents use Paperclip
credentials to request sends. Paperclip owns the provider key and
enforces access and task authority.

**Alternatives considered**

A general mailbox MCP connector does not provide durable task binding or
publication boundaries. A separate mailbox application duplicates task
collaboration. The board instead directs the agent through the normal
task conversation.

**Roadmap alignment**

This extends the existing experimental connections and task
infrastructure. Product scope and interaction design were reviewed with
the maintainer. Related connection authority work: #11831 and #11818.
The duplicate search found no competing task-based AgentMail
integration.

## What Changed

- Add AgentMail catalog data, shared contracts, company-scoped email
records, and an additive migration.
- Add vaulted setup, inbox assignment, access grants, trust guidance,
and provider-side allowlist guidance.
- Support WebSocket and signed-webhook intake through a shared durable
pipeline, deduplication, catch-up, and task wakeups.
- Queue explicit new conversations and replies with immutable send
intents, idempotency, delivery state, and uncertain-send resolution.
- Show inbound and outbound email cards in normal task conversations.
Keep internal messages internal.
- Add task-scoped CLI actions and the sandbox callback routes required
for Daytona execution.
- Provide a dedicated AgentMail skill automatically only to agents with
active authorized inbox assignments. Keep email instructions out of the
universal Paperclip skill.
- Advertise connector-owned `agentmail_inboxes`,
`agentmail_read_thread`, `agentmail_send`, and `agentmail_delivery`
tools only in eligible native sessions. Recheck live authority on
execution.
- Isolate Codex CLI connector skills by agent and skill revision.
Deliver the assigned skill in the run prompt for adapters that use
shared skill directories, including resumed turns. Keep automatic skills
out of manual persistent sync. Show them as read-only and document the
pattern in the connector playbook.
- Fix AgentMail health checks that entered local-stdio validation and
optional missing Codex credential cleanup in sandboxes.
- Add API, pipeline, authorization, sandbox, browser, and Storybook
coverage.

## Verification

- Live AgentMail testing covered WebSocket intake, signed webhooks,
restart catch-up, and a full receive → task → Daytona Codex CLI →
explicit reply → Delivered round trip. The reply was verified in the
other inbox. The normal task composer also initiated an outgoing email
child task.
- The connector-skill change was verified in the browser: AgentMail
appears once as an automatic, read-only skill with its assigned address.
Disabling experimental chat connections removes it; re-enabling restores
it. A regression test covers assignment data arriving after library
data.
- Connector regression coverage passed 178 runtime utility, email
integration, skill-route, and heartbeat tests. All 17 Codex execution
tests passed, including per-agent skill isolation, model identity,
revision changes, removal, and prompt delivery without shared skill
files.
- After rebasing onto master, all 44 focused email, heartbeat, and
native-authority tests passed. All 313 native-session executor tests
passed. The UI regression suite passed all 3 tests. These test sets
overlap earlier focused runs.
- Full workspace typecheck and build passed after the rebase. Token
gates passed. Earlier focused Playwright task/setup coverage and the
Storybook build also passed.
- Native connector tool execution uses deterministic integration tests.
Live Daytona qualification used the Codex CLI adapter; the new
shared-home prompt fallback has deterministic coverage.
- The full repository suite is run by CI. The earlier unsharded local
full-suite attempt was stopped after the equivalent CI suites passed and
is not reported as a completed local run. Greptile reviewed
`7e57dc267a8446d3c906e3cc5b8abc94fb8860eb` at 5/5 with no unresolved
threads. All server, workspace, serialized server, and browser suites
passed in CI. The build job hit a five-second timeout in a runner
transport test; both variants and the full 80-test file passed locally
with unchanged timeouts. The build passed on retry on the same commit
without code or timeout changes. All required CI gates, including the
final `ci / verify` and `ci / e2e` summaries, are green on
`7e57dc267a8446d3c906e3cc5b8abc94fb8860eb`.

## Risks

- Email from external senders can start normal agent work. Setup
recommends a low-trust agent and AgentMail sender controls. Sender
addresses never grant board membership.
- Provider timeouts can leave uncertain sends. Retries retain their
idempotency key; expired windows require reconciliation or operator
resolution.
- Connector skills and native tools are assignment-dependent and require
current access. Revocation denies retained calls; assignment changes
select a new runtime context.
- Activation remains behind the experimental-channel setting. The native
runner path has deterministic coverage; live Daytona qualification used
the Codex CLI adapter.
- Schema changes are additive. Inbox ownership is unique across
companies. Disconnect preserves provider inboxes and task history.

## Model Used

OpenAI GPT-6 (Codex). Used reasoning, repository tools, code execution,
and browser testing. The exact deployment model ID and context-window
size were not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-11 16:56:38 -05:00
Nicky LeachandPaperclip ad4f0b5867 Fix Codex API key authentication in tests and runs (#13260)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Agent runtime settings can bind organization secrets to an adapter
environment
> - Paperclip redacts plain environment values when it returns a saved
agent to the UI
> - A saved-agent test sent the redacted `CODEX_HOME` value back to the
server
> - Codex ACP also received the API key without an ACP API-key
authentication request
> - This pull request restores saved environment values for tests and
selects API-key authentication for Codex ACP runs
> - The benefit is that Codex agents can test and run with an
organization-scoped OpenAI API key

## Linked Issues or Issue Description

**What happened?**

Testing a saved Codex agent sent `***REDACTED***` as `CODEX_HOME`.
Secret normalization rejected that placeholder. Remote Codex ACP runs
received `OPENAI_API_KEY`, but session creation stopped with
`Authentication required`.

**Expected behavior**

Paperclip must use the saved `CODEX_HOME` value when it tests an
existing agent. Codex ACP must select API-key authentication when
`OPENAI_API_KEY` is available.

**Steps to reproduce**

1. Create an organization-scoped secret named `OPENAI_API_KEY`.
2. Give a Codex agent access to the secret.
3. Save the agent runtime settings.
4. Test the saved agent again.
5. Run the agent in a remote sandbox through ACP.

**Paperclip version or commit**

Reproduced on master before commit
`68c17709d7c051a804a416263e2e08920f1dfcb1`.

**Deployment mode**

Self-hosted server with a remote sandbox environment.

**Installation method**

Built from source.

**Agent adapter(s) involved**

Codex.

## What Changed

- Send the saved agent ID with adapter environment tests.
- Restore redacted plain environment values from the saved agent before
test-time secret resolution.
- Select the Codex ACP `api-key` authentication method when
`OPENAI_API_KEY` is present.
- Add focused regression coverage for saved-agent tests and remote ACP
launch configuration.

## Verification

- `pnpm --filter @paperclipai/adapter-utils exec vitest run
src/acpx-engine/execute.test.ts`
- `pnpm --filter @paperclipai/server exec vitest run
src/__tests__/agent-adapter-validation-routes.test.ts`
- `pnpm --filter @paperclipai/ui exec vitest run
src/lib/test-agent-setup.test.ts`
- `pnpm -r typecheck`
- `pnpm test:run`
- `pnpm build`
- `git diff --check`

## Risks

- Low risk. The test route reads saved configuration only when the
request supplies a compatible agent ID and the caller can update that
agent.
- The Codex ACP change applies only when `OPENAI_API_KEY` exists and no
explicit `DEFAULT_AUTH_REQUEST` exists.
- There are no schema migrations or telemetry changes.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- OpenAI Codex with `gpt-5`. The context-window size is not exposed in
this runtime. The model used reasoning, repository search, file editing,
command execution, and test execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-11 13:58:29 -07:00
DottaandPaperclip 7b829efdf6 feat: show tasks created from a task by project (#13241)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - A task can cause an agent to create more tasks.
> - Those tasks can belong to other projects or have another parent.
> - The subtask view does not show all work created from the current
task.
> - This pull request adds a Tasks tab with separate subtask and
creation groups.
> - Operators can follow created work without changing its parent or
project.

## Linked Issues or Issue Description

**Problem or motivation**

Operators need to see all work that an agent creates while running for a
task. Parentage alone does not describe this relationship. Legacy and
native runs must follow the same rules.

**Proposed solution**

Show all subtasks in one section. Separately group tasks created from
the source task by their current project, with a No project group when
needed. A created subtask appears in both sections. Use saved run
context and recorded creation activity to find the source task.

**Alternatives considered**

Making all created tasks children would change their hierarchy. Removing
overlap between sections would hide the creation relationship. This
change keeps the two memberships separate.

**Roadmap alignment**

This extends the existing Activity log & action attribution capability.
It does not add a new roadmap area. Related PR: #9727 adds a stored
source-task field and inbound attribution UI. This PR adds the outgoing
task list using existing run and activity records and does not require
that schema change.

## What Changed

- Add a company-scoped createdFromIssueId filter to issue lists.
- Save the actor run during task creation, including legacy child-helper
calls.
- Recover historical run attribution from creation activity when the
origin run is absent.
- Render the production Tasks panel with all subtasks and independently
grouped created work.
- Keep progress only for subtasks. Add folding, hover fades and project
links.
- Fetch all result pages and refresh on issue activity. Show load
failures with Retry.
- Add database, API, UI and pagination tests, design-guide examples and
Storybook pages.

## Verification

- Before rebase: 158 targeted tests passed. Workspace typecheck,
UI/server builds, Storybook build and token gates passed.
- After rebase: full workspace typecheck and build passed. The cursor
fix passes 26 focused tests and UI/server typechecks.
- The full local test command completed its general-server group with
10,563 passing tests and two failures: a missing native-runner fixture
and a concurrency-test timeout. Building the fixture and rerunning both
affected files passed all 41 tests. The local command stopped before its
remaining groups; all corresponding GitHub test shards passed on the
submitted head.
- GitHub checks on commit 3b4bf0bd8: 31 passed, including the aggregate
CI gate, all server/workspace/browser test shards, typecheck, build,
runner verification, canary dry run, policy and security checks. Two
conditional Storybook jobs were skipped by the workflow.
- Greptile reviewed commit 3b4bf0bd8 at 5/5 with no unresolved comments.
- Open Storybook at UX Labs / Tasks Created From a Task / Full Task
Page. Check that a created subtask appears in both sections. Fold each
group and use a project link. Check the No project group and first-task
arrival story.

## Risks

- Old tasks without a saved origin run or attributed creation activity
cannot be linked to a source. The code does not infer a source from a
shared creator or a comment.
- The new list filter reads run and activity records. Source, run,
activity and result stay within the requested company.
- Task creation now saves the actor run when no explicit origin run is
supplied. Existing explicit origins remain unchanged. No database
migration is required.

## Model Used

- OpenAI GPT-6 through Codex. The runtime identifies the model family as
GPT-6 but does not expose a more specific model ID or context-window
size. Used reasoning, repository tools, code execution and browser
inspection.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-11 12:39:08 -05:00
DottaandPaperclip 889947c238 feat: add experimental native chat connectors (#13038)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - People also ask agents for work in their existing chat tools.
> - Each external conversation needs one task and a current authorized
source.
> - Retries, Stop, and provider failures must not duplicate work or
expose private data.
> - The first chat PR establishes the opt-in provider and data
contracts.
> - This PR adds experimental channel integration and its durable
control plane.
> - Users can request work from connected channels and inspect delivery
in Paperclip.

## Linked Issues or Issue Description

Refs #13100 and #13092. This is the second of exactly two chat PRs.
Foundation #13100 is merged and changed 143 files. Runner prerequisite
#13092 is also merged. This PR changes 400 files against master, below
the 500-file review limit. It contains no wireframe images or HTML
galleries.

## What Changed

- Add native Slack, GitHub, Microsoft Teams, Telegram, and Discord chat
connections. Keep chat disabled unless the operator enables experimental
chat connectors. Preserve the production GitHub tool connection and its
normal setup path.
- Bind each provider bot identity to one immutable Paperclip agent. Bind
each admitted external conversation to one task. Paperclip owns tasks,
runs, permissions, and audit records.
- Add durable admission, per-conversation queues, questions, task
controls, progress, final replies, images, files, and delivery receipts.
Board comments remain internal unless explicitly sent to the channel.
- Check current identity, provider reach, resource access, credentials,
runtime generation, and exact source before provider effects. Keep
private responses private. Never send raw reasoning, private logs,
credentials, or tool arguments.
- Hold uncertain sends for explicit audited resolution. Make Board
Send-to-channel atomic and idempotent. Keep reconnect and setup
credentials in Paperclip secret storage.
- Preserve current native-runner authority across retries, lost
acknowledgements, and recovery. Keep immutable input and completion
contracts separate from newer user input. Receipt reconciliation cannot
launch a provider.
- Reconcile chat close/new ordering and provider-effect lock order.
Audit resource access changes in the same transaction. Submit only the
selected resource from each UI toggle so stale pages cannot undo
unrelated access changes.
- Drain Codex stdout before certifying process exit. Bound the drain
with the existing shutdown grace. Preserve observed terminal authority
without treating an undrained process as successful or reusable.
- Incorporate master `018ca5da` with its ACP Stop, mobile task layout,
runner packaging, and official lock changes. Preserve dedicated
chat-answer continuations in both directions when ordinary queued
comments are adopted after Stop.
- Fence late adapter readiness behind an earlier Stop for the same run.
Preserve verified cleanup for registered adapters. Handle single Stop,
agent pause, duplicate Stops, and failure release without creating a
false cancellation receipt.
- Incorporate master's `6dd48cad4` wake-queue extraction. Preserve exact
failed-chat retry authorization and lineage, retired question-source
suppression, and the block on generic recovery that would discard the
admitted source. Fresh deferred input retains its separate promotion
path.
- Incorporate master `2a05b5ed3` and its queue-admission extraction,
simplified transaction ports, and separate runner CI job. Preserve exact
durable receipts, actor separation, and dedicated-answer isolation
through the new module. A failed receipt insert rolls back the
accompanying deferred-wake merge.

## Verification

Current head: `afe19299d06253cb628eb398e91d1200ea9f412a`, incorporating
master `2a05b5ed3457ea33efd6895520447d1d97fe98d8`. The conflicts are
resolved. This successor fixes two test-harness boundaries exposed by
CI: per-case route-module preparation and actual durable-save completion
before intentional runner termination. Production code and all existing
test/turn deadlines are unchanged. [Exact-head Greptile
review](https://github.com/paperclipai/paperclip/pull/13038#issuecomment-5587250594)
is **5/5**, completed September 10 at 13:20:55 UTC, with no actionable
findings or open review threads. [Fresh exact-head
CI](https://github.com/paperclipai/paperclip/actions/runs/34481724341)
passes **all 24 jobs**, including Build and both required aggregates.
Normal exact-head guarded merge was attempted and rejected by the
remaining branch approval policy: CODEOWNER review is required and no
human approval is present. Normal **squash auto-merge is enabled** as of
September 10 at 13:36:26 UTC. Requested CODEOWNERS have been notified;
no approval bypass or self-approval was used. Earlier-head results below
remain historical evidence, not qualification of this successor.

- Final exact-head Linux evidence: 995/995 chat integration cases; 36/36
agent-skills routes; 35/35 runner live-session cases, including real
process kill/resume; 1948 runner Vitest cases with three existing
benchmark/platform guards; 870/870 API-authority cases; and 104 browser
cases with four existing optional skips. Rust, conformance/replay, full
repository build, typecheck, canary, all server/workspace shards, and
both required aggregates pass with normal CI concurrency. Earlier failed
attempts remain recorded below.

- Latest test-only qualification: 141/141
route/permissions/authentication cases pass in separate cold forks, with
plain server types and independent review clear. The real-runner suite
passes 35/35, with plain runner types and independent review clear. A
controlled premature-save acknowledgement fails as expected; matching
ownership/effect/process evidence, rejected saves, real turn outcome,
test abort, and pre-kill liveness are covered. No local reproduction of
the original CI scheduling failure is claimed. The preceding [CI
run](https://github.com/paperclipai/paperclip/actions/runs/34479680858)
passes 21/24 jobs, including all 995 Linux chat cases and browser
aggregate (104 passed, four existing optional skips); only Build, the
skills serialized shard, and the required verification aggregate fail.
Its exact-head Greptile review was 5/5. Both failed job logs are
retained.

- Final fixture qualification: all eight focused Discord cases and all
995 chat integration cases pass. The exact modal statement/PID is
observed before taking the real connection lock; the test then proves
its actual blocking relationship before mutation. Original SQL
execution, provider behavior, negative assertions, and 1s/15s timeouts
remain unchanged. Independent review is clear and test/production hashes
remain frozen. The preceding [CI
attempt](https://github.com/paperclipai/paperclip/actions/runs/34477184777)
passed 22 jobs, including Build/runner, typecheck, canary, all other
test shards, and browser aggregate (104 passed, four existing optional
skips); the two fixture failures and failed verification aggregate
remain recorded, not relabeled as a pass.

- Current queue-module composition: 308/308 recovery/batching/queue/Stop
tests; 995/995 full chat integration; 89/89 module tests, including real
PostgreSQL receipt-insert rollback; 24/24 workflow/module-boundary
tests; plain server and UI types. All four actual local process/ACP
browser paths pass in 1.4 minutes. Fresh databases, no skips or retries,
stable reviewed source hashes. The initial boundary failure is retained;
its no-op service wrapper was removed without changing recovery context
or weakening the check. An exploratory standalone test-directory
typecheck fails because its new upstream transformation config is not a
standalone typechecking project; standard CI/build does not invoke it,
and no configuration was weakened to suppress those diagnostics.

- The preceding head `e02a63d462ce5d47433b0aeb632bb6fd20aab1ba` passed
[all 24 CI
jobs](https://github.com/paperclipai/paperclip/actions/runs/34436462958)
and exact-head Greptile review at 5/5. Required CODEOWNER review
prevented its normal merge before master advanced again.

- Final extracted-module composition: 307/307 recovery, batching, queue
and Stop-control tests; 995/995 full chat integration; 49/49 module
tests including eight PostgreSQL adapter cases; and 19/19 issue-update
tests. Plain server types pass. All four actual local process/ACP
browser paths pass in 1.3 minutes. Fresh databases, no skips or retries
in these cohorts, frozen source hashes, and independent review clear.

- The preceding head `3e4e1c1c` passes [all PR CI
jobs](https://github.com/paperclipai/paperclip/actions/runs/34415826820),
including Build and required `ci / verify` and `ci / e2e`. Both the
original Rust failure and the previously load-sensitive lineage fixture
pass with unchanged Linux concurrency. Master advanced afterward and
required this reconciliation.
- Final master composition: 448/448 focused UI tests, 186/186 adapter
tests, 24/24 queue/control tests, and 11/11 packaging tests. Plain UI,
server, shared, and adapter types pass. Token gates and diff checks
pass. Independent server and UI reviews are clear.
- Stop-registration regression: both real-service cases fail against
exact `a95` source and pass with the fix. The full corrected
recovery/control suite passes 265/265. Duplicate-owner and failed-Stop
controls also pass. Plain server types pass. The readiness barrier
prevents provider startup without adding an acknowledgment to an already
terminal run.
- Final qualification strengthens terminal-field equality and repeats
both affected cases successfully on a fresh database. All four actual
local process/ACP browser paths pass again in 1.3 minutes, without skips
or retries. The final screenshot shows Cancelled, a paused subtree,
retained input, and no error toast.
- Two new actual-service regressions fail before the merge fix. They
prove that queued-comment adoption could consume a dedicated chat answer
or add unrelated input to that answer. The fixed four-case cohort
passes, including ordinary upstream continuation and adapter Stop
controls. Full recovery passes 257/257. All four actual local
process/ACP Stop browser flows pass in 1.4 minutes, without skips or
retries, on a fresh database.
- The unchanged runner artifact was qualified with 171/171 transport
tests, 870/870 API-authority tests, conformance 1/1, and replay 11/11.
Six controlled reader tests prove the exit/drain repair. Its local
serial Rust workspace passed 546 top-level cases plus two invoked
helpers; the later passing Linux CI supplies default-concurrency
evidence.
- Prior exact-source full chat integration passes 995/995. Settings
regressions cover concurrent stale pages, 501 destinations, pending
state, rejected updates, and explicit retry. These deterministic tests
do not prove live provider behavior.
- Retained failed attempts and their causes are in the [qualification
log](https://github.com/paperclipai/paperclip/blob/afe19299d06253cb628eb398e91d1200ea9f412a/doc/plans/chat-adapters/2026-09-08-chat-queue-and-webhook-repair.md).
The first merge adapter run timed out while macOS slept for 290 seconds.
Its unchanged repeat passed with a temporary sleep guard. No assertion,
deadline, or CI gate was weakened.

Review commands include `pnpm --filter @paperclipai/server exec vitest
run src/__tests__/heartbeat-process-recovery.test.ts
src/__tests__/issue-queued-comments-routes.test.ts` and `pnpm exec
playwright test --config tests/e2e/playwright.config.ts
tests/e2e/acp-stop-continuation.spec.ts`. Database suites require fresh
disposable databases. See the [browser
runbook](https://github.com/paperclipai/paperclip/blob/afe19299d06253cb628eb398e91d1200ea9f412a/doc/plans/chat-adapters/2026-09-04-chat-adapters-browser-e2e-runbook.md)
for provider setup and separate live acceptance steps.

## Risks

- This remains experimental. Deterministic tests and bounded live
evidence do not establish every provider feature, tenant, permission
layout, or media shape. Teams work-tenant qualification is still open.
- Failed and uncertain provider effects remain visible and can require
operator action. A transport receipt does not prove recipient
visibility.
- Native controller and runner artifacts must remain compatible.
Preserve lease ownership, terminal authority, source binding, and
quarantine during future changes.
- Access and audit rows commit together, but activity notifications
remain best-effort. This is not a new durable event outbox.
- The PR operation does not deploy a live server, replace its runner, or
change provider permissions. Remaining live qualification is documented
in the [temporary
handoff](https://github.com/paperclipai/paperclip/blob/afe19299d06253cb628eb398e91d1200ea9f412a/doc/plans/chat-adapters/2026-09-08-open-qualification-followups.md).

## Model Used

OpenAI Codex assisted with implementation, tool execution, testing, and
review. The work records `gpt-6-astra` assistance. The environment does
not report a context-window size. No private reasoning traces are
included.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-10 10:06:45 -05:00
DottaandPaperclip 8cfd30fb07 feat(ui): add composer Stop and simplify task controls (#13104)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The task composer is where operators direct running agents.
> - Operators need to stop work without leaving the conversation.
> - Existing pause controls already hold task trees and interrupt both
runner types.
> - This pull request connects the composer to those controls and
removes repeated feedback.
> - Operators can pause work quickly and still queue messages while
agents run.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

Task pause, resume, and cancellation in the task page and composer.

**Current behavior**

The empty composer cannot stop a running task. Task controls require
extra confirmation and reason text. Pause can show several notifications
for the task already on screen.

**Proposed behavior**

Show Stop while this task runs and the composer is empty. Text or
attachments switch it to Send. Stop and the menu use the same manual
pause hold. Parent pauses include descendants. Keep task cancellation in
the menu with a compact confirmation. Show one quiet pause row and gray
cancelled-run details.

**Reason and benefit**

Operators can interrupt execution with one click. Drafts and queued
messages keep their existing behavior. The UI waits for actual
termination, including native cancellation acknowledgment.

**Breaking changes**

No endpoint, schema, or task-status change. Pause no longer asks for
confirmation or a reason. Resume now honors the existing wake-agents
option. Task notifications are suppressed for the task and subtree
currently in view.

Related UI work: #8228 changes navigation and composer shortcuts. This
PR covers execution controls. No duplicate Stop-button PR was found. The
change improves existing controls and does not duplicate a roadmap
milestone.

## What Changed

- Add Stop, pending feedback, duplicate-click protection, and inline
errors to the composer.
- Share the pause mutation across the composer, active-run controls, and
menu.
- Poll affected runs after a pause request. Require native cancellation
acknowledgment.
- Remove pause confirmation and shared reason fields. Reduce cancel
confirmation to its task count and actions.
- Honor wake-agents for executable tasks only. Preserve the pause when
recovery review is needed; show partial wake failures inline.
- Preserve explicit legacy reconciliation decisions while their
continuation waits for dispatch.
- Suppress notifications for visible task trees. Use quiet pause and
cancellation feedback.
- Add interactive stories using production controls and native/legacy
end-to-end tests.

## Verification

- User reviewed the running feature and revised Storybooks in the
browser.
- Rebased focused checks passed: 295 original targeted tests, 161
updated route/page/notification/status tests, and 26 recovery
integration tests.
- Both isolated runner journeys pass on the final revision (1.7
minutes). Coverage includes queueing, parent and child interruption,
persisted holds, no automatic continuation, reconciled resume,
cancellation, terminal exclusions, and no Stop toast.
- Native coverage uses real runnerd with a deterministic provider
fixture. Legacy coverage checks actual process termination. Live
hosted-provider execution was not tested.
- Repository typecheck and build, Storybook build, and token gates
passed after rebase. The final server typecheck/build also passed.
- The broad local run completed its general-server stage with 7,219
passing tests, 48 skipped, and two failures from cached pre-fix source
and a stale native provider fixture. Both failed tests pass in fresh
final-head reruns after rebuilding the fixture; the script did not
continue to its later local stages. CI runs all test groups on the final
revision.
- Final revision: all 31 applicable CI checks passed; Storybook visual
regression was skipped by its workflow conditions. Greptile: 5/5, zero
unresolved comments.
- Review `Tasks / Execution Controls` in Storybook. Type and clear a
draft, stop a run, expand cancellation details, and test the menu on
desktop and mobile.

## Risks

- Stop pauses descendants for a parent task. This is the existing pause
contract.
- A held task can remain active if interruption fails. The UI shows an
error instead of claiming termination.
- Resume can start multiple assignees when wake-agents is selected.
Backlog, blocked, and terminal tasks stay excluded. Existing execution
reconciliation remains mandatory where required; Resume never invents
action-outcome evidence.
- Notification suppression uses the visible task and cached subtree.
Notifications for unrelated work remain enabled.

## Model Used

OpenAI GPT-6 through Codex. The exact runtime snapshot and
context-window limit are not exposed in this session. Used reasoning,
tool calls, code execution, and browser inspection.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-09 12:18:56 -05:00