mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-09 16:00:03 +02:00
codex/hermes-native-runner
2091
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
ab98588317 |
Merge bounded selection and verified cold CLI checks
Co-Authored-By: Paperclip <noreply@paperclip.ing> * codex/hermes-routines: fix(ui): bound manual agent selection to the switch transition |
||
|
|
30d57b2971 |
fix(ui): bound manual agent selection to the switch transition
Reproduce the page/sidebar mismatch after a real manual switch and browser history navigation. Dismiss covering announcements through the normal UI. Preserve dense ZIP bytes with a CRC table and bound cold CLI health inside the existing setup budget. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
295488bf19 |
Merge verified company isolation and runtime teardown fixes
Co-Authored-By: Paperclip <noreply@paperclip.ing> * codex/hermes-routines: fix(ui): scope canonical agents to their authorized company test(server): drain runtime writes and load routes in setup |
||
|
|
a7843be973 |
fix(ui): scope canonical agents to their authorized company
Keep alias caches and redirects on the agent company. Defer remembered UUID paths until canonical resolution and verify organization switching with a delayed public alias response. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
cf595b8aa3 |
Merge branch 'codex/hermes-routines' into codex/hermes-native-runner
* codex/hermes-routines: fix(ui): preserve exact retry feedback across agent aliases |
||
|
|
fb725383b9 | fix(ui): preserve exact retry feedback across agent aliases | ||
|
|
c834f9dd8a | fix(hermes): release credentials and reproduce relocatable runtime assets | ||
|
|
9e07748e91 |
feat(runner): add native Hermes ACP qualification candidate
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
dd777f4b73 |
feat(dot): complete onboarding and expand governed Runner capabilities (#15414)
## Thinking Path
> - Paperclip manages AI agents, their work, and their permissions.
> - Paperclip Runner supplies the same admitted tool authority to each
provider.
> - The Dot provider in #15402 needs reliable onboarding and useful
agent capabilities.
> - An idle Dot could not start work, assign a human, read skills, or
produce a workspace artifact.
> - Pairing also relied on a second configuration save before ordinary
admission could work.
> - This pull request adds governed idle admission and shared Runner
tools, and completes pairing atomically.
> - Operators can test Dot with fresh data while keeping normal company,
approval, budget, and run ownership checks.
## Linked Issues or Issue Description
**Subsystem affected**
Paperclip Runner, dedicated Dot MCP access, OAuth onboarding,
experimental settings, and empty worktree startup. This PR builds on
merged provider PR #15402. It reuses the merged MCP gateway from #14846
and assistant connection work from #14933 and #15380.
**Problem or motivation**
An idle Dot could see assignments but could not act on a conversation
request until someone created a task first. Its Runner catalog could not
assign tasks to humans or read pinned skills and workspace files.
First-time OAuth discovery and pairing also needed browser fixes, and a
completed pairing did not persist its binding reference on the agent.
**Proposed solution**
Keep Dot within the existing Runner. Admit a visible agent-authored
intake task for idle requests. Add people, human assignment, cross-task,
skill, and optional sandbox workspace tools through shared authority.
Relay assigned app calls through the configured MCP gateway. Add lease
renewal and follow-up references. Save pairing and its configuration
revision atomically. Show prerequisites and provide a complete copy
prompt.
**Roadmap alignment**
This extends the experimental provider in #15402. It uses the existing
governed gateway, task model, skills, artifact path, and native Runner.
It adds no separate execution subsystem.
## What Changed
- Add a standalone OpenAI Dot agent choice with an independent
experimental opt-in. It works with the general Runner option off.
Require Assistant connections (MCP), authenticated sign-in, and public
HTTPS for pairing.
- Use Dot’s own option for company package import. Allow an unpaired Dot
configuration to save after external billing acknowledgement; task
admission still requires pairing. Prepare the shared dev binary for
Dot-only opt-in.
- Label saved Dot agents as OpenAI Dot. Use prerequisite-check copy in
setup and runtime configuration.
- Route Dot creation directly to pairing after explicit external billing
acknowledgement. Hide local CLI, model, and harness setup for Dot.
Preserve the shared Runner implementation and canonical API type. Other
Runner providers still require their general opt-in.
- Add an empty worktree option with fresh signing keys and no production
data copy.
- Fix public-client OAuth negotiation, discovery compatibility, and
optional separate browser authorization origin.
- Add one-use pairing consent preview, clear copied setup instructions,
and atomic binding persistence and cleanup.
- Add idle request admission with stable request IDs and normal
scheduling, permissions, budgets, and task ownership.
- Add identity and people discovery, human task assignment and
reassignment, and authorized cross-task comments and documents.
- Read assigned skill files from pinned manifests. Relay assigned app
calls through the merged gateway without exposing credentials.
- Add an off-by-default workspace bridge. Constrain paths and writes.
Run commands in a deny-by-default OS sandbox with no network or injected
credentials. Reserve mutations before effects and never blindly repeat
uncertain work.
- Add rolling lease renewal, task pagination, bounded operation limits,
and deduplicated follow-up references without comment bodies in
webhooks.
- Add an off-by-default attachment reading setting. Restrict reads to
files on the current assigned task. Verify size and hash, cache bounded
verified copies per run, paginate text or binary bytes, and recheck live
authority before returning.
- Keep file grants operator-owned. Reject agent self-grants across
configuration routes. Preserve attachment consent in create/import
forms. Close generic API file bypasses while retaining current-run
response snapshots and permitted uploads.
- Keep provider limits explicit. Do not inherit a Dot binding,
attachment permission, or workspace permission when hiring another
agent.
- Include the required Markdown format in cross-task document writes and
validate the API title limit. Verify real creation and revision
persistence.
- Restrict command execution to Linux bubblewrap with descendant
containment. macOS retains workspace file tools and artifact publishing,
while refusing command calls. Explain the platform limit in setup.
- Exclude Paperclip instance state from workspace files, uploads,
artifact publication, and sandbox commands. Protect nested directories
and case variants. Fail closed when the directory protection scan
exceeds 4,096 directories.
- Add static UI compression for slow public tunnels. Document setup, the
complete tool inventory, and qualification limits.
## Verification
- Merge preparation on `00ca2c75b` integrates merged base #15402 and
master `fc6304dfe`. The ancestry commit preserves the reviewed follow-up
source tree. The subsequent security fix excludes instance state from
file tools, uploads, artifact publication, and sandbox commands. It
preserves private task authorization and task monitors. It regenerates
the combined tool catalog, seeded catalog digest, and protocol manifest.
Workspace typecheck, full build, and UI token gates pass. All sixteen
real Dot broker cases and 93 company import cases pass after the review
fixes and creator-attribution test correction. The exported
human-assignment catalog and Unicode page boundaries are also fixed. All
ten catalog tests and nine workspace/skill bridge tests pass. All 35
workspace bridge and authority tests passed after the instance-state
fix, including real macOS commands in that intermediate version. The
subsequent document and descendant-containment fixes pass 44 focused
tests across bridge, authority, and setup UI, with six Linux command
cases skipped on macOS. Real cross-task documents pass the route
validator and persist two revisions. macOS refuses command execution and
does not advertise the tool. Full workspace typecheck and build, changed
server/UI typechecks, and token gates pass again on the final commit.
Greptile rates final head
|
||
|
|
fc6304dfe5 |
feat(runner): add experimental OpenAI Dot provider over MCP Events (#15402)
## Thinking Path > - Paperclip manages AI agents, tasks, permissions, and execution budgets. > - Paperclip Runner gives each provider the same admitted task and tool authority. > - OpenAI Dot runs outside the local process tree and needs asynchronous work delivery. > - The merged MCP gateway supplies OAuth consent and signed event delivery. > - A personal assistant grant cannot safely stand in for an assigned agent. > - This pull request adds a separate Dot agent connection and a durable Rust Runner bridge. > - The operator can assign work to Dot and inspect its accepted work, tool receipts, and result. ## Linked Issues or Issue Description **Agent or provider** OpenAI Dot, as an experimental provider of the existing Paperclip Runner adapter. **Why this adapter is useful** An operator can assign normal Paperclip tasks to an existing Dot. Dot can read its mailbox, request work on an assigned task, use admitted task tools, and submit a result. Paperclip keeps company scope, checkout, approvals, known budget limits, and activity attribution. **How the agent is invoked** A dedicated `/mcp/runner` OAuth resource pairs one Dot grant with one agent. A signed MCP mailbox event wakes Dot. Dot explicitly accepts the assignment. The Rust Runner owns the durable turn and operation receipts. The first release supports self-hosted instances with a local Runner controller. **Additional context** This extends the merged public MCP gateway from #14846 and the assistant invitation and device-consent work from #14933. This also integrates the merged assistant tool and configuration expansion in #15380. Dot retains its dedicated agent resource and cannot receive personal configuration permission. The public assistant connection remains a personal connection. ## What Changed - Add a durable Rust Dot provider and its TypeScript Runner driver. - Add closed PRP v3 external-provider operations and native execution input v6. - Add company-scoped pairing, mailbox, assignment, and operation records. - Reuse merged browser/device consent, client metadata verification, webhook admissions, refresh, secret rotation, and warm-standby gates. - Keep Dot scopes, issuer, grants, event workers, and tool access separate from personal assistant access. - Add Dot configuration, pairing, readiness, and consent UI. Keep agent grants out of the personal Connections entry. - Regenerate the Dot-only migration after master. Preserve published gateway migrations. Make the new migration safe to reapply. - Document setup, recovery, accounting limits, evidence, and remaining account qualification. - Reverify reconnect callbacks and wake outstanding work with a fresh mailbox reference; preserve the existing assignment and operation receipts. - Clean up Dot bindings and waiting runs on OAuth revoke and refresh-token replay. Old grants cannot revoke replacement bindings. - Restore the pairing reference when an unsaved agent form is reopened; document board-only pairing routes in OpenAPI. - Accept a clean Rust exit after the acknowledged shutdown receipt. Unexpected exits still require recovery. - Clear the cached binding after a successful revoke so a failed connection refresh cannot restore it. - Add production-component Storybook states and screenshots for pairing and connection review. All preview account data is synthetic. - Persist normalized completion, serialize Dot turns and durable work admission, and poll subscription readiness. - Serialize mailbox writes and cursor reads; retain paused fence acknowledgement without task authority. - Authorize admitted review runs without changing the worker assignee. Include the fenced assignment ID in production stop notices. ## Verification - This PR integrates master `4a8178e9c`. Dot migration `0317_messy_famine.sql` follows the published history and is safe to reapply. The merge preserves the reserved migration connection, batch-commit handling, private task checks, task monitors, and native accounting. - Local workspace typecheck, full build, and UI token gates pass. The server typecheck passes after the review fixes. Database and native executor regressions pass. - All twelve real Rust/PostgreSQL Dot integration tests and twelve Dot driver tests pass. The tests cover native document writing and finalization, durable replay, queue admission, mailbox ordering, admitted reviews, stale authority, production stop references, and paused acknowledgements. - Current head `d0e7e0626` passes all 57 checks: 53 pass and four are intentionally skipped. This includes full typecheck, build, tests, Rust Runner verification, browser E2E, release verification, and Canary Dry Run. Greptile rates this exact head 5/5. All review threads are resolved. - The full local root test run is slower than the sharded CI run and has not completed. The full CI test gates pass on the current commit. Focused local regressions pass. - Real-account pairing and event delivery on this base commit remain unqualified. Live account and setup proof are recorded in the follow-up #15414. The following screenshots use synthetic preview data. They show the production pairing component and do not qualify a real account or the full agent setup journey.   ## Risks - This base adapter uses `PAPERCLIP_ENABLE_OPENAI_DOT=1` plus Public MCP and Paperclip Runner. The separate experimental-settings follow-up in #15414 replaces this environment flag with saved operator settings. - Dot does not expose provider token usage or cost. The operator must acknowledge external billing. Known Paperclip budget gates still apply. - Cancellation fences Paperclip authority. It does not confirm that Dot stopped all external activity. - Assigned skill files and third-party MCP bindings are unsupported and reject admission. There is no mounted workspace, model selector, or provider thread identifier. - Hosted agent-broker and remote controller deployments are not qualified. - The new migration follows the merged master history. Existing prototype databases still need the normal master migration history before this Dot-only migration. ## Model Used OpenAI Codex, based on GPT-6. The exact deployment ID and context window size are not exposed in this session. Capabilities used: reasoning, repository editing, code execution, and test inspection. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #123` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub references) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4a8178e9cf |
fix(ui): stabilize composer model selection and refine effort slider (#15444)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Task composers let users select an agent, a model, and an effort level. > - The model label changes width when users move the effort slider. > - This movement makes the control harder to use. > - This pull request keeps the label fixed while the picker is open. > - A larger slider gives clearer feedback during selection. ## Linked Issues or Issue Description **What existing behavior does this improve?** The shared model and effort picker in task composers. **Current behavior** The trigger shows each effort change while the picker is open. Different label lengths move the control. The slider is small. **Proposed behavior** The trigger shows “Select model” while any picker view is open. It shows the selected model and effort after the picker closes. The slider has a thicker track and a larger thumb with hover and press feedback. **Reason and benefit** Users can change effort without moving the trigger. The larger thumb gives clearer pointer feedback. ## What Changed - Keep the trigger text fixed while the picker is open. - Add slider size and feedback tokens. - Remove the gray input outline and thumb border. Show keyboard focus on the thumb. - Keep a thumb focus outline when Windows high contrast suppresses shadows. - Keep native keyboard and pointer controls. Respect reduced motion. - Test open and closed labels on desktop and mobile. ## Verification - Picker tests: 27 pass. Composer settings and new-task tests: 86 pass. - UI typecheck and UI build passed before the final CSS-only high-contrast fix. - Token gates pass. - Full typecheck and build stop at installed Runner API type mismatches outside these files. - Revised Chromium checks pass for desktop pointer drag, no input outline during drag, stable open label, keyboard input, and mobile rendering. The earlier rendered-image check confirmed reduced-motion behavior. - The earlier full local suite did not finish. The approved visual revision passed 113 focused tests. The final high-contrast fix passed 27 picker tests and token gates. - Chromium high-contrast and reduced-motion verification shows a visible thumb focus marker and working keyboard input. - Open the model picker. Change effort. Confirm “Select model” stays visible. Close the picker. Confirm the selected model and effort appear. - All CI gates pass on final head `7a64104e14251a354b06b7916d09d44692557c4f`, including full typecheck, tests, build, and browser E2E. Greptile: 5/5 with no findings. ## Risks - Low risk: shared UI only. Native range controls stay in place. - Browser thumb rendering can differ between engines. ## Model Used - OpenAI GPT-6 Codex. Code editing, tool use, and test execution. The runtime does not expose an exact model variant or context window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a7a244ab33 |
feat: add company decision models with permission and cost controls (#15473)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Optional product features need small, typed model decisions. > - Each company needs to choose which shared API connection pays for those decisions. > - Calls must retain user and task permissions, budget limits, and cost attribution. > - This pull request adds a managed decision service, setup UI, and request history. > - Features can check availability cheaply and keep their existing behavior when decisions are unavailable. ## Linked Issues or Issue Description **Subsystem affected** Server services, shared contracts, database accounting, Company Settings, and Costs. **Problem or motivation** Paperclip has no common decision-model service. Adding provider calls within each feature would duplicate credential access, permission checks, and billing rules. **Proposed solution** Let a connection manager configure one company decision model. Support OpenAI Decisions and Jev through OpenRouter. Provide a fixed setup test and metadata-only history. Default company-sponsored background decisions to on during setup, and preserve a saved off setting. **Alternatives considered** Per-feature credentials would duplicate existing connection management. Personal overrides and provider fallback chains add permission and billing complexity; they remain deferred. **Roadmap alignment** Reviewed ROADMAP.md and searched open PRs. This extends existing connection access and budget accounting. Product features that call the service remain outside this change. No matching decision-model service PR was found. ## What Changed - Add company settings, an internal `decisionModelService`, local availability checks, and trusted human, agent/run, and system contexts. - Pin Vercel AI SDK provider dependencies and adapt boolean, choice, and ordered-score decisions for both providers. Bound requests and time; disable paid retries. - Add durable invocation metadata and agentless decision ledger charges. Preserve fractional cents, pricing evidence, dispatch identity, and unresolved billing holds. - Share the company accounting lock and apply company, agent, and project budgets. Settle charges once, retain unknown holds, and recover interrupted calls without resubmission. - Reuse connection setup and management UI. Add a Decisions view under Costs, production-component Storybook coverage, database migration, and service documentation. ## Verification - Passed 166 current-code tests covering the decision service/provider, setup component, Costs, OpenAPI, and every failure from the earlier broad run. Coverage includes native SDK wire formats, refusals, billed malformed responses, permission and secret-rotation races, identity changes, concurrent budget admission, unresolved holds, agent/task deletion, and stale setup feedback. - Passed 170 existing connection, cost, budget, heartbeat-accounting, and profile regression tests. - Passed repository typecheck, production build, and design token gates after integrating master. Verified the generated migration on a fresh test database and upgraded the populated preview database from the branch's earlier migration without losing settings or usage. - Ran the required full `pnpm test:run`: its general phase completed with 16,354 passed and 10 failures across five files while this branch was still being updated. Every reported failure passes in the current-code rerun; the serialized phase did not run after that failure. The full GitHub CI suite passed on `b1b856a88`: general and serialized tests, browser shards, runner checks, typecheck, production build, packaging/canary, and policy gates. [CI evidence](https://github.com/paperclipai/paperclip/actions/runs/37673106360). - Passed the full-shell Storybook setup-to-history interaction test again after integrating master. Greptile rates the final revision 5/5 with zero unresolved threads. - Walked through the running app: empty setup, add each provider, save, reload, run all three sample questions, inspect fractional charges in history, switch provider while sponsorship is off, disable, and reconnect. Checked mobile settings. These tests used the actual UI, vault, server, SDKs, and database with simulated upstream responses. - Live paid setup tests remain unverified: this environment has no authorized OpenAI/OpenRouter credentials available. No mocked test is presented as live provider evidence. Reviewer journey: Company Settings → General → Decision model. Add/select a shared API connection, save, run the billed sample, open View usage, then disable decisions and verify Run test is disabled after reload. ## Risks - The SDK decision interface is experimental. Pinned versions and wire-format tests limit upgrade drift. - The migration allows agentless service charges and reservations. Existing agent cost-reporting APIs still require an agent, and decision receipts stay separate from run reconciliation. - Timeouts can have unknown provider charges. Holds remain until an audited accounting correction resolves them. - OpenAI prices use a versioned Decisions rate snapshot; OpenRouter costs use provider receipts. Unknown pricing is retained as unknown. - A configured company authorizes background spending by default. Setup explains this, and managers can turn it off. - Live provider account/model availability still needs the two credentialed acceptance checks. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository inspection, code execution, and browser tools. The exact serving revision and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f1c44b7b56 |
Preserve remote work when required workspace restore fails (#15479)
Persist exact source-retention obligations before run finalization and protect them across cancellation, restart and task changes. Require board-authorized repair evidence without replaying old work or changing current task ownership, state or locks. Hide an unavailable retry action and document operator recovery and retention costs. Validated with 674 scoped regressions, 224 combined integration tests, full typecheck/build, independent safety reviews and green CI with Greptile 5/5. No historical file recovery is claimed. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
572cec0344 |
fix(ui): present missing costs as an informational notice (#15459)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Costs dashboard shows spending and gaps in cost data. > - Historical usage can have token counts without a recorded price. > - The red notice implies that the user must fix those records. > - This pull request shows the missing costs as a neutral note below the totals. > - Users can see the limit of the totals without a false call to action. ## Linked Issues or Issue Description Refs #14997. **What happened?** The Costs dashboard shows a red notice when a selected period contains usage without prices. It also shows “0 runs await accounting” when no runs are pending. Historical pricing gaps can therefore look like current errors that need cleanup. **Expected behavior** Show a neutral explanation below the totals. State that totals include only known costs. Show pending runs separately and only when the count is positive. **Steps to reproduce** 1. Open Costs for a period with unpriced usage and no pending runs. 2. Observe the red notice above the totals. 3. Apply this change. Confirm that the notice is muted and below the totals, with no zero-count pending message. **Paperclip version or commit** Reproduced on master after #14997. **Deployment mode** Board UI in both the embedded Activity page and the standalone Costs page. ## What Changed - Move the cost-data notice below the summary tiles and use the muted text token. - Explain unavailable prices with “Totals include known costs only.” - Separate pending runs from missing prices. Omit zero counts and use singular or plural copy as needed. - Update the existing UI tests for both page variants, including missing-only, pending-only, and fully accounted states after refresh. ## Verification - `pnpm --dir ui exec vitest run src/pages/Costs.test.tsx`: 21 tests passed on the final commit. Both page variants cover mixed, missing-only, pending-only, and fully accounted states. - `pnpm check:token-gates`: passed on the final commit. - `pnpm build`: passed locally. - `pnpm -r typecheck`: passed locally. - Full Linux CI passed on `0c23544c0f17e6641cc0c8de6ed498a731748bde`, including browser tests, server tests, shared-package tests, build, and typecheck. [CI run](https://github.com/paperclipai/paperclip/actions/runs/37661883464). - Greptile Apex: 5/5 on the final commit, with no unresolved review threads. - Local full-suite limit: `pnpm test:run` first failed because the fresh checkout lacked embedded PostgreSQL library links. A runtime-skill fixture also selected an unrelated parent directory, and a load test timed out. After setup repair, the database suite (70 tests), skill suite (3 tests), email suite (39 tests), and load suite (4 tests) passed individually. The subsequent full local rerun was stopped after the complete Linux CI run passed. This is not a claim that a full local test invocation passed. ## Risks Low risk. This changes presentation only. The note remains visible and retains its accessible status role. Cost calculations, stored records, and budget enforcement do not change. The copy does not assume that all missing prices are historical. No schema migration or operator documentation change is needed. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, repository inspection, code editing, and test execution. The exact served model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ae6f95ed7a |
feat: add personal primary agents (#15470)
Add a personal primary agent per company and user. Initialize it from the first human-created agent, expose profile-only switching with confirmation, and use it after recent choices for task and Chat defaults. Persist authenticated preferences, preserve lifecycle and membership rules, keep selections out of shared audit events, and document the API contract. Include the reviewed Storybook surfaces and regression coverage for concurrent choices, onboarding, cross-device updates, and browser journeys. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
06484b3c41 |
fix: preserve conversation retries through execution cleanup (#15463)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Recovery schedules bounded retries after a provider disconnects. > - A stopped run can still hold its environment lease while cleanup runs. > - Retrying before that lease is released cancels the new run before it starts and spends another retry. > - Restoring the task can also send the worker repair instructions from an already resolved recovery action. > - This pull request preserves the waiting retry and removes settled recovery instructions from later wakes. > - The task can continue after cleanup without an operator repairing the same incident again. ## Linked Issues or Issue Description **What happened?** A legacy conversation run disconnected while it was doing ordinary work. Its environment cleanup took longer than the retry delay. Two retries were cancelled before dispatch with `execution_reconciliation_required`. Those cancellations exhausted the failure budget. After an operator restored the task, the wake still told the original worker to repair the runtime and hand the task back to itself. **Expected behavior** Cleanup waits preserve the pending attempt. The same retry can continue after ownership is released, subject to all current gates. Once a recovery action is resolved or cancelled, subsequent task wakes omit its repair instructions. **Steps to reproduce** 1. Fail a legacy conversation run while its environment lease remains in `pending_cleanup`. 2. Schedule a bounded retry and run promotion before cleanup releases that lease. 3. Repeat the scheduler sweep. Before this fix, retries promote and then cancel without starting. 4. Resolve a stranded-task recovery action and build the restored task wake with that action ID. Before this fix, the wake still includes the settled repair instructions. **Paperclip version or commit** Reproduced with database regressions against `ceabc3bc880` on master. Related: #15019 restores a skipped assignment handoff after lease release. #15235 filters stale handoff evidence in the recovery sweep. This change preserves an existing scheduled conversation retry and corrects restored wake content. It does not create a new handoff wake. ## What Changed - Keep an unstarted legacy conversation retry on the same durable row while prior execution ownership remains active. Recheck after 30 seconds without increasing retry accounting. - Return a queued retry to scheduled state if it encounters that hold at the claim gate. Retain its issue claim and publish the status change. - Record one local lifecycle diagnostic per blocking run. Remove that wait marker on promotion. - Include recovery action metadata only while the referenced action is active or escalated. - Return an explicit `waiting` response and the saved schedule when Retry now meets cleanup. Show the wait inline without a false success or disabled button. - Add database, rendered-prompt, route, and UI regressions. Document the execution and run-log contracts. ## Verification - Five cleanup and restored-wake regressions fail against the original production code. The Retry now route and UI regressions also fail before their correction. - Related retry, dispatch, stale-queue, and recovery suites: 321 tests pass across seven files. All ten focused cleanup/restored-wake cases pass after rebase. The final dispatch adjustment passes all 46 adapter tests. - Retry now routes and affected UI suites: all 46 tests pass. The tests cover repeated clicks, the saved schedule, unchanged accounting, promotion after release, and no false success or error state. - `pnpm -r typecheck`, `pnpm build`, `pnpm build-storybook`, and `pnpm check:token-gates` pass. - Full local `pnpm test:run` was started and then stopped after the final commit passed all GitHub CI test shards. No complete local full-suite result is claimed; CI supplies the complete test result for the final commit. - Final head `6f1058a332c039e33c4f002b296d20a5554e760e`: all 55 GitHub checks are green or intentionally skipped. Apex review is 5/5 after two reviews, with no unresolved threads. The PR has no merge conflicts. ## Risks The wait applies only to unstarted legacy conversation retries. Native runs and non-conversation execution keep their existing recovery rules. Cleanup must actually release ownership before execution can resume. The wait does not fix a cleanup service that never finishes. Promotion and dispatch still enforce cancellation, reassignment, pause, budget, and reconciliation gates. No schema change is required. The Retry now response adds a `waiting` outcome; the shared contract and all three UI controls handle it. ## Model Used OpenAI Codex based on GPT-6, with repository analysis, tool use, and local code execution. The runtime does not expose the precise serving model ID or context window. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused suites; complete final-head test coverage in CI) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0830737030 |
feat(ui): add CSV table previews to task file tabs (#15445)
## Thinking Path > - Paperclip helps people manage AI agents and inspect their work. > - Task file tabs show the files that agents produce. > - Markdown has a readable view, but CSV files show only source text. > - People need to scan CSV reports without a separate app. > - This pull request adds a table view to attachment and workspace file tabs. > - View buttons let people inspect the original source beside download. ## Linked Issues or Issue Description **Subsystem affected** ui/ — task file tabs. **Problem or motivation** CSV exports are difficult to scan as source text. Related work: https://github.com/paperclipai/paperclip/pull/14297 added text attachment tabs. **Proposed solution** Render CSV as a table by default. Add row numbers, sticky headers, and record counts. Keep rendered and raw icons beside download in task tabs. **Roadmap alignment** This is a small improvement to existing task file previews. It does not add a new core subsystem. ## What Changed - Add a shared CSV parser and table preview. - Handle quoted commas, escaped quotes, UTF-8 BOM, CRLF, and multiline cells. - Show at most 500 data rows and 100 columns. Show a notice for larger files. Keep the complete download. - Add CSV attachment and workspace previews. Put workspace tab view controls in the header. - Add parser and interaction tests, Storybook examples, and product documentation. ## Verification - 28 focused tests pass across the parser, attachment tabs, and file viewer. - UI typecheck, UI build, and token gates pass. - Playwright checked desktop and mobile rendered → raw → rendered flows using the real attachment preview component in Storybook. Screenshots are recorded on the driving task. - Repository-wide typecheck and build were attempted. Both require cargo, which is absent in this environment. - The earlier repository-wide local test run stopped after a server test race failure. The prior commit passed all GitHub CI jobs. - Review fixes add source-line navigation coverage and CSV boundary tests. All 22 file-viewer tests pass. UI typecheck and token gates pass. - All 54 active GitHub checks pass on the final commit. Two Storybook checks are skipped. Greptile gives 5/5 with no actionable findings. - The final version preserves HTML previews from master and uses the shared toolbar control. All 35 focused tests, UI typecheck, UI build, and token gates pass. ## Risks - The first CSV record supplies column headers. Files without headers use that same rule. - The table preview is bounded. Download and raw view retain the complete file. - No API, database, or dependency changes. ## Model Used - OpenAI Codex, GPT-6. The runtime does not expose a more precise model ID or context window. Used code editing, shell tools, and browser checks. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1640c5b6ab |
feat(ui): render HTML artifacts in a secure sandbox (#15447)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents deliver reports as task attachments and workspace files. > - The board already renders Markdown, but it shows HTML reports as source or rejects their preview. > - Reports can need inline scripts to build charts and tables. > - Artifact scripts must not read board cookies, storage, or the parent page. > - This pull request adds an opaque-origin HTML preview and view controls beside Download. > - Users can explore a report, inspect its source, and download the original file. ## Linked Issues or Issue Description **What existing behavior does this improve?** The task attachment panel and workspace file viewers. **Subsystem affected** ui/ and the server workspace file preview service. **Current behavior** HTML attachments show source text. Workspace HTML files cannot be previewed. **Proposed behavior** Render self-contained HTML reports in a sandboxed iframe. Place eye and code controls beside Download. Preserve the original file for raw view and download. **Reason and benefit** Users can read and filter agent reports in Paperclip without giving artifact scripts access to the board session. **Breaking changes** HTML previews now open rendered. Scripts and styles must be embedded in the report. Remote resources, API requests, forms, popups, and host navigation are blocked. Workspace HTML content uses the existing bounded UTF-8 JSON response. **Additional context** Related closed proposals: #3293 and #4857. This change uses the current attachment and workspace viewers and adds no report-serving endpoint. The Artifacts and Work Products roadmap item is complete; this improves its existing preview behavior. ## What Changed - Add a shared HTML iframe renderer with `sandbox="allow-scripts"` and no `allow-same-origin`. - Install a restrictive CSP before artifact markup. Keep inline report scripts and styles. - Add shared rendered/raw icon controls beside Download in the attachment panel, workspace panel, and file sheet. - Return workspace HTML as bounded text inside JSON. Keep download and path-access protections. - Add Storybooks for reports, security probes, workspace viewers, a task journey, and mobile layouts. - Document the security boundary and browser test command. ## Verification - `pnpm -r typecheck` and `pnpm build` pass. - Token gates and the static Storybook build pass. - Four browser tests pass. They cover report filters, raw mode, task entry, workspace viewers, cookies, storage, host DOM, resource requests, forms, popups, and host navigation. - The focused UI suite passes 26 tests. The file-resource server suite passes 36 tests. - Local full-stack acceptance passes with an actual 52 KB HTML report. Its chart and filter render. Raw mode preserves the source. Download bytes match the original file. - The complete local UI suite passes: 698 files and 7,754 tests. - `pnpm test:run` was attempted. Its server group recorded two unrelated timeouts and eight connector failures. All failed cases pass in isolated reruns, including the complete 388-test connector suite. The remaining local run was stopped after the full CI test coverage passed, to release test database resources. The full local command did not complete cleanly. - The separate local CLI group passes 511 tests. Three worktree database cases fail to start embedded PostgreSQL because this Mac has exhausted its shared-memory allocation. One isolated rerun fails at the same database startup step. No CLI code was changed. The corresponding CI test jobs pass. - All 56 PR checks are clean on `c771255f1a2e86bd8825f22eafb5d14a5e342cce`. Greptile gives 5/5 with no actionable findings. The security scan passes. The branch has no merge conflict. - Review the **HTML artifacts** section in Storybook. Run `pnpm exec playwright test --config tests/html-preview/playwright.config.ts` for the browser checks. ## Risks - Never add `allow-same-origin` to this iframe. Its opaque origin is the cookie, storage, and host-DOM boundary. - Reports that depend on CDN assets need to embed their dependencies. - The frame can navigate itself. Sandbox restrictions remain after navigation. This is not full network isolation. - Inline scripts can consume browser resources. This change does not isolate CPU or memory use. - No database migrations or API shape changes are required. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, repository editing, code execution, and browser tool use. The runtime does not expose a more specific model ID or context window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass — feature and UI checks pass; the full local run has the infrastructure limits recorded above. All CI test jobs pass. - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b315580640 |
feat: add private task creation and sharing controls (#14718)
Add private task creation and sharing controls, effective access explanations, private project management, and locked task references. Constrain the mobile composer plus menu to the viewport and explain named inherited privacy on hover, focus, or tap. Include 125 production-component Storybook stories and thirteen interactive user journeys. Final head passes all CI checks and Greptile Apex 5/5 with no findings. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
b67db12d90 |
feat: add durable storage for private tasks (#14717)
Add private task and project ownership, downward access grants, and immutable run/workspace provenance. Apply migration 0313 with bounded batch commits and concurrent indexes. Keep the enforcement and sharing changes in their dependent PRs. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
799e4d556f |
fix: make accounting durable and synchronize cost reporting (#14997)
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
99a9de9940 |
fix(mcp): personalize assistant connections and hide revoked grants (#15411)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Assistant connections let people use their organization from another assistant. > - Setup begins with an invitation and returns to a list of connected assistants. > - Client-only labels obscure the authorizing person, while revoked rows clutter that list. > - This pull request shows the person’s avatar, names connections by owner and client, and hides revoked rows. > - It also shortens the copied invitation while keeping the approval instructions. ## Linked Issues or Issue Description Related: #14933 and #15380. **What existing behavior does this improve?** Invitation copy and assistant connection management, in the organization’s Connections screen and the account-wide management page. **Subsystem affected** Cross-cutting: shared MCP connection types, server profile projection, UI and Storybook. **Current behavior** Connections are labeled only with a client name such as “Codex.” Revoked connections remain visible. The invitation includes an extra sentence about agent identity. **Proposed behavior** Show the authorizing person’s avatar and use names such as “Dotta’s Codex connection.” Hide revoked rows after successful revocation and when loading retained revoked grants. Failed revocation leaves the connection visible. Remove the extra identity sentence from invitation copy. **Reason and benefit** Make connection identity clear and keep the list focused on usable connections. **Breaking changes** The connection response adds optional `user` metadata with name and image. Older servers remain usable. Names and revoked-row visibility change in the UI; OAuth client identity, authorization and audit retention remain unchanged. ## What Changed - Shorten the shared invitation text. - Project the authorizing person’s name and avatar through the user-scoped connection endpoint, without returning email or credentials. - Reuse the existing Identity component and owner naming conventions across the connection page, catalog card and account-wide list. - Hide revoked grants and remove a successfully revoked row from the shared cache, even if the subsequent refresh fails. - Update documentation, regression tests and production-page Storybook fixtures and revocation journeys. ## Verification - `pnpm -r typecheck`, `pnpm build`, `pnpm build-storybook` and `pnpm check:token-gates` pass. Final UI type checks also pass. - Focused consent and connection UI tests: 31 pass, including company filtering, legacy metadata, custom client names, user-initial fallbacks, failed revocation, retained revoked rows and refresh failure after successful revocation. - Existing MCP regression suite: 6 pass; 71 database checks are skipped locally because embedded PostgreSQL cannot start on this machine. The added database check verifies user-profile isolation and retained revocation history; CI runs these checks. - Browser verification with Storybook fixtures: owner avatar and name render; revocation removes the selected row in both production pages, leaves other connections visible, and restores the empty state after the last revocation. - All 54 current-head CI checks pass, with two optional Storybook jobs skipped. CI includes database, browser, runner, typecheck, build and clean-install canary coverage. - Greptile reviewed commit `1368d79e1066b418712224378d89d64c2b11cb86`: 5/5, no actionable findings or unresolved threads. - The full local `pnpm test:run` was stopped after complete CI passed. Local database coverage remains unavailable because embedded PostgreSQL cannot start; no full local-suite pass is claimed. ## Risks Low risk. The additive profile field is optional for compatibility. Revoked grants are filtered only from management UI and retained for audit. Revocation failure does not hide an active connection. No authorization scopes, token handling or schema changes. ## Model Used OpenAI GPT-6 through Codex, with code editing, command execution and browser verification. The exact deployment ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a6306ba606 |
feat(runner): consolidate Cursor production integration (#15075)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The native Runner keeps provider sessions under company authority, approvals, budgets and durable recovery. > - Cursor work was spread across candidate branches. The published branch lacked later plan, permission and cleanup fixes. > - Production also needs public installation and matching runtime assets for local and Daytona execution. > - This pull request consolidates Cursor onto current mainline recovery behavior and completes that installation path. > - The installed v11 release passed focused local and Daytona qualification after the generic mode and lifecycle cleanup. The later model-selection correction and current mainline merge produce v14 artifacts that need matching release qualification. > - Cursor admission is enabled in source; publish only an artifact combination with matching qualification. Native AskQuestion and complete per-run dollar accounting remain excluded. ## Linked Issues or Issue Description Refs: #14435, #14631, #14669, #14699, #14724. This completes the Cursor implementation by @cryppadotta from combined source `22c78242a4e0c2369fecf0c2dc4e7600fbad6706`. It preserves newer mainline recovery, completion and warm-directory behavior. Pi and Copilot remain gated. ## What Changed - Generate named Rust and TypeScript ACPX release profiles from one manifest. Share runtime pins with packaging and server verification. Preserve vendor runtime versions; bind the updated ACPX patch to Cursor profile v14 and reject stale generated declarations at build/typecheck. - Remove ACPX model allowlists, including the former Codex and Pi restrictions and the duplicate developer test-drive gate. Send any explicit model ID unchanged to its provider and verify the effective selection before prompting. The bundled ACPX package forwards unlisted IDs, rejects mismatched acknowledgements, and restores the exact selection after session load. It does not expand Cursor model aliases. Provider rejection, mismatch, or missing model controls fails without a fallback. Model examples live in evaluation fixtures, outside runtime declarations. - Add pinned Cursor execution, contained instructions, exact model verification and Agent/Plan/Ask modes. - Carry an opaque generic `mode` identifier in shared native execution, sidecar, Rust and recovery contracts. The provider adapter owns supported modes, defaults, native translation and acknowledgement. - Keep native RPC recognition, accepted-plan interpretation and permission evidence behind provider adapters. Shared settlement and recovery verify normalized facts and their committed evidence. - Replace the Cursor-only warm-attachment branch with a runner-owned capability. Only Cursor opts into it. Move profile compatibility and optional usage parsing into provider metadata and adapters. - Write generic plan-wait receipts. Read exact historical Cursor receipts through a separate compatibility decoder. Reject mixed formats and preserve existing authority checks. - Carry native plans, semantic questions, todos, child activity, permission identities and partial usage diagnostics through the Runner. - Preserve durable response delivery, cancellation, warm ownership and process retirement. - Finish accepted planning runs successfully. Keep their tasks open for explicit direction. Acceptance does not start implementation. - Ship `paperclipai runtime setup cursor` and its provisioner through the public package. npm installation does not download Cursor. Setup uses the OS account's closure-keyed cache so system-wide npm packages can remain read-only. Run it as the Paperclip service account. - Include Cursor in normal provider packs and Daytona images for macOS ARM64/x64 and Linux x64. - Reject stale release packs by source revision and current ACPX/Cursor pins before assembly writes files. Verify current Cursor version/profile/closure again at runtime. - Ship all three daemon targets and the expected Linux image-pack identity. A macOS controller uses its packaged Linux daemon for Daytona. Image mismatches fail before provider launch. - Use the vendored Runner boundary for installed readiness probes. Verify the actual installed Cursor probe. - Verify compiled public Daytona plugins and their release versions in installed smokes. - Record exact artifacts, the acceptance matrix, retained failures, supported capabilities and rollback behavior in the [readiness report](https://github.com/paperclipai/paperclip/blob/codex/cursor-production-readiness/doc/plans/2026-10-03-cursor-production-readiness.md). ## Verification - Current head `f7ec5cc1f0e30c62a829c016ff2013a2a9d79037` merges mainline `faa8e452c73bae5e044dd6379179a00106abb131`. It keeps Cursor plan and cancellation guards alongside mainline historical-question filtering. The evaluation catalog includes both Cursor and expanded adapter accounting cases (683 total). Recursive typecheck, full build, 696 lifecycle/recovery tests, 45 fixture tests and fixture typecheck passed. Current-head CI passed: 56 successful checks, one neutral and four skipped. [Complete CI](https://github.com/paperclipai/paperclip/actions/runs/37557996535). The fresh Base Greptile review is 5/5 on this exact head, with 304 files reviewed, zero new comments and zero unresolved threads. The user authorized overriding the CODEOWNER review gate after checks passed; no failing checks are overridden. Prior results below retain their own head identities. - Corrective head `3d2b168366258036f6b6a6fccb382c49138cc601` fixes the post-merge Apex finding. Automatic-review and new-evidence reconciliation preserve pending child results and recheck delivery under the status lock before completing. Account repair now excludes unrelated secret consumers and requires the failed agent's identity. Regression coverage includes the commit race, delivery statuses, current-run/current-intent exclusions, repeated reconciliation, both database reconciliation paths, and credential consumer boundaries. All 184 affected tests, server typecheck and server build passed. Current-head Base Greptile review is 5/5, with 304 files reviewed, zero new comments and zero unresolved threads. Current-head CI passed: 56 successful checks, one neutral and four skipped. [Complete CI](https://github.com/paperclipai/paperclip/actions/runs/37535994724). This Base review is distinct from the earlier Apex review. - Merge head `5957c257a` reconciles mainline `b508a05c4`. It preserves both accepted-plan waits and pending-child-completion checks, current provider selectors, task-creation response identities, and mainline ACPX missing-file handling. The combined patch is bound to Cursor profile v14; historical records keep their original identities. - Merge head `5957c257a` passed recursive typecheck, full build, 43 installed ACPX/package contracts, 107 provider UI and plan/recovery tests, 593 database-backed lifecycle tests, 49 profile/native contract tests, 45 Product E2E fixture tests, fixture typecheck, token gates, three provider-free browser task-creation cases, and Runner conformance/replay checks. Its complete CI passed (55 successful checks, one neutral and four skipped), while Apex returned 2/5 with a child-delivery finding addressed below. - The local full-suite attempt again failed the unchanged Git streaming test (360-second timeout) and was stopped. The concurrent local Rust attempt failed four unchanged Codex process/deadline tests; all four passed serially without code changes in 7.29 seconds after removing the competing test load. These failed commands are retained and are not reported as full-suite passes; the fresh Linux CI runs are tracked separately. - The previous head `907bdb2a2778c7ffeb4a662a91460c9d1ddfc9c5` earned Apex 5/5 with zero comments after fixing all three findings: per-user install cache, stale release-pack rejection, and public Linux smoke account/home handling. Its real built installer passed from read-only public packages on macOS ARM64 and Linux x64. All 137 release-registry checks and 64 ACPX package contracts passed. That review does not cover this mainline reconciliation. - Prior `beadd3654` passed the full CI matrix; its one unchanged chat test failure and successful single retry remain in the [CI history](https://github.com/paperclipai/paperclip/actions/runs/37521449327). Historical results below remain attributed to their original builds. - Fixture follow-up `dd59d7e82b103a88b7cbd7d2c38b612c0fbbff7a` removes provider-specific model choices from generic offline ACPX tests. The fake sidecar preserves the model and session identity selected at open through suspension. Affected verification passed: 106 Rust tests and 73 TypeScript tests. This commit changes test code only; the production-code checks below retain their recorded identities. Its CI and Greptile review later passed; those results belong to that historical head. - Model-selection cleanup `9a070808b48960a41fdfd369ae0636b95af82459`: 252 focused Runner tests passed (six platform skips), covering all six ACPX agents, native model acknowledgement, rejected selections, installation integrity and recovery identity. The merged branch passed recursive typecheck, full build, token gates, server admission (19 tests), and the Product E2E catalog (45 tests). The acceptance catalog passed all four tests. The full Rust suite passed: 643 tests, 2 ignored. It verifies sidecar acknowledgement of unlisted models and rejection of model mismatches. The final commits only update Rust tests; production sources match the verified build at `65ec3279ac50185e3cda109b5cfd9b4f56105de0`. No new paid provider calls were made. - The merge preserves both Cursor and the new mainline public-MCP fixture cases. Auto-merge remains disabled; the latest follow-up status is recorded above. The local `pnpm test:run` attempt hit the unchanged Git streaming test's 300-second timeout and was interrupted before merging mainline. The broad Runner attempt found obsolete single-model assertions plus three macOS fixture-path failures caused by a `/private/tmp` override. The assertions are corrected; affected TypeScript checks passed with the standard macOS temporary directory, and the complete Rust suite passed. Neither interrupted command is a full-suite pass. - Earlier declaration-cleanup head `6f4a5e9e2` passed recursive typecheck, build, Rust and focused tests. Its CI later exposed a test expecting duplicated Grok digest literals. The current source fixes that assertion to compare launcher bytes with the shared manifest. Historical successes and failed attempts are retained; no new live provider qualification is claimed. - Previous head `e75fde6098b0ddd8cec765bfb6ecaeecb88a26a6` passed complete CI (56 successful checks, one neutral, four skipped) and Greptile 5/5. [Historical complete CI](https://github.com/paperclipai/paperclip/actions/runs/37489112305). Those results are not claimed for the cleanup head. - Frozen live application: `d7b696f9b8f79095233e9e3d56d23e6a6018dd48`. Public package version: `0.0.0-cursor-verify.3d0c9b7761c6`. The declaration cleanup preserves release pins and does not relabel that tested artifact as a build of the new source. Mainline through `e34abee670` was reconciled while preserving accepted-plan waits, provider-capacity handling, and both Cursor and public-MCP fixtures. - Clean normal installation, explicit Cursor setup and daemon resolution passed on macOS ARM64, macOS x64 under Rosetta, and Linux x64. npm lifecycle hooks ran without silently downloading Cursor. - Historical v11 live matrix: **18/18 passed with cleanup** (nine local, nine Daytona) after the generic mode and lifecycle cleanup. The campaign has 23 attempts; all five failures and their diagnoses remain recorded. Exact case identities, hashes and limits are in the readiness report. All provider calls are real, use the explicit Luna model and company-bound credentials, and run without qualification or runtime-asset overrides. - The immutable Daytona image is `ghcr.io/paperclipai/paperclip-daytona-runner@sha256:d6259b6bba094702c13fc2283bd85550849c1c53145b656fb2746778f9fa1747`. The public Daytona plugin is installed independently and its version is checked. - Recursive typecheck, full build, token gates and Runner contract/conformance/replay checks passed on the frozen application. Its complete Linux CI suite passed. The duplicate local full-suite command was incomplete after timing failures; affected repeats passed, but that command is not reported as a clean pass. - Qualification fixtures passed typecheck, 1,675 Vitest tests (one skip), 128 Node checks, three provider-free browser tests, and 150 focused lifecycle tests after the final diagnostic correction. The affected legacy Cursor command file also passed all five tests after removing its shorter 10-second override; it now inherits the suite’s standard 15-second timeout. Greptile is 5/5 on `e75fde609` with no unresolved review threads. CI results above are recorded separately from historical build results. ## Risks - Cursor v14 includes the updated ACPX dependency patch and release identity. The v11 live matrix and image below remain historical evidence. They do not certify new v14 package/image artifacts. - ACPX accepts models beyond the qualification fixtures. Availability and entitlement depend on the provider. Successful configuration is not a claim of live qualification for every model. - Shared mode is an opaque identifier. Provider adapters own its meaning. Incompatible historical sessions remain fenced; exact committed plan waits and task history remain inspectable. - Native AskQuestion is excluded. Paperclip semantic questions are supported. Authoritative per-run dollar accounting is unavailable; partial counters remain diagnostics and unknown cost is not zero. - Image input, detailed native diffs, deeper child transcripts and native plan-file export remain follow-ups. - macOS x64 has clean-install and daemon-startup proof under Rosetta, not a separate live campaign on Intel hardware. - Release only the tested package/image combination. Merging this PR does not publish npm packages or deploy that image. Later builds need their own release verification. Rollback disables new Cursor admission while preserving records and recovery inspection. - A model can fail an exact instruction: one cancelled-plan attempt returned the wrong summary marker despite correct cancellation. The unchanged repeat passed; both results remain in the report. > ROADMAP.md was checked. This completes existing native Runner/Cursor work; it does not add an independent core feature proposal. ## Model Used OpenAI Codex, GPT-6. The exact serving variant and context window are not exposed in this session. The agent used reasoning, repository inspection, code execution, protocol tests and browser-backed Product E2E tools. Cursor acceptance uses the explicit `gpt-5.6-luna[context=272k,reasoning=medium,fast=false]` model. That is the evaluated provider model. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass — affected suites passed; full CI and the retained local failed attempts are recorded separately above. - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green — 56 successful checks, one neutral and four skipped on `f7ec5cc1f0e30c62a829c016ff2013a2a9d79037` - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups — fresh Base review passed on `f7ec5cc1f0e30c62a829c016ff2013a2a9d79037`; zero new comments and no unresolved threads. The earlier Apex finding remains fixed. - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
2d0c138122 |
Expand direct assistant MCP tools for work and configuration (#15380)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - People also use assistants in Codex, Claude, and other MCP clients. > - The existing assistant connection can read work and create tasks or comments. > - It cannot edit tasks, exchange files, or manage normal agent and project settings. > - These operations must retain the person's permissions and Paperclip's execution rules. > - This pull request adds an explicit operation registry and separately consented configuration access. > - Assistants can manage work without receiving credentials or runner authority. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: assistant MCP, domain routes, consent UI, storage, and Product E2E. **Problem or motivation** A connected assistant cannot update tasks, maintain documents, attach files, or configure existing agents, projects, and skills. Users must leave the assistant for these routine actions. **Proposed solution** Add named tools and a restricted API registry to direct connections. Require separate configuration consent. Reuse domain routes and retry receipts. File uploads save the attachment when the byte transfer succeeds. **Alternatives considered** Arbitrary REST forwarding would expose administration and credential operations. Runner impersonation would bypass execution ownership. Separate upload completion calls add unnecessary client state. **Roadmap alignment** Checked ROADMAP.md and related MCP pull requests. This extends the human-authorized connection from #14933. It does not replace the runner or introduce agent impersonation. Companion Cloud routing and directory isolation: https://github.com/paperclipai/paperclip-cloud/pull/678. ## What Changed - Add task editing, finish/block, documents/revisions, deliverables, agent settings/instructions, projects/repositories, and skills/files. - Add an allowlisted API search/call registry with identical field restrictions, scopes, and retry identities. - Add unchecked configuration consent. Existing write grants retain their current authority. - Add hashed, expiring file transfer tickets and atomic upload receipts. No completion call is required. - Preserve company boundaries, human attribution, active native execution ownership, and execution review gates. - Add protocol/domain tests, consent stories, and eight paid Product E2E workflows. - Repair two CI fixture races: await cold route setup before assertions, and wait for asynchronously loaded connection copy. Both fixture suites pass (24 + 48 tests). ## Verification - Consent revision: one write-access checkbox controls requested work and configuration permissions in browser and device flows. All 16 consent tests, UI typecheck/build and token gates pass. Updated interactive stories cover default approval, opt-out and viewer restrictions. The paid browser helper uses the new exact label. Real GPT-5.4 Mini Product E2E passes 2/2 at `64f96373118eb190f8cba1c2ab17cb979555f3ad` (configuration + permission denial), campaign `local-2026-10-07T00-51-14-337Z`, no automatic retries, cleanup passed; $0.04149375 estimated assistant cost plus unpriced worker usage. Raw results, usage and source fingerprints are retained in the worktree. UI and Product E2E typechecks pass. - Prior head `2f246d4b74f1f98c75ebcb37ae6753a748237fac`: all 52 checks pass; two optional Storybook checks skip. Greptile 5/5 on that head, no unresolved review threads. Final consent head `64f96373118eb190f8cba1c2ab17cb979555f3ad` also has all checks passing and Greptile 5/5 with no unresolved threads. The unchanged Cursor sandbox test had one 10-second timeout, passed in local isolation, and passed its single CI rerun; the failed attempt remains in [the CI run](https://github.com/paperclipai/paperclip/actions/runs/37554106934). The existing chat retry-denial browser test had one visibility failure; its single rerun passes, and the failed attempt remains in [the CI run](https://github.com/paperclipai/paperclip/actions/runs/37542735691). - Full workspace `pnpm -r typecheck` and `pnpm build` pass at final runtime source `b2196fae1`. UI token gates pass. - 139 MCP/OAuth/transfer/privacy tests and 76 grader calibration tests pass, including one-connection PostgreSQL OAuth and concurrent upload retries. - Paid Product E2E: all eight expanded cases qualified across Mini, Haiku and Sonnet. A merged-source repeat passed 23/24; one Haiku cell timed out before application startup. Final affected-case qualification passes 9/9 on all three models with grader v16, including the failed cell. Automatic retries disabled; failures, costs, source hashes and independent durable-state/file assertions are retained in [the verification record](doc/plans/2026-10-06-expanded-assistant-mcp-verification.md). - Actual Codex CLI, Claude Code and OpenCode clients completed local reads/mutations. Codex wrote a report, Claude updated it in a later conversation, and OpenCode uploaded/downloaded a file with matching SHA-256 and registered the attachment. Revoking the CLI grant rejects subsequent bridge initialization. - Butter staging is verified on final runtime `b2196fae1` ([deployment](https://github.com/paperclipai/paperclip-cloud/actions/runs/37538432138)). A fresh OpenCode workspace fetched the copied invitation, configured remote MCP, started OAuth and reached real consent with configuration unchecked. Invalid transfer tickets return 403 through Cloud. Human approval for the new persistent staging grant is pending; hosted task/file success is not yet claimed. The final transaction fix is deployed. - Full local `pnpm test:run` passed 15,614 general-server tests but stopped on two macOS timeouts. The heartbeat test passed in isolation; the existing 40,000-file Git stress fixture timed out again. Its Linux CI lane passes. Later local full-suite phases did not run after the timeout; this is not an all-green local full-suite claim. - Instructions and security limits are in `doc/public-mcp.md`; the saved plan is `doc/plans/2026-10-06-expanded-assistant-mcp-tools.md`. ## Risks - This expands the experimental direct MCP surface. Explicit schemas and domain permissions must stay synchronized. - Migration 0311 adds transfer tickets and upload receipts. Expired orphan cleanup must not remove committed attachments. - Configuration requires a new consent request containing that scope; the single write-access choice controls it alongside work mutations. Refreshing an old grant does not add it. - The public directory keeps its original ten tools through the companion Cloud change. - Hosted consent/work proof remains the final delivery gate. The PR stays draft while approval of the new staging grant is pending; code checks and review are green. Merging is a separate action. ## Model Used OpenAI Codex (GPT-6, tool use and code execution). The exact serving model ID and context window are not exposed in this session. Paid evaluation models: gpt-5.4-mini, claude-haiku-4-5-20251001; claude-sonnet-4-6. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused suites; full-suite macOS limitation disclosed above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
892b0b3606 |
fix: scope quota reports to authorized subscription accounts (#14994)
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
3ebd7bc0c9 |
test: isolate accounting fixtures and include all adapter suites (#14989)
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
ac8f3eb143 |
feat(ui): remove star and leave buttons from agent index (#15400)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The Agents index page (`/agents/all`) shows one row per agent in both the streamlined and production sides of the UI > - Each row had a star toggle and a leave/join button that write per-user state through the resource-memberships API (`agent_memberships` table) > - The default streamlined sidebar no longer lists agents, so the star and leave actions on the index no longer feed a visible pinned list > - This pull request removes the star toggle and the leave/join button from agent index rows and org-tree nodes > - The backend resource-memberships endpoints and the `agent_memberships` table stay live because the agent detail page and the legacy sidebar still use them > - The benefit is a cleaner index page with no redundant actions for state that is not surfaced on the page itself ## Linked Issues or Issue Description No public GitHub issue exists for this change; the description follows the feature request issue template inline. **Subsystem affected** ui/ — React + Vite board UI. The change is presentational only; no backend or schema code is touched. **Problem or motivation** The agent index rows show a star toggle and a leave/join button. Those two buttons write per-user membership state that the default streamlined UI no longer surfaces anywhere on the page itself. The default sidebar shows a plain "Agents" nav item and does not list agents. The buttons are now redundant on the index. **Proposed solution** Remove both buttons from the agent index list rows and org-tree nodes in both UI variants. Keep the data layer. The agent detail page still renders the header star toggle and the join/leave banner, and the legacy sidebar still renders star and leave rows for instances that opt out of the streamlined UI. **Alternatives considered** - Remove the backend endpoints too. Rejected: the agent detail page and the legacy sidebar still read and write the same state. - Keep the buttons hidden behind a setting. Rejected: no product need, and the plan approved the straightforward removal. **Roadmap alignment** This is a small presentational change. It does not overlap any section in `ROADMAP.md`. **Additional context** Rows whose membership state is `left` keep their dimmed styling on the index. ## What Changed - Removed the `StarToggle` component from agent index list rows in `ui/src/pages/Agents.tsx` and `ui/src/pages/Agents.production.tsx` - Removed the `MembershipAction` (Leave/Join) component from agent index list rows in both files - Removed both actions from the org-tree nodes in both files - Dropped now-unused imports and derived values (`StarToggle`, `MembershipAction`, `isStarred`, `useResourceMembershipMutation`, and the per-row pending/starred variables) - Kept `useResourceMemberships` and the dim-on-left row styling so rows whose membership is `left` remain visually identified - Updated `ui/src/pages/Agents.test.tsx`: a new test asserts both modes omit star and leave/join actions from list and org-chart views; the dim-on-left tests are kept ## Verification - `pnpm --filter @paperclipai/ui typecheck` passes - `pnpm --filter @paperclipai/ui build` passes - UI Vitest suite for `Agents.test.tsx` passes (21 tests) - `pnpm check:token-gates` is clean - Manual check on `/agents/all`: no star and no Leave/Join buttons in the list view or the org-chart view; star/leave still work on the agent detail page ## Risks Low risk. This is a presentational UI change with no backend or schema changes. Star/leave remain available on the agent detail page and the legacy sidebar. No migration is involved. ## Model Used - Provider: DeepSeek (via the Paperclip opencode_local adapter, model `openrouter/~deepseek/deepseek-v4-flash-latest`) - Model: deepseek-v4-flash (OpenRouter `openrouter/~deepseek/deepseek-v4-flash-latest`) - Context window: 128K; used with tool use in the repository - Assisted with the code change and this PR body; the plan and scope came from the tracking work item ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: devinfoley <139239+devinfoley@users.noreply.github.com> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f77fcbf4bf |
feat(apps): add Telem.AI web search connection (#15379)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Research agents need current web information. > - The Apps catalog connects agents to remote MCP tools through the normal access rules. > - Telem.AI supplies web search and page reading through one API key. > - This PR adds its catalog entry, optional search settings, artwork, and setup guide. > - The connection keeps an operator's saved header policy when they reconnect. ## Linked Issues or Issue Description Refs #15302. Related catalog work: #13881. This PR continues #15302 by Yifei Ai (@aiwen324). Thank you for the connector and its review fixes. All seven original commits are preserved. GitHub denied the attempt to push to the contributor's fork, so this branch retains the repair commit and merges current master. Master now includes the same test fix. For a squash merge, keep the original author in the final commit message: ```text Co-Authored-By: Yifei Ai <aiwen324@users.noreply.github.com> Co-Authored-By: Paperclip <noreply@paperclip.ing> ``` A search found no separate public Telem issue or competing Telem PR. The existing Apps path matches the roadmap. **Agent or provider** Telem.AI provides web search and page reading through a hosted MCP server. **Why this adapter is useful** Agents can use multiple search providers through one governed connection. Operators can set the search tier, auto routing, and provider lists. **How the agent is invoked** The remote MCP server uses Streamable HTTP at `https://mcp.telem.ai/mcp`. The API key uses an `Authorization: Bearer` header. See the [official MCP guide](https://docs.telem.ai/integrations/mcp/). ## What Changed - Add the Telem.AI definition, research entry, permission review, and generated registry entry. - Add four optional settings. Unset settings send no request header. - Add official light and dark artwork, source records, and a setup guide. - Forward company, issue, agent, run, project, and correlation IDs by default. Preserve a saved policy, including disabled forwarding, on reconnect. - Add catalog and connection tests. ## Verification Current head: `b4164477fb1b312a504789bf17b51c963244c0fd`. Merged master: `228f0e2807c5b59d2aa129cf2d80b9777ebabf07`. - Resolved five shared catalog conflicts after the Superagent connection merged. - Keep both providers in the research ledger, generated registry, generator, branding manifest, and connection guide index. - Correct the combined catalog totals: 52 self-serve candidates, 55 research entries, and 68 Apps entries. - The published Git tree exactly matches the tested local resolution. - Catalog and Apps UI suites: **295 tests pass** after the catalog count fixes. - Connection service suite: **387 tests pass** in the full run. Its only failure was the old catalog count. That test passes on a focused rerun after the fix. This gives **388 passing service tests** across the two runs. - Total focused coverage: **683 passing tests**. The first runs exposed four fixed-count assertions that needed the combined totals. - Shared package build and plugin SDK compile pass. Token gates and whitespace checks pass. - Generation with `--definitions-only` reproduces the Telem definition and registry. The unrelated AgentMail and Linear drift remains excluded. - Local UI typecheck ended with exit 137 at the container memory limit. Full local typecheck, test, and build are not claimed. Earlier runs also recorded missing Cargo and Node development headers. - GitHub reports a clean merge state against master `228f0e280`. - All 54 checks are complete: **52 passed and two Storybook checks skipped**. No check failed or remains pending. - [CI](https://github.com/paperclipai/paperclip/actions/runs/37545412406) passes on this head. This includes typecheck, build, tests, browser shards, Runner checks, and Canary Dry Run. - [Greptile](https://github.com/paperclipai/paperclip/pull/15379#issuecomment-6024519537) is **5/5 on this head**. There are no review threads, open P2s, recommendations, or follow-ups. - [Superagent](https://github.com/paperclipai/paperclip/runs/112548142930) passes. - Final recovery checks confirm all seven original commits and current master remain in history. Token gates and whitespace checks pass. - The final recovery run makes no source change. It verifies the published repair and retains the local check limits below. - [Commitperclip](https://github.com/paperclipai/paperclip/actions/runs/37545407943) passes with no failures. Its only informational note asks the merger to keep the author trailer above. - All seven original contribution commits remain in history. The diff against master contains the same 14 Telem files. It adds no dependency, lockfile, schema, or workflow change. - The managed GitHub CLI capability was missing in this run. The installed GitHub connection applied the base files, merged master, then restored the tested combined catalog. No history was rewritten. The previous head `2d32a0094` passed all remote gates and had Greptile 5/5. Those results do not verify this new head. The original PR reports live setup, discovery, settings headers, gateway calls, and context-header forwarding. This repair does not repeat those account-bound checks. The permission record still marks maintainer live qualification as outstanding. ## Risks - Telem.AI receives the six context IDs by default. The saved header policy controls forwarding. Search use is billed to the account that owns the key. - All agents on a connection share its search settings. - The merge uses master's route-test setup unchanged. The company-boundary assertions remain intact. - No schema, dependency, or workflow change is included. - Live provider evidence is attributed to the original contributor. Maintainer live qualification remains outside this CI repair. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Original contribution: Anthropic Claude Opus 5.5 (`claude-opus-5-5`), 1M-token context, through Claude Code with shell, editing, and test tools, as disclosed in #15302. - CI repair and review: OpenAI `gpt-6-astra`, through Codex with reasoning, shell, editing, and GitHub tools. The runtime does not expose the context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Yifei Ai <aiwen324@gmail.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Co-authored-by: devinfoley <139239+devinfoley@users.noreply.github.com> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
228f0e2807 |
fix(ui): make pull request review waits visible (#15375)
## Thinking Path > - Paperclip lets people supervise agent work and review its outputs. > - A task can wait for a human to review a pull request while an agent schedules checks. > - The PR work product and detected external links use separate display paths. > - A private PR can disappear from the useful sidebar view when its provider lookup fails. > - The monitor countdown says when the next check runs but omits the saved review request. > - This change shows the saved PR and requested review in the task's existing surfaces. ## Linked Issues or Issue Description **What happened?** A task kept checking whether a private GitHub PR had merged. Its work product requested board review, but the Properties sidebar used detected external objects instead. The PR URL was stored only in metadata, which the PR refresh path did not read. The composer showed a generic monitor countdown with no PR link or review action. **Expected behavior** Show saved PR links even if provider access fails. During a GitHub monitor wait, make outstanding review requests visible with the next check time and a status-check action. Stop requesting review when a PR merges or closes. **Steps to reproduce** 1. Register a PR work product with `reviewState: needs_board_review` and its URL in `metadata.url`. 2. Schedule an external-service monitor for GitHub. 3. Open the task with external-object lookup unavailable or unable to access the private repository. 4. Inspect Properties and the composer wait strip. Related work: #14469 added rich artifact cards. #8759 addresses attention on monitored task blockers. This change uses the existing work products and monitor action; it adds no blocker or approval mechanism. ## What Changed - Show saved PRs in Properties, including metadata-only links, independent of external-object availability. - Deduplicate equivalent GitHub PR URLs while preserving the saved navigation link, and put explicit review requests first. - Keep provider status and freshness on the combined PR row; keep private PR links usable when lookup fails. - Show the saved PR review request above the composer and in the existing monitor banner during a GitHub monitor wait. - Label the existing monitor action `Check status` and show check failures inline. - Suppress review prompts for merged, closed, or archived PRs even if their review flag is stale. - Refresh PR metadata using `metadata.url` and the existing `repository` alias. - Share saved work-product reads across the thread, Properties, and Artifacts. Refresh GitHub in a separate query and enrich only matching PR versions. - Start a fresh saved-row request on live invalidation so late provider responses cannot hide new artifacts or changed review requests. - Cover stalled GitHub lookups in both panels so refresh latency cannot hide saved work. - Scope monitor-check mutation state to the task so failures and late responses do not leak across navigation. - Refresh GitHub status when the displayed run finishes, including when saved PR rows have not changed. - Clarify that PR review and external release handoffs need a saved human-input interaction with an agent assigned for continuation; a flag, monitor, or handoff comment alone does not create that card. ## Verification - 386 tests pass across eleven affected UI and server suites on `184c7d8746`. Regressions cover saved links, provider status, cold-cache loading, panel reopening, monitor errors during navigation, and live updates during provider refresh. - Three live-update regressions and two run-completion regressions fail before their fixes and pass afterward. These tests use the real API client's GET coalescing and abort handling. - Full repository typecheck and build passed on the merged parent `1a49112dd2`. The latest UI changes pass UI typecheck and token gates; CI also verifies the latest build. Capability contract/inventory checks pass. - Storybook build and browser checks passed before the query race fix. Browser checks cover desktop and 390px mobile review waits, plus video, mixed-file, and empty artifact galleries after background refresh. - The earlier full local `pnpm test:run` was stopped under disk pressure. It reported failures outside the changed suites; an isolated skill-cache run reproduced three existing macOS permission failures. The full local suite was not rerun for this follow-up. - Latest-head CI passes on `184c7d8746`, including the browser shards and canary dry run. Apex is 5/5 with no actionable findings and no unresolved threads. All eight reported findings are addressed. ## Risks - This uses saved PR review state. When GitHub access fails, the saved state can remain stale until the agent updates it. The link remains visible and the status check remains available. - `Check status` wakes the existing monitor owner. It does not merge a PR, accept an approval, or mark the task done. - No schema, permissions, or scheduler behavior changes. ## Model Used OpenAI GPT-6 via Codex, with reasoning, repository analysis, code editing, and test execution. The exact deployment model ID and context window were not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass — 277 affected tests - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8cbd21b3e7 |
feat(apps): add Superagent connection (#15394)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents reach external services through the Apps catalog. Each catalog entry is a reviewed `AppDefinition` that connects a provider's hosted MCP server to Paperclip's shared vault, grants, policies, gateway, and audit trail. > - Superagent (`superagent.sh`) is a security platform. Its hosted MCP server lets agents read and triage security findings, start red-team reports, run Contributor Trust and dependency update jobs, and score web pages, email, files, skills, MCP repositories, and packages before an agent trusts them. > - Superagent is not in the catalog. Operators must use the generic "Connect your own MCP server" flow. That flow has no branding, no key guidance, and no warning about billable or destructive tools. > - The connector playbook supports this provider with existing definition fields. The server accepts an organization API key as a bearer header. It publishes no OAuth authorization-server metadata, so browser sign-in is not possible. > - This pull request adds the Superagent definition, its artwork, the research and permission-review ledger rows, documentation, deterministic tests, and one reviewed risk-classification rule. > - The benefit is a branded, governed Superagent connection with clear key guidance and a warning about tools that cost credits or delete data. ## Linked Issues or Issue Description **Problem or motivation** Teams that use Superagent for PR security, red teaming, and agent guardrails want their Paperclip agents to read findings, return structured reports, and score content before they use it. Superagent is not in the Apps catalog. Operators must paste the MCP URL and an `Authorization` header into the generic remote-MCP flow. That flow gives no branding and no provider guidance. It also does not tell the operator that the key reaches the whole organization, or that some tools consume credits or permanently delete findings. **Proposed solution** Add a catalog-only Superagent connection that follows the connector playbook. It has one method: a customer organization API key (`sk_live_...`), sent as an `Authorization: Bearer` header to `https://www.superagent.sh/mcp`. The field helper text explains that Superagent keys are not scoped. The method warning tells operators to set billable and destructive actions to Ask first before agents run unattended. All discovered tools stay governed by the normal per-action policies. **Alternatives considered** A browser sign-in method was not added. The server's protected-resource metadata names `https://superagent.sh` as its authorization server, but that origin publishes no `oauth-authorization-server` or `openid-configuration` document, so Paperclip cannot discover OAuth endpoints. A plugin was not needed because the connection needs no custom UI, tables, workers, or webhooks. Relying on the generic risk classifier was not enough. Several Superagent mutations (`triage_finding`, `scan_*`, `restore_agent_builtin_rule`) use names that it reads as reads, so a narrow reviewed Superagent rule was added instead. **Roadmap alignment** This extends the existing self-serve remote-MCP connection catalog. It does not overlap planned core work. ## What Changed - Added the `superagent` row to `packages/shared/src/self-serve-mcp-research.json` (API-key auth, risk tier S4). - Added the `superagent` provider to `scripts/ingest-app-definitions.mjs` (category, key placement and placeholder, console links, guidance, description). Regenerated `packages/shared/src/app-definitions/superagent.json` and the generated registry. - Added the `superagent/mcp-api-key` permission review to `doc/connections/tool-method-permission-reviews.json`, with key-permission text and evidence links. - Added Superagent's official mark (`ui/public/brands/apps/superagent.png`, the 460×460 avatar of the official `superagent-ai` GitHub organization) and the brand manifest entry. - Added gallery copy for the Superagent card. - Added a reviewed Superagent rule to `classifyRisk` in `server/src/services/tool-access.ts`. Only `list_*` and `get_*` tools, and tools that Superagent marks read-only, are reads. `delete_*` and `revoke_agent_client` are destructive. All other tools are writes, so billable and rule-changing tools can be set to Ask first. - Put the Ask-first advice in the API-key helper text, because the key form shows helper text and not method warnings. - Added `doc/connections/SUPERAGENT.md` (transport and auth, why there is no OAuth, administrator setup, capabilities and policy, manifest, brand provenance, validation hook). Linked it from the connections README and the permission audit. - Tests: definition shape, store visibility and artwork, URL recognition, the bearer header on discovery with the key kept out of connection config, read/write/destructive classification of fixture tools, the Superagent risk rule (including `triage_finding` and `restore_agent_builtin_rule`), the visible Ask-first advice and API-key gating of the connect form, and the pinned catalog counts. ## Verification - `pnpm exec vitest run packages/shared/src/app-definitions.test.ts packages/shared/src/app-definitions-url.test.ts ui/src/pages/apps/AppsConnect.test.tsx ui/src/pages/apps/Browse.test.tsx ui/src/lib/app-brand-assets.test.ts ui/src/pages/apps/AppLogo.brand-assets.test.tsx`: 317 passed. - `pnpm exec vitest run server/src/__tests__/tool-access-service.test.ts`: 385 passed. - `node scripts/check-app-brand-assets.mjs` and `node --test scripts/app-brand-validation.test.mjs`: passed. - `pnpm --filter @paperclipai/shared typecheck`, `pnpm --filter @paperclipai/server typecheck`, and `pnpm --filter @paperclipai/ui typecheck`: clean. - Manual: in a local instance, open Apps → Browse and confirm the Superagent card and icon. Open `/apps/connect?source=superagent`. Confirm the single API-key method, and confirm that Connect enables only after a key is entered. - Live metadata probe on 2026-10-06: an unauthenticated `initialize` on `https://www.superagent.sh/mcp` returns 401 with `resource_metadata="https://www.superagent.sh/.well-known/oauth-protected-resource"`. That document returns 200. No authorization-server metadata exists at the named issuer. ## Risks - Low risk to existing providers. The change is additive catalog data plus tests. The generated registry only gains one import. The new risk rule runs only for Superagent connections. - A Superagent key reaches its whole organization. Some tools consume credits (`create_*_report`, `triage_finding`) or delete data permanently (`delete_finding`). Every action starts Allowed under the current product default. The key helper text tells operators to set these actions to Ask first. - The permission-review ledger records live proof as not run. No Superagent account was used. The lifecycle checklist in `doc/connections/SUPERAGENT.md` needs a documented pass before the entry is fully qualified. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Claude Opus 5.5 (`claude-opus-5-5`, 1M context) in Claude Code, with extended thinking and tool use (shell, file editing, web fetch, browser checks). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
ee3c728b95 |
fix(ui): keep No project available in new composer (#15382)
Keep the no-project action first and available during project search, while preserving matching-project keyboard selection. Add focused regression coverage for ordering, projectless submission, and filtered Enter behavior. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
582911ba74 |
fix(access): give Operators default company editing permissions (#15377)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Human company members receive permission grants from their role preset. > - Human invitations use the Operator role by default. > - The old Operator preset only granted task assignment, so ordinary members could not edit agents or manage connections. > - This pull request gives Operators company editing and audit access while keeping join approval and member-permission management separate. > - Members can configure their agents and accounts without an Owner role. ## Linked Issues or Issue Description **What existing behavior does this improve?** The default permission grants for human Operators and the invitation role description. **Subsystem affected** server/ permission presets and invitation authorization, plus ui/ invitation and tool access gates. **Current behavior** Operators receive only `tasks:assign` as an explicit default grant. Agent configuration and managed account setup fail their permission checks. **Proposed behavior** Operators receive agent creation/configuration, skills, environments, invitations, task assignment, pipelines, connections, tool management/use, and audit grants. The preset excludes `joins:approve` and `users:manage_permissions`. Operators can invite Operators and Viewers. Inviting a role with either excluded power requires that power, so invitations cannot bypass the restriction. **Reason and benefit** Ordinary company members can edit company work and connect accounts through the existing routes. **Breaking changes** The Operator preset grants more permissions. The existing startup and Cloud sign-in seeding paths can insert these missing grants for existing Operators. They retain custom grant scopes. Explicit invitation grants still take precedence. Inviting Admins now requires join approval; inviting Owners also requires member-permission management. This PR adds no migration or new backfill path. Related: #5945 describes missing role grants in another membership entry point. This PR changes the preset and does not change that endpoint. ## What Changed - Expanded the Operator preset to 14 company editing, invitation, tool, and audit grants. - Kept join approval and member-permission management out of the preset. - Required those powers when an invitation's selected human role includes them, preventing Operators from delegating the excluded powers through Owner/Admin invitations. - Updated the invitation role copy and the product specifications. - Allowed active Operators (including legacy Member roles) through the invite shortcut and advanced tool/profile UI gates. Kept inactive memberships, Viewers, and other companies excluded; API grants remain authoritative. - Added tests for the default Operator grants, both excluded actions, agent editing, company boundaries, and invitation role authorization. - Simplified adapter-route test imports so authorization errors and the error handler use one module graph; retained the upstream mock-initialization fix. - Kept Owner/Admin/Viewer presets and explicit invitation grants unchanged. Added no EE code or database migration. ## Verification - Permission and invitation tests: 38 passed across access service, invitation defaults, and invitation creation routes. - UI access tests: 61 passed across invitation shortcuts, Operator tool/profile access, and invitation UI. - Adapter route tests: 15 passed after resolving the upstream test setup conflict. - `pnpm -r typecheck`, `pnpm build`, and `pnpm check:token-gates`: passed on the final branch. - The initial local `pnpm test:run` general-server batch reported three tool-access failures and eight upload-helper timeouts. Both complete affected suites passed in isolation on the final branch: 403 tests. The initial full local run was not green. - Both complete tool-access and adapter route suites also passed on an untouched snapshot of base `88ff98b83`: 387 tests. The earlier tool failures did not reproduce there. - [All CI gates passed](https://github.com/paperclipai/paperclip/actions/runs/37527422346) for `068be153b9c2a064f2aaccc0627fdf458e73747e`, including the previously failing serialized adapter lane, general tests, E2E, typecheck, build, Runner checks, and canary dry run. - Greptile scored that exact commit 5/5. Superagent passed. Both invitation review threads are resolved. ## Risks - Operators gain broad company editing and invitation access by design. - Existing default seeding can add the new missing grants during startup or Cloud sign-in. It does not replace existing scopes. - Company boundaries and the two excluded permission checks still apply. - An Admin without member-permission management can no longer invite an Owner. Operators retain invitation access for Operator/Viewer roles and agent-only invites. - This PR changes company permissions. Cloud workspace invitation rules remain separate. ## Model Used OpenAI Codex, GPT-6. The exact deployment model ID and context window were not exposed in this session. Used repository inspection, code editing, shell execution, and test tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting a merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
2ca0d26a99 |
fix(connections): recover missing personal AI credentials in chat (#15376)
Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
e38d6d16b6 |
feat(connections): add advanced provider setup and live browser qualification (#15341)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Users connect accounts and choose an agent harness and model. > - The runtime change in #14970 supports custom providers on those connections. > - Normal setup must stay simple while advanced users can choose a compatible gateway. > - Shared connector rows and access controls keep these choices consistent. > - This pull request refines the agent setup UI and adds review stories and repeatable browser qualification. > - The qualification checks real tools and downloaded outputs, not only a successful run status. ## Linked Issues or Issue Description Refs #14970, #37, #13083, #14104, #14565, #12692. The core implementation in #14970 is merged. This branch incorporates its squash commit and targets `master`. Both PRs contain our implementation. #14016 is a reference only and is not a dependency. This PR has 96 changed files. ## What Changed - Complete model-provider connector presentation beside other connectors. Each row uses the existing Connect action and connection list. Tags are stored without category UI. The base PR includes the provider forms and routes. - Show persistent Subscription, API Key, and Advanced choices. Label Advanced as Custom Gateway. Reuse provider logos, connection lists, and permissions controls. Default access to the organization and all agents when permitted; keep narrowing controls under Advanced. - Keep Configure reachable before subscription sign-in, so users can select a supported environment when the default cannot sign in. Testing and saving still require a connection. Show the execution environment in Configure. Preserve the confirmed Connect choice. Editing a method, credential, saved account, or advanced choice requires that current choice to connect before testing or saving. Use matching model and thinking-effort dropdowns and retain connection icons in selected values. - Preserve the new harness model default when switching an existing OpenCode agent to Codex or Claude, and resolve user-selected model names with the effective harness. - Load popular OpenRouter models through the shared connection-model discovery path. Keep explicit model lists and manual model entry available. - Group onboarding, connection setup, agent runtime, management, recovery, and production-component stories under AI Connections / Provider routing. - Add an explicit-only provider-connections browser suite for managed local or existing local/staging targets. Use private browser profiles and credential handoffs. Support human-assisted subscription sign-in without sharing passwords or tokens in reports. - Verify persisted connection identity, runtime probes, tool execution, exact artifact bytes, completion, and context-dependent follow-up. Retain source/model provenance, cost bounds, closed error diagnostics, original failures, and cleanup evidence. - Add Gemini startup-model and skill-root fixes, Grok private-history detection, ACP filesystem regression fixtures, selected-workspace handling for local Hermes, and artifact-helper workspace fallback. - Keep managed Grok runtime homes disposable. Remove host-side transcript retention/restoration because private file modes do not isolate same-user agent processes. Ignore earlier development archives and use a fresh task handoff when history is unavailable. Verify the absence of restored transcripts with a separate same-user process. - Capture stopped-run diagnostics before deleting an attached-company fixture agent. Track creation and owned sign-in receipts; revoke only this attempt's accounts and never adopt a concurrent campaign's newly created account. Preserve failure signals and final status through cleanup. - Require the requested environment in the saved agent and every run, including follow-ups. Reject a forced incompatible target. Keep one cancellation state through startup, every cell, reporting, and teardown for SIGINT, SIGTERM, and SIGHUP. Stop further paid cells after interruption. Document qualification limits. ## Verification - Current head `b3bb3e94d577d43d9965a6b9daba039f599b2e49` includes master `d9f600043`. The security fix in `a758fde31` passes full workspace typecheck, production build, and 119 connection/Grok regressions. The unchanged UI passes all 126 configuration/model-discovery tests and token gates. The final published-guide correction passes Grok adapter typecheck. Earlier head `eebd8225c` passed the complete deterministic runner suite (1,404 Vitest tests and 128 Node tests) and all CI jobs. Current-head CI run `37520147514` passed all 47 jobs, including the full sharded Vitest and browser matrix, production build, and canary dry run. All 55 checks completed: 53 successes and two expected skips. The current-head security scan passed, Greptile is 5/5, and no review threads remain open. - A separate same-user process reproduced reading a restored Grok transcript before the security fix. The regression now finds no transcript. Existing fresh-session fallback and ordinary session metadata behavior pass. - The final account-choice and cleanup fixes pass 85 setup tests and 26 qualification-harness tests. Regressions verify that editing a connection invalidates confirmation, Configure remains reachable before sign-in, diagnostics are captured before fixture deletion, and concurrent campaigns cannot adopt or revoke each other's accounts. UI and E2E typechecks pass. - The Storybook build and actual Chromium production-component stories passed during this change. Review the neighboring AI Connections / Provider routing stories, regular connector rows, three connection modes, model discovery, and the single execution-environment control in Configure. - Cancellation smoke verified authenticated cleanup before browser close for SIGINT, SIGTERM, and SIGHUP. Regressions cover interruption during startup and reporting, missing-file ACP resource errors, and preserved permission denials. Both ACP runtime versions and 54 ACPX/Grok regressions passed. The deterministic connection-intent browser suite passed two tests. - Historical local qualification retained 43 passing API/gateway cells out of 46, with downloaded outputs and follow-up receipts. These attempts span earlier builds; they do not qualify this exact commit or staging. Subscription combinations, Gemini overloads, and the unresolved follow-up failure remain recorded rather than counted as passing. - Use `pnpm test:e2e:runner -- --list --suite provider-connections` to inspect the matrix. Follow `tests/runner-e2e/PROVIDER-CONNECTIONS.md` for credentials, target URL, sign-in assistance, budget, evidence, and cleanup. Paid live tests remain opt-in. ## Risks - The core implementation in #14970 is merged. This PR adds no database migration of its own. - Subscription login needs an interactive provider session. Dedicated accounts and staging qualification remain follow-up work; this PR does not certify every login combination for production. - Managed Grok transcript resume is deferred until provider history has an OS isolation or authorized broker solution. Follow-ups start fresh with Paperclip task context; earlier live Grok results do not qualify this behavior. - Gemini CLI 0.58.0 has an upstream ACP new-file error conversion defect. Live overloads and one unresolved follow-up timeout remain recorded. The stock CLI is unchanged, and those cases are not marked as passing. - Real-provider tests spend credits and use private credential/evidence directories. The launcher requires explicit selection and checks target ownership. It must not attach to a developer's database by accident. - OpenClaw Gateway, Hermes Gateway, Claude Managed, AWS AgentCore, Process, HTTP, and legacy ACPX local remain outside custom provider setup. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, code execution, and browser testing. The exact deployment model ID and context window size were not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a7a2ed63a0 |
fix(ui): restore wheel and touch scrolling in task selectors (#15374)
Restore wheel and touch scrolling in nested task selectors and keep mobile picker controls reachable above the keyboard. Stabilize the affected server route test harnesses for Vitest 5. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
88ff98b83d |
feat(apps): add read-only Enterpret MCP with OAuth and org tokens (#13906)
## Thinking Path > - Paperclip agents need governed access to Enterpret's hosted MCP tools for customer feedback. > - The official Enterpret MCP is read-only. Enterpret Agent’s beta write MCP is a separate service and is outside this PR. > - This PR supports both organization auth tokens and personal browser OAuth for the official MCP. > - Enterpret committed to fixing OAuth scope behavior and reducing its token-revocation cache window. The provider fixes will follow after connector deployment and do not gate this release. Paperclip functional validation remains required. New tools stay quarantined on refresh. > - This PR combines the original connector and Vivek's token-path work on current master, with live self-hosted QA recorded in the validation guide. ## Linked Issues or Issue Description No public issue covers this connector. This PR incorporates the work from #14737. The independent OAuth scope-provenance fix is #14059; it is not required for the organization-token path. #13890 is a merged connector precedent. #13692 provides the connector-authoring skill. **Problem or motivation** The generic MCP setup can reach Enterpret, but it has no reviewed Enterpret entry, method-specific setup, artwork, or policy defaults. The original connector branch had also fallen behind master. **Proposed solution** Publish the Enterpret catalog entry with an organization auth token as the selectable method. Support browser OAuth through instance-local DCR on the same read-only endpoint. Default `run_graph_query` to Ask first and quarantine tools newly found on refresh. ## What Changed - Added the Enterpret app definition, generated registry, official icon, Browse card, setup copy, tests, and Storybook states. - Combined #14737 and resolved the current-master conflicts, including the connector permission audit. - Preserved catalog quarantine on token reconnect and updated expiry guidance to follow the token's dashboard expiry. The QA token displayed three months, contradicting the former six-month copy. - Enabled personal OAuth for the official read-only MCP. Removed the write-capability inference and OAuth draft-only setup restriction. The beta Agent endpoint is excluded. - Preserved tool quarantine during token reconnect, with a regression test for existing quarantined and newly discovered tools. - Added OAuth scope and revocation-cache disclosures and an interactive Storybook retry check. - Merged master `d9f600043` into the existing PR branch without rewriting history. Resolved four conflicts, regenerated the combined registry, and retained the model-provider assertions with a catalog count of 66. - Recorded live QA and its limits in [doc/connections/ENTERPRET.md](https://github.com/paperclipai/paperclip/blob/950bacc52/doc/connections/ENTERPRET.md). ## Verification Current head: `6c73d3084202ccbb25080e7a4e9c3467bc1341dd`. All 513 focused connector, discovery, OAuth, policy, shared-definition, and UI-handoff tests passed across 11 suites. Full repository typecheck, build, Storybook build, token gates, module boundaries, Node policy, no-git-push policy, and typecheck-build-gaps passed. Release-registry tests passed 129/129 with modern Bash. The four failures with macOS Bash 3.2 also reproduced on untouched master. The full test suite passed in CI on this exact head, including every general-test and serialized-server shard and all eight E2E shards. The duplicate local `pnpm test:run` was stopped after these remote test lanes passed; it did not complete locally. The combined registry was regenerated. Regeneration also reproduces existing Linear and AgentMail definition drift on untouched master; those unrelated changes were excluded. On an isolated self-hosted Paperclip instance, an organization token connected and discovered eight actions. `get_organization_details` returned the intended organization before and after reconnect. Catalog refresh, invalid-token failure and valid-token recovery, activity logging, local disable, and local removal were exercised. No feedback quotes or graph query were requested. Paperclip revoked both QA connection secrets and archived the QA connections. The token-path agent access preview correctly showed an action Off, but the board Test endpoint still invoked it as the board user. This shared Test-path mismatch is documented; it is not evidence that a real agent can bypass policy. An actual token-path agent-session denial, a real agent process, expiry recovery, server/VPS, and Cloud remain unverified. Enterpret's dashboard confirmed the QA token was revoked, but the same token immediately reconnected and completed a safe read. Vivek explained this as a 24-hour validity cache and committed to reducing the window. The new maximum delay and deployment remain unverified. The earlier OAuth check reported `mcp:write` and `email` for an `mcp:read` request. This is a scope mismatch; it did not prove access to write tools. On 2026-10-01 the account holder relayed Vivek’s clarification that the official MCP is read-only and the beta Agent write MCP is separate. ## Risks The organization token can read the organization's customer feedback, and provider-side revocation did not immediately stop access. Operators should check each token's displayed expiry and remove Paperclip's local connection when retiring access. The account holder confirmed the provider fixes are fast follows after deployment. Release can proceed with the current scope reporting and up-to-24-hour revocation delay documented, once fresh OAuth authorization/refresh, real-agent allow/deny, CI, and review pass. Retest the reduced revocation window after Enterpret deploys it. Scope labels alone must not be used to claim write access. Greptile is 5/5 on head `6c73d3084` with no actionable findings. All current-head CI checks pass, including the canary dry run. Live agent-session denial and fresh OAuth refresh remain explicit deployment QA items. The board Test panel does not prove agent authorization. Provider scope-reporting and revocation-cache fixes are accepted fast follows after deployment. ## Model Used Original connector and validation record: Claude Opus 5 through Claude Code. Integration, current-master reconciliation, and token-path QA: OpenAI Codex, GPT-6 family (exact host variant not exposed), using repository tools and an isolated runtime. Contributor commits retain their authors. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked existing issues or described the issue in-PR - [x] I have not referenced internal or instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal issue ID - [x] I have run relevant local tests and checks - [x] I have updated relevant documentation - [x] I have considered and documented the risks above - [x] All current-head CI gates are green - [x] Greptile is 5/5 with no actionable follow-ups If squash-merging, preserve contributor credit in the squash message: Co-Authored-By: Vivek Kaushal <kaushalvivek@users.noreply.github.com> --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: Vivek Kaushal <vivek@enterpret.com> |
||
|
|
854af7df19 |
build(deps-dev): bump vitest from 4.1.11 to 5.0.3 (#12969)
Bumps [vitest](https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest) from 4.1.11 to 5.0.3. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/vitest-dev/vitest/releases">vitest's releases</a>.</em></p> <blockquote> <h2>v5.0.3</h2> <h3> 🐞 Bug Fixes</h3> <ul> <li>Isolate <code>result.status</code> between <code>repeats</code> runs - by <a href="https://github.com/hi-ogawa"><code>@hi-ogawa</code></a>, <strong>Hiroshi Ogawa</strong> and <strong>Codex (GPT-6)</strong> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11218">vitest-dev/vitest#11218</a> <a href="https://github.com/vitest-dev/vitest/commit/5dbebe9e3"><!-- raw HTML omitted -->(5dbeb)<!-- raw HTML omitted --></a></li> <li>Don't print an interceptor warning in browser mode - by <a href="https://github.com/sheremet-va"><code>@sheremet-va</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11377">vitest-dev/vitest#11377</a> <a href="https://github.com/vitest-dev/vitest/commit/15cc006aa"><!-- raw HTML omitted -->(15cc0)<!-- raw HTML omitted --></a></li> <li>Don't retry when <code>test.fails</code> expectedly failed - by <a href="https://github.com/hi-ogawa"><code>@hi-ogawa</code></a>, <strong>Hiroshi Ogawa</strong> and <strong>Codex (GPT-6)</strong> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11219">vitest-dev/vitest#11219</a> <a href="https://github.com/vitest-dev/vitest/commit/b24585f08"><!-- raw HTML omitted -->(b2458)<!-- raw HTML omitted --></a></li> <li>Scope cache key generators to projects - by <a href="https://github.com/ecoyoung"><code>@ecoyoung</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11281">vitest-dev/vitest#11281</a> and <a href="https://redirect.github.com/vitest-dev/vitest/issues/11301">vitest-dev/vitest#11301</a> <a href="https://github.com/vitest-dev/vitest/commit/92ba7fc1d"><!-- raw HTML omitted -->(92ba7)<!-- raw HTML omitted --></a></li> <li><strong>browser</strong>: <ul> <li>Delay server <code>listen</code> until tests start running - by <a href="https://github.com/sheremet-va"><code>@sheremet-va</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11366">vitest-dev/vitest#11366</a> <a href="https://github.com/vitest-dev/vitest/commit/7d8ed3e9b"><!-- raw HTML omitted -->(7d8ed)<!-- raw HTML omitted --></a></li> <li>Check mock path boundaries - by <a href="https://github.com/saryn17"><code>@saryn17</code></a>, <strong>Ryosei Sato</strong> and <a href="https://github.com/sheremet-va"><code>@sheremet-va</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11361">vitest-dev/vitest#11361</a> and <a href="https://redirect.github.com/vitest-dev/vitest/issues/11362">vitest-dev/vitest#11362</a> <a href="https://github.com/vitest-dev/vitest/commit/1c3888bce"><!-- raw HTML omitted -->(1c388)<!-- raw HTML omitted --></a></li> <li>Keep config of browser-consumed environments - by <a href="https://github.com/kasperpeulen"><code>@kasperpeulen</code></a> and <strong>Claude</strong> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11378">vitest-dev/vitest#11378</a> <a href="https://github.com/vitest-dev/vitest/commit/aafc0996f"><!-- raw HTML omitted -->(aafc0)<!-- raw HTML omitted --></a></li> <li>Ignore page crash while cancelling - by <a href="https://github.com/sheremet-va"><code>@sheremet-va</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11386">vitest-dev/vitest#11386</a> <a href="https://github.com/vitest-dev/vitest/commit/7c36748fa"><!-- raw HTML omitted -->(7c367)<!-- raw HTML omitted --></a></li> <li><code>toMatchScreenshot</code> uses wrong reference on retried tests - by <a href="https://github.com/macarie"><code>@macarie</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11393">vitest-dev/vitest#11393</a> <a href="https://github.com/vitest-dev/vitest/commit/c22aba992"><!-- raw HTML omitted -->(c22ab)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>cache</strong>: <ul> <li>Revalidate imports of cached modules - by <a href="https://github.com/sheremet-va"><code>@sheremet-va</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11381">vitest-dev/vitest#11381</a> <a href="https://github.com/vitest-dev/vitest/commit/38f98855f"><!-- raw HTML omitted -->(38f98)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>deps</strong>: <ul> <li>Pin <code>why-is-node-running</code> to <code>3.2.1</code> to avoid users running into <code>ERR_PNPM_TRUST_DOWNGRADE</code> - by <a href="https://github.com/AriPerkkio"><code>@AriPerkkio</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11403">vitest-dev/vitest#11403</a> <a href="https://github.com/vitest-dev/vitest/commit/f6c9a4977"><!-- raw HTML omitted -->(f6c9a)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>expect</strong>: <ul> <li>Pass current equality testers to <code>expect.extend</code> asymmetric matchers - by <a href="https://github.com/hi-ogawa"><code>@hi-ogawa</code></a>, <strong>Hiroshi Ogawa</strong> and <strong>Claude</strong> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11401">vitest-dev/vitest#11401</a> <a href="https://github.com/vitest-dev/vitest/commit/3e794a96b"><!-- raw HTML omitted -->(3e794)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>jsdom</strong>: <ul> <li>Support Blob on jsdom 30.1 - by <a href="https://github.com/sheremet-va"><code>@sheremet-va</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11379">vitest-dev/vitest#11379</a> <a href="https://github.com/vitest-dev/vitest/commit/6c49b7197"><!-- raw HTML omitted -->(6c49b)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>pool</strong>: <ul> <li>Preserve unique pool ids when <code>groupOrder</code> is set - by <a href="https://github.com/mtorp"><code>@mtorp</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11392">vitest-dev/vitest#11392</a> <a href="https://github.com/vitest-dev/vitest/commit/50312ebb4"><!-- raw HTML omitted -->(50312)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>ui</strong>: <ul> <li>Split-pane handle overlapping iframe - by <a href="https://github.com/macarie"><code>@macarie</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11221">vitest-dev/vitest#11221</a> <a href="https://github.com/vitest-dev/vitest/commit/f91db0dfd"><!-- raw HTML omitted -->(f91db)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>vitest</strong>: <ul> <li>Remove root temp dir on close - by <a href="https://github.com/abhinav-phi"><code>@abhinav-phi</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11248">vitest-dev/vitest#11248</a> <a href="https://github.com/vitest-dev/vitest/commit/7c7119cf7"><!-- raw HTML omitted -->(7c711)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>vm</strong>: <ul> <li>Do not optimize deps from index.html - by <a href="https://github.com/ezefernandezyf"><code>@ezefernandezyf</code></a>, <strong>Hiroshi Ogawa</strong> and <strong>Codex (GPT-6)</strong> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11329">vitest-dev/vitest#11329</a> and <a href="https://redirect.github.com/vitest-dev/vitest/issues/11360">vitest-dev/vitest#11360</a> <a href="https://github.com/vitest-dev/vitest/commit/caf2887de"><!-- raw HTML omitted -->(caf28)<!-- raw HTML omitted --></a></li> <li>Don't reuse scripts across vite environments - by <a href="https://github.com/MO2k4"><code>@MO2k4</code></a>, <strong>Martin Oehlert</strong> and <strong>Claude</strong> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11395">vitest-dev/vitest#11395</a> <a href="https://github.com/vitest-dev/vitest/commit/346d3896b"><!-- raw HTML omitted -->(346d3)<!-- raw HTML omitted --></a></li> </ul> </li> </ul> <h5> <a href="https://github.com/vitest-dev/vitest/compare/v5.0.2...v5.0.3">View changes on GitHub</a></h5> <h2>v5.0.2</h2> <h3> 🐞 Bug Fixes</h3> <ul> <li>Bind <code>process</code> in case global is overwritten - by <a href="https://github.com/AriPerkkio"><code>@AriPerkkio</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11343">vitest-dev/vitest#11343</a> <a href="https://github.com/vitest-dev/vitest/commit/0b79231ad"><!-- raw HTML omitted -->(0b792)<!-- raw HTML omitted --></a></li> <li><strong>detect-async-leaks</strong>: <ul> <li>Ignore <code>process.stdio</code> handles - by <a href="https://github.com/AriPerkkio"><code>@AriPerkkio</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11333">vitest-dev/vitest#11333</a> <a href="https://github.com/vitest-dev/vitest/commit/0fd6b9790"><!-- raw HTML omitted -->(0fd6b)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>expect</strong>: <ul> <li>Fix <code>toMatchObject</code> with asymmetric matchers - by <a href="https://github.com/ShreeBohara"><code>@ShreeBohara</code></a>, <strong>Claude Opus 5</strong>, <a href="https://github.com/hi-ogawa"><code>@hi-ogawa</code></a>, <strong>Hiroshi Ogawa</strong> and <strong>Codex (GPT-5)</strong> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11100">vitest-dev/vitest#11100</a> <a href="https://github.com/vitest-dev/vitest/commit/42523289e"><!-- raw HTML omitted -->(42523)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>jsdom</strong>: <ul> <li>Fix <code>Request</code> with <code>Blob</code> body on jsdom 28+ - by <a href="https://github.com/harshit-d3v"><code>@harshit-d3v</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11295">vitest-dev/vitest#11295</a> <a href="https://github.com/vitest-dev/vitest/commit/d1c3ecc93"><!-- raw HTML omitted -->(d1c3e)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>reporter</strong>: <ul> <li><code>agent</code> to respect <code>--silent</code> - by <a href="https://github.com/Raj4478"><code>@Raj4478</code></a> and <a href="https://github.com/AriPerkkio"><code>@AriPerkkio</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11271">vitest-dev/vitest#11271</a> <a href="https://github.com/vitest-dev/vitest/commit/5b95efb6d"><!-- raw HTML omitted -->(5b95e)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>reporters</strong>: <ul> <li>Handle concurrent <code>createReport</code> calls - by <a href="https://github.com/7rulnik"><code>@7rulnik</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11278">vitest-dev/vitest#11278</a> <a href="https://github.com/vitest-dev/vitest/commit/e8e556ff7"><!-- raw HTML omitted -->(e8e55)<!-- raw HTML omitted --></a></li> <li><code>hanging-process</code> to use ESM entrypoint - by <a href="https://github.com/AriPerkkio"><code>@AriPerkkio</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11316">vitest-dev/vitest#11316</a> <a href="https://github.com/vitest-dev/vitest/commit/4e91e5668"><!-- raw HTML omitted -->(4e91e)<!-- raw HTML omitted --></a></li> </ul> </li> <li><strong>spy</strong>: <ul> <li>Fix stack overflow when spying <code>Set.prototype.add</code> - by <a href="https://github.com/fengmk2"><code>@fengmk2</code></a> in <a href="https://redirect.github.com/vitest-dev/vitest/issues/11299">vitest-dev/vitest#11299</a> <a href="https://github.com/vitest-dev/vitest/commit/a0a939653"><!-- raw HTML omitted -->(a0a93)<!-- raw HTML omitted --></a></li> </ul> </li> </ul> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/vitest-dev/vitest/commit/33cadea62e8763c455c7fca38d9ab1dda87c5f75"><code>33cadea</code></a> chore: release v5.0.3 (<a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11409">#11409</a>)</li> <li><a href="https://github.com/vitest-dev/vitest/commit/346d3896b65c3c907174447a035807342799f346"><code>346d389</code></a> fix(vm): don't reuse scripts across vite environments (<a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11395">#11395</a>)</li> <li><a href="https://github.com/vitest-dev/vitest/commit/f6c9a4977ad3363f796a737834572e54c6ad5c18"><code>f6c9a49</code></a> fix(deps): pin <code>why-is-node-running</code> to <code>3.2.1</code> to avoid users running into `...</li> <li><a href="https://github.com/vitest-dev/vitest/commit/062c75d8b63519211d951d8293ea81b5a9e3c124"><code>062c75d</code></a> chore: fix standalone docs build, update exports maps (<a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11394">#11394</a>)</li> <li><a href="https://github.com/vitest-dev/vitest/commit/caf2887dee8987a60118d53933f6e9cabd6b3e2a"><code>caf2887</code></a> fix(vm): do not optimize deps from index.html (fix <a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11329">#11329</a>) (<a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11360">#11360</a>)</li> <li><a href="https://github.com/vitest-dev/vitest/commit/50312ebb4eca6a98f6d0b2b61d5d9d38cbbabcef"><code>50312eb</code></a> fix(pool): preserve unique pool ids when <code>groupOrder</code> is set (<a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11392">#11392</a>)</li> <li><a href="https://github.com/vitest-dev/vitest/commit/7c36748fad1eae9687312f2f7ceadce6ec88b5df"><code>7c36748</code></a> fix(browser): ignore page crash while cancelling (<a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11386">#11386</a>)</li> <li><a href="https://github.com/vitest-dev/vitest/commit/92ba7fc1df16a4fa5bbee3f198c582fbd56689d8"><code>92ba7fc</code></a> fix: scope cache key generators to projects (fix <a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11281">#11281</a>) (<a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11301">#11301</a>)</li> <li><a href="https://github.com/vitest-dev/vitest/commit/38f98855fa9cd7fd376afb84094eba0fda256a74"><code>38f9885</code></a> fix(cache): revalidate imports of cached modules (<a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11381">#11381</a>)</li> <li><a href="https://github.com/vitest-dev/vitest/commit/b24585f08f2ea267746a2d6ca0e43edcbb29726f"><code>b24585f</code></a> fix: don't retry when <code>test.fails</code> expectedly failed (<a href="https://github.com/vitest-dev/vitest/tree/HEAD/packages/vitest/issues/11219">#11219</a>)</li> <li>Additional commits viewable in <a href="https://github.com/vitest-dev/vitest/commits/v5.0.3/packages/vitest">compare view</a></li> </ul> </details> <br /> --------- Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: Priya Raman <priya.raman@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
a590ab769d |
Give managed agents persistent cryptographic identities (#15352)
Give agents persistent Ed25519 identities encrypted with the existing instance master key. Create keys transactionally for new agents and lazily before supported managed runs, expose public identities in the API and agent UI, and protect private material during runtime delivery and output persistence. Preserve identities in recovery backups while giving imported and development-cloned agents fresh keys. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
9f7057e122 |
feat(connections): configure custom model providers across agent harnesses (#14970)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents use a harness, a model, and a credential to run tasks. > - Connections already store credentials and control who can use them. > - Custom providers also need an endpoint and a supported API format. > - A per-agent endpoint would duplicate credentials and access rules. > - This pull request stores routing on the connection and projects it into the harness. > - Isolated credentials and protocol checks keep the selected connection authoritative. ## Linked Issues or Issue Description Refs #37, #13083, #14104, #14565, #12692. #14016 is a reference only. This PR has its own schema, vault persistence, routing validation, runtime projection, and tests. None of the five commits in #14016 is an ancestor of this branch. We do not depend on or plan to merge it. #14967 addresses task-pinned account pools. #14422 addresses another provider integration. This is the first of two linked PRs. Merge this connection change before #15341, which refines agent setup and adds the qualification harness. The split keeps each review below 100 changed files. Provider catalog entries have usable setup forms in this PR. Local browser subscription sign-in is included. ## What Changed - Store non-secret routing metadata on AI connections. Vault provider API keys, including Bedrock bearer API keys. Reject general AWS access keys. - Enforce company, owner, human audience, agent access, connection status, and protocol checks before resolving credentials. Keep reconnect destinations immutable and retain connection identity during key rotation. - Project OpenRouter and compatible custom endpoints into Codex, Claude, OpenCode, and local Hermes. Carry these settings through both legacy and native runner transports. Clear conflicting host credentials and redact keys from diagnostics. - Preserve older OpenRouter accounts and native personal defaults. Add Google API-key accounts and migration `0306` for the two provider-default constraints. - Run local Claude and Codex subscription sign-in behind the existing browser sign-in card. Use private attempt homes and owner-bound completion instead of a copied terminal command. - Seed isolated Gemini authentication and preserve OpenCode workspace permissions. Keep the selected connection authoritative. The independent Gemini and Grok workflow fixes are in #15341. - Keep native OpenCode custom gateway keys in a runner-owned selected-model proxy; the harness config contains only a session-scoped capability. Honor runtime outgoing proxy and certificate settings. Preserve streamed responses and revoke the proxy on close or startup failure. - Allow ordinary members to connect native personal accounts before an agent exists. - Repair routed accounts from task cards using the saved provider destination, protocol, model aliases, and connection identity. - Add provider catalog definitions, model discovery, pinned logos, and complete native and routed setup forms. Allow a personal routed connection before a new agent exists. Keep endpoint authentication keys out of Hermes terminal children. - Recover cancelled or restarted browser sign-in with a clear restart action. Support no-auth endpoints without a vault credential. Add isolation and recovery regressions and runtime documentation. ## Verification - Updated with `origin/master` at `22a3ea341`. Migration `0306` follows the new master migration and passes migration and snapshot checks. - The integrated connection regressions passed 152 tests and 50 native OpenCode driver tests, including key-free child-shell configuration reads, authenticated/no-auth forwarding, streaming, model/path restrictions, cancellation, outgoing proxy routing, and NO_PROXY bypass. Provider setup has 14 passing tests. The pinned real OpenCode 1.18.34 executable also completed a turn through the proxy against a local synthetic provider; the reusable key was absent from its config. A second real-executable smoke passed with an HTTPS CONNECT proxy and runtime-specific synthetic certificate trust. Certificate-file and certificate-directory regressions pass. - Task-card repair passed 48 tests, including OpenRouter, Bedrock, and custom gateway reconnect cases. UI typecheck and token gates passed. - The prior core regression set passed 133 tests across new-agent setup, provider forms, browser sign-in, routing projection, and connection authorization. Token gates and UI typecheck passed. - Full workspace typecheck and production build passed again after the latest integration and credential-proxy fix. The merged deterministic runner E2E suite passed 1,400 Vitest tests and 128 Node tests. - Full workspace typecheck passed on the prior linked combined implementation. Production build, Storybook build, 1,316 browser-harness Vitest tests, and 128 Node tests passed. Head `9d964c8d9` includes the latest master integration and regenerated migration. This exact head passed 54 remote checks with four expected skips and Greptile 5/5; no review threads remain open. An unchanged server fixture had a random six-character issue-prefix collision on its first attempt. All 245 tests passed locally and the single CI retry passed. - A provider-free terminal check used the cited supported Hermes source and dummy keys. Gateway and OpenRouter terminal children could not read the selected key. - The broad local Vitest attempt passed 15,442 tests but was not green. It had an embedded-Postgres startup failure, an HTTP logger timeout, an origin socket error, and a browser cancellation wait timeout. The cancellation wait was corrected. The relevant connection tests and the full origin test file passed separately. Latest-head CI must pass before merge. - Prior credential-backed acceptance exercised task creation, tool use, artifact delivery, completion, and context-dependent follow-up. Claude legacy and native runners passed Bedrock with `us-east-1` and `us.anthropic.claude-sonnet-4-6`. - Historical local qualification retained 43 passing API/gateway cells out of 46. Those attempts span earlier builds. They do not qualify this exact commit or staging. All subscription combinations and staging remain unqualified. - Verify native subscription and API-key setup. Connect a regular provider catalog row. Verify an incompatible harness and a changed reconnect URL are rejected. Use #15341 for the complete browser campaign. ## Risks - Migration `0306` changes two check constraints. It preserves rows and is safe to reapply. It takes normal constraint-change locks. - Credential projection touches several harnesses. CLI upgrades can change provider configuration and session behavior. - The native OpenCode proxy adds a loopback hop, pins requests to the selected model, limits request bodies to 16 MiB, rejects redirects, and expires at session close. It prevents reusable keys in the child configuration; it is not an OS isolation boundary against a process debugger running as the same user. - Custom endpoints must be reachable from the agent environment. Saving a connection does not prove connectivity. Bedrock keys require rotation before expiry. - Gemini CLI 0.58.0 has an upstream ACP new-file error conversion defect. Provider overloads and an unresolved follow-up timeout also affect live Gemini qualification. We have not patched the installed CLI or marked those cases as passing. - OpenClaw Gateway, Hermes Gateway, Claude Managed, AWS AgentCore, Process, HTTP, and legacy ACPX local are excluded. Vertex, ambient AWS identity, arbitrary auth headers, and custom routing for other harnesses are excluded. - These PRs do not establish production or staging qualification for every provider and login method. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, code execution, and browser testing. The exact deployment model ID and context window size were not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
22a3ea3414 |
Invite assistants from Connections with scoped browser and device consent (#14933)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Public MCP lets people use their organization from an external assistant. > - Operators need a visible control for this experimental access. > - Hosted users should select an organization once and then approve its permissions. > - This pull request adds the setting, invitation-first setup, and browser or device consent. > - Connections provides a copyable invitation with public instructions that grant no access. > - Users reach browser consent from their assistant and return to inspect or revoke access. ## Linked Issues or Issue Description Builds on merged foundation #14846. This PR now targets master. Related settings convention: #13905. **Current behavior** The preview uses an environment variable to enable MCP. Hosted consent repeats organization selection. Assistant access has no entry in Connections, so users must already know the endpoint and how to reach consent. **Proposed behavior** An administrator enables Settings → Experimental → Assistant connections (MCP). A hosted connection shows the selected organization and its icon, then asks for permissions. Requested write access starts checked when the user’s role permits it; the user can opt out before connecting. Direct instance connections show an organization picker with the first available organization selected. The selection stays fixed across refetches and still requires an explicit Connect action. Connections includes Assistant Connection (MCP). Its setup page explains the canonical endpoint, client configuration, browser authentication, and connected access. It connects as the current person and does not select or impersonate an agent. **Reason and benefit** Operators manage access with the other experiments. Users select one organization, and both the UI and server enforce that choice. **Breaking changes** The old enable variable has no effect. Preview operators must enable the setting once. Apply the additive consent-request migration before deploying the tenant, then deploy the compatible Cloud broker. Existing direct requests and grants keep their behavior. ## What Changed - Simplify OAuth and device consent: show the Paperclip logo beside “Connect {client} to Paperclip”, fall back to “your assistant”, and show the identifying origin plus its favicon below, with the callback URL also visible when different. Remove the hosted-organization creation action. Default to the first available organization without silently changing it on refetch; preserve company restrictions and write opt-outs. Keep the button row contained on narrow screens. - Make Copy invitation the primary action, using the shared animated AgentSetupPrompt and a collapsed manual setup section with icon-labeled line tabs. Remove redundant link actions, copy-status text, the extra first-prompt well and revocation explanation from the setup page. Serve shared version-aware HTML and Markdown instructions without private organization data. - Support guarded Client ID Metadata Documents alongside dynamic registration, and include authorization response issuer identification. - Add RFC 8628 device authorization with separately hashed codes, expiry, shared request quotas, persistent polling backoff and atomic redemption. Reuse human consent, role checks, scoped grants, audit and revocation. - Add CLI device login and a local stdio bridge. Store credentials separately with private permissions and serialize rotating refreshes. - Add device consent stories and five cold-start paid Product E2E cases with independent grant, configuration and durable-work assertions. - Add `enablePublicMcp` to the settings validator, normalizer, feature catalog, and toggle UI. Check it live for OAuth, tools, subscriptions, and event delivery. Keep connection management and revocation available while disabled. - Default the MCP origin to the existing auth public URL, with strict validation and an explicit override. - Persist the optional OAuth `company_id` restriction. Describe only that company and reject approval for any other company, even if the person belongs to both. Keep active-membership and role checks. - Show the Paperclip icon and a large organization icon during consent. Return the saved company logo through the company-scoped request response and reuse the standard fallback icon. Use the requested concise permission labels: “Read all of your Paperclip data” and “Allow write access and creating tasks as me”. Use concise permission copy, retain a compact client and callback-origin disclosure, and remove the footer link. - Default requested write access on for eligible roles. Preserve opt-out across organization changes and refetch, reset defaults for a new request, and submit read-only access when the request or role does not allow writes. Align the shared checkbox with its label. - Use organization wording in consent, management, settings, and walkthroughs. Keep the organization fixed for hosted requests and retain direct-instance choice. - Add an Assistant Connection (MCP) card to the Connectors catalog, a setup page in the app shell, and a return link from Experimental settings. Include Codex, Claude Code, OpenCode, and generic remote MCP instructions. - Read the live gate and canonical server URL through authenticated setup metadata. Show only the current person’s grants for the selected organization, refresh after consent, and support revocation. Surface catalog status failures with an explicit retry action; do not present them as an empty connection list. Opening setup grants no authority. - Start the eight guided chapters in Connections. Keep presenter notes and chapter controls around real product pages in the app shell. Explain the terminal, consent, delegation, retrieval, and revocation handoffs. Mark conversation examples as illustrative. Cover first use, client setup, connected, loading, and error states. Keep the existing consent and management stories. - Keep the paid-eval setup and browser helper aligned with the setting and consent button. ## Verification - Warm-standby integration fix `ec64ea05e`: public MCP ingress now follows the Cloud claim guard; MCP and discovery paths return 503 instead of SPA HTML while unclaimed. Event polling checks the in-memory claim before reading the persisted experimental setting. All 97 focused OAuth/Cloud tests and server typecheck pass, including new request and timer regressions for idle-before-claim and resume-after-claim behavior. Fresh review is 5/5 with no unresolved threads, and all security scans pass on this final head. All browser shards, typecheck, build, canary installation and other test groups passed on the first attempt. The unchanged Cursor sandbox default-command test timed out at 10 seconds; the exact test passed locally without edits in 587 ms. The single failed-job retry passed, with the original failure retained in workflow 37500711895. All 54 final-head checks pass on `ec64ea05e9a03e2179d4e2f84c2de03761f7ce26` (two optional Storybook jobs are intentionally skipped). - Final master integration `8457828fc`: merged foundation #14846 and current master, preserving the invitation changes and all 33 files from the two newer upstream changes. No migration renumbering was required. All 95 focused OAuth/Cloud integration tests, full recursive typecheck and token gates pass. All CI gates passed on that integration head; review identified the warm-standby issue fixed above. - Security-review fix `8c1d0b696`: commit shared global/per-source admission before outbound CIMD work, preserve failed-attempt receipts, and validate resource/scope before fetching. Added migration `0305_chubby_vin_gonzales.sql` and six concurrent/adversarial regression cases. All 69 OAuth/metadata tests, 26 migration checks, full recursive typecheck and production build pass. The security scanner passed that commit. Follow-up `87f9658e7` limits only actual cache-miss fetches; 18 authorization requests sharing one proxy across two service instances use just two fetches. All 70 OAuth/metadata tests and server typecheck pass after that refinement. Final follow-up `8ebeae84c` reports admission-storage failures as retryable HTTP 503 instead of invalid client metadata. Its regression proves no outbound request before admission and successful retry after storage recovers. All 71 OAuth/metadata tests and server typecheck pass. Final-head security scanning passes; Greptile is 5/5 with no unresolved findings. CI passed all browser shards, typecheck, build, token gates and canary installation. One unchanged adapter-utils bridge test raced a response-file write (expected a JSON error, received the safe file-changed error). The exact test passed locally without edits. The single failed-job retry passed; the original failure is retained in workflow 37490609192. All 54 checks now pass on final head `8ebeae84ca77c0cf7ac12c2006f0f8743fe50e0b`, with security scan and fresh Greptile 5/5 and no unresolved threads. Foundation #14846 subsequently merged as `e34abee670069cca84afb2efb86041bce7dccbec`; the final integration above now targets master. - Integration with current master: preserved the new Connections source filters and pagination, kept all eval suites, and regenerated the consent/device snapshots as migrations 0303/0304. All four MCP migration SQL hashes are unchanged from the staging versions. Full recursive typecheck and production build, 132 focused UI tests (including catalog filtering), 89 server authorization/settings tests, 26 migration tests, 120 eval calibration tests and token gates pass. Review follow-up `4820ce74c` also keeps active assistant grants in Installed, with pending/error recovery and revocation/company-isolation coverage. All 76 setup/catalog tests, UI typecheck and token gates pass after that fix. The unchanged signoff browser test timed out waiting for a heartbeat in CI at `4820ce74c`; the exact test passed locally without code changes, and the preceding CI head passed that shard. That same unchanged test failed at the reviewer stage in the next CI run. All five signoff tests passed three times locally (15/15), without test changes. All eight browser shards pass at final head `8ebeae84c`; no browser-test edits or failed-browser-job retries were needed. - Setup-page refinement at `9ab009178`: all 17 focused setup/consent tests pass, along with UI typecheck, production build, Storybook build and token gates. Browser exercised the shared prompt preview and client tab switching, and the updated InvitationCopied Storybook interaction checks its clipboard fixture. All final-head CI checks pass at `9ab009178`, with no unresolved review findings. Deployed successfully to Butter in https://github.com/paperclipai/paperclip-cloud/actions/runs/37475189524. Verified the actual page, tab switching and line styling, removed actions/copy, and successful native copy/paste of the complete Butter invitation into a local-only test field. The existing Claude grant was left intact. - Consent follow-up at `dc8e9fd11`: all 10 consent tests and token gates pass. UI typecheck and production build passed again at `4e4d5e4d9`; Storybook build and eval-helper typecheck passed for `28101cf91`. Follow-ups let the primary button wrap on narrow screens, preserve a distinct callback URL, and use only bundled icons to avoid pre-consent requests to client-selected sites. Browser-verified the real consent component in desktop and 320px mobile stories, including default selection, write access and preserved opt-out. Updated E2E heading/default-selection helpers. All CI checks passed at `dc8e9fd11`, with review 5/5 and no unresolved threads. The Butter preview publication needed a retry because npm initially accepted the DB package before making it visible; the retry succeeded and `dc8e9fd11` deployed. Verified a fresh, unapproved native Codex CIMD request on Butter: default organization/write selection, known-client heading and icon, distinct callback origin, and removed creation action. No grant was approved for this UI check. Prior paid runs below retain their exact source provenance; this UI-only follow-up did not rerun paid qualification. - Source-pinned paid matrix at `2992ef2710f47230e7f484c709c6ba02524f884c`: **15/15 passed**, five cases each on GPT-5.4 Mini, Claude Haiku and Sonnet. Campaign `local-2026-10-06T02-41-14-462Z`. Covers cold start, existing config, unavailable host, denied consent and reconnect/later retrieval, with independent configuration/grant/task/run/document assertions. Original failures, transcripts, source fingerprints and billing remain retained. - Final instruction follow-up `cda8178af`: **3/3 cold starts passed** on Mini, Haiku and Sonnet. Campaign `local-2026-10-06T02-58-30-041Z`. Latest `0637b9f1c` shares that same guidance across HTML, Markdown and manual UI after review; generated Markdown is verified byte-identical to the paid-evaluated version. Shared build, server/UI typechecks, token gates and 63 auth/metadata tests passed again. Every CI gate passed at prior HEAD `0637b9f1c`, with review 5/5 and no unresolved threads. - Other focused checks: 11 CLI credential/refresh-lock tests, 120 eval calibration tests, server/UI/eval typechecks, token gates and Storybook build passed. Full recursive typecheck and production build passed during implementation; CI also passed them at `2992ef271`. - Local full-suite limitations: a large-file Git streaming test times out on this Mac, and broader CLI/route runs hit DB hook timeouts. Fresh MCP reruns passed, and the corresponding CI groups passed. No claim that the local full suite is green. - Actual clients: Codex 0.153.4 and Claude Code 2.1.245 reach CIMD consent; device CLI reaches verification/consent. New grants await human approval. Existing local OpenCode retrieved a saved result in a fresh conversation through its previously approved grant. - Fresh OpenCode 1.18.17 on Butter: started with no MCP config, received the exact copied invitation, read public setup, configured its server and started PKCE consent. Its shell command timed out; background retry reached the client's own callback deadline while approval remained pending. Latest instructions cover that handoff. **No completed Butter read/delegation/result retrieval is claimed.** - Cloud companion https://github.com/paperclipai/paperclip-cloud/pull/672 passes checks/review and deployed. Anonymous setup and device-protocol routing verified. Core `2992ef271` deployed successfully and the actual Claude web flow now reaches consent. Its extra JWT-bearer metadata is filtered to implemented grants; unsupported token grants remain rejected. Final `0637b9f1c` deployed successfully to Butter in https://github.com/paperclipai/paperclip-cloud/actions/runs/37409195300; live HTML and Markdown both contain the final guidance. The superseded instruction-only build was canceled before deployment. This is a core-only staging preview; private Cloud plugins are omitted. ChatGPT web is signed out, so browser connector use is unverified. - Screenshot gallery begins at Butter's dashboard and distinguishes real setup/pending consent from local reuse and fixtures. It records the timeout finding. New persistent access needs human confirmation before the remaining actual-client acceptance work. - Manual path: Connectors → Assistant Connection (MCP) → Copy invitation → paste into assistant → configure and start authorization → sign in and approve → verify `paperclip_connection` → delegate → retrieve the saved report later. - Plan and instructions: `doc/plans/2026-10-05-assistant-invitations.md` and `doc/public-mcp.md`. ## Risks - Apply additive, replay-safe migration `0304_curvy_shadow_king.sql` before using device authorization. The public setup link carries no credential. Device codes and tokens stay private; neither sharing instructions nor installing a plugin authorizes access. - Apply additive migration `0305_chubby_vin_gonzales.sql` before deploying the shared metadata admission gate. It retains at most 60 short-lived, hashed-source receipts per instance and rejects excess attempts with 429. - CIMD metadata fetching is a new external-input boundary. It requires HTTPS, exact client ID and redirect validation, bounded responses and guarded DNS/network access. Client names remain self-reported. - Device support is per-instance. The central Cloud broker retains its existing grant support. Host installation and tool reload capabilities vary by client; instructions describe manual settings and restart requirements. - Consent names the registered client in its heading and displays its identifying origin below. Known-origin icons are bundled; all other origins show a neutral site icon without contacting client-selected sites. Client names are self-reported; the callback origin is the recipient check. The Cloud chooser also displays the original client and receiving origin before tenant handoff. - A user who accepts the preselected write permission can create tasks and comments. Task creation and comments can start or wake agents and use execution budget; the consent label uses the concise wording explicitly requested by the maintainer. Scope requests, role checks, and the final Connect action still apply. - Migration `0303_supreme_garia.sql` adds one nullable UUID column with `IF NOT EXISTS`. Requests without a company restriction keep the direct-instance picker. The binding stays recorded if its company is deleted; consent then fails closed. - Deploy tenant support before the Cloud broker sends `company_id`. Unknown or inaccessible organizations must never fall back to a different company. - The setting defaults off. Disabling access does not cancel work already delegated. Existing tokens and unexpired subscriptions can resume when enabled again; revocation remains separate. - The catalog entry is visible for discovery while the feature is off. Setup instructions, OAuth, and tool execution remain gated. No access is granted by viewing the entry. - Assistant sign-in starts in the external client so it owns PKCE and callback state. Client command syntax can change and links to official setup documentation are included. - An authenticated instance and valid public URL are required. Hosting, paid execution, and store publication remain separate rollout steps. ## Model Used OpenAI GPT-6 in Codex, with tool use and code execution. The exact serving model version and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused checks pass; unrelated local full-suite timeouts are explicitly recorded above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e34abee670 |
feat(mcp): connect assistants to a team with user OAuth (#14846)
## Thinking Path > - Paperclip gives teams durable tasks, agent execution, budgets, and approvals. > - People also use assistants in Codex, Claude, and other MCP clients. > - Those assistants need a scoped connection that preserves the person’s permissions and attribution. > - Delegating a task must not turn the assistant into the assigned agent. > - This PR adds opt-in user OAuth, ten first-party tools, browser consent, and workflow packages. > - Paid product evals verify the resulting tasks, documents, attribution, retries, and access boundaries. > - The team keeps working after the assistant conversation ends. ## Linked Issues or Issue Description **Problem or motivation** A person cannot connect an external assistant to an existing team through browser consent and safely delegate durable work as themselves. **Proposed solution** Expose an opt-in `/mcp/paperclip` endpoint with individually described first-party operations. Bind every connection to a person, client, company, resource, and scopes. Reuse domain authorization and scheduling. Package shared team-review, delegation, and follow-up workflows for OpenAI/Codex and Claude. **Alternatives considered** Related PRs #9393 and #12549 cover earlier remote MCP and board-operator approaches. This change uses user OAuth and a bounded public catalog. It does not expose a generic executor, operator administration, static shared board credentials, or external agent execution. Registry listing work in #9851 is a separate distribution step. **Roadmap alignment** This maintainer-requested implementation extends the governed MCP gateway, activity attribution, durable work products, and hosted deployment direction in `ROADMAP.md`. It implements the first release of the saved design plan; external agent participation and granted third-party tools remain later releases. ## What Changed - Add MCP 2.0 discovery and task status/comment/document Events on the same authenticated endpoint. Persist subscriptions and delivery receipts, verify HTTPS callbacks, sign Standard Webhooks, encrypt callback material, recheck permissions/Cloud membership, and bound retries/expiry. Older MCP clients keep their existing tools. - Add discovery, dynamic client registration, S256 PKCE, resource validation, rotating refresh tokens, revocation, and company consent. Store credentials as hashes and recheck membership at execution. - Add tools for connection identity, agents/projects, task search/read/create, human comments, documents/deliverables, and pending-approval links. Preserve current domain permissions and scheduling. - Add durable mutation receipts across reconnects. Matching retries replay results; uncertain outcomes keep the same request ID and require inspection. - Add consent and connection-management pages, OAuth log redaction, shared plugin workflows, and separate OpenAI/Codex and Claude package outputs. - Add eight paid Product E2E cases across three models, independent durable-state grading, usage evidence, cleanup, and report integration. Add task-document guidance and regenerate the runner capability inventories. - Add migrations 0301 and 0302, the dated implementation plan, result notes, and direct-client setup instructions in `doc/public-mcp.md`. ## Verification - Merge integration `e180b1948`: resolved conflicts with current master, preserved both eval registries, regenerated capability catalogs, and regenerated migrations as 0301/0302 while keeping the original replay-safe SQL byte-identical. Local migration safety/snapshot tests (26), MCP/OAuth tests (38), redaction/OpenAPI tests (71), and eval catalog/grading tests (198) pass. Token and capability gates pass. Full recursive typecheck passed. Fresh Greptile review is 5/5 with no unresolved findings. CI is green on this exact head (55 successes, two intentional skips, one neutral result): one unchanged Cursor sandbox test timed out at 10 seconds, then passed locally in 856 ms. A single retry of that failed shard and the aggregate workflow passed. Merge remains blocked on the repository code-owner approval rule. Earlier checks passed at `6aa0962d4fb715f2190bb7bb22efacab2e58495d`: 55 successes, two intentional skips and one neutral result. [The earlier CI run](https://github.com/paperclipai/paperclip/actions/runs/36901592350) includes all test shards, browser tests, typecheck, build and canary dry run. Greptile was 5/5 on that commit with no unresolved review threads. GitHub still requires code-owner review under the repository merge rules; passing checks do not bypass that approval. Paid source fingerprints remain separate below and in the dated result note. - Paid Events qualification passes **3/3**: GPT-5.4 Mini, Claude Haiku 4.5 and Claude Sonnet 4.6. Each uses a real public HTTPS callback, signature verification and report retrieval in a fresh conversation. A final Mini regression passes after the quota/status fixes. All evidence validates. Bounded tunnel startup retries occur before provider calls and remain visible; failed earlier attempts retain their original grades. - The earlier complete seven-case matrix passes **21/21**, with a separate **3/3** delegation regression. Two preceding matrices also passed 21/21 each. A complete 24-cell matrix including Events has not been run. [The dated results](doc/plans/2026-10-01-public-mcp-paid-eval-results.md) retain exact source fingerprints, failures, model IDs and partial costs. - Node 24: repository-wide `pnpm -r typecheck` and `pnpm build` pass after merging master. Server typecheck passes after the final quota/status changes. Eval typecheck and all 892 eval-support tests pass. - All 33 real MCP/OAuth tests pass. The preceding combined MCP, redaction, private-address and DNS-rebinding run passed 129 tests; two later MCP regressions cover quota reuse and unchanged-status suppression. All 28 adjacent issue-tree/stale-lock route tests pass. CI then found a null checkout result in the existing concurrent-workspace path; logging now uses optional status access. All 12 closed-workspace tests and all 33 MCP tests pass after that correction. The exact-start event calibration exposed a timestamp gap; scanning now includes the subscription start, with all 33 MCP tests and server typecheck passing. These two narrow corrections follow the paid regression. - A real Core → Cloud → Core authority round trip passes OAuth, MCP 2.0 subscription/delivery, current membership loss, unsubscribe, legacy SDK tools, refresh and revocation. Its callback transport is a fixture with independent HMAC verification. The paid Events campaigns separately prove public HTTPS delivery. - Earlier component qualification passed UI 7,117 tests, CLI 502, shared 832, skills catalog 20, database 160 and OpenAPI 10. Token gates, module boundaries, migration order and plugin regeneration passed. CI covers general/serialized suites, eight browser shards, runner checks, typecheck, build and canary dry run. - **Local full-suite limitation:** the earlier monolithic run was not clean. It encountered overlapping schema rebuilding, Mac database shared-memory limits and isolated CLI/fixture failures. Targeted reruns passed. The existing >32 MiB Git filename stress test still hit its 300-second Mac timeout. The additional serialized sweep stopped after 62 passing suites once CI passed. Original failures and partial logs remain; this PR does not claim a wholly green local monolithic run. - Local Codex CLI and Claude Code OAuth login and MCP SDK interoperability were verified. Public-store installation, actual ChatGPT Work Cloud Events UI, staging HTTPS client behavior and hosted newcomer provisioning remain release gates. Enablement is moving to **Settings → Experimental → Assistant connections (MCP)** in the stacked follow-up [#14933](https://github.com/paperclipai/paperclip/pull/14933). Merge both for the intended setup experience. This foundation branch alone still uses `PAPERCLIP_PUBLIC_MCP_ENABLED=true`. After deployment, set `PAPERCLIP_PUBLIC_URL` to the authenticated instance's HTTPS origin, and connect to `/mcp/paperclip`. Select a team and allow writes in browser consent. Configure an available agent and budget, then delegate and retrieve results later. For Events, rescan the deployed plugin catalog in ChatGPT Work Cloud; the host supplies its webhook credentials when the user asks to watch a task. See [the setup runbook](doc/public-mcp.md). ## Risks - Events are at-least-once and may arrive out of order. No replay cursor is advertised. Clients must refresh finite subscriptions, read current state and avoid comment feedback loops. Callback material uses the instance secrets master key; hosted subscriptions require the updated Cloud broker and are bounded to five minutes/the access proof expiry. - ChatGPT Work Cloud/dot event UI, plugin rescan and a hosted staging subscription remain deployment gates. Local signed-webhook and paid model evidence does not claim those surfaces have been exercised. - Disabled by default. Merging adds schema and opt-in code; it does not deploy a public endpoint, publish a store listing, create a team, or start paid agents. - Migrations 0301 and 0302 are additive and idempotent. Their SQL is unchanged from the earlier preview numbers, so hash-aware upgrade reconciliation preserves prior staging applications. Normal instance upgrades must apply it before enabling MCP. - Task creation and comments can schedule paid agent work. Consent and tool descriptions disclose that effect. Revocation blocks future calls but does not undo delegated work. - Public deployments need edge rate limits and credential-safe logging. Internal dispatch is restricted to the closed catalog and carries a request-local verified actor. - Hosted onboarding requires the companion Cloud broker, encryption-key configuration, and tenant rollout. Self-hosted direct connections can use this PR alone. - Store acceptance and agent-mode participation are not claimed. Checked-in plugin endpoints are development defaults; rebuild packages for a real deployment before installation. ## Model Used OpenAI GPT-6 in Codex, with reasoning, tool use, and code execution. A more specific serving version and context-window size were not exposed by the session. Paid eval models: `gpt-5.4-mini-2026-03-17`, `claude-haiku-4-5-20251001`, and `claude-sonnet-4-6`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (targeted/component checks; full local-run limitations are recorded above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
202c2d307e |
fix(server): leave unclaimed warm Cloud databases idle (#15314)
## Thinking Path > - Paperclip manages work performed by AI agents. > - Managed deployments prepare empty applications before an owner claims them. > - Those applications start database pollers even though no company can have work. > - Health probes also query SQL, so idle databases cannot remain suspended. > - This pull request adds an explicit standby marker for empty, unclaimed Cloud apps. > - The existing signed, durable claim resumes normal processing without restarting the app. ## Linked Issues or Issue Description **What happened?** An empty, unclaimed warm application runs recurring chat, email, plugin, heartbeat, cleanup, and reconciliation queries. Its health route also opens the database. This prevents idle database compute from suspending. **Expected behavior** An explicitly marked unclaimed application should keep its HTTP process and sandbox provider plugins ready while leaving the database idle. A successful signed claim should resume normal API behavior and background processing. Claimed and self-hosted instances should keep their current behavior. **Steps to reproduce** Start an empty Cloud-managed application and leave it unclaimed. Observe database activity while repeatedly requesting `/api/health`. Before this change, periodic queries continue without company data. Related: #15153 reduces allocation during chat polling. This change suppresses polling only for explicitly marked, empty, unclaimed Cloud apps. ## What Changed - Add `PAPERCLIP_CLOUD_WARM_STANDBY=1`. Check company emptiness once after restoring the persisted Cloud runtime identity. Missing Cloud configuration, existing data, or a persisted claim leaves normal processing active. - Gate recurring database pollers with an in-memory predicate. Keep startup preparation and sandbox provider plugin loading intact. - Serve unclaimed health probes without session or database reads and report `warmStandby: true`. Serve standby pages/assets directly from the UI router, bypassing session, bearer, tenant, and dynamic handlers. Refuse API requests and all WebSocket upgrades before authentication can query SQL or seed company data. - Exit standby after the existing signed identity assertion commits. Normal timers resume at their next tick; a restart restores the claim even with stale provider variables. - Document the marker, readiness semantics, rollout checks, and rollback. ## Verification - `pnpm -r typecheck` passed. A final server typecheck also passed after adding tests. - `pnpm build` passed. - Focused standby, signed claim, restart, health, static/Vite routing, hostname, HMR, and live-events suites: 72 passed after the review fixes. Includes real HTTP upgrade admission before/after claim. - `pnpm test:run` was attempted locally; both superseded runs were stopped after encountering checkout/platform failures. A clean-checkout rerun eliminated ancestor skill-directory lookup failures. The company-skills/runtime-cache families encounter macOS read-only-directory rename failures (`EACCES`); all three company-skills failures reproduce on unmodified base `bf14f803d5`. The initial full run also reported one native runner API test failure; an isolated comparison on both revisions was blocked by local embedded PostgreSQL startup failures. The full [Linux CI run](https://github.com/paperclipai/paperclip/actions/runs/37423263993) passed on final commit `5016c415ea`, including all server and workspace test shards, browser suites, typecheck, build, and release canary. This is not a claim that the full local suite passed. - Isolated full server with local PostgreSQL: after startup and connection expiry, 70 health probes, 70 page requests carrying valid synthetic tenant credentials, and 70 rejected WebSocket upgrades over 70 seconds observed zero app database connections. The signed claim completed in 62 ms and normal polling resumed (475 database transactions over 12 seconds). Restart with stale provider variables restored the durable claim. The latency is local-only, not a provider wake measurement. - Apex review: **5/5** on `5016c415ea`, both earlier threads resolved, no open recommendations. - No live-provider test or production deployment was performed. An actual database suspension/resume canary remains required before enabling the control-plane switch. ## Risks - Standby health reports HTTP readiness rather than current database connectivity. The signed claim still requires a durable database write; claimed health checks retain the SQL probe and 503 failure behavior. - Pollers resume at their usual intervals. A suspended database may add claim latency. Validate the real provider before enabling the marker. - Startup preparation and sandbox plugins remain loaded. New plugins or background loops must respect the same standby contract. - The marker is off by default. Remove it or set it to `0` and restart to roll back. No schema migration or claimed-workspace inactivity policy changes. ## Model Used OpenAI Codex, based on GPT-6. The exact serving snapshot and configured context-window size are not exposed in this session. Assistance included source review, TypeScript changes, command execution, and PostgreSQL tests. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass — targeted tests pass; full local suite limitations are documented above, and full Linux CI is green - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f2715e02bb |
fix(native): surface model capacity errors and retry automatically (#15347)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Native runs preserve provider events and results, then decide task state. > - Codex can fail a turn because the selected model is temporarily at capacity. > - Paperclip displayed a generic native-session failure and left the task in recovery. > - A known capacity failure needs a clear message and a durable delayed retry. > - This pull request uses committed terminal evidence, the status-effect ledger, and existing dispatch gates. > - The task can continue automatically without an unbounded retry loop or a model switch. ## Linked Issues or Issue Description Related: #13993 adds launcher capacity deferrals before provider startup. This change handles a committed native Codex `serverOverloaded` terminal after provider work has started. **What happened?** A native Codex turn ended with `codexErrorInfo: serverOverloaded` and “Selected model is at capacity. Please try a different model.” Paperclip saved the result, showed a generic failure, and required recovery instead of waiting and retrying. **Expected behavior** Show the capacity error directly. Schedule bounded retries after a short delay. Preserve saved work and honor current execution gates. **Steps to reproduce** 1. Run a native Codex task. 2. Commit a `turn.failed` event with `serverOverloaded`, then accept a failed result for the same turn. 3. Inspect the task error and its recovery state. The regression test reproduces this without a paid provider call. **Paperclip version or commit** Reproduced against master `16b7db35ffa0f9a95913c8cbdeea3d595435691f`. ## What Changed - Classify capacity failures from committed runner events and the pinned execution identity. Preserve accepted results and display the specific capacity error. - Atomically persist one scheduled successor with the status decision. Retry after one minute, then two minutes. Share the existing failure budget and stop after two automatic retries. - Reuse dispatch gates for ownership, task holds, dependencies, budget, and locks. Wait for predecessor execution, finalization, and cleanup before claiming a retry. - Preserve pending reviewer authority and suppress retries after reassignment or a successor claim. - Consume the failed run's resume receipt and delivered wake input. Rebuild ordinary continuation from the failed run so explicitly resumed tasks can retry without borrowing one-run authorization. - Label scheduled retries “Model at capacity.” Add a Storybook example and avoid duplicate punctuation in the existing retry card. - Add regression coverage and document the runtime contract. ## Verification - Targeted recovery and UI suites: 125 tests passed. Additional final cleanup and lock regressions passed. - Latest-head continuation and authorization regressions: 57 tests passed, including all four capacity integration tests against an isolated PostgreSQL database. - Rechecked all four capacity integration tests with the full runner's isolated `PAPERCLIP_HOME`, config, temporary directory, and serial fork settings: passed. - `pnpm -r typecheck`: passed. Final server typecheck passed. - `pnpm build`: passed. - `pnpm check:token-gates`: passed. - Browser: checked the real Storybook card. It shows the capacity message, automatic retry time, and existing Retry now action. - `pnpm test:run`: attempted and restarted after an interruption. The resumed run started before the final review correction and was stopped after recorded workspace/native test failures and the five-minute Git streaming timeout already documented on master. Exit 130; no complete local full-suite pass is claimed. Current-head isolated capacity tests and the complete CI suite pass. - A combined local recovery-suite attempt also hit PostgreSQL initialization failures at this macOS host's global shared-memory limit (32 slots). The focused recovery suites passed separately; final-head capacity tests also pass with the full runner's isolated environment settings. - Latest-head GitHub checks: all 56 checks green, including general and serialized test shards, browser E2E, typecheck, build, Runner verification, and canary dry run. No merge conflicts. - Greptile: 5/5 on `813a470b8c18c05aeb7e31e63e573c3a3a2f5cac`, with zero unresolved findings after fixing the resumed-task continuation issue. ## Risks - Capacity retries can repeat a task turn after partial work. They start a fresh provider session with task history and wait for predecessor cleanup. They do not resume the failed turn. - Retries retain the configured model unless an operator changes configuration. Persistent overload consumes the existing failure budget and then needs an explicit retry or model change. - Only run-bound native Codex `serverOverloaded` failures qualify. Usage-limit exhaustion, model/account incompatibility, unbound text, and unknown provider failures retain their current recovery behavior. - No schema migration is required. ## Model Used - OpenAI GPT-6 through Codex. The exact model ID and context-window size are not exposed to this session. Used reasoning, repository tools, code execution, and browser verification. No subagents. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9b3fe260ba |
fix(tasks): surface Codex ChatGPT model rejection (#15299)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Native Runner runs report provider errors and can also save a structured failure result. > - Codex rejects a model that the user's ChatGPT account cannot use. > - Master now diagnoses this rejection, but the task's compact run data omits its message. The thread can still call it a generic run failure. > - Users need the account restriction and a clear step to repair the model selection. > - This pull request shows the existing diagnosis as “Model unavailable” on the task. ## Linked Issues or Issue Description **What happened?** Codex returns HTTP 400 with `invalid_request_error` and the message `The 'gpt-6.1-sol' model is not supported when using Codex with a ChatGPT account.` Master now stores an actionable diagnosis for this rejection. The task thread still labels it “Run failed” and does not receive its error text in compact run data. **Expected behavior** Show the account restriction on the task and in the run error. Tell the user to choose a supported model or clear the task's model override before retrying. **Steps to reproduce** 1. Use a Native Runner Codex agent signed in with a ChatGPT account. 2. Select a model that produces the rejection above and start a task. 3. Let the runner save its generic failed result. Inspect the task's failure marker and recovery notice. **Additional context** Refs: #15304. That merged PR diagnoses the provider failure and preserves worker and review recovery rules. This PR adds its task-facing message and guidance without changing that diagnosis or those rules. Refs: #13134. That PR improves model discovery for ChatGPT accounts. This PR exposes the rejection when a configured model still fails at execution time. ## What Changed - Return a bounded model rejection message in compact issue-run data. - Show “Model unavailable” and model-change guidance in the task thread and recovery notice. - Add database and UI regression tests, including both the original rejection text and master's fixed diagnosis. Document the new failure message. ## Verification - Focused merged-branch validation: 279 tests passed across activity service, native provider failure observation and PostgreSQL integration, TaskChatThread, and ExecutionBlockerNotice. - `pnpm -r typecheck`, `pnpm build`, and `pnpm check:token-gates` passed on the merged source tree. - Greptile reviewed conflict-resolution commit `1db39c4e636a92e2e76ba224b71c6f1e2f556c81` at 5/5 with no findings or inline comments. - All 55 check runs completed without failure on that commit. The two optional Storybook jobs were skipped. The legacy Snyk status passed. The branch has no merge conflicts. - UI regression assertion: the task's failure marker says “Model unavailable” and retains the account restriction. It no longer says that this failure happened after a final response. ## Risks - Provider recognition and recovery are owned by the existing master implementation. This PR exposes only bounded error text for failed runs with `native_provider_model_rejected`. - Historical runs with the generic `adapter_failed` code are not reclassified. No stored run is rewritten. - Existing retry and reconciliation gates remain in place. No schema migration or model configuration change is required. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, code execution, and browser inspection. The exact deployment model ID and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
858094ba81 |
Fix browser polling for unsaved agent chats (#15298)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent Chat uses the shared task surface and browser panel. > - An unsaved chat has a temporary ID instead of a stored task UUID. > - The browser poll sent that ID to a PostgreSQL UUID query and returned a server error. > - This pull request waits for a saved task and validates task IDs before the query. > - Saved chats keep their browser tools and access checks. ## Linked Issues or Issue Description Refs #14331. That report describes the same draft-ID problem in email requests. This change covers the browser routes; it does not change email behavior. Opening an unsaved Agent Chat sends `chat:<agent-id>` to `/issues/:issueId/browsers` every three seconds. PostgreSQL rejects this as a task UUID. The draft is only a UI view model. The first send or upload creates its stored task. ## What Changed - Disable browser polling for unsaved chat IDs. Resume polling when the chat has a task UUID. - Return the existing task-not-found response for malformed task IDs before database access. Keep board authentication first and company and credential checks unchanged. - Test all six task browser routes, draft-to-saved polling, normal tasks, saved chats, and denied access. Keep canonical UUID v4, v7, and nil values queryable. - Document the draft and saved-chat browser contract. ## Verification - `pnpm install --frozen-lockfile` passed with pnpm 9.15.4 and Node 24.21.0. - `pnpm exec vitest run server/src/__tests__/browser-use-route-scope.test.ts server/src/__tests__/browser-use-connection.test.ts` passed all 35 tests. - `pnpm exec vitest run ui/src/hooks/useTaskBrowsers.test.ts` passed all 5 tests. - Independent review passed 14 hook and route tests plus both real task and saved-chat authorization cases. - `pnpm check:token-gates` and `git diff --check` passed. - Full workspace `pnpm -r typecheck` and `pnpm build` passed. - Full Linux CI passed on `dafff44a32`: 53 successful checks and two intentional Storybook skips. This includes all general and serialized test groups, UI and browser tests, typecheck, build, and the canary dry run. No CI retries were needed. - The full local `TMPDIR=/private/tmp pnpm test:run` started but did not complete. It was stopped after full Linux CI passed. Before the stop, three company-skills cache cases failed on macOS. A focused rerun reproduced all three as `EACCES` while renaming immutable cache staging directories. The test, company-skills service, and runtime cache files are byte-identical to the baseline previously reproduced on `e99854249c` for #15291. Other local groups were not reached; the full Linux CI run provides aggregate coverage. - Greptile scored the exact head `dafff44a32` 5/5 with no findings or unresolved review threads. ## Risks - Malformed task IDs now return 404 instead of reaching PostgreSQL and returning 500. The route does not map draft IDs to agents or tasks. - No schema, browser provider, task creation, or authorization policy changes. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository editing, code execution, and independent agent review. The exact deployment model ID and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub references) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused suites; full local run limitations are reported above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
81cd03b3b5 |
test(ui): keep initial reasoning with the saved reply (#15107)
## Thinking Path > - Paperclip helps people manage AI agents and their work. > - The issue thread shows saved replies and agent run history. > - History that appears after the reply can move the page. > - Master now waits for initial run history before it shows the thread in #15228. > - This pull request checks that reasoning and the saved reply appear together in the first visible frame. > - The test protects that behavior from a later change. ## Linked Issues or Issue Description No public issue tracks this report. This is a bug regression test. **What happened?** The issue thread showed a saved answer before its reasoning and tool history loaded. The later history moved the page. **Expected behavior** The first visible thread contains the initial reasoning and the saved reply in order. **Steps to reproduce** 1. Open an issue with an agent run and a saved reply. 2. Delay initial run-history hydration. 3. Observe the saved reply before its reasoning appears. Related work: #15228 now contains the reveal gate. This PR adds a direct regression assertion for that gate. #14667 takes a different loading approach and remains open. ## What Changed - Add a saved agent reply to the initial-history test. - Assert that the native run's reasoning appears before that reply when the thread becomes visible. - Assert that the thread stays inert until history is ready. ## Verification - `vitest run ui/src/components/TaskChatThread.test.tsx ui/src/pages/IssueDetail.test.tsx` passed: 353 tests. - `git diff --check origin/master...HEAD` passed. - GitHub CI checks are in progress for the rebased head. ## Risks - Low risk. This PR changes only a test. - The production fix is already in #15228. This PR must not restore the older code from the previous branch head. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex CLI assisted with tool use and code execution. The exact model ID and context window were not exposed to this agent session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My remote branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (no user-facing documentation changed) - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
0e0b63e5a5 |
feat(connections): add experimental task-pinned AI routing (#14967)
## Thinking Path > - Paperclip manages AI agents and their work. > - AI Connections separate account access from models and harnesses. > - A pool must act as one connection while retaining each task’s account. > - Core must enforce member access and preserve session and recovery rules. > - A plugin supplies rotation policy without receiving credentials. > - This change adds durable routing and native connector setup and management. ## Linked Issues or Issue Description **Subsystem affected** AI Connections, Connectors, plugins, run dispatch, and session compatibility. **Problem or motivation** Operators need to rotate new tasks across saved accounts while each task keeps its account and session. Pool setup must fit the existing connector catalog and account workflow. **Proposed solution** Add an experimental router binding, a capability-gated plugin hook, and transactional task pins. Plugins declare native pooled connectors through `aiConnectionRouter`. Core hosts the existing-account picker, ordering step, and account settings. Related usage contract: #14936. Companion private plugin: https://github.com/paperclipai/paperclip-cloud/pull/643. **Roadmap alignment** This extends Apps and AI Connections. Core supplies generic enforcement and native connector UI; the private plugin owns rotation and quota policy. The prior duplicate search found no matching router implementation. ## What Changed - Add a router binding without changing existing concrete bindings. Keep the instance flag and new pools disabled by default. Require manual operator configuration. Show no routing toggle in Experimental settings on either open-source or Cloud installs, even after routing is enabled. - Persist company-scoped pools, one shared cursor per pool, and pins keyed by company, pool, agent, and task. Commit pins and cursor advances together with revision checks and bounded retries. Persist run-ID affinity before allocation. - Pass only authorized metadata and normalized usage to plugins. Core retains credential handling, member access checks, runtime qualification, and recovery evidence. Probe outside locks with a shared 15-second budget and freshness cache. - Resolve routing before credential preparation and backend selection. Preserve pins through turns, session resets, removed members, and quota waits. Retain admitted recovery after disable or uninstall. - Separate credential session epochs from token generations. Verified refresh preserves the epoch; reconnect and manual replacement change it. Include the credential slot ID in session and usage-cache identity, so reconnecting an indexed legacy account invalidates its old session even when both epochs are zero. - Validate pool member installations before accepting saved-agent bindings and recheck compatibility when the harness changes. Install only authorized members in the new-agent transaction and record their IDs in local activity. Pool membership cannot install a restricted shared connection. - Preserve pool bindings when agents hire teammates through either creation API or native caller runtime inheritance. Block stale manager credential references; retain explicit child authentication precedence and reject incompatible inherited pools. - Add native connector registration through plugin metadata. Reuse the Connectors catalog, setup header, account header, sidebar, dialogs, and usage display. Setup selects and orders saved connections. Advanced settings hold usage rules and member runtime defaults. New-account setup opens in another tab. - Use revision-checked pool archival from the Connectors catalog and account page. Keep task pins, cursors, recovery evidence, and underlying connections. Reject ordinary connection updates or removals that bypass pool revisions. - Add pool selectors, composer models, override notes, quota status, run details, activity records, and local run-log records. Keep session-adoption copy minimal. - Show **Used by** below the pool connections. List current company agents with shared avatars and profile links. Include paused agents; exclude terminated agents and agents using another pool. - Add Core stories for the generic connector workflow and runtime surfaces. Cloud stories reuse these production routes and tokens through a preview-only alias. ## Verification - Final head `73cb953bca30ed83e4505dd820edd9b5edffd28b`: full workspace `pnpm -r typecheck`, `pnpm build`, and `pnpm check:token-gates` pass locally. - All 422 focused connector/settings/shared-contract/migration tests and all 156 database-backed AI connection, hiring, reconnect, and durable-routing cases pass (69 hiring cases rerun after the final auth-precedence fix). The merged shared contract retains connection instructions and pool metadata. The pool migration is generated at sequence 0299 after the latest upstream migrations; this PR makes no lockfile changes. - All four full-app Playwright tests pass on the final head after a cold restart and migration, against the installed private plugin and isolated database, with no route or pool-API mocks. They cover hidden routing controls after manual opt-in, native pool creation, ordering, membership edits, rename, paused defaults, enabling/save/refresh persistence, stale edits, cancellation/removal, preserved underlying accounts, unavailable routers, and Used by avatars and profile links. Exact command: `PAPERCLIP_CONNECTION_POOL_E2E=1 AI_CONNECTIONS_TEST_COMPANY_ID=a37b9625-5ecf-4e29-8081-04df3d6e7d6f AI_CONNECTIONS_TEST_URL=http://127.0.0.1:3108 pnpm exec playwright test --config tests/ai-connections-app/playwright.config.ts connection-pools.spec.ts`. - [Native setup, ordering, and management screenshots](https://github.com/paperclipai/paperclip/pull/14967#issuecomment-6006976278) address the review follow-up. [Earlier selector, quota, and run-detail screenshots](https://github.com/paperclipai/paperclip/pull/14967#issuecomment-5971537316) show the runtime surfaces. Core previews: `pnpm --filter @paperclipai/ui storybook`, then **Connectors / Pool host** or **AI Connections / Connection pools**. Cloud owns its host-backed plugin stories; both repositories’ Operator Setup Required story assertions pass. - Live acceptance used OpenAI/Codex and Anthropic/Claude ACPX, resumed both exact sessions after restart, preserved pinned accounts through explicit reset and controlled quota deferral/recovery, and committed only two allocations across fourteen runs. A later UI-created task test again rotated OpenAI then Anthropic and resumed OpenAI through follow-up/restart/quota recovery. That later Anthropic execution was blocked by its saved OAuth token expiring (provider 401). No live usage probes ran. - The full local `pnpm test:run` was attempted earlier and did not complete because of macOS embedded PostgreSQL bootstrap/shared-memory failures and the 40,000-file Git fixture timeout. The focused database suites above now pass; full-suite verification is provided by the split CI lanes. The preceding CI run had one runtime readiness timeout; it passes locally both alone and inside the larger runtime suite. That larger local suite also encountered an embedded PostgreSQL setup failure and two macOS temporary-path alias assertions; those two assertions pass with canonical TMPDIR=/private/tmp. All final-head CI checks are terminal green, including full general/serialized server suites, Runner checks, browser E2E shards, canary verification, build, and typecheck. Greptile is 5/5 on that exact head with no unresolved threads. ## Risks - The migration adds routing tables and a credential epoch column. Install the private plugin only with the compatible Core contract. - Routing and each pool require opt-in. Production distribution and fleet defaults remain unchanged. - Unknown usage stays eligible. Known pinned exhaustion waits; revoked access requires operator repair. - Legacy adapters require compatible members. Runner model and effort overrides remain limited by qualified backend support. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository editing, code execution, and browser testing. The exact deployment model ID and context window are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes:` / `Closes:` / `Refs:` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (targeted suites; full-suite limitations are reported above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4857799a88 |
feat(connections): deliver saved instructions to authorized agent turns (#15216)
Persist optional connection instructions and deliver authorized snapshots to agent execution prompts. Keep provider templates with each app definition, preserve edits and opt-outs, and replace sessions when guidance or access changes. Use shared production settings across setup and Permissions, with source visibility in agent Instructions. Add the initial memory-provider defaults and managed Honcho workspace configuration. Include migration 0298 and regression coverage for generic providers, runtime delivery, authorization, and catalog regeneration. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
2c43b39167 |
refactor(ui): share task composer and project/worktree controls (#15091)
Use the shared composer for new tasks, including project and worktree selection, rich mentions and slash commands, and remembered task settings. Save project preferences after successful task updates. Show known agent models by name and use Default for unknown runtime defaults. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
ad0c4f0767 |
fix(ui): keep mobile task output above the composer (#15282)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - People read agent output in the task conversation. > - Mobile tasks scroll the document and keep the composer above the footer. > - The document body keeps a fixed height while the thread grows. > - The old observer misses late thread growth, and a layout scroll event can cancel follow mode. > - This pull request observes the thread and preserves follow mode until the reader scrolls up. > - The benefit is visible new output without repeated manual scrolling. ## Linked Issues or Issue Description **What happened?** New task output can move behind the mobile composer while the reader is already at the bottom. Late content layout does not resize the fixed-height body. A scroll event after layout growth can also clear follow mode before the resize observer runs. **Expected behavior** Keep new output visible above the composer while the reader follows the bottom. Hold the reading position after an upward scroll. Resume following when the reader returns to the bottom. **Steps to reproduce** 1. Open a long task conversation on mobile and scroll to the bottom. 2. Let a streamed row or delayed content grow after the parent renders. 3. Check whether the latest output stays above the composer. 4. Scroll up to read earlier output, then return to the bottom. **Paperclip version or commit** Reproduced by regression tests against master at |