mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 21:05:21 +02:00
243430f76eee6ed84b9fb855df8524d4e1dc318d
58
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
14867bd186 |
test(release-smoke): follow the mission-less onboarding reorder (#12135)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The nightly release lane publishes only after the release smoke suite passes against the newest canary > - Onboarding was reordered: step 1 now creates the company and routes straight to the agent step, and the mission step is gone (collected later in the tenant app, deliberately writing no goal) > - The smoke spec still walked the removed mission step, so the scheduled nightly has been red since the reorder shipped > - This pull request updates the spec to the current flow and asserts the deliberate empty goal list > - The benefit is a green nightly lane and an unblocked beta promotion from current master ## Linked Issues or Issue Description **What happened?** The scheduled `Release` nightly run fails in `smoke_nightly / smoke` since 2026-08-23 (runs 32630184811, 32710905212): `docker-auth-onboarding.spec.ts` waits for the `Define your mission` heading after step 1, but the wizard now routes 1 → 3 with no mission step (the step buttons literally skip from 1 to 3). The retry then fails on step 1 because the first attempt's company persists. **Expected behavior** The smoke passes against canaries carrying the reordered wizard, and the nightly lane publishes again. **Steps to reproduce** Run `scripts/docker-onboard-smoke.sh` with `PAPERCLIPAI_VERSION=2026.824.0-canary.7` and `pnpm run test:release-smoke` against it. **Paperclip version or commit** `2026.824.0-canary.7` Related (not duplicates): #11565 updated this same spec for the chat-first rewrite; this is the follow-up for the mission-less reorder. ## What Changed - Remove the mission-step interaction; step 1's "Next" now creates the company and the spec goes straight to the agent step. - Replace the mission-goal API assertion with the truthful one: onboarding deliberately writes no goal, so a fresh company's goal list is empty. - Update step comments to match the shipped flow. ## Verification - Local run of the exact CI harness against `paperclipai@2026.824.0-canary.7`: 1 passed (6.8s), exit 0. - The suite's remaining API assertions (company, CEO agent, seeded task assignment, landed issue URL, assignment-sourced heartbeat run) pass unchanged. ## Risks - Low risk: test-only. The spec remains copy-coupled to the wizard — this is the third drift in two weeks; stable `data-testid` hooks in the wizard remain the durable fix and can follow separately. ## Model Used Claude Fable 5 (Claude Code) ## Pre-submission checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template |
||
|
|
10d2781a29 |
feat(sandbox): add the duplex bridge broker, gated transport selection, and fixed observability (#11769)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Sandbox adapters provide controlled execution for untrusted provider environments. > - The sandbox channel needs one persistent duplex transport with strict host control. > - The transport must remain off unless the instance setting and provider capability both allow it. > - The host must detect loss, bound resource use, and expose only safe telemetry. > - This pull request adds the broker, gated selection, kill-switch wiring, fixed observability, and real-process proof. > - The benefit is safer sandbox execution with bounded failure behavior and inspectable transport results. ## Linked Issues or Issue Description No public issue exists for this change. The related pull requests are #11738 and #11750. **Problem or motivation** The sandbox duplex channel needs a host-controlled broker, strict transport gates, bounded provider input, and safe loss telemetry. Without these controls, a provider can cause replay, resource growth, unsafe endpoint selection, or data exposure through telemetry. **Proposed solution** Add a host broker with nested time limits, request limits, one-shot loss, and per-id deduplication. Select duplex transport only when the instance setting and provider capability both equal true. Assign the endpoint and nonce on the host. Reject invalid readiness data and use the file bridge on failure. Add fixed redacted telemetry and a real-process end-to-end test harness. **Alternatives considered** Keep the file bridge as the only transport. This avoids new channel behavior but does not provide persistent duplex operation for supported sandbox providers. **Roadmap alignment** This change supports the Cloud / Sandbox agents section in ROADMAP.md. ## What Changed - Add the duplex bridge broker with bounded forward, response, and gateway wait budgets. - Bound concurrent requests, lifetime requests, and request-id bytes before retention or forwarding. - Select duplex transport only when both required gates are true. - Assign the loopback port and nonce on the host and enforce a liveness-only READY frame. - Fall back to the file bridge after invalid readiness, contamination, bind failure, or timeout. - Carry the kill switch through the server, acpx engine, and six local adapters. - Add fixed, redacted duplex telemetry with a provider allowlist. - Add a real-process end-to-end harness for readiness, round trips, loss, and teardown. - Add regression coverage for limits, loss, UTF-8 splits, concurrency, and telemetry dimensions. ## Verification - Adapter-utils, server, and Daytona typechecks pass locally. - Adapter-utils tests pass, including the codec, broker, execution-target sandbox, and real-process harness. - Server kill-switch tests pass. - Live Daytona tests pass with the required provider key and skip without that key. - The root pnpm-lock.yaml file has no diff. - The branch contains ten commits after origin/master. ## Risks - Duplex transport remains disabled unless both gates equal true. - A provider remains an untrusted boundary and needs least-privilege credentials and quotas. - The server telemetry recorder stays deferred; the default recorder does nothing. - A provider that pre-binds the host port causes a fail-closed fallback to the file bridge. - The change adds no database migration and changes no root lockfile. ## Model Used OpenAI GPT-5, exact model family GPT-5, large context window, reasoning, and tool use. The model assisted with Git handoff validation and PR preparation. The implementation commits came from the engineering worktree. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with Fixes: # / Closes # / Refs # OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub #NNN / github.com/paperclipai/paperclip URLs) - [x] My branch name describes the change (e.g. docs/... or fix/...) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge - [x] I searched the GitHub PR list for similar PRs and confirmed this is not a duplicate --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3ff636bc48 |
Drop the mission step from the wizard arc (#11935)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - New customers arrive through an onboarding arc that spans Paperclip Cloud and the tenant app > - Cloud's naming screen stopped asking for the company mission, but the tenant wizard still decided its first step by asking whether the company had one > - Every Cloud-created company therefore looked mission-less on arrival, so every walk detoured through a "Define your mission" screen the design had already removed > - This pull request removes that step from the arc, and makes step 1 create the company itself > - The benefit is a shorter arc that matches the design, and a three-step progress strip that now counts the steps that exist ## Linked Issues or Issue Description No public issue exists. The problem was found by walking staging end to end. **What happened:** A new customer signs in, names their organization, and waits for it to build. The tenant wizard then asks "Define your mission" before it asks for the first agent. Cloud no longer collects a mission, so this screen appears for every new customer. **Expected behavior:** The wizard asks for the first agent, the model, and a review. The progress strip counts three steps. **Actual behavior:** The wizard asks for the mission first. The progress strip counts five segments, because the run does not enter on the agent arc. **Additional context:** Three merged pull requests built the mission-based step choice this change removes: #11352, #11416 and #11429. The mission is now collected later, inside the tenant app, so onboarding does not ask for it at all. ## What Changed - `onboardingStepForCompany` always returns the agent step. The `companyHasMission` parameter is removed, because it cannot change the answer. - `resolveRouteOnboardingOptions` no longer accepts `companyHasMission`. - The dashboard no longer waits for the goal lookup before it opens the wizard. That wait only chose a step, and the step is now fixed. - Step 1 creates the company in a new `handleCreateCompany`. Company creation used to sit at the end of `handleConfirmMission`. - No company goal is written during onboarding. - The three-step strip now shows on the agent, model and review steps, because every Cloud-first run enters on the arc. - The full-length bar drops its second segment. No run can fill it. - The grow path keeps its step 2 questionnaire. Only the create path skips ahead. - Back from the agent step goes to the screen the run came from. - Four end-to-end specs no longer drive the wizard through the mission step. ## Verification Run the tenant test suite: ``` cd ui && npx vitest run ``` - 4356 tests pass. 471 files pass. - `npx tsc --noEmit` reports no errors. - Fault injection: forcing `skipsMissionStep` to `true` fails the grow questionnaire test. Removing the Back rule fails the Back test. Both tests fail on the exact defect they guard. - The three-step strip is asserted by an existing test. It checks `Step 1 of 3` and `aria-label="Create your first agent"`. Manual check on staging after the paired Cloud change: 1. Open a new incognito window. 2. Sign in with a new account. 3. Name the organization. 4. Confirm the wizard shows "Create your first agent" and "Step 1 of 3". ## Risks - **Behavioral change.** Onboarding no longer writes a company goal. An agent hired during onboarding starts without a seeded mission. This is intended. The mission moves to the tenant app. - **Dead code.** `ONBOARDING_MISSION_STEP` and the mission screen stay in the codebase, but nothing in the app opens them. They wait for the surface that collects the mission later. - **Grow path.** The grow path is unchanged, but it shares step 2 with the removed screen. New tests cover it. - **Superseded work.** #11352, #11416 and #11429 tuned the mission-based step choice. This change removes the branch they tuned. ## Model Used Claude Opus 5 (`claude-opus-5`), extended thinking, with tool use and code execution through Claude Code. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket id - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
3d366ba15f |
Rebuild the onboarding agent arc on the prototype's step design (#11905)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The onboarding wizard in `ui/` hires that first agent. It runs three steps: create the agent, connect a model, and review > - A standalone prototype holds the agreed design for these steps. #10786 ported that prototype, but #11067 reverted it in full because the port deleted `OnboardingWizard.tsx` while four pull requests were editing that file > - Those four pull requests have since merged. The revert said the port can "re-land incrementally", and this is that re-land > - This pull request takes the presentational layer from the prototype only. It keeps master's wizard as the source of behaviour, so the eight onboarding fixes merged since the revert stay in place > - The benefit is that the three agent steps match the agreed design, and no merged fix is lost to get there ## Linked Issues or Issue Description Refs #10786 — the first attempt to land this design. Refs #11067 — the revert that asked for it to re-land in smaller steps. No public issue exists for the re-land. The problem is described below. **Subsystem affected** The `ui` package. The change touches the onboarding wizard, the agent capsule, and one Storybook story. It adds four small presentational components under `ui/src/components/onboarding/`. **Current behavior** The wizard's agent steps do not match the prototype. Each step shows a small heading beside an icon, above a form. The agent capsule sits below that heading and does not animate. The agent gets a name but no role, so every first agent is created as `ceo`. The wizard also shows a five-segment progress bar on these steps. A walker who enters on the agent step cannot reach the first two segments, so two of the five can never be filled. **Proposed behavior** The three steps use the prototype's card, its centred display heading, and its footer. One capsule sits above the heading and stays mounted across all three steps, so it reads as one object being built rather than three screens that each show their own. A three-segment strip counts these steps for a walker who enters on them. The full-length bar stays for a walker who starts at step one, so that count never restarts partway. The agent step gains a role. The options come from the agent role enum, not from the prototype's mock list. **Reason and benefit** The design is agreed and already built once. Re-landing it presentation-first keeps the behaviour that master gained after the revert. Sourcing roles from the enum matters. The prototype offers "Coder", which is not a valid role — the enum uses `engineer` — so a walker who picked it would fail validation at hire time. **Breaking changes** None. The wizard keeps its routes, its draft format, and its hire call. The draft gains one optional field, `agentRole`. A draft saved before this change loads without it and falls back to the default. ## What Changed - Add `ui/src/components/onboarding/`: `Stepper`, `OnboardingCard`, `OnboardingHeading`, `FooterNav`, `AgentPreview`, and shared motion constants - Rebuild wizard steps 3–5 on those parts: one card, the capsule above a centred heading, and one footer - Hold one `AgentCapsule` across the three steps. It springs in once, then morphs from dashed slot to traced outline to filled - Add `strokeDraw` to `AgentCapsule`. It traces the outline instead of cross-fading it. The dashed layer holds until the trace ends - Add a role select to the agent step. Choosing a role fills the name, unless the walker typed one - Show one progress indicator per run, not two - Label strip segments by destination, not by number - Add `motion` to the `ui` package - Add a Storybook story for the strip and the capsule states ## Verification Run the tests: ``` pnpm --filter @paperclipai/ui exec vitest run pnpm --filter @paperclipai/ui exec tsc -p tsconfig.json --noEmit ``` 4235 tests pass. The typecheck is clean. To see the steps, start the app and open `/<PREFIX>/onboarding` for a company that has a company-level goal. The wizard opens on the agent step. Step three requires a hire. Three absence assertions were checked by fault injection. Each one fails when the old behaviour returns: - put the step counter back, and the "shows no step counter" test fails - default `strokeDraw` to true, and the cross-fade test fails - restore the timer gate on the strip, and the indicator test fails ## Risks Low to medium. `motion` is one new dependency in `ui`. #11067 gave dependency weight as one of three reasons to revert #10786, so this branch carries the smallest set that works. `motion` drives the step transitions and the capsule choreography, and three files import it. An earlier revision of this branch also added `three` and `@types/three`. Both are removed. They existed for the 3D backdrop, which belongs to the auth and welcome screens rather than to these three steps, so nothing on this branch imported them. The role select changes what the wizard sends. Before this change every first agent was hired as `ceo`. Now the walker chooses. The values come from the enum, so the server accepts all of them. Steps 1 and 2 keep the older design. They do not run on the Cloud-first path, where the company already exists. ## Model Used Claude Opus 5 (`claude-opus-5`), with extended thinking, tool use, and code execution. Used for the code, the tests, and this description. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
a7e689b3c3 |
feat(ui): rename "Agent mode" to "Auto mode" and show full work-mode labels (#11866)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - The task composer and the New Task dialog show a work-mode chip (how
the agent will run a task)
> - The default mode was labeled "Agent mode", and the New Task dialog
chip abbreviated every mode to one word ("Auto", "Plan", "Ask")
> - "Agent mode" is confusing because every mode runs an agent, and the
abbreviated chip hid what the label means
> - This pull request renames the mode to "Auto mode" and makes every
mode chip show the full label
> - The benefit is a clearer, consistent mode name in every place the
user selects a work mode
## Linked Issues or Issue Description
No public GitHub issue exists for this change. Description follows the
enhancement template:
**What existing behavior does this improve?**
The work-mode selector chips in the task composer and in the New Task
dialog.
**Subsystem affected**
UI (`ui/src/lib/work-mode-meta.ts`,
`ui/src/components/NewIssueDialog.tsx`).
**Current behavior**
The default work mode is labeled "Agent mode". The New Task dialog chip
shows a shortened label ("Auto", "Plan", "Ask") from a separate
`shortLabel` field.
**Proposed behavior**
The default work mode is labeled "Auto mode". Every chip shows the full
label ("Auto mode", "Plan mode", "Ask mode"). The `shortLabel` field is
removed so no surface can fall back to the short form.
**Reason and benefit**
"Agent mode" does not describe the behavior — all modes use an agent.
"Auto mode" states what the mode does. One label field keeps every
surface consistent.
**Breaking changes**
None. This is a display-string change only. No API, storage, or mode-key
changes.
## What Changed
- Renamed the `standard` work-mode label from "Agent mode" to "Auto
mode" in `ui/src/lib/work-mode-meta.ts`, the single source for all mode
chips.
- Changed the New Task dialog mode chip to render the full `label`
instead of `shortLabel`.
- Deleted the `shortLabel` field from `WorkModeMeta` so nothing can
silently regress to the short form.
- Updated unit tests and fixtures to pin the full labels.
## Verification
- Run `pnpm --filter @paperclipai/ui test --
src/lib/work-mode-meta.test.ts src/components/NewIssueDialog.test.tsx
src/components/IssueChatThread.test.tsx
src/components/task-chat/TaskChatComposer.test.tsx`. All tests pass. The
tests assert the labels are exactly "Auto mode", "Plan mode", and "Ask
mode".
- Manual: start the dev server, open the board, press `c` to open the
New Task dialog, and press Cmd+Period to cycle modes. The chip reads
"Auto mode", "Plan mode", then "Ask mode". The composer chip on an open
task shows the same labels.
## Risks
- Low risk. Display strings only. The chip is a few pixels wider in the
New Task dialog; no layout overflow was observed in any of the three
modes.
## Model Used
- Claude Fable 5 (Anthropic, model ID `claude-fable-5`), extended
thinking with tool use, run inside a Claude Code agent session.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
1c366a9059 |
fix(server): reject invalid agent credentials instead of downgrading to the local user actor (#11589)
<!-- Write all pull request text in Simplified Technical English (ASD-STE100): short sentences, one instruction per sentence, simple approved vocabulary, and the active voice. --> ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The server authenticates each agent request in `actorMiddleware` before it attributes chat comments > - When an agent bearer token failed verification, the middleware called `next()` with no error and the request continued without an agent actor > - The request then fell back to the local user actor, so the server stored agent replies as user comments > - The task chat UI renders user comments in blue bubbles, so agent messages appeared as blue user bubbles > - This pull request rejects invalid agent credentials with 401 instead of a silent downgrade > - The benefit is that agent messages keep agent attribution, and broken credentials fail loudly with a clear retry message ## Linked Issues or Issue Description **What happened?** A user cancelled an onboarding question card. The agent posted a follow-up reply. The reply appeared in a blue bubble, which the UI reserves for human messages. The agent run held an expired local agent JWT. The auth middleware could not verify the token, called `next()` without an actor, and the request fell back to the local user identity. The server stored the agent comment as a user comment. **Expected behavior** Agent messages always render as agent bubbles. A request with invalid agent credentials must fail with 401 so the adapter can refresh credentials and retry. It must not post content under a human identity. **Steps to reproduce** 1. Start a local Paperclip instance. 2. Give an agent run an expired or malformed agent JWT. 3. Let the agent post an issue comment through the API bridge. 4. Before this change: the comment is stored with the local user identity and renders as a blue bubble. After this change: the request fails with 401 and a message that tells the caller to obtain fresh credentials. ## What Changed - `server/src/middleware/auth.ts`: a bearer token that fails verification now produces a 401 `unauthorized` error instead of a silent fall-through to the anonymous/local-user actor. - The 401 message states the cause: expired token, unverifiable token, empty bearer token, missing agent record, agent record in another company, terminated agent, or agent pending approval. - The API-key path now also rejects an agent record whose company does not match the key. - `packages/adapter-utils/src/execution-target.ts`: the bridge proxy now writes a `comment id: <id>` marker to the run log for each posted issue comment, so misattributed comments can be traced to a run. - `ui/src/components/task-chat/task-chat-adapter.test.ts`: a regression test asserts that a recovered `local-board` comment with a derived agent author renders as an agent bubble, not a user bubble. - `server/src/__tests__/agent-auth-middleware.test.ts` and `packages/adapter-utils/src/execution-target-sandbox.test.ts`: new tests cover each rejection path and the log marker. ## Verification - Run `pnpm vitest run src/__tests__/agent-auth-middleware.test.ts` in `server/` — 14 tests pass. - Run `pnpm vitest run execution-target-sandbox` at the repo root — 44 tests pass. - Run `pnpm vitest run src/components/task-chat/task-chat-adapter.test.ts` in `ui/` — 4 tests pass. - Manual check: post an issue comment with an expired agent JWT; the API returns 401 with a retry message and no comment is stored. ## Risks - Behavioral shift: requests that previously continued as anonymous or local-user actors after a failed agent-token verification now receive 401. Any caller that relied on the silent downgrade must refresh its credentials. This is the intended fix, and the adapters already handle 401 with a credential refresh. - No schema or migration changes. Low risk otherwise. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Claude (Anthropic), model ID `claude-fable-5`, via Claude Code with extended thinking and tool use (agent harness with shell, file, and git tools). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
4af55ba6bd |
test(release-smoke): follow the onboarding wizard's chat-first rewrite (#11565)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release channel system publishes a nightly build only after the release smoke suite passes against the newest canary > - The scheduled nightly run has failed every night since August 12, so no nightly, and therefore no beta candidate, has shipped for six days > - The failures are stale test locators, not a product regression: the chat-first onboarding rewrite (#11101) changed wizard copy and the post-launch destination > - This pull request updates the smoke spec to match the current wizard > - The benefit is a green nightly lane and an unblocked beta promotion ## Linked Issues or Issue Description **What happened?** The scheduled `Release` nightly run fails in `smoke_nightly / smoke` every night since 2026-08-12. The failing spec is `tests/release-smoke/docker-auth-onboarding.spec.ts`. Four assertions no longer match the product after the chat-first onboarding rewrite (#11101): - The step-1 heading is now "Name your organization", not "Name your company". - The step-4 hire button is now "Connect", not "Give it a heartbeat". - The seeded first task is now titled "Paperclip onboarding". - A successful launch navigates to the seeded task's thread (`/issues/<ref>`), not `/dashboard`. **Expected behavior** The smoke suite passes against a canary that contains the current onboarding wizard, and the nightly lane publishes again. **Steps to reproduce** Run `.github/workflows/release-smoke.yml` against `paperclipai@canary` (any version at or after the rewrite), or dispatch `release.yml` with `channel: nightly`. Example red runs: 32014452506 (Aug 17), 31938284401 (Aug 16). **Paperclip version or commit** `2026.817.0-canary.12` Related (not duplicates): #11190 updated this same spec for the mission-first wizard; this PR is the follow-up for the chat-first rewrite that landed after it. ## What Changed - Update the step-1 wizard heading locator to "Name your organization". - Update the step-4 hire button locator to "Connect" and reword the step comment. - Update `FIRST_TASK_TITLE` to "Paperclip onboarding" (the wizard's current `DEFAULT_TASK_TITLE`). - Assert the post-launch URL is the seeded task's thread (`/issues/`), not `/dashboard`. ## Verification - Local run of the exact CI harness: `scripts/docker-onboard-smoke.sh` with `PAPERCLIPAI_VERSION=2026.817.0-canary.12`, then `pnpm run test:release-smoke` against the container — 1 passed. - The suite's later API assertions (company, CEO agent, mission goal, seeded issue assignment, assignment-sourced heartbeat run) all pass unchanged against the current canary. ## Risks - Low risk: test-only change; no product code is touched. - The spec remains copy-coupled to the wizard. If wizard copy churn continues, a follow-up could add stable `data-testid` hooks to the wizard so the smoke spec stops breaking on wording changes. ## Model Used Claude Fable 5 (Claude Code) ## Pre-submission checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template |
||
|
|
a09d7dcc06 |
feat(ui): bounce cold arrivals off archived company URLs, add Unarchive (#11302)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Archiving a company hides it from the sidebar switcher, but remembered last-visited paths, browser history, bookmarks, and restored tabs keep depositing users onto its URLs long after archiving > - Since the selection ping-pong fix (#11300) those arrivals render, but the user is stranded inside a workspace the sidebar refuses to show — and unarchiving had no UI anywhere, so the only way back was a hand-typed settings URL > - This pull request bounces cold arrivals at archived company URLs to an active company (with a toast naming why), lets deliberate visits stick, and adds an Unarchive action to the companies list > - The benefit is that stale URLs stop stranding users in retired workspaces, and archived companies become restorable from the one page that still lists them ## Linked Issues or Issue Description Follow-up to #11300. No existing issue for the remaining gap; description follows the enhancement template: **What happened?** After #11300, opening an archived company's URL (stale tab, history, bookmark, remembered path) renders that company's pages — but the sidebar switcher does not list it, so the user is stranded in a workspace they retired, and every stale URL pulls them back in. Separately, unarchiving a company has no UI: the archive button lives in company settings, which becomes unreachable through normal navigation once the company is archived. **Expected behavior** Arriving cold at an archived company's URL lands the user in an active workspace, with a toast explaining the redirect. Explicitly choosing the archived company (from the companies list) still works, so its pages remain reachable. Archived companies can be restored from the companies list. **Steps to reproduce** 1. Create two companies; archive one. 2. Open `/{archivedPrefix}/dashboard` directly — before: renders the archived workspace with no sidebar presence; after: bounces to the active company's dashboard with a toast. 3. On the companies list, open the archived company's row menu — before: no restore action anywhere; after: Unarchive. ## What Changed - `ui/src/lib/company-selection.ts`: `resolveArchivedCompanyBounce` — pure policy: bounce when the URL names an archived company that is not the current selection and an active company exists; prefer the currently selected active company as the destination. - `ui/src/components/Layout.tsx`: the route-sync effect applies the bounce (toast + selection + `replace` navigation) before syncing selection from the route. - `ui/src/pages/Companies.tsx`: Unarchive action (`PATCH status: "active"`) in the row menu for archived companies. - Tests: unit cases for the bounce policy; the e2e now drives all three behaviors (direct-load bounce with toast, re-arrival bounce, deliberate visit sticks) on top of the existing crash regression. ## Verification - `pnpm vitest run src/lib/company-selection.test.ts src/context/CompanyContext.test.tsx src/pages/Companies.test.tsx` in `ui/` — 20 tests pass. - `npx playwright test --config tests/e2e/playwright.config.ts archived-company-url` — passes, covering bounce, toast, and deliberate-visit paths. - `pnpm typecheck` in `ui/` — clean. ## Risks Low risk. The bounce only fires for archived-company URLs when the archived company is not already selected and an active company exists; all-archived instances render as before. Deliberate selection from the companies list is unaffected (selection equals the matched company, so no bounce). Unarchive reuses the existing `PATCH /api/companies/:id` status transition the server already supports. ## Model Used - Claude (Anthropic), Claude Fable 5 (`claude-fable-5`), via Claude Code CLI with extended thinking and tool use (code search, edit, test execution, Playwright e2e). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
61a5b7c6f9 |
fix(ui): stop the selection ping-pong on archived company URLs (#11300)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The UI keeps a selected company in `CompanyProvider` with two writers: a bootstrap effect that repairs invalid selections, and a Layout route-sync effect that selects the company the URL prefix names > - The route-sync matches the URL against the full company list (archived included), while the bootstrap resolver only accepted companies from the sidebar-filtered non-archived list > - On any archived company's URL the two effects overwrite each other's selection in a synchronous loop until React throws error #185 ("Maximum update depth exceeded") and unmounts the root to a blank page — armed by remembered last-visited paths, back/forward navigation, or bookmarks, on first load and client navigation alike > - This pull request makes an already-selected company only need to exist, keeping the sidebar filter for fresh-boot resolution where no explicit selection exists > - The benefit is that archived company URLs render instead of blanking the entire app ## Linked Issues or Issue Description No existing issue. Description follows the bug template: **What happened?** Opening (or back-navigating to) a URL whose company prefix belongs to an archived company blanked the whole app with `Minified React error #185`. Console in dev mode: "Maximum update depth exceeded. This can happen when a component calls setState inside useEffect…". A workspace whose first/seeded company was archived hit this on every load of its remembered URL. **Expected behavior** An archived company's URL renders its pages (the company still exists and its API routes serve data). The sidebar simply does not feature archived companies, and fresh boots still land on a non-archived company. **Steps to reproduce** 1. Create two companies; archive one (`PATCH /api/companies/:id` with `status: "archived"`). 2. Navigate to `/{archivedPrefix}/dashboard` — direct load or client-side back-navigation. 3. Before this fix: React #185 and an unmounted blank page (reproduced deterministically by the new e2e test). ## What Changed - `ui/src/context/CompanyContext.tsx`: `resolveBootstrapCompanySelection` keeps an explicitly selected company that exists in the full company list; stored-id and default resolution still prefer sidebar (non-archived) companies. - `ui/src/context/CompanyContext.test.tsx`: resolver keeps an archived-but-existing selection; a truly deleted selection is still replaced. - `tests/e2e/archived-company-url.spec.ts`: end-to-end regression driving both field shapes (direct load and back-navigation onto an archived company URL); it failed with the exact #185 console errors before the fix and passes after. ## Verification - `pnpm vitest run src/context …` in `ui/` — 122 tests pass (includes the new resolver cases). - `npx playwright test --config tests/e2e/playwright.config.ts archived-company-url` — fails before the fix (captured "Maximum update depth exceeded" console errors), passes after. - `pnpm typecheck` in `ui/` — clean. ## Risks Low risk. The only behavioral change is that a selection naming an archived-but-existing company survives the bootstrap repair — previously that state was unreachable without crashing. Boots with no valid selection behave exactly as before (non-archived preferred), covered by the existing and new resolver tests. ## Model Used - Claude (Anthropic), Claude Fable 5 (`claude-fable-5`), via Claude Code CLI with extended thinking and tool use (code search, edit, test execution, Playwright-driven crash reproduction). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
9c941169a6 |
fix(ui): keep new task dialog visible above mobile keyboard (#11281)
<!-- Write all pull request text in Simplified Technical English (ASD-STE100): short sentences, one instruction per sentence, simple approved vocabulary, and the active voice. --> ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Operators create tasks in a dialog that includes the assignee and project fields. > - Mobile browsers reduce and offset the visual viewport when the on-screen keyboard opens. > - The dialog used layout viewport units, so its upper fields could move off-screen while the user typed. > - This pull request makes the dialog follow the live visual viewport and keeps the focused editor visible. > - The benefit is that operators can see the task context and the field they edit on mobile devices. ## Linked Issues or Issue Description **What happened?** On mobile browsers, opening the keyboard in the new-task dialog could move the assignee and project fields above the visible screen. The active editor could also become difficult to see. **Expected behavior** The full dialog must stay inside the visible browser area. The active editor and task controls must remain reachable while the on-screen keyboard is open. **Steps to reproduce** 1. Open Paperclip on a mobile browser. 2. Open the new-task dialog. 3. Focus the title or description editor to open the on-screen keyboard. 4. Observe that the upper fields can move outside the visible viewport. **Paperclip version or commit** Reproduced before commit `838cdbb325` on `master`. **Deployment mode** Local dev (`pnpm dev`) in a mobile browser viewport. ## What Changed - Read `window.visualViewport` while the dialog is open. - Apply token-based dialog geometry when the visual viewport is constrained. - Keep the focused editor visible after viewport resize and scroll events. - Add unit coverage for visual viewport updates and focus scrolling. - Add Playwright coverage for mobile, tablet, desktop keyboard, and unconstrained desktop layouts. ## Verification - `pnpm exec vitest run ui/src/components/NewIssueDialog.test.tsx` — 27 tests passed. - `pnpm --filter @paperclipai/ui typecheck` — passed. - `pnpm check:token-gates` — passed with all gates clean. - `pnpm --filter @paperclipai/ui build-storybook` — passed. - `pnpm exec playwright test tests/storybook-visual/new-issue-dialog-viewport.spec.ts --config tests/storybook-visual/playwright.config.ts` — 4 tests passed. ## Risks - Low risk. The custom geometry only activates when `visualViewport.height` is less than `window.innerHeight`. - Browsers without the Visual Viewport API keep the existing dialog primitive behavior. - The browser test checks hit targets and visible bounds at mobile, tablet, and desktop widths. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5. The session used reasoning, repository tools, shell execution, and browser automation. The service did not expose the context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
2494a2a0fe |
perf: add repeatable issue-detail baseline rig (#10409)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The issue detail page is a core operator surface where perceived latency directly affects task navigation > - Performance work needs repeatable evidence so later optimizations can be compared against the same scenarios > - The page did not expose stable user-timing marks for its header or first useful content > - There was also no isolated seeded browser rig that measured warm navigation, cold deep links, waterfalls, or server time > - This pull request adds the instrumentation and a one-command Playwright baseline harness > - The benefit is that issue-page performance changes can be validated with reproducible median measurements instead of anecdotes ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: `ui/`, `server/`, and browser performance tooling. **Problem or motivation** The issue detail page performs a large client bootstrap and request fan-out, but the repository lacks stable user-timing boundaries and a repeatable benchmark. That makes performance changes difficult to compare and allows regressions to be judged from anecdotes instead of consistent evidence. **Proposed solution** Add stable header/content paint measures, development/QA-only lifecycle vital reporting, aggregate server timing for the issue endpoint, and a seeded Playwright command that runs warm/cold scenarios under throttled and unthrottled profiles with N≥5 median reporting. **Alternatives considered** Ad hoc DevTools recordings were rejected because they are not repeatable or reviewable. Production telemetry was rejected because this baseline should not change production data collection. A unit-only harness was rejected because it cannot capture browser bootstrap, rendering, and network waterfall costs. **Roadmap alignment** The roadmap calls for agent performance to be measurable over time. This change applies that evidence-first principle to a core operator page and does not duplicate a listed roadmap deliverable. **Additional context** The generated report includes warm and cold medians, TTFB/FCP/LCP where applicable, request and byte totals before first useful content, JavaScript bytes, and issue endpoint server timing. ## What Changed - Added `issue-detail:navigate→header-paint` and `issue-detail:navigate→content-paint` user-timing measures to the issue detail page. - Added development/QA-only TTFB, LCP, and INP console reporting without production telemetry delivery. - Added `Server-Timing` for `GET /api/issues/:id`. - Added `pnpm exec playwright test --config tests/perf/issue-detail/playwright.config.ts`, which seeds an isolated instance and runs N≥5 warm/cold samples under unthrottled and Fast 4G/4x CPU profiles. - Added Markdown, raw JSON, and Chrome-trace outputs with median baseline tables and waterfall data. ## Verification - `pnpm --filter @paperclipai/ui typecheck` - `pnpm --filter @paperclipai/server typecheck` - `pnpm check:token-gates` - `npx playwright test --config tests/perf/issue-detail/playwright.config.ts --list` - `pnpm exec playwright test --config tests/perf/issue-detail/playwright.config.ts` — passed 20 samples in 9.4 minutes (5 runs × 2 scenarios × 2 profiles) for the baseline; post-review integrity reruns also exercised the corrected paths, while this shared runner intermittently killed Chromium processes, so the rig now performs one bounded browser-crash retry per sample. - Baseline medians: warm unthrottled 278/447 ms header/content; cold unthrottled 646/646 ms; warm throttled 1240/2060 ms; cold throttled 3932/3933 ms. ## Risks - Low product risk: the new browser measurements are development/QA tooling and the UI timing work does not change visible layout. - `Server-Timing` exposes only aggregate handler duration, not query contents or private identifiers. - Native INP reporting uses supported browser event timing entries and silently no-ops where unsupported. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5.4, tool-assisted coding and browser execution with reasoning enabled; context-window size is not exposed in this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Dev Agent <dev@paperclip.ing> |
||
|
|
815e49bb7c | feat: make chat-style tasks the default experience (#11101) | ||
|
|
5a0985f80a |
test(release-smoke): update onboarding spec for the mission-first wizard (#11190)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release subsystem's nightly lane gates every nightly publish on the release smoke suite, which drives real onboarding in a browser against the published artifact > - With the harness fixed (#11187, #11189), the gate reached the Playwright suite for the first time in CI — and the spec still walks the old onboarding wizard, so it fails at "Create your first agent" on every current build > - The wizard was redesigned to a mission-first five-step flow, and the spec rotted silently because the suite never ran in CI before > - This pull request rewrites the spec to drive the current wizard end to end > - The benefit is a smoke gate that actually tests today's product, verified against a real published canary ## Linked Issues or Issue Description **Subsystem affected** Release smoke testing: `tests/release-smoke/docker-auth-onboarding.spec.ts`. **Problem or motivation** Nightly run 31431273139 failed in the smoke Playwright suite: the spec expects the old wizard step "Create your first agent", but current builds show the redesigned mission-first flow (front door → company → mission → team lead → connect model → review). The page snapshot in the run artifact shows the "Define your mission" step where the spec expected the agent step. Both retries failed identically — this is deterministic spec drift, not flake. **Proposed solution** Rewrite the spec for the current flow: fill the company name, define the mission directly (confirming creates the company), name the team lead, hire it through the adapter step — the adapter environment probe reports unhealthy in the CLI-less smoke container by design and must not block the hire — then launch to the dashboard. Assert the company, the ceo-role agent, and the company goal through the API. The first-task and assignment-run assertions are removed together with the wizard flow that created them. ## What Changed - `tests/release-smoke/docker-auth-onboarding.spec.ts`: rewritten for the mission-first wizard; sign-in and wizard-opening helpers and the company-name step are unchanged ## Verification - Full local run against the real nightly candidate: launched the smoke container for `paperclipai@2026.810.0-canary.3` via `scripts/docker-onboard-smoke.sh` (with the #11189 bind fix), then ran `pnpm run test:release-smoke` against it — 1 passed (4.5s) - After merge: dispatch `release.yml` with `channel: nightly` to run the full gate in CI ## Risks - Low. Test-only change. The spec now asserts less about first-task creation because the wizard no longer creates a first task; if a first-run trigger returns to onboarding, the spec should grow that assertion back ## Model Used Claude Fable 5 (`claude-fable-5`, Anthropic) in Claude Code, with extended thinking and full tool use (CI artifact forensics, UI source tracing, local Docker + Playwright reproduction and verification). All changes model-authored under human direction. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending — will confirm before merge) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending — will confirm before merge) - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
384e5f6178 |
revert(ui): back out onboarding port (#10786) (#11067)
This reverts commit |
||
|
|
0a511ed1b0 |
feat(apps): support multiple provider connections (#11060)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Apps subsystem connects company tools through governed provider connections. > - A company can need more than one account for the same provider. > - The current database constraint and Apps flow assume one named connection per company. > - New quarantined actions also need an explicit review decision before activation. > - This pull request supports multiple provider connections and complete action review decisions. > - The benefit is safer access control and a clear multi-account Apps workflow. ## Linked Issues or Issue Description Refs: #11040 **Subsystem affected** Cross-cutting. This change affects the Apps UI, the tool access API, the shared request contract, and the database schema. **Problem or motivation** The connection name constraint prevents a company from keeping more than one connection for a provider. The Apps UI also reuses an existing OAuth connection when a user asks to connect another account. Action review can enable selected entries without recording a decision for every quarantined action. **Proposed solution** Remove the company and connection name uniqueness constraint. Let users open, count, edit, and create multiple provider connections. Require the finish request to cover every quarantined action exactly once before the server activates reviewed entries. **Alternatives considered** The UI could generate unique internal names and keep the database constraint. This would preserve a one-connection assumption in the data model and would make display names part of identity. The server could also infer review decisions from enabled actions. This would not distinguish a reviewed disabled action from an action that the user did not review. **Roadmap alignment** This change extends the completed MCP Tool Gateway and Apps milestone. It also supports the Connected Apps roadmap item. It follows the navigation and connection management work in #11040. ## What Changed - Remove the company-scoped connection name uniqueness index with an ordered and idempotent migration. - Add a reviewed action list to the finish-app contract and reject incomplete or duplicate review decisions. - Activate reviewed entries and keep unreviewed quarantined entries blocked. - Enable a completed connection and preserve the company and connection scope in all updates. - Show provider connection counts and open the provider setup page from Browse. - Let users edit existing connections or connect another account without reusing an active OAuth connection. - Update focused server and UI coverage for multiple connections and action review. ## Verification - Ran the focused Apps UI suite. All 116 tests passed in 11 files. - Ran the focused server and CLI suite. All 276 tests passed in 3 files. - Ran `pnpm --filter @paperclipai/db check:migrations`. The migration safety check passed. - Ran `pnpm -r typecheck`. All projects passed. - Ran `pnpm build`. All projects built successfully. - Ran `pnpm test:run`. It passed 3,735 tests and skipped 4 tests. One worktree-safety assertion failed because the execution workspace reloads its worktree marker. The same test passed with an isolated non-worktree marker. - Ran `pnpm check:token-gates`. It reports 12 existing violations in the unchanged `PaperclipOrbit3D.tsx` file from the target branch. - Started the six affected Playwright specifications. Chromium could not start because the host does not provide `libatk-1.0.so.0`. The GitHub e2e jobs will verify these specifications. - GitHub Actions passed every final-head CI gate, including all three e2e shards and the aggregate `e2e` and `verify` jobs. - Greptile reviewed final commit `9af9200426` at 5/5 with zero review threads. ## Risks - Removing the name uniqueness index permits duplicate display names. Stable connection IDs and UIDs remain unique within a company. - The finish-app endpoint accepts the new review field as optional for backward compatibility. When clients send it, the server requires a complete decision for all quarantined actions. - Multiple OAuth connections depend on the explicit new-connection route flag. Focused tests cover active and draft connection reuse. - The migration is ordered after migration 0210. Its `DROP INDEX IF EXISTS` statement is safe to repeat. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex with the `gpt-5.6-sol` model assisted this change. The agent used repository tools, code execution, test execution, and agentic reasoning. The Codex runtime manages the context window. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b18b0fc39b |
feat: refine app connections and legacy worktree startup (#11040)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Apps UI manages app discovery and app connections. > - The managed worktree runtime starts agent work in repository worktrees. > - The Apps routes do not match the main discovery flow, and the connections view lacks a delete action. > - Legacy managed worktrees can also start before their pending seed operation runs. > - This pull request makes app discovery the main Apps route and makes connection management explicit. > - It also seeds legacy managed worktrees before runtime startup and makes the CLI read the repository-local config. > - The benefit is a clearer Apps workflow and a safer managed-worktree startup path. ## Linked Issues or Issue Description **What existing behavior does this improve?** This improves the Apps navigation, app connection management, managed git-worktree startup, and CLI worktree selection. **Subsystem affected** Cross-cutting. The change affects `ui/`, `server/`, `cli/`, and development documentation. **Current behavior** The `/apps` route opens the connections list while discovery uses a nested route. The connections list has no delete action. Some legacy managed worktrees can start runtime work before their pending seed operation runs. The CLI can also read an ambient Paperclip config instead of the repository-local config. **Proposed behavior** The `/apps` route opens Browse, and `/apps/connections` opens the connection list. Users can delete a connection after confirmation. Runtime startup seeds legacy managed worktrees when required. The CLI resolves the current worktree from the repository-local `.paperclip/config.json` file. **Reason and benefit** Users can discover apps from the canonical Apps route and can manage existing connections from a dedicated route. Legacy worktrees receive their required repository content before agent runtime starts. CLI worktree selection stays scoped to the current repository. **Breaking changes** The `/apps` and `/apps/browse` route behavior changes. Old Browse links redirect to `/apps`. The change does not modify an API schema or database schema. ## What Changed - Make Browse the canonical `/apps` page and move the connection list to `/apps/connections`. - Align Apps navigation, redirects, attention links, empty states, and connection actions with the new routes. - Add connection deletion with confirmation and clear failure feedback. - Seed legacy managed git worktrees before runtime startup when their seed status is pending. - Read the CLI worktree selection from the repository-local Paperclip config. - Update focused UI, server, CLI, and development documentation coverage. ## Verification - Ran 202 focused UI, server, and CLI tests. All tests passed. - Ran `pnpm -r typecheck`. All projects passed. - Ran `pnpm build`. All projects built successfully. - Ran `pnpm test:run`. The server and UI stages passed 7,168 tests. The CLI stage found one environment-sensitive secrets test because this workspace injects static AWS credentials. The isolated CLI file passed all 8 tests after those injected variables were unset. - Ran `pnpm check:token-gates`. It reports 12 existing color-token violations in the unchanged `PaperclipOrbit3D.tsx` file from the target branch. This pull request does not modify that file. - Ran focused regression coverage for repository-root CLI config resolution and connection deletion state. All tests and affected package typechecks passed. - Collected all 27 tests in the six changed Playwright specifications successfully. - GitHub Actions passed every latest-head CI gate, including all three e2e shards and the aggregate `e2e` and `verify` jobs. - Greptile reviewed the final commit at 5/5 with zero unresolved threads. ## Risks - Existing bookmarks for `/apps/browse` redirect to `/apps`. - Connection deletion changes visible connection state and requires user confirmation. - The legacy seed path runs only for managed git worktrees with pending seed state. Tests cover the startup condition. - The rebase preserves the target branch's direct OAuth policy for the Notion connection flow. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex with the GPT-5 model family assisted this change. The agent used reasoning, repository tools, code execution, and test execution. The runtime does not expose the exact model snapshot or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
11e56654f8 |
feat(ui): port onboarding flow from prototype; add cloud + local variants (#10786)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - First-run onboarding is the subsystem that turns a brand-new install
into a working company: it creates the company, its goal, a lead agent,
and that agent's first task
> - The existing `OnboardingWizard` carried all of that wiring
correctly, but its UI had drifted from the current design direction, and
a separate design prototype (`paperclip-onboard`) existed as a
standalone visual mock with no backend
> - Porting the prototype's *logic* would have thrown away working,
well-tested backend orchestration; leaving the two apart meant the
design never shipped
> - Separately, cloud and local (self-hosted) installs need meaningfully
different first runs — local has no sign-in and must let the user pick a
locally-installed CLI adapter — so a single linear wizard could not
serve both
> - This pull request rebuilds the presentational layer from the
prototype on top of the existing backend orchestration, and splits it
into two thin flow containers over a shared core
> - The benefit is that the shipped onboarding matches the intended
design, cloud and local can diverge without duplicating logic, and each
can later ship to a different app version while sharing one set of step
components
## Linked Issues or Issue Description
No existing issue — describing inline (feature request).
**What problem does this solve?**
Onboarding is the first thing a new user sees, and the shipped wizard
had drifted from the current design. In parallel, cloud and local
installs need different first-run paths: local has no hosted sign-in,
and its agent runs on a CLI adapter installed on the user's machine,
which the cloud path never has to ask about. There was no way to express
that difference without either forking the whole wizard or bolting
conditionals onto a single linear flow.
**Proposed solution**
Extract the onboarding step views and shell into a shared core, then
compose two thin flow containers (cloud and local) over it. Keep all
backend orchestration in the existing `useOnboardingFlow` hook so no
working logic is rewritten.
**Alternatives considered**
- *Single flow with a `variant` prop* — most DRY, but the two flows are
intended to ship on different app versions, and a shared file would have
to be split later anyway.
- *Two fully independent copies* — simplest per-flow, but every shared
refinement (spacing, motion, copy) would have to be made twice and would
drift.
## What Changed
- **Shared core** under `ui/src/components/onboarding/`:
`OnboardingScaffold` owns the full-screen shell and the single
`AnimatePresence` step crossfade, so both flows transition identically;
step views (Start / Company / Agent / Task), `FooterNav`, `AgentPreview`
and the motion constants are extracted for reuse.
- **`CloudOnboardingFlow`** — `start → company → agent → task`; mounted
in the real app via `OnboardingWizardVariant`. Behaviour matches the
retired wizard, including `previewMock` and the existing-company ("add
an agent") entry point.
- **`LocalOnboardingFlow`** — skips sign-in and adds an optional email
ask (with a privacy assurance), a local model/adapter step that hires
with `requireEnvProbe: true`, and a "star us on GitHub" interstitial
before completing. **Harness-only for now** — the real app still mounts
the cloud flow.
- **Deleted `OnboardingWizard.tsx`** (1,786 lines); updated its
Storybook stories and the `OnboardingWizardVariant` test to the new
components.
- **Orbiting 3D paperclip backdrop** behind the auth and welcome screens
(`three`), code-split so it only downloads on those screens; honours
`prefers-reduced-motion` and disposes its GL context on unmount.
- **`motion`** added for step transitions and the agent-capsule
choreography.
- Visual values routed through design tokens per `DESIGN.md`; `Stepper`
generalized to take a step total (backward compatible); `/design-guide`
page and the component index updated.
- **Standalone preview harness** (`ui/onboarding-preview.html`) with
`?flow=` and `?step=` for backend-free review, wired as a second Vite
rollup input.
- **Adapter env probe bound to the adapter it ran against.**
`hireLeadAgent` reused `adapterEnvResult` for any adapter, so when a
hire failed and the user picked a *different* local adapter and retried,
the previous adapter's verdict satisfied the `requireEnvProbe` guard
while the hire posted the new adapter's config — hiring it unprobed. The
cache is now keyed on the adapter type plus the exact config posted to
the test endpoint, the config is built once and shared by probe and
hire, a failed probe clears the cache, and `clearAdapterEnvResult()`
(called on adapter change) stops the step displaying a stale verdict.
Cloud is unaffected — it hires with `requireEnvProbe: false`. Reported
by Greptile.
- **E2E specs re-pointed at the new flow.** Four specs still drove the
deleted wizard (`onboarding`, `conference-room-typing-intro`,
`planning-mode-visual-verification`, `nux-phase4-screenshots`) and
failed with `element(s) not found` on `"Name your company"` /
`input[placeholder="Acme Corp"]`. Rather than repeat the new drive
sequence four times, `tests/e2e/onboarding-flow.ts` adds one driver per
step (`startCloudOnboarding`, `completeCompanyStep`,
`completeAgentStep`, `completeTaskStep`, `completeCloudOnboarding`) and
the specs import it, so the next flow change touches a single file. Two
now-dead `**/test-environment` route stubs went with it — the cloud flow
hires with `requireEnvProbe: false`, so that probe never fires.
## Verification
- `pnpm --filter @paperclipai/ui typecheck` — clean.
- `npx vitest run` over the onboarding suites
(`OnboardingWizardVariant`, `AgentCapsule`, `onboarding-launch`,
`onboarding-goal`, `onboarding-route`, `onboarding-adapter-config`) — 33
tests pass.
- `pnpm --filter @paperclipai/ui build` — succeeds; the three.js chunk
splits out separately (522 kB raw / 133 kB gzip) rather than entering
the main bundle.
- Both flows driven end-to-end in the preview harness in `previewMock`
(no database writes), plus the cloud flow rendered in the real
authenticated app at `/onboarding` to confirm the mount swap.
- The four re-pointed e2e specs pass locally against the new flow.
- New `ui/src/hooks/useOnboardingFlow.test.tsx` — 4 cases pinning the
adapter-probe cache (switch-adapter retry, cold path, explicit clear,
and the cloud flow's `requireEnvProbe: false`). Verified non-vacuous:
the switch-adapter case fails against the pre-fix code.
- Rebased onto current `master`; `pnpm-lock.yaml` is deliberately
**not** committed — `.github/workflows/pr.yml` regenerates it when a
manifest changes and shares it with downstream jobs as the `pr-lockfile`
artifact.
## Risks
- **Deleting `OnboardingWizard.tsx` is the one change that alters
existing app behaviour.** The cloud flow is intended to be
behaviour-equivalent, and its entry points are covered by the updated
`OnboardingWizardVariant` test, but this is the area to review most
closely.
- **Conflict risk with open PRs that touch the old wizard**: #9900,
#9501, #8982 and #6636 all modify
`ui/src/components/OnboardingWizard.tsx`, which this PR removes.
Whichever lands second will need its change re-applied to the new step
components. Flagging so ordering can be decided deliberately.
- **New dependencies**: `motion` and `three` (+ `@types/three`). `three`
is large, so it is lazily imported and code-split — it does not affect
the main bundle. Both are MIT.
- The **local flow is not reachable in the app** yet (harness/canary
only), so it carries no runtime risk today; wiring it up is a follow-up.
- The auth screens remain **presentational only** — they are not wired
to real auth, unchanged from before this PR.
- **Pre-existing, not introduced here:** `OnboardingWizardVariant`
renders outside `<Routes>` in `App.tsx`, so its `useParams()` never
resolves `:companyPrefix` and `/{prefix}/onboarding` opens the welcome
screen instead of jumping to the agent step. `master` has the identical
structure, so this PR faithfully ports existing behaviour; the working
"add an agent" entry is the launcher card behind the overlay, which is
what the screenshot spec drives. Worth a separate fix.
## Model Used
Claude Opus 5 (`claude-opus-5`) via Claude Code, with extended thinking
and tool use (repo search/edit, local test + build execution, and
browser-driven visual verification of the rendered flows). Portions of
the session also ran on `claude-opus-4-8` and `claude-fable-5`.
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [ ] I will address all Greptile and reviewer comments before
requesting merge
🤖 Generated with [Claude Code](https://claude.com/claude-code)
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
5858ccb981 |
feat: make in-app features cloud-aware (#10850)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Operators use the same board application in self-hosted and Paperclip Cloud deployments. > - A Cloud tenant contains one company, so an in-app company switch does not change the active Cloud stack. > - Cloud operators need the sidebar and company surfaces to use the signed-in user's stack portfolio. > - The server must derive Cloud identity and links from trusted instance context instead of client input. > - This pull request adds canonical Cloud context, a trusted stack portfolio proxy, and Cloud-aware navigation. > - The benefit is consistent stack switching on Cloud while self-hosted company behavior stays unchanged. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: server REST routes and the React board UI. **Problem or motivation** A Cloud-managed instance contains one company. The existing company switcher could only switch records inside that tenant. It could not move the operator to another Cloud stack. The existing header also gave long organization names too little width. **Proposed solution** Expose a canonical public Cloud context in health data. Add a trusted server proxy for the current user's stack portfolio. Use that data in the board UI to switch stacks with top-level navigation. Keep the existing company behavior on self-hosted instances. Move search into the navigation and keep long organization names inside the sidebar panel. **Alternatives considered** An in-app `/stacks` route was rejected because Cloud tenant hosts reserve that path and stack selection must wake or authenticate another tenant. Client-supplied user identity was rejected because the server can derive the trusted Cloud actor. **Roadmap alignment** This change advances the Cloud deployments milestone. It keeps the product local-first and Cloud-ready without changing the self-hosted mental model. ## What Changed - Added canonical Cloud instance context and public health metadata. - Added a Cloud-only stack portfolio proxy with trusted actor forwarding and per-user caching. - Prevented normal company creation on Cloud-managed instances. - Switched the sidebar and Companies page from company actions to stack actions on Cloud. - Added full-page stack navigation and Cloud create-stack links. - Moved search into the sidebar navigation so the organization name keeps more width. - Added truncation and hover recovery for long organization and stack names. - Added server and UI regression coverage for Cloud and self-hosted behavior. - Updated the implementation specification for the Cloud contracts. ## Verification - `node scripts/check-token-gates.mjs` passed. All three token gates are clean. - `pnpm --dir server exec vitest run src/__tests__/health.test.ts src/__tests__/cloud-instance.test.ts src/__tests__/cloud-routes.test.ts src/__tests__/company-cloud-floor.test.ts src/__tests__/company-portability-routes.test.ts` passed: 5 files and 66 tests. - `pnpm --dir ui exec vitest run src/components/SidebarCompanyMenu.test.tsx` passed: 1 file and 11 tests. - Pre-PR QA report `7da87ca7` passed all 8 acceptance criteria with real HTTP route factories and real Chromium screenshots in Cloud and self-hosted modes. - Security reviews passed for the canonical Cloud context and stack portfolio proxy. ## Risks - Cloud stack switching depends on the configured Cloud application and tenant portfolio URLs. - The new health `cloud` block is public by design, but it contains only canonical public instance metadata. - The stack proxy fails closed on self-hosted instances and derives the user identity from the trusted actor. - Self-hosted navigation and company creation retain their existing paths and behavior. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, model `gpt-5`. The run used reasoning, repository tools, shell execution, and GitHub integration. The deployment did not expose its context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f91a6e27c0 |
feat(issues): contain cross-issue agent side effects (#10837)
## Thinking Path > - Paperclip is the control plane that coordinates autonomous agent work. > - Agents need to collaborate on issues beyond their current assignment. > - Cross-issue comments and updates are useful, but an unbounded run can create cascading side effects. > - The control plane must preserve company-wide collaboration while containing each run's influence. > - Comment attribution must also show the responsible user and the acting agent in audits. > - This pull request adds run-bound cross-issue containment, attribution, and agent-class wake rules. > - The benefit is safer collaboration without restoring issue-assignee ownership restrictions. ## Linked Issues or Issue Description **What existing behavior does this improve?** Agent-authenticated issue comments, updates, reopen behavior, and assignee wake routing. **Subsystem affected** Cross-cutting: server routes and services, shared contracts, database schema and migration, and implementation documentation. **Current behavior** An authenticated agent can collaborate across company issues, but one heartbeat run has no per-run side-effect boundary. Comment records also do not persist the responsible user separately from the acting agent. **Proposed behavior** Require a valid heartbeat run for agent cross-issue comments and updates. Audit each attempt and cap a run at 20 cross-issue effects. Keep the cap in log-only mode until it automatically changes to enforcement at 2026-08-11 00:00 UTC. Preserve same-issue writes. Use agent-class wakes for agent comments. Keep same-run completion comments from reopening completed work. Record the responsible user on agent-authored comments and activity. **Reason and benefit** Agents can collaborate on other issues without an assignment gate, while each run has an atomic and inspectable side-effect limit. Operators can identify both the acting agent and the responsible user. **Breaking changes** After 2026-08-11 00:00 UTC, the twenty-first cross-issue comment or update from one heartbeat run returns a containment error. Agent cross-issue writes without valid run context are rejected. The migration is additive and backfills existing agent-authored comment attribution where the source data is available. ## What Changed - Added an atomic per-run counter for cross-issue agent comments and updates. - Added audit events for allowed and rejected cross-issue effects. - Added the automatic log-only to enforcement flip at 2026-08-11 00:00 UTC. - Added responsible-user attribution to agent-authored comments, activity records, shared types, and validators. - Added an additive migration and migration coverage for existing comments. - Updated reopen, resume, and wake behavior so agent comments create agent-class wakes and same-run completion comments remain inert. - Updated the implementation specification and regression coverage. ## Verification - `pnpm exec vitest run server/src/__tests__/cross-issue-influence-limit.test.ts server/src/__tests__/issue-comment-attribution-audit-routes.test.ts server/src/__tests__/issue-comment-reopen-routes.test.ts packages/db/src/issue-comment-on-behalf-migration.test.ts` — 97 tests passed. - `pnpm -r typecheck` — passed, including migration safety checks. - `pnpm test:run` — server batch: 3,364 passed and 2 skipped; UI batch: 3,504 passed. One unrelated CLI doctor test warned because this agent runtime injects static AWS credentials. - `env -u AWS_ACCESS_KEY_ID -u AWS_SECRET_ACCESS_KEY pnpm exec vitest run cli/src/__tests__/secrets.test.ts` — 8 tests passed and confirmed the CLI failure was ambient-environment sensitive. - `pnpm build` — passed. ## Risks - The fixed enforcement timestamp changes production behavior automatically on 2026-08-11 00:00 UTC. Audit logs before that time provide rollout visibility. - The per-run counter serializes on the heartbeat-run row. This prevents concurrent attempts from racing past the cap but adds a small lock scope for cross-issue writes. - Existing comments can only be backfilled when their acting run or agent attribution is recoverable. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-5 in the Codex agent runtime. The runtime did not expose a context-window size. Reasoning, shell tools, code editing, and test execution were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
42d0ddcb86 |
test(e2e): deflake applications Connections list against the health sweep (#10763)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The applications page shows connected services and their status > - The e2e suite covers the Connections row on that page > - The server runs a periodic connection health sweep during the test > - The sweep can change the connected row label and action after the first render > - This pull request accepts both connected states on that row > - The benefit is the test checks the real user path without a sweep race ## Linked Issues or Issue Description No public issue exists. This pull request fixes a race in the applications Connections e2e test. **Bug:** The test pinned the exact connected-state pill and action. **Expected:** The test should accept both connected states for the same connection row. **Impact:** The health sweep can change the label between assertions. **Fix:** The test now accepts either label and action on the connected row. ## What Changed - Allowed the connected row pill to match `Healthy` or `Needs attention`. - Allowed the connected row action to match `Open` or `Reconnect`. - Kept the not-connected row exact. ## Verification - `git diff --check origin/master..origin/test/deflake-applications-crud-health-sweep` - `git show --stat --summary --oneline 9785785a5d41cd13ee5a0f8aeb73f389cdbdac2e` - Local Playwright e2e did not run in this worktree. - CI on this pull request should provide the full proof. ## Risks - Low risk. The change only widens the expected labels for the connected row. - If the UI adds a new state, the test may need another update. ## Model Used OpenAI GPT-5, tool use, context window not exposed in this shell session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
9c1f8e7887 |
feat(decisions): add first-class propose mode (#10010)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents can currently perform many mutations directly, while humans often need a durable review point before cross-issue or destructive actions occur > - Existing approvals and issue-thread interactions do not provide a standalone, reusable object for presenting options, collecting typed inputs, detecting stale targets, and auditing effect execution > - The control plane therefore needs a first-class propose mode that separates an agent's recommendation from the governed mutation it may cause > - This pull request adds Decisions v1 across the database, shared contracts, server execution and telemetry, agent skill guidance, and operator UI > - The benefit is that agents can propose multi-option actions safely while operators get explicit provenance, fail-closed execution, per-effect results, and a focused attention workflow ## Linked Issues or Issue Description ### Subsystem affected Cross-cutting: `packages/db`, `packages/shared`, `server`, and `ui`. ### Problem or motivation Agents need a governed way to propose consequential work without immediately mutating issues, especially when one choice can affect several issue trees. Existing approvals and issue-thread interactions do not provide a standalone object with typed options, target snapshots, effect-level authorization, expiration, execution outcomes, and reusable attention-feed presentation. ### Proposed solution Add first-class Decisions that store options and typed inputs, surface open proposals in the operator attention feed, validate target freshness and the origin-agent/operator authorization intersection at decision time, execute a bounded set of auditable effects, and retain terminal outcomes. Decisions v1 supports comments, status and assignee changes, follow-up issue creation, blocker resolution, and issue-tree cancellation, plus bundle grouping, expiration/dismissal, rule-key telemetry, and agent-facing API guidance. ### Alternatives considered - Extend approvals with arbitrary effects: rejected because approvals represent governed yes/no actions and would become an unsafe generic mutation envelope. - Model every proposal as an issue-thread interaction: rejected because decisions can span several targets and need independent lifecycle, telemetry, idempotency, and effect results. - Let agents perform the mutation and ask for retrospective review: rejected because it removes the pre-execution governance boundary this feature is meant to provide. ### Roadmap alignment Aligns with `ROADMAP.md` sections **Agent Reviews and Approvals**, **Enforced Outcomes**, **MCP Tool Gateway & Apps (governed tool access)**, and **Activity History** by making explicit decisions, authorization gates, auditable execution, and terminal outcomes first-class control-plane objects. ### Additional context This does not replace existing approvals or issue-thread interactions, and it does not add an unrestricted generic mutation effect. ## What Changed - Added company-scoped decision, option, target, and effect-execution schema plus migration and shared TypeScript/Zod contracts. - Added decision routes and services for propose, list/get, decide, dismiss, cancel, target freshness checks, authorization intersection, idempotency, activity logging, and execution auditing. - Added rule-key decision telemetry and attention-feed metadata so open decisions are visible and measurable. - Added agent skill documentation for proposing and resolving decisions through the Paperclip API. - Added the Decisions UI: API client, query keys, inline attention resolver, bundle grouping, target-issue strip, terminal history, destructive confirmation, and per-effect result rendering. - Added server service coverage, DecisionCard state tests, and Storybook stories for the supported visual states. ## Verification - `pnpm -r typecheck` — passed. - `pnpm test:run` — 2,876 passed, 1 skipped, with one unrelated cross-suite cleanup-order failure in `heartbeat-responsible-user-invariant.test.ts`; the failing file passes in isolation (`6/6`). - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/heartbeat-responsible-user-invariant.test.ts` — passed. - `pnpm --filter @paperclipai/ui exec vitest run src/components/DecisionCard.test.tsx` — passed (`9/9`). - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/authz-existence-oracle-guard.test.ts src/__tests__/openapi-routes.test.ts` — passed (`5/5`). - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/decisions-service.test.ts` — passed (`16/16`). - `pnpm --filter paperclipai exec vitest run src/__tests__/company-import-export-e2e.test.ts` — passed (`1/1`). - `pnpm --filter @paperclipai/server typecheck` and `pnpm --filter paperclipai typecheck` — passed. - `pnpm build` — passed. - Rebased-head focused suite — passed (`6` files, `88` tests): shared decision contracts, Decisions service, OpenAPI routes, startup feedback export, DecisionCard states, and attention helpers. The follow-up stale-secondary-target regression passes in the DecisionCard suite (`10/10`). - Rebased-head scoped typechecks — passed for `@paperclipai/shared`, `@paperclipai/db`, `@paperclipai/server`, and `@paperclipai/ui`. - Rebased-head migration numbering and safety checks — passed after renumbering the additive migration to `0193` and making it replay-safe for environments that applied the earlier feature-branch number. - `pnpm check:token-gates` — passed with all gates clean. - GitHub PR workflow and Greptile review for `1f9f7645882d05dfdd9c99377c03a1f53f20e8be` — running after the stale-secondary-target fix and PR metadata refresh on July 27, 2026. - `pnpm --filter @paperclipai/ui build-storybook` exposes an existing Storybook version mismatch (`storybook` 10.4.6 vs `@storybook/addon-docs` 10.5.0); Decisions stories were validated with the docs addon temporarily disabled and the tracked config remains unchanged. ## Risks - **Migration:** Adds replay-safe migration `0193`; migration numbering and safety checks pass. The new tables and indexes are additive. - **Authorization:** Effect execution intersects the proposing agent's permissions with the responsible user context and fails closed; mistakes could reject a valid proposal rather than silently over-authorize it. - **Concurrency:** Target snapshots and idempotency keys protect against stale or duplicate execution, but reviewers should focus on mixed-effect partial outcomes and retry behavior. - **UI:** Decisions are integrated into the existing attention feed rather than a separate navigation surface, reducing routing risk but increasing the importance of attention-item metadata compatibility. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex CLI using `gpt-5.6-sol` for final PR preparation, review fixes, and verification; repository tools and code execution were enabled, and context-window size is not exposed in this runtime. - Anthropic Claude Opus 4.8 with 1M context assisted with the Decisions UI implementation, as recorded in the relevant commits. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
91c79d28cb |
test(e2e): retry run-lock 409 in signoff-policy agentPatch (#10386)
## Thinking Path > - Paperclip is the control plane people use to manage AI-agent work across companies > - This repository's e2e suite verifies the execution and approval paths that keep the control plane reliable > - The signoff-policy flow uses a helper that runs a heartbeat and then PATCHes the issue with that run id > - That PATCH can race the heartbeat's run-lock ownership and intermittently receive a transient 409 > - When that happens, a single-positive transition fails the shard even though the underlying behavior is only a lock contention race > - This pull request makes the helper retry the 409 path by re-reading the current lock and re-PATCHing under the winning run id > - The benefit is a stable e2e shard without weakening the negative-path assertions that protect the contract ## Linked Issues or Issue Description No public GitHub issue exists for this change. This PR addresses a flaky signoff-policy e2e transition where the helper can lose a run-lock race and receive a transient 409 while the issue is still assigned to the acting agent. ## What Changed - Added bounded retry/backoff handling in the signoff-policy `agentPatch` helper for transient run-lock 409 responses. - Re-read the issue's current lock before retrying so the helper can re-PATCH with the winning run id. - Kept the retry guarded so non-participant rejection and missing-comment 400s still surface unchanged. - Preserved the existing positive-path behavior without adding Playwright retries or weakening assertions. ## Verification - Reviewed the diff shape for a single-file change in `tests/e2e/signoff-policy.spec.ts`. - Verified the pushed commit matches the authorized submit SHA from the handoff. - Confirmed the branch contains only the expected commit and no unrelated history. - The handoff notes record deterministic harness evidence showing the pre-fix helper fails on the 409 race and the post-fix path passes. ## Risks - Low functional risk: the retry is narrowly scoped to the transient 409 lock-contention path. - If the lock semantics change server-side, the helper may need a follow-up adjustment. - The change only affects the e2e helper and does not alter production API behavior. ## Model Used OpenAI Codex (GPT-5, tool-using coding agent) ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
c8c2ae82a3 |
test(e2e): wait for connections fetch before asserting applications status pill (#10384)
## Thinking Path > - Paperclip relies on end-to-end tests to catch regressions in the operator UI > - The applications list status pill is rendered from more than one data fetch, so the visible label can lag behind the row itself > - A default Playwright assertion timeout can expire before the connections fetch finishes under CI load > - That creates a flaky failure without changing the underlying product behavior > - This pull request extends the two status-pill assertions to use the longer timeout already used elsewhere in the spec > - The benefit is the same UI coverage with fewer false negatives ## Linked Issues or Issue Description There is no public GitHub issue linked to this repo-local fix. This PR addresses a flaky Playwright assertion in `tests/e2e/applications-crud.spec.ts`, where the status pill can render after the default assertion timeout because it depends on the connections fetch as well as the applications fetch. ## What Changed - Increased the wait window for the two status-pill visibility assertions in `tests/e2e/applications-crud.spec.ts`. - Kept the assertions anchored to the exact expected labels (`Healthy` and `Not connected`) so the test still verifies the same behavior. ## Verification - The spec is still discovered cleanly by Playwright. - The diff stays limited to `tests/e2e/applications-crud.spec.ts`; there are no product, dependency, or lockfile changes. - The full browser-backed e2e signal remains CI for this change. ## Risks - Low risk. This only changes assertion timing in a test file and does not alter runtime product behavior. ## Model Used OpenAI Codex, GPT-5, tool-using coding agent. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7e00f67138 |
feat(connections): add v3 schema core (#9958)
## Thinking Path > - Paperclip is the open source control plane people use to manage AI-agent companies and their governed access to external systems. > - Connected Apps build on the existing Apps and MCP gateway substrate so companies can configure reusable, auditable integrations. > - The current connection record does not yet have a stable public address, explicit ownership/auth method fields, or subject-specific credential grants. > - Without that schema core, later OAuth, per-user authorization, token brokering, triggers, and connector-service phases cannot enforce tenant and subject boundaries consistently. > - This pull request adds the forward-compatible Connections v3 schema core while preserving the existing connection lifecycle and directly migrating the remote MCP transport name. > - The benefit is a company-scoped, least-privilege foundation for one-click integrations without bypassing Paperclip secrets, profiles, rules, or audit controls. ## Linked Issues or Issue Description No matching public issue was found. **Problem** Paperclip's current app connections need a durable identity and authorization substrate before Connected Apps can safely support multiple setup methods, per-user credentials, provider tenants, and managed connector services. The existing schema only models a single connection-level credential set and uses legacy transport terminology. **Proposed solution** Add a stable company-scoped connection UID, explicit ownership/auth/transport fields, a subject-aware `connection_grants` table, and multi-key credential annotations. Backfill existing connections and workspace grants in a reversible migration, then update shared/server/UI contracts to the new `mcp_remote` transport name. **Related work** - Related foundation: #9534 - Roadmap: Connected Apps (one-click integrations) ## What Changed - Added company-scoped connection `uid`, `ownership`, `authKind`, and canonical transport fields across database, shared contracts, validators, services, and UI fixtures. - Added `connection_grants` with workspace/user subject rules, provider tenant metadata, credential secret refs, revocation state, company scoping, and uniqueness constraints. - Added migration `0182_connections_v3_schema_core` to backfill stable UIDs, rename `remote_http` to `mcp_remote`, infer auth kinds, create default workspace grants, and support rollback coverage. - Added multi-key credential annotations and updated gateway/access services without changing the existing lifecycle behavior. - Updated the connection glossary, connector playbook, and security threat model for the new identity, grant, and relay boundaries. - Added explicit test UIDs to direct database fixtures so the new non-null invariant is exercised across affected server suites. ## Verification - `pnpm --filter @paperclipai/shared typecheck` - `pnpm --filter @paperclipai/db typecheck` - `pnpm --filter @paperclipai/server typecheck` - `pnpm --filter @paperclipai/ui typecheck` - `pnpm exec vitest run server/src/__tests__/tool-access-service.test.ts server/src/__tests__/tool-gateway-service.test.ts server/src/__tests__/tool-gateway.test.ts server/src/__tests__/heartbeat-runtime-skills.test.ts server/src/__tests__/tool-oauth-legacy-backfill.test.ts server/src/__tests__/tool-access-policy-service.test.ts server/src/__tests__/heartbeat-runtime-mcp-servers.test.ts packages/db/src/connections-v3-schema-core-migration.test.ts packages/shared/src/validators/tool-access.test.ts --config vitest.config.ts` — 9 files, 218 tests passed. - Latest-head GitHub Actions: build, typecheck, general/serialized suites, backup/worktree restore coverage, both e2e shards, canary, policy, and security scans pass. - Greptile: 5/5 with zero unresolved threads. - `pnpm check:token-gates` remains red only on five pre-existing `#9627` color literals outside this change. ## Risks - **Migration risk:** UID backfill and default-grant creation touch every existing connection. The migration uses company-scoped uniqueness, deterministic legacy UIDs with ID suffixes, and seeded up/rollback coverage. - **Authorization risk:** Grant rows carry credential references. Constraints enforce workspace-vs-user subject shape, company/connection lookup indexes, one default grant per connection, and one user grant per connection/subject. Security review is requested specifically for this design. - **Compatibility risk:** `remote_http` is renamed directly to `mcp_remote`; all repository call sites and fixtures are updated in the same change. - **Future-phase risk:** Subject-bound token issuance, triggers, and connector-service relay verification remain fail-closed requirements documented for later phases; this PR does not expose those capabilities. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex CLI coding agent. The runtime did not expose an exact underlying model ID or context-window size; capabilities used include repository inspection, code editing, shell execution, test execution, Git/GitHub CLI operations, and structured reasoning. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3db2e6bdd2 |
feat(mcp) [split 8/8]: add e2e coverage and operator docs (#9563)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Governed MCP access spans contracts, runtime enforcement, adapters, UI surfaces, and operator verification > - The parity reference PR #9534 is too large for effective automated or human review > - The feature therefore needs a linear stack whose individual diffs stay below the 100-file review limit > - This pull request is split 8/8 and focuses on end-to-end coverage, operator docs, evals, and release notes > - The benefit is a standalone, testable review boundary while preserving byte-for-byte parity at the top of the stack ## Linked Issues or Issue Description - Related parity reference: #9534 - Problem: The complete stack needs discoverable browser scenarios, operator guidance, threat modeling, eval coverage, and a parity proof before merge. - Proposed solution: Adds MCP user-story and Smoke Lab e2e suites, docs/evals/release notes, the skill update, and the root e2e driver script registration. - Alternatives considered: keeping #9534 as one 403-file review, or rewriting the feature to manufacture seams; both were rejected in favor of path extraction plus compile-driven boundary moves. - Roadmap alignment: this advances the existing governed MCP/tool-access work already represented by #9534; it does not introduce a separate roadmap initiative. - Stack position: base branch is `pap10341-split/07-ui-apps-activation`. - Merge policy: merge bottom-up, in order, only after the complete eight-PR stack has been reviewed and the top-of-stack parity gate remains empty. - Requested review: QA for flag audit and e2e/browser acceptance; Greptile on every PR. ## What Changed - Adds MCP user-story and Smoke Lab e2e suites, docs/evals/release notes, the skill update, and the root e2e driver script registration. - Keeps this PR below 100 changed files and independently typecheckable. - Preserves the final tree from #9534 when combined with the other seven stack levels. ## Verification - `pnpm typecheck` - `node --check scripts/e2e-mcp-user-stories.mjs` - `pnpm exec playwright test --config tests/e2e/playwright.config.ts --list` — 43 tests discovered - `git diff pap10341-split/08-e2e-docs 6b40e3876d9297105d4ec306e47e46d351c86172` — empty (0 bytes) ## Risks - Browser suites depend on runtime services and environment setup; this PR validates discovery locally while QA owns full flag-on/flag-off execution. - Stack risk: merging out of order can expose incomplete layers; mitigate by following the documented bottom-up merge policy. - Parity risk: later edits to an intermediate branch can drift from #9534; mitigate by re-running the empty top-of-stack diff before merge. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, exact model ID `gpt-5.4`; runtime-managed context window; medium reasoning with repository, shell, Git, GitHub CLI, and code-execution tools enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] Internal references are omitted except the execution-plan link explicitly required for this coordinated split stack - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge ## Stack Coordination - Internal execution plan: [PAP-13874](/PAP/issues/PAP-13874#document-plan) - Parity reference: #9534 - Stack: #9556 → #9557 → #9558 → #9559 → #9560 → #9561 → #9562 → #9563 - Merge bottom-up only after full-stack review and an empty parity diff at #9563. --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b13eb5b2b5 |
Skill Studio: three-pane skill IDE with sandboxed test runs (#9241)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Skills Manager gives operators a reusable skill layer, but iteration still required manual edits, ad hoc prompts, and indirect run inspection. > - Skill authors need a focused workflow for editing skill files, saving representative test inputs, and running those inputs through an agent without exposing harness tasks as normal company work. > - The backend therefore needs durable test inputs, reusable run templates, hidden harness issues, scoped run execution, retention metadata, and read-containment rules around hidden work. > - The frontend needs a three-pane Studio that keeps skill files, saved inputs/templates, and run output/history visible together while preserving the existing design system and token rules. > - This pull request ships that Skill Studio surface end to end: database migrations, shared contracts, server APIs/services, hidden harness execution behavior, UI routes/components, and focused tests. > - The benefit is faster and safer skill iteration, with inspectable outputs and fewer ways for internal harness work to leak into normal task lists, costs, or adjacent read APIs. ## Linked Issues or Issue Description No public GitHub issue exists for this feature. Feature request summary: - Problem: Skill authors need to edit and test company skills in one place instead of switching between the skill detail page, task creation, run output, and manual prompt history. - Proposed solution: Add a Skill Studio workbench with saved inputs, reusable templates, hidden sandboxed test runs, live run status, output inspection, run history, rerun/delete controls, and frontmatter-aware editing. - Expected users: Paperclip operators and agent-company maintainers who create, fork, import, and tune skills. - Related public PRs: Supersedes #9205, which was replaced so the public PR branch name follows contributor policy. - Duplicate search: searched public GitHub issues and PRs for "Skill Studio"; no other active public issue or PR directly covers this feature. ## What Changed - Added database migrations for Skill Studio test inputs, test runs, test run retention, and reusable run templates. - Added shared Skill Studio types, validators, route helpers, frontmatter utilities, and status handling. - Added server services and routes for saved inputs, test runs, templates, reruns, terminal-run deletion, hidden harness issue execution, and run-detail hydration. - Strengthened hidden-issue read containment across issue-adjacent routes and cost rollups used by skill test harness work. - Added the Skill Studio UI with skill file editing, frontmatter editing, saved inputs, templates, run creation/cancel/rerun/delete flows, output rendering, history, route support, and responsive pane behavior. - Added focused backend, shared, and UI tests for the new APIs, routing logic, editor/run behavior, hidden-issue containment, and migration safety. - Rebased onto current `master`, removed the generated lockfile diff from the PR, and verified no workflow files are changed. ## Verification - [x] `pnpm --filter @paperclipai/db check:migrations` - [x] `pnpm check:token-gates` - [x] `pnpm exec vitest run server/src/__tests__/company-skills-service.test.ts server/src/__tests__/company-skills-routes.test.ts server/src/__tests__/company-skill-test-runs-service.test.ts ui/src/lib/skill-studio.test.ts ui/src/pages/SkillStudio.test.tsx` — 5 files, 132 tests passed - [x] Greptile review on the latest PR head - [x] GitHub PR checks on the latest PR head ## Risks - Medium risk because this is a broad feature touching database schema, server orchestration, issue visibility, and a large UI surface. - Hidden harness issue containment is security-sensitive; this PR includes regression coverage for adjacent read paths and cost rollups. - The new migrations are additive and use idempotent guards where applicable, but deployed databases that previously tested draft migration numbers should still be checked carefully. - The UI depends on a new resizable panels package in `ui/package.json`; the lockfile is intentionally left to repository automation. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5 coding agent with shell, git, and GitHub CLI tool use. Earlier feature commits include assistance from other Paperclip coding agents; this PR preparation, rebase, cleanup commit, and PR body were completed by OpenAI Codex in a Paperclip worktree. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
c90e66bdd6 |
feat(ui): design-system component convergence — Card/Badge adoption, multiplicative radius ladder, unified list surfaces (#9240)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Its UI is governed by a design system (`DESIGN.md` + the token layer in `ui/src/index.css`, merged in #9134) whose first principle is "one way to say each thing" — one Card, one Badge, one nav row > - After the token extraction landed, ~35 files still hand-rolled card containers, ~45 files hand-rolled pill spans, the sidebar agents section duplicated the nav-row chrome, and the inbox and tasks lists rendered the same task rows two subtly different ways > - Each divergence is a place where a future design change (radius, hover language, status vocabulary) silently misses surfaces, defeating the "edit tokens + run checks" model the design system exists for > - This pull request converges those surfaces onto the shared primitives, codifies the radius scale as the modern multiplicative shadcn ladder, and unifies the row/hover/tree-guide language across the inbox and tasks lists — every visible delta was human-reviewed screen-by-screen against a live instance across nine feedback rounds > - The benefit is that the app's look is now steerable from single knobs (one `--radius` anchor, one Card, one Badge, one row component), and the visual regression suite covers the result (514 snapshots including a new AgentDetail page story) ## Linked Issues or Issue Description No public GitHub issue exists for this work; describing per the feature template: - **Problem**: after the design-token foundation (#9134), component-level drift remained — hand-rolled cards/pills, duplicated sidebar row chrome, and two different renderings of task rows (inbox vs tasks list) meant design changes had to be applied per-surface and frequently missed spots (e.g. status glyphs rendered 16px in the inbox but 20px in the tasks list because a slot override silently beat the component default). - **Proposed behavior**: all card-shaped containers render via `Card`, all label pills via `Badge`, sidebar rows via `SidebarNavItem`, and both task-list surfaces via one `IssueRow` configuration; the radius scale is a single multiplicative ladder anchored at `--radius: 0.5rem`. - **Alternatives considered**: converting interactive `<button>`/`<Link>` cards to `Card` divs (rejected — breaks semantics; documented inline with `design-allow` comments instead); keeping the legacy additive radius ladder (rejected in favor of the standard shadcn multiplicative mapping). ## What Changed - `Card` adoption across ~35 files (settings pages, auth/board flows, dashboards, list containers, KPI tiles); non-adoptable sites (interactive cards, `<li>` rows, class-string props, chart tooltip) carry documented `design-allow(card-pattern)` comments - `Card` gains an `interactive` prop — one quiet hover affordance for clickable cards (cursor, border darken, shadow lift, focus ring), applied to skills tiles, artifact cards, and the company selector; cards carry no resting shadow - `Badge` adoption for 113 hand-rolled pill spans across ~45 files; `PropertyChip` wraps `Badge` internally; `StatusBadge`, external-object chips, and match chips stay bespoke by documented decision (WCAG-tuned status mechanics) - Radius ladder becomes the multiplicative shadcn mapping (`sm/md/lg/xl/2xl/3xl/4xl = 0.6/0.8/1.0/1.4/1.8/2.2/2.6 × --radius`, anchor `0.5rem`); every card surface unifies on `rounded-lg`; the orphaned 8px literal token is deleted - Sidebar: agent rows render via `SidebarNavItem` (new additive props: `iconNode`, `active`, `trailing`, `liveAccessory`); live dots use `--status-agent-running`; one row rhythm and inset rounded pill highlight; right-aligned trailing badges; every labeled section is collapsible - Inbox + tasks lists unified: md status glyphs, `accent/50` rounded row hovers, vertical tree guides under parent rows (opaque underlay so dark-mode translucent borders don't stack), no horizontal dividers under expanded parents; the swipe-to-archive reveal layer shows only mid-swipe; board toggle uses the `SquareKanban` glyph - Kanban: every column carries a status-hued tint; lanes default expanded (including empty); compact mode collapses empty lanes to labeled rails (fixes a clipped, label-less empty-column state) - Storybook: new AgentDetail page story (realistic fixtures, light+dark) joins the visual suite; suite captures with `reducedMotion: 'reduce'` and the ux-lab reasoning ticker honors `prefers-reduced-motion`; a stale lexical alias in `storybook/main.ts` is fixed (build was broken since the lexical 0.46 bump) - Keyboard navigation, from live review of the unified lists: inbox navigation keys work on every tab (archive keys stay scoped to the archivable tab); keyboard-driven scrolling no longer hands the selection to whatever row lands under the stationary cursor (hover selects only after real pointer movement); the tasks list view gains the same j/k / arrows / Enter selection model as the inbox; and the `g` then `i` go-to-inbox chord works app-wide instead of only on the issue detail page - Token gates restored to 3/3 CLEAN (tokenized a post-#9134 regression in the recovery card); decisions recorded in `doc/design/DECISION-SHEET.md` and `doc/design/COMPONENT-INVENTORY.md` (investigation verdicts: FileTree vs WorkspaceFileBrowser and the four entity pickers stay separate — evidence included) ## Verification - `pnpm check:token-gates` → 3/3 CLEAN - `pnpm typecheck` → green (all packages) - `cd ui && npx vitest run` → 2106/2106 (assertions updated in lockstep where they documented superseded decisions; new tests for the global go-to-inbox chord) - `pnpm --filter @paperclipai/ui build` → exit 0 - `pnpm build-storybook` → succeeds (also fixes the lexical-alias break on master) - Visual regression: 514-snapshot Playwright suite green against the updated baseline (zero diffs from the keyboard-navigation round — those changes are purely behavioral). Note: baselines live outside git per the suite design and the baseline-manifest archive is not yet published, so CI cannot run this suite — it was run locally throughout; every visible delta was reviewed screen-by-screen in a live instance across nine review rounds. Review evidence (before/after triplets) intentionally kept out of the repo for size; available on request. - Manual: exercised dashboard, tasks (list + board), inbox, agents, skills, costs, settings, and artifact surfaces in light and dark themes ## Risks - Wide but shallow visual surface: most changes are class-string substitutions with behavior preserved (props, handlers, roles, test ids). The riskiest areas — dnd-kit card refs (React 19 ref-as-prop), inbox swipe-to-archive, and sidebar overlays — are covered by existing unit tests (all green) and were manually exercised. - Intentional visual deltas (rounded cards, tinted kanban columns, md status glyphs, unified hovers) are design decisions recorded in `doc/design/DECISION-SHEET.md`; each maps to a re-baselined snapshot set locally. - The visual suite cannot yet run in CI (unpublished baseline archive — pre-existing condition from #9134); until that lands, snapshot coverage is local-only. ## Model Used - Claude Fable 5 (`claude-fable-5`, Anthropic) via Claude Code — agentic coding with tool use (file editing, test execution, Playwright screenshot verification); extended thinking enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (will confirm once CI runs) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending first review) - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
d3919713bc |
[codex] Document Storybook visual baseline platform lock (#9216)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Storybook visual baselines protect UI surfaces from unintended visual drift. > - Pixel-perfect screenshot baselines are sensitive to OS, font rasterization, and browser environment. > - The suite already stores external baseline artifacts and has an opt-in CI path. > - Local runs on non-matching platforms can report false-positive diffs unless the platform lock is explicit. > - This pull request documents the Linux/Ubuntu baseline constraint and makes the local static server port explicit. > - The benefit is clearer visual-review guidance and more predictable Playwright web server startup. ## Linked Issues or Issue Description No public GitHub issue exists. ### What happened? The Storybook visual baseline suite requires a matching Linux capture environment for pixel-exact comparisons, but the docs did not clearly warn local users that non-Linux environments can produce false-positive diffs. The Playwright web server command also relied on the static server's default port instead of passing the configured port explicitly. ### Expected behavior Developers should see clear Linux/Ubuntu baseline guidance before running the visual suite locally, and Playwright should start the Storybook static server on the same explicit port that the test config expects. ### Steps to reproduce 1. Review the Storybook visual docs before this PR. 2. Run or inspect the Storybook visual Playwright config. 3. Notice the missing platform guidance and implicit static server port coupling. ### Paperclip version or commit Reproducible on `master` before this branch. ### Deployment mode Local dev (pnpm dev) / built from source. ## What Changed - Documents the Linux/Ubuntu-only baseline limitation in the developer docs and visual-suite README. - Adds `--port` parsing and validation to the Storybook static server helper. - Adds regression coverage for `--port` followed by another flag. - Passes the Playwright web server port explicitly from the Storybook visual config. ## Verification - Passed: `node --check scripts/serve-storybook-static.mjs` - Passed: `node --test scripts/__tests__/serve-storybook-static.test.mjs` - Passed: `node --test scripts/__tests__/storybook-visual-baseline.test.mjs` - Greptile: 5/5 with no unresolved review threads after commit `94a649755a2ae7c4a34a3e8a1f16ec4d26d738fd`. - Not run: full `pnpm test:storybook-visual`, because it builds Storybook and runs the browser visual suite; this PR only changes docs plus server port plumbing. ## Risks Low risk. The server still defaults to port 6106 when no explicit port is provided, and invalid port values now fail fast with a clear error before the Playwright server waits for an unreachable URL. ## Model Used OpenAI GPT-5 Codex coding agent with local command execution and repository editing tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
c07e650cd7 |
feat(ui): single-source design tokens, visual regression suite, and theme retune (#9134)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Its UI is the operator's daily surface: task lists, boards, budgets, agent status — all built on shadcn components and Tailwind > - Visual values (colors, spacing, type sizes, radii) were hardcoded at ~1,600 call sites: the same "small gray label" was 9/10/11px depending on the file, charts disagreed with chips about status colors, two toggle-switch implementations coexisted in two greens, and there was no visual regression coverage > - This made the UI drift-prone and made any restyle a hundreds-of-files project, which discourages design iteration > - This pull request extracts visual values into a single token layer in `ui/src/index.css`, adds a Storybook visual regression suite backed by external immutable baseline archives, and then applies a deliberate retune reviewed change-by-change on screenshot diffs > - The benefit is that Paperclip's look becomes a config surface: retheming is a token edit reviewed as a snapshot diff, drift is blocked by a token gate, and future UI PRs can prove exactly what changed visually without committing hundreds of PNGs ## Linked Issues or Issue Description No existing public issue covers this work (searched "design tokens", "visual regression", "design system" across issues and PRs). Related in spirit: Refs #8982 (theming a hardcoded panel — a one-off instance of the same problem class this PR addresses systematically). **Problem (feature-request form):** UI visual values are hardcoded per call site with no source of truth and no regression coverage; consistency depends on reviewer memory, and restyling requires mass file edits. **Proposed solution (this PR):** a single token layer + enforcement gate + externally stored visual snapshot suite, then an intentional restyle on top of that foundation. ## What Changed - **Token extraction (zero visual change, machine-verified during development):** committed codemods (`scripts/codemod-*.mjs`) moved ~1,600 hardcoded color/type/spacing/radius/shadow/misc values into named tokens in a non-inline `:root` block of `ui/src/index.css`. - **Visual regression suite:** `pnpm test:storybook-visual` covers 255 stories × light/dark = 510 Playwright screenshots at `maxDiffPixels: 0`, plus new primitive-coverage stories and deterministic-render fixes. - **External visual baselines:** committed PNG snapshots were removed. `tests/storybook-visual/baseline-manifest.json` pins an immutable archive URL/hash/size/count, and `scripts/storybook-visual-baseline.mjs` handles `download`, `verify`, `pack`, and trusted maintainer `upload` flows. - **Opt-in visual CI artifacts:** added a `Storybook Visual` workflow that runs on manual dispatch or PRs labeled `storybook-visual`, downloads/verifies the baseline, runs Playwright, and uploads Playwright report/test-result artifacts for review. Normal PR runs do not mutate baseline objects. - **Token gate:** `pnpm check:token-gates` — zero hex literals, zero arbitrary bracket values, zero raw font-sizes in `ui/src/components/**` and `ui/src/pages/**`, with a documented inline allowlist for legitimate opt-outs. - **Theme retune (intentional, snapshot-reviewed):** new base theme values; radius ladder derived from a single `--radius` knob; micro-type cluster collapsed to a named ladder (`--text-nano/micro/compact` + Tailwind `text-xs`/`text-sm`); letter-spacing collapsed to named steps. - **One status-color vocabulary:** charts, quota/budget bar fills, RUNNING/live chips, and liveness indicators all use the canonical `--status-*` hues. Light-mode legibility fixes for red alert surfaces that used dark-tuned text classes. - **One switch:** `ToggleSwitch` restyled to the registry capsule form, second hand-rolled implementation removed, and all call sites unified. - **Docs:** `DESIGN.md` is the design contract; `doc/design/` holds audit reports, decision logs, and updated guidance for external baseline review/update workflows. - Dead code removed (`agentStatusBadge` duplicate map), byte-identical contrast constants consolidated, semantic renames (`--project-seed`/`--project-none`, `--liveness-blue`). ## Verification - `pnpm check:token-gates` — 3/3 gates CLEAN during the design-system run - `pnpm typecheck` && `pnpm --filter @paperclipai/ui build` — green during the design-system run - `node --test scripts/__tests__/storybook-visual-baseline.test.mjs` — pass after external-baseline rework - `pnpm exec tsc --noEmit --pretty false --module NodeNext --moduleResolution NodeNext --target ES2022 --types node,@playwright/test tests/storybook-visual/playwright.config.ts tests/storybook-visual/storybook-visual.spec.ts` — pass after external-baseline rework - `git diff --check origin/pr/9134..HEAD` — pass after external-baseline rework - `find tests/storybook-visual -type f -name '*.png' -print | wc -l` — `0` - `node scripts/storybook-visual-baseline.mjs verify` — intentionally fails closed until the first trusted maintainer publishes the baseline archive and updates `baseline-manifest.json` ## Risks - **Large but shallow:** the PR still touches many UI files due to mechanical token extraction and retune work, but committed PNG snapshot churn has been removed from the branch. - **Baseline publication required before the visual suite can pass in clean clones:** the manifest currently has placeholder archive metadata. A trusted maintainer must publish the first immutable archive, then update `baseline-manifest.json`. - **Rendering platform variance:** the external baseline should be captured in the documented Linux/Chromium environment. Future CI runs verify against the pinned archive and fail closed on checksum/count mismatch. - **Visual CI is opt-in while stabilizing:** add the `storybook-visual` label or dispatch the workflow manually to produce downloadable Playwright report/test-result artifacts. - **Scheduled follow-ups, deliberately out of scope:** Tailwind palette classes map to semantic tokens in a dedicated pass; card/pill component consolidation; ESLint ratchet. Tracked in `doc/design/DECISION-SHEET.md`. ## Model Used Claude Fable 5 (Anthropic, `claude-fable-5`, Mythos-class tier) with extended thinking, running in Claude Code with tool use; mechanical phases delegated to Claude Sonnet subagents. Follow-up external-baseline rework assisted by OpenAI Codex (`gpt-5` coding agent with repository, terminal, and GitHub tool use). All bulk rewrites executed via deterministic, idempotent scripts committed in `scripts/`; intentional visual changes were human-reviewed on screenshot contact sheets. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run targeted local verification and documented the intentional baseline-publication failure above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green *(pending new CI run after this rework)* - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups *(pending review)* - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) and OpenAI Codex --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Dotta <bippadotta@protonmail.com> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f42a66bf04 |
test: port pipelines tutorial e2e (#9149)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Pipelines are a control-plane workflow surface where operators need confidence that setup, intake, board movement, review, and learning flows keep working. > - The tutorial flow is broad enough that regressions can slip through unit tests when UI state, route behavior, and workflow fixtures drift apart. > - End-to-end coverage gives reviewers a realistic smoke path through the pipelines tutorial experience. > - This pull request ports the pipelines tutorial flow into Playwright and updates the variable flow covered by the scenario. > - The benefit is stronger browser-level regression coverage for a high-value onboarding/workflow path. ## Linked Issues or Issue Description No public GitHub issue exists. Inline feature request: **Subsystem affected** Cross-cutting: `tests/e2e`, server startup, and the pipelines UI workflow. **Problem or motivation** The pipelines tutorial flow lacked focused browser-level regression coverage for setup, intake, board movement, detail/review surfaces, and learning flow behavior. This left a broad user-facing workflow dependent on narrower unit coverage. **Proposed solution** Add a Playwright spec that boots a throwaway Paperclip instance and exercises the tutorial flow across setup, intake, board movement, item detail, review queue, learnings, agent fan-out, drift acknowledgement gates, child-terminal gates, and stale approvals. **Alternatives considered** Unit tests alone are faster but do not validate the integrated browser workflow. A release-smoke-only path would be broader than necessary for this tutorial-specific coverage. **Roadmap alignment** This supports the roadmap theme of easier onboarding and stronger first-run confidence without adding a new core feature surface. **Additional context** The local host used for this PR is missing Chromium’s `libatk-1.0.so.0` dependency, so the full browser run needs CI or a machine with Playwright system dependencies installed. ## What Changed - Added a Playwright e2e spec covering the pipelines tutorial flow. - Covered agent fan-out, drift acknowledgement gates, child-terminal gates, stale approvals, setup/intake, board movement, item detail, review queue, and learnings. - Added Playwright config support needed by the new tutorial flow. - Updated the tutorial variable flow assertions in the ported spec. ## Verification - Attempted `/srv/paperclip/home/paperclipai/paperclip/node_modules/.bin/playwright test --config tests/e2e/playwright.config.ts tests/e2e/pipelines-tutorial-flow.spec.ts` - Local result: the throwaway Paperclip server booted and `covers agent fan-out, drift acknowledgement gates, child-terminal gates, and stale approvals` passed. - Local blocker: the second Chromium test failed before executing because this host is missing the Playwright system library `libatk-1.0.so.0`. CI should run this on an image with browser dependencies installed. ## Risks Medium risk for CI duration/flakiness because this adds browser-level coverage over a broad workflow. The test is intentionally scoped to one spec file and uses the dedicated e2e config/server path. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5.5 coding agent with repository tool use and local shell execution. Context window was not surfaced by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ad961227f5 |
feat(secrets): add user-specific runtime secrets (#8825)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent runs often need provider credentials, API tokens, and other environment-bound secrets. > - Company-level secrets work for shared credentials, but they do not model values that should differ by human operator. > - Without a user-scoped model, a run can dispatch without knowing whether the responsible human has supplied the needed value. > - Paperclip also needs run attribution to make those user-scoped runtime checks deterministic and auditable. > - This pull request adds user-specific secret definitions, per-user values, environment bindings, responsible-user attribution, and runtime resolution gates. > - The benefit is that teams can define the secret once, let each user provide their own value, and block runs before dispatch when required user secrets or active definitions are unavailable. ## Linked Issues or Issue Description Refs #224 Refs #6057 This PR implements user-specific secret support as a core secret-management capability rather than a one-off adapter setting. It is related to existing public work on company secrets UI and runtime secret refs, but is distinct because the value is owned by the responsible user and resolved at run dispatch time. Related PR search before opening found existing secrets work such as #1550, #8256, #8614, #8634, and #8647; none of those add the full user-secret definition/value/runtime gate covered here. ## What Changed - Added user-secret definitions and per-user "My secrets" values, keeping stored values out of access metadata. - Added `user_secret_ref` environment bindings and UI affordances to pick them alongside existing secret refs. - Added responsible-user runtime resolution so user-secret refs resolve against the human responsible for the run. - Added pre-dispatch missing-secret gates so runs fail before adapter dispatch when required user values are absent or definitions are inactive. - Added low-trust allowlist hardening for user-secret runtime access. - Added issue, routine, run, and agent API key responsible-user attribution and fail-closed dispatch behavior when attribution cannot be resolved. - Added denial-copy mapping so responsible-user authorization failures surface as actionable run outcomes instead of opaque setup failures. - Added OpenAPI documentation for the user-secret routes. - Rebases cleanly on current `master`; migrations were renumbered incrementally as `0128_user_specific_secrets`, `0129_agent_api_key_responsible_user`, and `0130_run_responsible_user_invariant` after upstream `0126`/`0127` migrations. - Removed previously committed local design screenshots so the PR contains code/docs/tests only. ## Verification - PASS: PR head `2527febd106bcf3ca264ca0da7fca491084192d6` is based on `paperclipai/paperclip:master`. - PASS: `git diff --check` - PASS: `git diff --name-only public/master...HEAD | rg '^(pnpm-lock\\.yaml|\\.github/workflows/|screenshots/)' || true` produced no files. - PASS: migration journal audit confirmed unique indexes through `130` with tail entries `0126_issue_comment_derived_attribution`, `0127_environment_custom_images_instance_scoped`, `0128_user_specific_secrets`, `0129_agent_api_key_responsible_user`, and `0130_run_responsible_user_invariant`. - PASS: `pnpm --filter @paperclipai/ui typecheck` - PASS: `pnpm --filter @paperclipai/server typecheck` - PASS: `pnpm --filter @paperclipai/server exec vitest run src/__tests__/heartbeat-responsible-user-invariant.test.ts` - PASS: `pnpm --filter @paperclipai/server exec vitest run src/__tests__/heartbeat-active-run-output-watchdog.test.ts src/__tests__/heartbeat-stale-queue-invalidation.test.ts src/__tests__/heartbeat-workspace-finalize-branch.test.ts src/__tests__/issue-monitor-scheduler.test.ts` - PASS: `pnpm --filter @paperclipai/server exec vitest run src/__tests__/heartbeat-comment-wake-batching.test.ts src/__tests__/heartbeat-retry-scheduling.test.ts src/__tests__/heartbeat-accepted-plan-workspace-refresh.test.ts src/__tests__/heartbeat-plugin-environment.test.ts` - PASS: `pnpm --filter @paperclipai/server exec vitest run src/__tests__/low-trust-red-team-routes.test.ts` - PASS: `pnpm --filter @paperclipai/server exec vitest run src/__tests__/secrets-service.test.ts` (55 tests) - PASS: `pnpm vitest run server/src/__tests__/secrets-routes.test.ts server/src/__tests__/secrets-service.test.ts` (89 tests after final Greptile cleanup fixes) - PASS: `pnpm --filter @paperclipai/server exec vitest run src/__tests__/heartbeat-issue-liveness-escalation.test.ts` (17 tests after the final rebase CI fix) - PASS: focused server Vitest batches covering heartbeat recovery, project env, plugin env, routines, low-trust, pipelines, monitors, watchdog, and stale queue paths. - PASS: GitHub checks are green on `2527febd106bcf3ca264ca0da7fca491084192d6`, including Typecheck + Release Registry, Build, General tests, serialized server suites, e2e, Canary Dry Run, verify, security checks, and Greptile Review. - PASS: Greptile Review completed successfully on `2527febd106bcf3ca264ca0da7fca491084192d6` with Confidence Score 5/5, and GraphQL review-thread audit returned zero unresolved non-outdated threads. ## Risks - Runtime behavior now depends on a run having a correct responsible user; missing or incorrect responsibility assignment can block runs before adapter dispatch. - `user_secret_ref` bindings intentionally expose metadata without values, but UI/API callers may need to handle the new binding kind explicitly. - External secret providers and IAM policies are not automatically provisioned by this PR; operators still need to configure provider-side access for non-local vaults. - The PR is broad across db/shared/server/UI/runtime paths, so release validation should include both API and UI secret workflows before merge. - The migration renumbering is intentionally incremental after upstream migrations; the branch migrations use guarded column/table/index/constraint creation so users who tested the older draft numbering should not hit duplicate DDL for the existing objects. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5-based coding agent (`gpt-5`), Codex local adapter with shell/tool use and code execution. Context window and internal reasoning mode are not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
c79d347abe |
Update company creation copy (#8653)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The onboarding and company switcher UI are the first places users create or select an organization > - Some of that surface still used older team/workspace wording even though the product model is company-centric > - Mixed wording makes the setup path feel inconsistent and can make users wonder whether they are creating a team, workspace, or company > - This pull request updates the affected UI copy to consistently say company > - The benefit is a clearer first-run and navigation experience without changing behavior ## Linked Issues or Issue Description No public issue exists for this small UI polish change. ### Problem or motivation The company creation and switcher surfaces used mixed team/workspace/company wording for the same concept, which makes the setup path feel inconsistent. ### Proposed solution Update the visible copy, accessibility label, inline comment, e2e expectations, and matching test expectations to use company-centric language consistently. ### Alternatives considered Leave the existing wording alone, but that preserves inconsistent terminology in a high-traffic setup path. ### Roadmap alignment This is focused UI polish and does not overlap with a roadmap-level core feature. ### Additional context The create action keeps its trailing ellipsis because it opens the onboarding wizard rather than completing immediately. ## What Changed - Updated front door and onboarding wizard labels from team-oriented copy to company-oriented copy. - Updated the sidebar company menu from workspace/team wording to company wording, including the trigger accessibility label and empty fallback text. - Kept the sidebar create action ellipsis for the dialog/wizard affordance. - Updated component and Playwright test expectations for the new copy. ## Verification - `pnpm exec vitest run ui/src/components/SidebarCompanyMenu.test.tsx` - Attempted `npx playwright test --config tests/e2e/playwright.config.ts tests/e2e/onboarding.spec.ts tests/e2e/nux-phase4-screenshots.spec.ts tests/e2e/planning-mode-visual-verification.spec.ts tests/e2e/conference-room-typing-intro.spec.ts`; local browser launch is blocked by missing host Chromium dependencies (`libatk1.0-0t64`, `libatspi2.0-0t64`, `libxcomposite1`, `libxdamage1`, `libxfixes3`, `libxrandr2`, `libgbm1`, `libasound2t64`). - Screenshots intentionally omitted because this is a copy-only change and no design screenshots are needed for review. ## Risks Low risk. This is copy-only UI polish plus matching test updates; no data model, API, migration, workflow, lockfile, or behavior changes are included. ## Model Used OpenAI GPT-5 Codex (`gpt-5`) via the Paperclip Codex agent, with tool-assisted repository inspection, GitHub CLI usage, and local command execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
841742fc1a |
[codex] Graduate experimental conference room defaults (#8628)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The board UI has been graduating experimental conference-room task experiences into the default issue and onboarding flows > - Several task UI improvements were still coupled to the Conference Room Chat experimental flag even though they are useful outside chat itself > - That coupling meant disabling chat also reverted unrelated defaults such as work-mode labels, task status colors, team creation copy, and the graduated issue thread > - This pull request keeps chat-specific gating scoped to chat while making the graduated task UI the default experience > - The benefit is that operators can use the newer task workflows without needing to enable the separate chat experiment ## Linked Issues or Issue Description No public GitHub issue exists for this exact change. ### Problem or motivation The Conference Room Chat experimental flag was controlling unrelated task UI defaults, which made non-chat workflows regress when chat was disabled. ### Proposed solution Remove that flag from task-thread, work-mode, onboarding, status, and team-creation presentation paths while leaving chat-specific behavior separately gated. ### Alternatives considered Keeping the flag as a broad umbrella until chat graduates would avoid a behavior change, but it keeps unrelated UI improvements hidden behind the wrong capability switch. ### Roadmap alignment Checked `ROADMAP.md`; this is focused graduation/polish for existing UI surfaces rather than a new roadmap-level core feature. Related search: - Searched open PRs and issues for `conference room chat experimental flag`: no matches. - Searched open PRs and issues for `graduated issue thread`: no matches. ## What Changed - Removes Conference Room Chat flag branching from task-thread rendering, work-mode labels, task status colors, and team creation copy. - Makes the onboarding completion path create/reuse an onboarding project, create the first assigned task, and send the user to the dashboard instead of chat. - Deletes the legacy onboarding wizard and classic task-thread files now that the graduated flow is the default. - Updates focused UI tests and affected E2E specs for the default task experience. - Adds a user-visible onboarding error if restored state is missing the company or agent required for launch. ## Verification - `pnpm run preflight:workspace-links && pnpm exec vitest run ui/src/components/NewIssueDialog.test.tsx ui/src/components/OnboardingWizardVariant.test.tsx ui/src/components/RunChatSurface.test.tsx ui/src/components/SidebarCompanyMenu.test.tsx ui/src/components/StatusBadge.test.tsx ui/src/lib/agent-order.test.ts ui/src/lib/onboarding-launch.test.ts ui/src/lib/work-mode-meta.test.ts ui/src/pages/IssueDetail.test.tsx` - `pnpm --filter @paperclipai/ui typecheck` - `npx playwright test --config tests/e2e/playwright.config.ts --list tests/e2e/conference-room-typing-intro.spec.ts tests/e2e/planning-mode-visual-verification.spec.ts` - Attempted targeted Playwright execution locally, but this host is missing Chromium system libraries (`libatk1.0-0t64`, `libatspi2.0-0t64`, `libxcomposite1`, `libxdamage1`, `libxfixes3`, `libxrandr2`, `libgbm1`, `libasound2t64`). CI runs the specs in the proper Actions environment. ## Risks - Medium UI behavior risk: this intentionally changes the default experience for users who have not enabled Conference Room Chat. - Medium onboarding risk: completion now creates/reuses a project and creates the first task instead of only navigating. - Low migration risk: no database schema or migration changes are included. - The PR avoids `pnpm-lock.yaml` and `.github/workflows` changes. ## Model Used OpenAI Codex, GPT-5-class coding model, tool-enabled local repository workflow with shell, git, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7069053a1f |
[codex] Add ask issue work mode (#8334)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Issue work mode controls how a task starts and how the conversation composer frames the operator's intent. > - Paperclip already supports standard agent execution and planning mode, but there is no lightweight mode for asking a question without immediately implying execution or plan drafting. > - That gap makes low-commitment clarification workflows look like normal task execution. > - This pull request adds an explicit Ask mode and threads it through shared contracts, server heartbeat context, and the issue composer UI. > - The benefit is that operators can create or switch a task into a question-oriented mode while preserving existing agent and planning flows. ## Linked Issues or Issue Description No public GitHub issue exists for this change. Inline feature request follows the repository feature request template. ### Subsystem affected Cross-cutting: `packages/shared`, `server/`, and `ui/`. ### Problem or motivation Issue conversations currently distinguish standard agent work from planning work, but question-first conversations do not have a clear public mode in the shared contract or UI. Operators who want to ask an agent a focused question have to use standard mode, which can imply normal task execution, or planning mode, which asks for a plan rather than an answer. ### Proposed solution Add Ask as a first-class issue work mode. It should be selectable from issue creation and issue chat, cycle alongside Standard and Planning from the keyboard shortcut/menu, appear distinctly in composer styling, and be included in heartbeat context so agents know to answer directly instead of executing or drafting a plan. ### Alternatives considered - Keep using standard mode for questions: rejected because it does not communicate answer-only intent to the agent or the UI. - Reuse planning mode for questions: rejected because planning mode asks for a plan and is semantically different from asking a question. - Add only local UI copy: rejected because the mode needs to be represented in the shared contract and server heartbeat context to be reliable. ### Roadmap alignment This is a focused issue-workflow improvement. `ROADMAP.md` was checked and no duplicate planned core work was found. ### Additional context Related public searches performed before opening this PR: - GitHub PR search for `"ask mode" repo:paperclipai/paperclip` - GitHub issue search for `"ask mode" repo:paperclipai/paperclip` - GitHub PR search for `"work mode" "ask" repo:paperclipai/paperclip` No duplicate PR was found. ## What Changed - Added `ask` to the shared issue work-mode contract and validation coverage. - Included issue work mode in heartbeat context summaries so agents can see standard, planning, and ask state. - Added Ask mode metadata, styling, composer tone handling, and selection/cycling behavior in the issue chat/new issue UI. - Updated focused tests for shared validators, heartbeat context, and affected UI work-mode flows. ## Verification - `NODE_ENV=test pnpm exec vitest run ui/src/components/ChatComposer.test.tsx ui/src/components/IssueChatThread.test.tsx ui/src/components/NewIssueDialog.test.tsx ui/src/lib/work-mode-meta.test.ts` - `NODE_ENV=test pnpm exec vitest run packages/shared/src/validators/issue.test.ts server/src/__tests__/heartbeat-context-summary.test.ts server/src/__tests__/issues-service.test.ts ui/src/components/ChatComposer.test.tsx ui/src/components/IssueChatThread.test.tsx ui/src/components/NewIssueDialog.test.tsx ui/src/lib/work-mode-meta.test.ts ui/src/pages/IssueDetail.test.tsx` The broader targeted command passed 8 test files / 245 tests. Visual reference for Standard/Planning/Ask composer states: https://gist.github.com/cryppadotta/714d8590bac55500a65e7e16de5bb4b8 It emitted an expected warning from an existing server test fixture about a missing run-log fixture while verifying derived issue comment metadata. ## Risks Low to moderate risk. This adds a new enum value that crosses shared, server, and UI contracts. Existing standard and planning modes are preserved, but any downstream code assuming only two non-terminal work modes may need to handle `ask`. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-5 Codex coding agent in Paperclip CodexCoder mode, with shell, git, GitHub connector, and local test execution tools. Context window and exact hosted model snapshot are not exposed in this runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`, `feat/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] If this change affects the UI, I have included before/after screenshots - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
6f9801a46b |
feat(ui): NUX rework behind enableConferenceRoomChat experimental flag — capsule onboarding, conference-room chat, unified composer (#8000)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - The first-run experience (onboarding wizard) and the chat surfaces
(conference-room/board chat, task threads, composers) are the product's
front door — they decide whether a new operator understands "hire
agents, give them work, review results" in the first five minutes
> - Today those surfaces feel ticket-y and form-like: the wizard is a
static multi-step form that ends in an anticlimactic "Launch" screen,
the task composer and board chat behave differently from each other, and
agent-feed issue quicklooks misbehave (multiple flyouts open at once,
cards jump on hover)
> - We wanted to iterate toward a conversational, team-centric NUX — but
without risking the workflows of everyone already running Paperclip
> - This PR reworks the NUX behind a new default-OFF
`enableConferenceRoomChat` experimental flag: a capsule-motif onboarding
wizard that builds your team as you answer, a conference-room chat
surface, one shared ChatComposer across surfaces, brand-accurate status
chips, and feed-quicklook fixes — with the pre-existing UI
fork-and-frozen as `*Classic` components that flag-OFF users keep
> - The benefit is a complete, testable modern NUX that anyone can opt
into from Settings → Experimental, with zero default behavior change and
a clean path to either graduate or drop the experiment
## Linked Issues or Issue Description
No pre-existing GitHub issue — feature description per
`feature_request.yml`:
- **Problem / motivation:** Paperclip's onboarding wizard and chat
surfaces grew up as separate ticket-centric forms. New users get a
form-filling experience rather than the feeling of standing up a team;
the board chat and task threads use different composers with different
affordances; the agent feed's issue quicklook can stack multiple
popovers and shifts cards on hover.
- **Proposed solution:** A coherent NUX experiment behind one
experimental flag (`enableConferenceRoomChat`, Settings → Experimental,
default OFF): capsule onboarding wizard with an evolving team capsule,
conference-room chat, unified `ChatComposer`, team-centric copy, brand
status chips, quicklook single-flight fix. Flag-OFF users get the exact
pre-experiment UI via frozen `*Classic` forks, verified by an on/off
parity test matrix.
- **Alternatives considered:** (a) incremental unflagged restyling —
rejected: the changes interlock across surfaces and would drip risk into
every release; (b) a separate app shell / route for the new NUX —
rejected: too much divergence, the flag + classic-fork pattern keeps the
diff reviewable and reversible.
- **Roadmap alignment:** `ROADMAP.md` lists **CEO Chat** ("a
lighter-weight way to talk to leadership agents... should still resolve
to real work objects"). This experiment is groundwork in that direction
(conference-room chat resolves to issues/tasks via the same composer
used in task threads) and does not change the core task-and-comments
model.
Related PRs found in the dedup search (same area, none duplicate this
work — they target the classic wizard, which this PR intentionally
leaves intact and mergeable):
- #5385 — Coach-driven onboarding: conversational entry +
agent-companies package import
- #5378 — Onboarding wizard: reusable adapter picker + probe card
- #6636 — ui(onboarding): friendly error surface + retry for the wizard
- #7005 — fix(onboarding): explicitly await first-task wake
- #2616 — fix: restore workspace directory config in onboarding wizard
## What Changed
- **Experimental flag plumbing** — `enableConferenceRoomChat` in shared
types/validators, server instance-settings service + API, Settings →
Experimental card with explicit enable/disable copy
- **Onboarding wizard** — classic wizard forked and frozen
(`OnboardingWizardClassic`); flag-ON variant is a 5-step capsule wizard
with a persistent evolving `AgentCapsule` (gradient/glow motif),
team-centric reframed copy, and a typing-dots intro (hardened with
fake-timer tests)
- **Conference-room chat** — flag-ON board-chat surface with agent
bubble name/icon headers and copy/vote/timestamp action rows
(`AgentBubbleActionRow`)
- **Unified composer** — shared `ChatComposer` adopted across surfaces;
translucent surface + scroll-mask removal; "Agent mode"/"Plan mode"
relabels; no-assignee confirmation `AlertDialog` (new
`ui/alert-dialog.tsx` primitive); `@task` reference picker +
linkification in mentions
- **Agent feed** — single-flight issue-quicklook store (one popover at a
time), flyouts open to the left, removed hover translate-y jitter
- **Status chips** — brand-accurate task status chips behind the flag
(light/dark, 1px borders per paperclip.ing/brand)
- **Tests** — flag on/off parity matrix across IssueDetail,
NewIssueDialog, Sidebar, wizard, gate components; component tests for
all new pieces
- **Merge with `master`** — one conflict in
`ui/src/components/IssueChatThread.tsx`, resolved by keeping master's
new `AssigneeChip`/`HandoffWakeRow`/`RunStatusBadge` components inside
the flag-gated metadata-row chrome (details in commit `21a5642a`);
post-merge fixes: vitest 4 mock typing in `MarkdownEditor.test.tsx`,
flag hook made safe for provider-less mounts (master's new isolated
component tests)
- **Branch hygiene** — internal design wireframes/mockups stripped
before the PR (they live in the Paperclip issue threads)
- No user-facing documentation changes required: the flag is
intentionally experimental and self-described in the Settings card; no
existing docs reference the affected surfaces
## Verification
- `pnpm run typecheck` — green across the workspace (ui, server, shared,
plugins)
- Full UI suite (`vitest run` in `ui/`, clean worktree at this HEAD):
**1593/1595 passing, 223/224 files** — the 2 remaining failures are in
`src/components/artifacts/ArtifactCard.test.tsx` and **fail identically
on pristine `origin/master`** (pre-existing upstream, unrelated to this
branch)
- Full server suite (`vitest run` in `server/`, same clean worktree):
results in PR checks; flag plumbing covered by instance-settings tests
- Targeted post-merge resolution check: `IssueChatThread`,
`IssueChatThreadSystemNotice`, `IssueDetail`, `Sidebar`,
`ConferenceRoomChatGate`, `OnboardingWizardVariant`, `NewIssueDialog`,
`InstanceExperimentalSettings`, `MarkdownEditor` — 172/172 passing
- Manual walkthrough: flag OFF (default) → onboarding wizard, task
thread, board chat, composer all render the classic UI; flag ON via
Settings → Experimental → capsule wizard, conference-room chat, unified
composer, status chips active
- Screenshots: see below
**Flag on/off screenshots** (committed on this branch under
`screenshots/PR-8000-*`):
| Surface | Flag OFF (classic, default) | Flag ON (experimental) |
| --- | --- | --- |
| Settings → Experimental | 
| 
|
| Task thread | 
| 
|
| Home / nav | 
| 
|
| Conference Room (flag-ON only surface) | — | 
|
Capsule onboarding wizard walkthrough screenshots (flag ON) are attached
to the Paperclip design/implementation threads; the wizard requires a
fresh instance so it is captured via the e2e harness
(`tests/e2e/nux-phase4-screenshots.spec.ts`).
## Risks
- **Large surface, but gated:** all new behavior sits behind
`enableConferenceRoomChat`, default OFF; flag-OFF rendering is locked by
frozen `*Classic` forks plus an on/off parity test suite
- **Classic forks are frozen at the fork point (`e3aada1d`):** master
features added to the live thread component after that point (assignee
handoff chips, run status badge, composer mention coach) render in the
flag-ON path; the flag-OFF task thread keeps the fork-point behavior
until the experiment graduates (forks deleted) or is dropped (forks
restored as canonical). Called out for reviewer attention.
- **Merge-conflict resolution in `IssueChatThread.tsx`** (commit
`21a5642a`) deserves reviewer eyes: master's new handoff/run-status
components were kept; the base toast-style no-assignee flow remains
replaced by the AlertDialog flow introduced on this branch
- Schema/server changes are additive (one optional boolean instance
setting); no migrations of existing data
## Model Used
- Claude (Anthropic) via Claude Code running in the Paperclip agent
harness (agent: ClaudeCoder)
- Branch implemented across multiple agent sessions on Claude Opus-class
models with extended thinking + tool use (file edits, shell, Playwright
screenshots); merge/PR session model ID as reported by the harness:
`claude-fable-5` (Claude Code CLI)
- All code was agent-authored and board-reviewed through Paperclip issue
threads (plans, wireframes, confirmations) before merging
## Checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] If this change affects the UI, I have included before/after
screenshots
- [x] I have updated relevant documentation to reflect my changes (none
required — experimental flag, self-documenting Settings card; noted
above)
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green (run 3 on `8af3041a`: all 16
gates SUCCESS, incl. e2e and all 4 serialized-suite shards)
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
(re-review verdict: Confidence 5/5, “Safe to merge”; all 4 round-1
findings fixed + confirmed resolved; both summary notes addressed in
`8af3041a`)
- [x] I will address all Greptile and reviewer comments before
requesting merge
---------
Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
50bff3b274 |
feat(ui): add collapsible sidebar rail and takeover panes (#7824)
## Thinking Path > - Paperclip is the open source control plane people use to manage AI agents, work, and company context. > - The board UI sidebar is the main way operators keep orientation across companies, projects, agents, issues, and settings. > - The existing fixed expanded sidebar competes with route-specific navigation, especially company settings and plugin routes that bring their own contextual sidebar. > - A collapsible primary rail preserves global navigation while giving contextual pages more horizontal room. > - This pull request adds a persisted collapsed rail, hover/focus peek, keyboard toggle, and a secondary sidebar takeover model for settings and plugin `routeSidebar` surfaces. > - The benefit is a denser board shell that keeps the app rail available without replacing it when a route needs its own navigation. ## Linked Issues or Issue Description Paperclip issue: PAP-10638 Create collapsible sidebar branch. Related GitHub PR found during duplicate search: #3838 (`feat/collapsible-sidebar`) covers a similar sidebar area but is a different head branch and implementation. This PR intentionally packages the work from `PAP-10638-collapsable-sidebar` into one reviewable branch. Problem description: The board shell needs a first-class collapsed sidebar mode. Contextual surfaces such as company settings and plugin route sidebars should not replace the global app sidebar; they should collapse the app sidebar to a rail and render their contextual navigation beside it. ## What Changed - Added desktop collapsed/sidebar-peek state to `SidebarContext`, including persisted user pins, route collapse requests, and forced collapse for secondary-sidebar routes. - Replaced the old resizable sidebar pane with `SidebarShell`, which supports a fixed 64px rail, persisted expanded width, keyboard/pointer resizing, and hover/focus peek overlay behavior. - Updated `Sidebar`, sidebar nav items, project/agent sections, badges, and account/company menu presentation for expanded, collapsed, and peeking states. - Added `RequestCollapsedSidebar` and `SecondarySidebar` so routes and plugin `routeSidebar` slots can request contextual sidebar layouts without replacing the primary app sidebar. - Wired company settings and plugin route sidebars into the secondary-pane takeover model. - Added focused Vitest coverage for sidebar state precedence, shell sizing, nav item rail rendering, keyboard shortcuts, layout takeover behavior, and route collapse requests. - Updated plugin authoring docs/spec references for route sidebar behavior. ## Verification Targeted local verification passed: ```sh NODE_ENV=test pnpm run preflight:workspace-links && NODE_ENV=test pnpm exec vitest run ui/src/context/SidebarContext.test.tsx ui/src/components/SidebarShell.test.tsx ui/src/components/Sidebar.test.tsx ui/src/components/Layout.test.tsx ui/src/components/RequestCollapsedSidebar.test.tsx ui/src/components/SidebarNavItem.test.tsx ui/src/components/SidebarAgents.test.tsx ui/src/components/SidebarProjects.test.tsx ui/src/components/KeyboardShortcutsCheatsheet.test.tsx ui/src/hooks/useKeyboardShortcuts.test.tsx ``` Result: 10 test files passed, 88 tests passed. Additional follow-up verification passed after review fixes: ```sh NODE_ENV=test pnpm run preflight:workspace-links && NODE_ENV=test pnpm exec vitest run ui/src/components/Layout.test.tsx ui/src/context/SidebarContext.test.tsx && pnpm --filter /ui typecheck ``` Result: 2 test files passed, 28 tests passed, and UI typecheck passed. Latest PR-head remote checks: Paperclip PR workflow, Snyk, Socket, and Greptile are green; commitperclip `review` is cancelled in its security-gate step after filing a non-blocking neutral `security-review` check. Notes: - A direct run without `NODE_ENV=test` loads React's production build in this workspace, where `act` is unavailable; the command above matches the repo stable runner's test environment. - I did not run Playwright/browser e2e or full workspace build/typecheck in this PR-creation heartbeat. - QA screenshots are attached in https://github.com/paperclipai/paperclip/pull/7824#issuecomment-4661968387 for expanded, collapsed rail, hover peek, and settings secondary-sidebar states. ## Risks - Medium UI layout risk: this changes the board shell and primary sidebar composition across many routes. - Local storage migration risk is low: new collapsed state uses a new key and existing width storage remains scoped to the sidebar width. - Plugin route risk: plugin `routeSidebar` slots now render as secondary panes on desktop, so plugin authors should confirm their route sidebar content fits a 240px contextual pane. - Mobile risk appears low because mobile keeps the drawer model and gates collapsed/peek behavior to desktop. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex coding agent based on GPT-5, with local shell/git/GitHub CLI tool use. Exact service-side model identifier and context window were not exposed in this runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] If this change affects the UI, I have included before/after screenshots - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
2e74d32871 |
PAP-10430: split Issue-to-Task copy migration (#7651)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The board UI is the operator surface where users create, assign, monitor, and review work items. > - The product language is moving toward "tasks" for user-facing work items while the internal API and database still use "issues". > - PR #7543 bundled this copy migration with broader information-architecture work, which made the branch too large for Greptile review. > - This pull request peels the Issue-to-Task copy migration into a smaller, independently reviewable change. > - The benefit is clearer user-facing terminology, less agent confusion via the Paperclip skill note, and a smaller PR that Greptile can review. ## Linked Issues or Issue Description Refs #7645 Refs #7543 Refs PAP-10430 This PR was split out of #7543 so the Issue-to-Task copy migration can be reviewed separately and the remaining IA PR can fall under Greptile's file limit. ## What Changed - Preserves Scott Tong's original `PAP-57` copy-only commit, with author and co-author credit intact, to rename user-facing "Issues" copy to "Tasks" across the UI while keeping routes/API/internal symbols as `issue`. - Updates onboarding and release-smoke browser selectors from `Create & Open Issue` to `Create & Open Task`. - Adds a terminology note to `skills/paperclip/SKILL.md` clarifying that task and issue refer to the same Paperclip work item. - Resolves the only cherry-pick conflict by keeping current search artifacts support and changing visible search copy to "tasks". ## Verification - `pnpm --filter @paperclipai/ui build` passed. - `NODE_ENV=test pnpm exec vitest run ui/src/components/IssuesList.test.tsx ui/src/components/Sidebar.test.tsx ui/src/components/NewIssueDialog.test.tsx ui/src/pages/IssueDetail.test.tsx` passed: 4 files, 66 tests. - `git diff --check origin/master...HEAD` passed. - Diff is 80 files, below Greptile's 100-file limit. - Before/after UI copy examples: "Issues" -> "Tasks", "New Issue" -> "New Task", "Create & Open Issue" -> "Create & Open Task". ## Risks - Medium copy-risk: this intentionally changes user-facing terminology broadly while keeping internal issue identifiers and routes unchanged. - Some docs and APIs still say `issue`; the skill note clarifies this so agents do not treat task and issue as separate entities. - Browser-level visual validation is expected from CI because this local container is missing usable browser dependencies. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used Scott Tong authored the original `PAP-57` copy migration, assisted by Claude Opus 4.8 and Paperclip agents per the preserved commit metadata. Codex / GPT-5-class coding agent with shell, GitHub CLI, and repository access performed the PR split, conflict resolution, skill note, and verification. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] If this change affects the UI, I have included before/after screenshots - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: scotttong <scott.tong@gmail.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
524e18b060 |
ci: use runner Chrome for headless workflows (#6967)
## Thinking Path > - Paperclip relies on CI browser suites to protect control-plane workflows, so a stalled browser bootstrap is a release blocker even when app code is unchanged. > - The failing signal on [PAPA-457](/PAP/issues/PAPA-457) was specific to the PR e2e lane timing out before tests started, which pointed at environment setup rather than assertions. > - The first shell-only Chromium attempt reduced download size, but the GitHub Actions log showed Playwright still hanging inside its install step after the headless shell download finished. > - That means the real problem is the Playwright browser-install path itself on the hosted Ubuntu runner, not just the size of the downloaded artifact. > - GitHub's Ubuntu runners already ship Google Chrome, and Playwright can target that binary through the `chrome` channel without downloading its own Chromium bundle. > - The safer workflow fix is therefore to remove the Playwright install step from the affected headless jobs and make the Playwright configs optionally use runner Chrome only when CI opts into it. > - This keeps local defaults unchanged, removes the failing browser-download dependency from CI, and preserves headless coverage for PR, standalone e2e, and release-smoke workflows. ## What Changed - Updated `.github/workflows/pr.yml`, `.github/workflows/e2e.yml`, and `.github/workflows/release-smoke.yml` to stop downloading Playwright browsers and instead verify the runner's preinstalled `google-chrome`. - Passed `PAPERCLIP_PLAYWRIGHT_CHANNEL=chrome` into the headless PR, standalone e2e, and release-smoke test steps so those jobs explicitly use runner Chrome. - Updated `tests/e2e/playwright.config.ts` and `tests/release-smoke/playwright.config.ts` to honor `PAPERCLIP_PLAYWRIGHT_CHANNEL` while keeping the default local/browser-bundle behavior unchanged when the env var is absent. ## Verification - Investigated the failed PR run log and confirmed the prior `Install Playwright` step stalled after `chromium-headless-shell` reached 100% download. - `PLAYWRIGHT_BROWSERS_PATH="$(mktemp -d)" PAPERCLIP_PLAYWRIGHT_CHANNEL=chrome PAPERCLIP_E2E_SKIP_LLM=true pnpm run test:e2e` Result: `7 passed (21.1s)` with an empty temporary Playwright browser cache, proving the e2e suite runs without any Playwright browser download when the `chrome` channel is selected. - `git diff --check` ## Risks - This assumes GitHub's Ubuntu runner continues to ship `google-chrome`; if that image contract changes, these workflows would need a dedicated Chrome install step. - The `chrome` channel can differ slightly from Playwright-managed Chromium, so the config gate is intentionally env-scoped to CI workflows that need the hosted-runner path. ## Model Used - OpenAI Codex, GPT-5-based coding agent running through Paperclip's `codex_local` adapter with tool use, shell execution, and repository editing enabled. The exact internal snapshot/version string is not exposed in-session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have run tests locally and they pass - [ ] I have added or updated tests where applicable - [ ] If this change affects the UI, I have included before/after screenshots - [ ] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
32605b71ad |
Remove planning badge from inbox issue rows (PAP-9691) (#6269)
## Summary - Removes the amber "Planning" pill from inbox / issue-list rows in `IssueRow` - Updates the focused IssueRow test to assert the badge is no longer rendered - Per [PAP-9691](https://paperclip.ing/PAP/issues/PAP-9691): user just doesn't want to see the badge in list rows The underlying `issue.workMode === "planning"` data, the issue detail composer toggle, and the server/plugin/heartbeat work-mode contract introduced in #5353 are all untouched. Planning mode still functions; the list-row indicator is just gone. ## Test plan - [x] `pnpm exec vitest run --project @paperclipai/ui ui/src/components/IssueRow.test.tsx` (11 passed) - [ ] Visual: open `/PAP/inbox` with a planning-mode issue assigned and confirm no amber Planning pill on the row 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
a1b30c9f35 |
Add planning mode for issue work (#5353)
## Thinking Path > - Paperclip is a control plane for autonomous AI companies. > - Issues are the core unit of work, and issue comments are how board users and agents coordinate execution. > - Some issue conversations need to produce plans and approvals instead of immediate implementation work. > - The existing issue contract did not distinguish standard execution comments from planning-oriented issue work. > - This pull request adds an issue work-mode contract and board UI affordances for standard vs planning mode. > - The benefit is that planning-mode issues can be created, displayed, discussed, and carried through agent heartbeat context without losing the normal issue workflow. ## What Changed - Added `standard` / `planning` issue work-mode contracts across DB, shared validators/types, server issue flows, plugin protocol, and adapter heartbeat payloads. - Added an idempotent `0081_optimal_dormammu` migration for `issues.work_mode`, ordered after current `public-gh/master` migrations. - Updated heartbeat/context summaries and issue-thread interaction behavior so planning work mode is preserved when creating suggested follow-up issues. - Added UI support for planning-mode issue creation, issue rows, detail composer styling, and composer work-mode toggles. - Added focused server/shared/UI tests plus a Playwright visual verification spec for planning-mode surfaces. - Rebased the branch onto current `public-gh/master` and added durable planning-mode screenshots under `doc/assets/pap-3368/`. ## Verification - `pnpm --filter @paperclipai/db run check:migrations` - `pnpm exec vitest run --project @paperclipai/shared packages/shared/src/validators/issue.test.ts` - `pnpm exec vitest run --project @paperclipai/server server/src/__tests__/heartbeat-context-summary.test.ts server/src/__tests__/issue-thread-interactions-service.test.ts server/src/__tests__/issues-goal-context-routes.test.ts --pool=forks --poolOptions.forks.isolate=true` - `pnpm exec vitest run --project @paperclipai/ui ui/src/components/IssueChatThread.test.tsx ui/src/components/NewIssueDialog.test.tsx ui/src/components/IssueRow.test.tsx ui/src/pages/IssueDetail.test.tsx` - `pnpm exec vitest run --project @paperclipai/adapter-utils packages/adapter-utils/src/server-utils.test.ts` - `PAPERCLIP_E2E_SKIP_LLM=true npx playwright test --config tests/e2e/playwright.config.ts tests/e2e/planning-mode-visual-verification.spec.ts` ## Screenshots Desktop planning detail:  Desktop planning row:  Desktop staged standard toggle:  Mobile planning detail:  Mobile planning row:  ## Risks - Medium migration risk: this adds a non-null issue column. The migration uses `ADD COLUMN IF NOT EXISTS` so installations that applied an older branch-local migration number can still apply the final numbered migration safely. - Medium contract risk: issue payloads, plugin payloads, and adapter heartbeat payloads now include work mode; compatibility is handled by defaulting missing values to `standard`. - UI risk is moderate because composer controls changed; focused component tests and visual e2e coverage exercise standard vs planning display and toggle behavior. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5 coding agent in a local Paperclip worktree, with shell/tool use. Exact context-window size is not exposed in this runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] If this change affects the UI, I have included before/after screenshots - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7a329fb8bb |
Harden API route authorization boundaries (#4122)
## Thinking Path > - Paperclip orchestrates AI agents for zero-human companies. > - The REST API is the control-plane boundary for companies, agents, plugins, adapters, costs, invites, and issue mutations. > - Several routes still relied on broad board or company access checks without consistently enforcing the narrower actor, company, and active-checkout boundaries those operations require. > - That can allow agents or non-admin users to mutate sensitive resources outside the intended governance path. > - This pull request hardens the route authorization layer and adds regression coverage for the audited API surfaces. > - The benefit is tighter multi-company isolation, safer plugin and adapter administration, and stronger enforcement of active issue ownership. ## What Changed - Added route-level authorization checks for budgets, plugin administration/scoped routes, adapter management, company import/export, direct agent creation, invite test resolution, and issue mutation/write surfaces. - Enforced active checkout ownership for agent-authenticated issue mutations, while preserving explicit management overrides for permitted managers. - Restricted sensitive adapter and plugin management operations to instance-admin or properly scoped actors. - Tightened company portability and invite probing routes so agents cannot cross company boundaries. - Updated access constants and the Company Access UI copy for the new active-checkout management grant. - Added focused regression tests covering cross-company denial, agent self-mutation denial, admin-only operations, and active checkout ownership. - Rebased the branch onto `public-gh/master` and fixed validation fallout from the rebase: heartbeat-context route ordering and a company import/export e2e fixture that now opts out of direct-hire approval before using direct agent creation. - Updated onboarding and signoff e2e setup to create seed agents through `/agent-hires` plus board approval, so they remain compatible with the approval-gated new-agent default. - Addressed Greptile feedback by removing a duplicate company export API alias, avoiding N+1 reporting-chain lookups in active-checkout override checks, allowing agent mutations on unassigned `in_progress` issues, and blocking NAT64 invite-probe targets. ## Verification - `pnpm exec vitest run server/src/__tests__/issues-goal-context-routes.test.ts cli/src/__tests__/company-import-export-e2e.test.ts` - `pnpm exec vitest run server/src/__tests__/plugin-routes-authz.test.ts server/src/__tests__/adapter-routes-authz.test.ts server/src/__tests__/agent-permissions-routes.test.ts server/src/__tests__/company-portability-routes.test.ts server/src/__tests__/costs-service.test.ts server/src/__tests__/invite-test-resolution-route.test.ts server/src/__tests__/issue-agent-mutation-ownership-routes.test.ts server/src/__tests__/agent-adapter-validation-routes.test.ts` - `pnpm exec vitest run server/src/__tests__/issue-agent-mutation-ownership-routes.test.ts` - `pnpm exec vitest run server/src/__tests__/invite-test-resolution-route.test.ts` - `pnpm -r typecheck` - `pnpm --filter server typecheck` - `pnpm --filter ui typecheck` - `pnpm build` - `pnpm test:e2e -- tests/e2e/onboarding.spec.ts tests/e2e/signoff-policy.spec.ts` - `pnpm test:e2e -- tests/e2e/signoff-policy.spec.ts` - `pnpm test:run` was also run. It failed under default full-suite parallelism with two order-dependent failures in `plugin-routes-authz.test.ts` and `routines-e2e.test.ts`; both files passed when rerun directly together with `pnpm exec vitest run server/src/__tests__/plugin-routes-authz.test.ts server/src/__tests__/routines-e2e.test.ts`. ## Risks - Medium risk: this changes authorization behavior across multiple sensitive API surfaces, so callers that depended on broad board/company access may now receive `403` or `409` until they use the correct governance path. - Direct agent creation now respects the company-level board-approval requirement; integrations that need pending hires should use `/api/companies/:companyId/agent-hires`. - Active in-progress issue mutations now require checkout ownership or an explicit management override, which may reveal workflow assumptions in older automation. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI Codex, GPT-5 coding agent, tool-using workflow with local shell, Git, GitHub CLI, and repository tests. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] If this change affects the UI, I have included before/after screenshots - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
16b2b84d84 |
[codex] Improve agent runtime recovery and governance (#4086)
## Thinking Path > - Paperclip orchestrates AI agents for zero-human companies. > - The heartbeat runtime, agent import path, and agent configuration defaults determine whether work is dispatched safely and predictably. > - Several accumulated fixes all touched agent execution recovery, wake routing, import behavior, and runtime concurrency defaults. > - Those changes need to land together so the heartbeat service and agent creation defaults stay internally consistent. > - This pull request groups the runtime/governance changes from the split branch into one standalone branch. > - The benefit is safer recovery for stranded runs, bounded high-volume reads, imported-agent approval correctness, skill-template support, and a clearer default concurrency policy. ## What Changed - Fixed stranded continuation recovery so successful automatic retries are requeued instead of incorrectly blocking the issue. - Bounded high-volume issue/log reads across issue, heartbeat, agent, project, and workspace paths. - Fixed imported-agent approval and instruction-path permission handling. - Quarantined seeded worktree execution state during worktree provisioning. - Queued approval follow-up wakes and hardened SQL_ASCII heartbeat output handling. - Added reusable agent instruction templates for hiring flows. - Set the default max concurrent agent runs to five and updated related UI/tests/docs. ## Verification - `pnpm install --frozen-lockfile` - `pnpm exec vitest run server/src/__tests__/company-portability.test.ts server/src/__tests__/heartbeat-process-recovery.test.ts server/src/__tests__/heartbeat-comment-wake-batching.test.ts server/src/__tests__/heartbeat-list.test.ts server/src/__tests__/issues-service.test.ts server/src/__tests__/agent-permissions-routes.test.ts packages/adapter-utils/src/server-utils.test.ts ui/src/lib/new-agent-runtime-config.test.ts` - Split integration check: merged this branch first, followed by the other [PAP-1614](/PAP/issues/PAP-1614) branches, with no merge conflicts. - Confirmed this branch does not include `pnpm-lock.yaml`. ## Risks - Medium risk: touches heartbeat recovery, queueing, and issue list bounds in central runtime paths. - Imported-agent and concurrency default behavior changes may affect existing automation that assumes one-at-a-time default runs. - No database migrations are included. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5.4 tool-enabled coding model, agentic code-editing/runtime with local shell and GitHub CLI access; exact context window and reasoning mode are not exposed by the Paperclip harness. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] If this change affects the UI, I have included before/after screenshots - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b9a80dcf22 |
feat: implement multi-user access and invite flows (#3784)
## Thinking Path > - Paperclip is the control plane for autonomous AI companies. > - V1 needs to stay local-first while also supporting shared, authenticated deployments. > - Human operators need real identities, company membership, invite flows, profile surfaces, and company-scoped access controls. > - Agents and operators also need the existing issue, inbox, workspace, approval, and plugin flows to keep working under those authenticated boundaries. > - This branch accumulated the multi-user implementation, follow-up QA fixes, workspace/runtime refinements, invite UX improvements, release-branch conflict resolution, and review hardening. > - This pull request consolidates that branch onto the current `master` branch as a single reviewable PR. > - The benefit is a complete multi-user implementation path with tests and docs carried forward without dropping existing branch work. ## What Changed - Added authenticated human-user access surfaces: auth/session routes, company user directory, profile settings, company access/member management, join requests, and invite management. - Added invite creation, invite landing, onboarding, logo/branding, invite grants, deduped join requests, and authenticated multi-user E2E coverage. - Tightened company-scoped and instance-admin authorization across board, plugin, adapter, access, issue, and workspace routes. - Added profile-image URL validation hardening, avatar preservation on name-only profile updates, and join-request uniqueness migration cleanup for pending human requests. - Added an atomic member role/status/grants update path so Company Access saves no longer leave partially updated permissions. - Improved issue chat, inbox, assignee identity rendering, sidebar/account/company navigation, workspace routing, and execution workspace reuse behavior for multi-user operation. - Added and updated server/UI tests covering auth, invites, membership, issue workspace inheritance, plugin authz, inbox/chat behavior, and multi-user flows. - Merged current `public-gh/master` into this branch, resolved all conflicts, and verified no `pnpm-lock.yaml` change is included in this PR diff. ## Verification - `pnpm exec vitest run server/src/__tests__/issues-service.test.ts ui/src/components/IssueChatThread.test.tsx ui/src/pages/Inbox.test.tsx` - `pnpm run preflight:workspace-links && pnpm exec vitest run server/src/__tests__/plugin-routes-authz.test.ts` - `pnpm exec vitest run server/src/__tests__/plugin-routes-authz.test.ts server/src/__tests__/workspace-runtime-service-authz.test.ts server/src/__tests__/access-validators.test.ts` - `pnpm exec vitest run server/src/__tests__/authz-company-access.test.ts server/src/__tests__/routines-routes.test.ts server/src/__tests__/sidebar-preferences-routes.test.ts server/src/__tests__/approval-routes-idempotency.test.ts server/src/__tests__/openclaw-invite-prompt-route.test.ts server/src/__tests__/agent-cross-tenant-authz-routes.test.ts server/src/__tests__/routines-e2e.test.ts` - `pnpm exec vitest run server/src/__tests__/auth-routes.test.ts ui/src/pages/CompanyAccess.test.tsx` - `pnpm --filter @paperclipai/shared typecheck && pnpm --filter @paperclipai/db typecheck && pnpm --filter @paperclipai/server typecheck` - `pnpm --filter @paperclipai/shared typecheck && pnpm --filter @paperclipai/server typecheck` - `pnpm --filter @paperclipai/ui typecheck` - `pnpm db:generate` - `npx playwright test --config tests/e2e/playwright.config.ts --list` - Confirmed branch has no uncommitted changes and is `0` commits behind `public-gh/master` before PR creation. - Confirmed no `pnpm-lock.yaml` change is staged or present in the PR diff. ## Risks - High review surface area: this PR contains the accumulated multi-user branch plus follow-up fixes, so reviewers should focus especially on company-boundary enforcement and authenticated-vs-local deployment behavior. - UI behavior changed across invites, inbox, issue chat, access settings, and sidebar navigation; no browser screenshots are included in this branch-consolidation PR. - Plugin install, upgrade, and lifecycle/config mutations now require instance-admin access, which is intentional but may change expectations for non-admin board users. - A join-request dedupe migration rejects duplicate pending human requests before creating unique indexes; deployments with unusual historical duplicates should review the migration behavior. - Company member role/status/grant saves now use a new combined endpoint; older separate endpoints remain for compatibility. - Full production build was not run locally in this heartbeat; CI should cover the full matrix. ## Model Used - OpenAI Codex coding agent, GPT-5-based model, CLI/tool-use environment. Exact deployed model identifier and context window were not exposed by the runtime. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] If this change affects the UI, I have included before/after screenshots - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] I will address all Greptile and reviewer comments before requesting merge Note on screenshots: this is a branch-consolidation PR for an already-developed multi-user branch, and no browser screenshots were captured during this heartbeat. --------- Co-authored-by: dotta <dotta@example.com> Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> |
||
|
|
7f893ac4ec |
[codex] Harden execution reliability and heartbeat tooling (#3679)
## Thinking Path > - Paperclip orchestrates AI agents for zero-human companies > - Reliable execution depends on heartbeat routing, issue lifecycle semantics, telemetry, and a fast enough local verification loop to keep regressions visible > - The remaining commits on this branch were mostly server/runtime correctness fixes plus test and documentation follow-ups in that area > - Those changes are logically separate from the UI-focused issue-detail and workspace/navigation branches even when they touch overlapping issue APIs > - This pull request groups the execution reliability, heartbeat, telemetry, and tooling changes into one standalone branch > - The benefit is a focused review of the control-plane correctness work, including the follow-up fix that restored the implicit comment-reopen helpers after branch splitting ## What Changed - Hardened issue/heartbeat execution behavior, including self-review stage skipping, deferred mention wakes during active execution, stranded execution recovery, active-run scoping, assignee resolution, and blocked-to-todo wake resumption - Reduced noisy polling/logging overhead by trimming issue run payloads, compacting persisted run logs, silencing high-volume request logs, and capping heartbeat-run queries in dashboard/inbox surfaces - Expanded telemetry and status semantics with adapter/model fields on task completion plus clearer status guidance in docs/onboarding material - Updated test infrastructure and verification defaults with faster route-test module isolation, cheaper default `pnpm test`, e2e isolation from local state, and repo verification follow-ups - Included docs/release housekeeping from the branch and added a small follow-up commit restoring the implicit comment-reopen helpers that were dropped during branch reconstruction ## Verification - `pnpm vitest run server/src/__tests__/issue-comment-reopen-routes.test.ts server/src/__tests__/issue-telemetry-routes.test.ts` - `pnpm vitest run server/src/__tests__/http-log-policy.test.ts server/src/__tests__/heartbeat-run-log.test.ts server/src/__tests__/health.test.ts` - `server/src/__tests__/activity-service.test.ts`, `server/src/__tests__/heartbeat-comment-wake-batching.test.ts`, and `server/src/__tests__/heartbeat-process-recovery.test.ts` were attempted on this host but the embedded Postgres harness reported init-script/data-dir problems and skipped or failed to start, so they are noted as environment-limited ## Risks - Medium: this branch changes core issue/heartbeat routing and reopen/wakeup behavior, so regressions would affect agent execution flow rather than isolated UI polish - Because it also updates verification infrastructure, reviewers should pay attention to whether the new tests are asserting the right failure modes and not just reshaping harness behavior ## Model Used - OpenAI Codex coding agent (GPT-5-class runtime in Codex CLI; exact deployed model ID is not exposed in this environment), reasoning enabled, tool use and local code execution enabled ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] If this change affects the UI, I have included before/after screenshots - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
be82a912b2 | Fix signoff e2e for auto-checked out issues | ||
|
|
b0b85e6ba3 | Stabilize onboarding e2e cleanup paths | ||
|
|
cb705c9856 | Fix signoff PR follow-up tests | ||
|
|
42b326bcc6 |
fix(e2e): harden signoff policy tests for authenticated deployments
Address QA review feedback on the signoff e2e suite (86b24a5e): - Use dedicated port 3199 with local_trusted mode to avoid reusing the dev server in authenticated mode (fixes 403 errors) - Add proper agent authentication via API keys + heartbeat run IDs - Fix non-participant test to actually verify access control rejection - Add afterAll cleanup (dispose contexts, revoke keys, delete agents) - Reviewers/approvers PATCH without checkout to preserve in_review state Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
97d4ce41b3 |
test(e2e): add signoff execution policy end-to-end tests
Covers the full signoff lifecycle: executor → review → approval → done, changes-requested bounce-back, comment-required validation, access control, and review-only policy completion. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
036e2b52db |
Merge pull request #1326 from paperclipai/ci/pr-e2e-tests
ci: run e2e tests on PRs |