Commit Graph
37 Commits
Author SHA1 Message Date
TonioandClaude Fable 5 141c5b1340 feat(onboarding): Figma pass over the tenant arc (#12726)
Takes the four tenant onboarding steps to the design, and makes the model
choice explicit.

**Nothing is preselected on the connect step.** It arrived with a source
already chosen, which made the row read as a confirmation rather than a
question and let a customer pass the step having touched none of it. The
step now opens unanswered and cannot advance until a source is selected in
the visible row.

Two defects fell out of that, both found by Greptile in review:

- A saved draft can name an adapter the row does not show — one the registry
  dropped, or one in the advanced list. The gate asked `sourcePicked`, which
  only means "a draft named something", so Next stayed live on a step that
  had visibly asked nothing and would hire against the hidden name. It now
  asks `sourceSelected`, read off the visible row, and the snap that replaces
  an unofferable adapter clears the pick rather than presenting its
  replacement as chosen.
- Cmd+Enter bypassed that gate entirely, because the condition was written
  out twice and drifted. Both paths now ask one predicate.

Also: selection is a fill rather than a border; the input canvas opens only
for a chosen source; the close control is gone from every step (nothing
downstream of the connect step works until a model is connected); the arc
draws to the design's 424px column with a filled 44px name field; the CTA
reads "Next" through the arc and "Get started" at the end; the login spinner
reads "Preparing..."; the OAuth panels number their fields and show the URL
above the code.

Storybook: the arc stories waited for a button named "Connect" and had been
silently rendering the wrong step since the CTA was renamed. They now wait on
the destination heading. Three e2e specs had the same fault and now select a
source before advancing.

One test was removed rather than replaced — `hydrates again when the same
company comes back through onboarding` drove the close control, and three
substitutes each passed against a wizard with the behaviour deleted. The
reason is recorded where it stood.

Known and not fixed here: SidebarCompanyMenu opens the wizard at step 1 for
"create a new organization", and with no exit an existing user who changes
their mind is trapped.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-09-02 12:19:15 -07:00
TonioandClaude Fable 5 889f5853f9 Onboarding: count the whole walk on a cloud tenant (#12706)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - A cloud customer meets it as one walk that crosses two applications:
Cloud names the organization, then hands off to the tenant for the
agent, the model and the review
> - The tenant's progress strip counted only its own three steps, so the
count restarted at the crossing — four dots on the naming screen, then
three beginning again at one
> - The strip that spans the whole walk already existed and simply was
not chosen on that path, which made the walk read as two products rather
than one
> - This pull request picks the four-step strip when the run came from
Cloud, so the count runs 1, 2, 3, 4 straight through
> - The benefit is that the hand-off stops announcing itself

## Linked Issues or Issue Description

**What existing behavior does this improve?**
The onboarding progress strip on a cloud tenant. It restarted its count
at the hand-off from Cloud, so a customer went from "1 of 4" to "1 of 3"
mid-walk.

**Subsystem affected**
Tenant onboarding wizard — `ui/src/components/OnboardingWizard.tsx`.

**Current behavior**
`showsAgentArcStepper` is `isAgentArcStep && entryStep >= 3`. Every run
entering on the agent step gets the three-step strip, including a cloud
tenant whose organization was named one screen earlier in Cloud.

**Proposed behavior**
A cloud tenant gets the four-step strip, positioned 2, 3 and 4. A
self-hosted run entering on the same step keeps three.

**Reason and benefit**
The count describes the walk the customer is on rather than the half of
it this application happens to render.

**Breaking changes**
None. No API or schema change; the self-hosted path is unchanged.

## What Changed

- Adds `enteredFromCloud`, read from `enableManagedSandboxOnly` — the
cloud-tenant shape the connect step already resolves its login
environment through.
- Excludes that case from `showsAgentArcStepper`, so it falls to the
existing four-step strip. `ONBOARDING_STEP_LABELS` and
`onboardingStepPositionFor` already produce 2, 3 and 4; neither needed
changing.
- Widens the experimental-settings query from step 4 to steps 3–5, so
the strip can read it on every step of the arc.
- Adds tests for both strip lengths.

## Verification

- `pnpm vitest run src/components src/adapters storybook` in `ui/` —
2309 tests pass.
- `pnpm typecheck` in `ui/` — clean.
- Storybook, `Onboarding/Agent arc`: its fixture is cloud-shaped, so the
three steps now announce "Step 2 of 4", "Step 3 of 4" and "Step 4 of 4"
with four dots.
- The two new tests fail without the change — the cloud case reports
`expected 'Step 1 of 3' to be 'Step 2 of 4'`.

## Risks

- **The two runs entering on the agent step are genuinely different, and
only one signal separates them.** A cloud tenant walked a naming screen
in Cloud; a self-hosted company with no agents walked nothing before
this. Getting it backwards would either restart the count or credit a
step nobody walked. Both directions are now tested, which they were not
before — the suite covered `agentArcStepFor` but nothing asserted which
strip renders.
- `enableManagedSandboxOnly` is being read as a proxy for "came from
Cloud". It is the same signal the connect step already trusts for its
environment, so this does not introduce a new dependency, but it is
inference rather than a fact passed at the hand-off. If the step count
ever needs to vary, Cloud should seed the position explicitly instead.
- The widened query means the value can be unresolved on first paint,
and the strip briefly renders three dots before settling to four. The
query key is shared and usually already cached, so this is rare rather
than routine — worth watching on staging, and the fix if it shows is to
hold the strip until the value resolves.
- Cosmetic only. Nothing here changes which steps run, what they
collect, or where the walk goes.

## Model Used

Claude Opus 5 (`claude-opus-5`), via Claude Code with extended thinking,
tool use, and browser-driven verification of the rendered strip.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [ ] I will address all Greptile and reviewer comments before
requesting merge

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-09-01 18:06:38 -07:00
TonioandClaude Opus 5 42c6f8a424 Onboarding: model source tiles, one input canvas, and Storybook coverage for the agent arc (#12613)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - New customers meet it through onboarding, whose last three steps run
inside the tenant: create an agent, connect a model, review
> - The connect step is the one that decides whether the agent can run
at all, and it had drifted — three contributors changed it in parallel,
and its visual language no longer matched the rest of the flow
> - It also could not be looked at without a provisioned stack, so
defects in it were only found by walking a real signup, and the review
step behind it could not be reached at all when it failed
> - This pull request brings the visual work onto the sign-in behaviour
that already shipped, and adds Storybook coverage for all three steps
> - The benefit is that the step is easier to read, and that it can now
be inspected and driven before it ships rather than after

## Linked Issues or Issue Description

**What existing behavior does this improve?**
The onboarding connect-a-model step. It presented the model choice as a
dropdown plus an "Advanced settings" disclosure, and put each credential
type in a different place, so the controls below moved whenever the
choice changed.

**Subsystem affected**
Tenant onboarding wizard (`ui/src/components/OnboardingWizard.tsx`) and
its Storybook coverage.

**Current behavior**
The step offered every registered adapter through a disclosure.
Credential entry appeared in a different shape per source. None of the
three agent-arc steps could be rendered outside a provisioned cloud
stack, so the sign-in panel and the review step were only reachable by
walking a real signup.

**Proposed behavior**
Two brand tiles for the recommended sources, a link that switches
between subscription and API-key credentials, and one canvas that holds
whichever input the current choice needs. Storybook stories mount the
real wizard against fixtures and walk it forward, so every step and its
states can be inspected locally.

**Reason and benefit**
The step reads as one decision rather than three scattered ones, and its
furniture stays still while the choice changes. The stories mean a
regression in it is visible before release instead of during a signup.

**Breaking changes**
No API or schema change. One behavioural narrowing, described under
Risks.

## What Changed

- Replaces the adapter dropdown and "Advanced settings" disclosure with
`ModelSourceTiles` — brand tiles for Claude Code and Codex.
- Adds `CredentialModeLink`, a text toggle between subscription sign-in
and API keys, replacing the disclosure.
- Adds `ConnectInputCanvas`: one surface that holds the sign-in panel or
the API-key field and resizes between them, so the Connect button below
does not move.
- Keeps the existing sign-in behaviour unchanged. `AgentConfigForm`
changes are presentation only — the provider name in the title, the CTA
wording, and `space-y` to `gap`. No change to the login mutations,
queries, or session handling.
- Restores the sleep marks on the dormant agent for the two steps before
the hire.
- Adds Storybook stories for all three agent-arc steps, with fixtures
for the environments, auth signal, both adapters' login flows, and the
hire.
- Copy: names the provider being signed in to ("Sign in to
Anthropic"/"Sign in to OpenAI"), and drops "Clippy" from the agent-name
helper text.

## Verification

- `pnpm vitest run src/components storybook` in `ui/` — 139 tests over
the touched suites, 1986 across `src/components`.
- `pnpm typecheck` in `ui/` — clean.
- Storybook, `Onboarding/Agent arc`: walk each story. Step 1 has no Back
button, steps 2 and 3 do.
- The sign-in gate: on `Connect a model`, press Connect without signing
in. It holds on step 2 and reports "No working authentication was
found." On `Review`, which fixtures an authenticated signal, Connect
reaches the review step.
- Both providers' login flows: press Sign in on the Claude tile for the
authorization URL and browser-code field, and on the Codex tile for the
device URL and code.
- The Claude sign-in was also walked end to end on a staging tenant,
including the OAuth redirect and pasting the code back.

## Risks

- **Onboarding now offers two model sources instead of every registered
adapter.** `ModelSourceTiles` is fed the `recommended` set, which is
`claude_local` and `codex_local`; Gemini, Cursor, Grok, Kimi, OpenCode
and Paperclip Runner are no longer selectable *during onboarding*. This
is deliberate. The full list is unchanged in agent settings, which is
where an adapter can still be switched after the agent exists, and
adding a source back is one `recommended: true` in
`adapter-display-registry.ts`. Flagging it because it is the one
behavioural narrowing here and it is not visible from the diffstat.
- The API key entered on this step is held in component state and
deliberately never written to the onboarding draft, because that draft
is `localStorage`. A customer who leaves mid-step re-enters the key;
that is the intended trade.
- Storybook-only risk: the fixtures now answer the environment test from
the story's auth state. If a future change moves the hire's gate off the
`adapter_auth_missing` check code, the stories would keep passing while
the product regressed. The gate is asserted in the adapter packages' own
tests, not here.
- Motion changes are low risk and reversible: the input canvas animates
its contents only, and its container was deliberately left unanimated
after an animated wrapper clipped the sign-in panel.

## Model Used

Claude Opus 5 (`claude-opus-5`), via Claude Code with extended thinking,
tool use, and browser-driven verification of the Storybook stories.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [ ] I will address all Greptile and reviewer comments before
requesting merge

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-01 09:57:46 -07:00
TonioandClaude Fable 5 4436cf00a2 Render the onboarding agent arc's real steps in Storybook (#12369)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - New tenants meet it through an onboarding wizard: create your agent,
connect a model, review
> - Those three screens only render for a signed-in account that owns a
provisioned stack, so the only way to look at one was to walk a real
signup
> - Round 4 redesigned all three and shipped them without anyone seeing
them render; when the connect step then failed on a live stack, the
review step behind it could not be reached at all
> - Storybook already mounts the app's provider stack and stubs `/api`,
and already has an `Onboarding/Agent arc` story — but that story
previews `AgentCapsule`, which the wizard stopped using in round 4
> - This pull request mounts the real wizard in Storybook at each step,
and replaces the stale capsule stories with the component the arc
actually renders
> - The benefit is that these screens can be reviewed, and regressions
seen, without provisioning anything

## Linked Issues or Issue Description

No existing issue. Describing it inline, following
`.github/ISSUE_TEMPLATE/enhancement.yml`.

**What existing behavior does this improve?**

Reviewing the tenant onboarding wizard. Today it can only be seen by
signing up for a real account and provisioning a real stack.

**Subsystem affected**

`ui` — the onboarding wizard and its Storybook coverage.

**Current behavior**

`OnboardingWizard` renders only for a signed-in account that owns a
company. There is no route, harness, or story that mounts it, so no
screen in the agent arc can be looked at in isolation. The existing
`Onboarding/Agent arc` story previews `AgentCapsule` in
`slot`/`configured`/`online` and describes it as what "the wizard holds
in one tree slot" — round 4 replaced that with `PillGuy` and
`dormant`/`alive`, so the story documents a component the arc no longer
renders.

**Proposed behavior**

Stories that mount the real `OnboardingWizard` at each of the three
steps against the existing Storybook API fixtures, plus stories for
`PillGuy` and its transition.

**Reason and benefit**

Round 4 shipped three redesigned screens that nobody could see. The
connect step then failed on staging, which made the review step
unreachable even with an account — reviewing it required hand-editing
`localStorage`. Stories remove that whole class of problem.

**Breaking changes**

None. Storybook-only; no product code is touched.

## What Changed

- `CreateYourAgent`, `ConnectAModel`, `Review` — the real wizard, per
step.
- `PillStates` and `PillMorph` — the two states, and the transition on a
loop with a toggle. The morph is the arc's payoff and the hardest thing
to judge from a still.
- Replaces the `AgentCapsule` stories in this file with `PillGuy`.
`AgentCapsule` is still used by `DesignGuide` and keeps its coverage
there.
- Four routes added to the Storybook fetch fixtures:
`/api/instance/settings`, `…/environments`, `…/adapters/:type/models`.
The empty environment list is the cloud-tenant shape, and also the state
that produces the "no managed sandbox environment is available" notice —
worth being able to look at rather than only meeting it on a live stack.

## Three properties of the wizard the stories had to respect

Each of these cost a debugging cycle, so they are documented at the call
site:

1. **The draft is seeded during render, not in an effect.** Roughly
twenty `useState(saved?.x ?? default)` initializers read the restored
blob exactly once, so a draft written after mount arrives too late.
2. **Nothing mounts until the companies list settles.** The wizard's own
mount gate waits on `isFetching`, but that query is *disabled* until the
account settles, and a disabled query is not fetching. Mounting straight
away gets an inner wizard that reads a null draft, falls back to
`initialStep`, and then persists that back over the seed. A real session
never hits this because the dashboard has already loaded the list.
3. **The review step opens with no `initialStep`.** An explicit option
takes precedence over saved state by design, so passing one clamps 5 to
4 and lands on Connect.

## Verification

- `npx vitest run src/components/OnboardingWizard.test.tsx
src/components/OnboardingWizard.step.test.tsx` — 44 tests, all passing.
- `npx tsc -p tsconfig.json --noEmit` — no new errors.
(`src/lib/sentry.*` reports pre-existing missing-type errors for
`@sentry/browser` on master.)
- Each of the three step stories loaded in a running Storybook and read
back:
- `create-your-agent` → "STEP 1 OF 3 / Create your first agent / Name /
Next"
- `connect-a-model` → "STEP 2 OF 3 / Connect a model / Paperclip works
with your existing subscription or API keys. / Claude Code / Codex /
Advanced settings / Connect"
- `review` → "STEP 3 OF 3 / Let's get started... / Darnold is ready to
work! / Get started"

One difference from a real walk, worth knowing before treating a story
as ground truth: the review story has no Back button, because
`entryStep` is 5 there and back-navigation is bounded by where the run
entered.

## Risks

Low. Storybook-only — no product code, routes, or bundles change. The
four added fetch fixtures are inside the Storybook mock and cannot
affect the app.

The one thing to watch: these stories mount the real wizard, so a future
change to how it restores drafts or decides its initial step can break
them. That is arguably the point — the stories would be the first place
it shows — but it does mean they are coupled to internals rather than to
a prop surface, and the three notes above are what a maintainer needs.

## Model Used

Claude Fable 5 (`claude-fable-5`), extended thinking, with tool use:
file editing, shell, and browser automation for loading and reading back
each story.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [ ] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-27 16:53:12 -07:00
TonioandClaude Opus 5 eb86fcd498 The agent, drawn as itself (#12274)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - A new customer's first session ends in the tenant's agent arc:
create an agent, connect a model, review
> - Walking it turned up questions the arc had no business asking — a
role picker using a vocabulary the customer has not been given, a model
picker asking them to judge models they have not met — and chrome
restating what they had just watched happen
> - Each one costs a first-session customer attention at the exact
moment they are deciding what this product is
> - This pull request cuts the arc to what it must ask, and draws the
agent as itself so the arc has a visible subject
> - The benefit is three steps that each ask one thing, ending on an
agent that is visibly ready

## Linked Issues or Issue Description

No public issue exists. The changes come from walking the sign-up arc
end to end.

**What happened:**
The agent step asks for a role from a fixed enum before asking for a
name. The model step shows two "Recommended" badges (on both options),
an "Adapter type" eyebrow, and a model picker. The review step lists a
three-row checklist of work the customer just performed. The progress
strip is a full-width segmented bar.

**Expected behavior:**
The agent step asks for a name. The model step offers the two harnesses
and hides the rest behind advanced settings. The review step says the
agent is ready. The strip counts three discrete steps.

**Steps to reproduce:**
1. Sign up and enter the tenant wizard on the agent arc.
2. Observe the role select above the optional name field.
3. Continue to the model step: both options carry a "Recommended" badge,
and a model picker sits below.
4. Continue to review: a checklist restates the organization name,
agent, and model.

**Additional context:**
The brand pill assets (`pill-1-dormant.svg`, `pill-1-alive.svg`) are
transcribed verbatim into a component rather than approximated. The role
removal exposed a latent silent-failure path — see Risks.

## What Changed

- `PillGuy` renders the brand pill in two states; the arc holds one
instance, dormant through create and connect, alive on review.
- The agent step asks for a name only. The name is required; the role
picker is gone.
- `DEFAULT_AGENT_ROLE` (`general`) backs every onboarding hire, and
`agentRole` now defaults to it rather than empty.
- The model step drops both "Recommended" badges, the "Adapter type"
eyebrow, and the model picker; "More Agent Adapter Types" becomes
"Advanced settings"; the sub-line becomes "Paperclip works with your
existing subscription or API keys."
- The review step drops its checklist; the heading becomes "Let's get
started..." with "[name] is ready to work!".
- The progress strip renders three left-aligned dots at the previous
gap.
- Five e2e specs and both wizard unit suites migrate off
`#onboarding-agent-role`.

## Verification

Run the tenant suite:

```
cd ui && npx vitest run
```

- 4398 tests pass across 474 files; `npx tsc --noEmit` clean.
- Walked live in a local instance: agent step (dots, dormant pill, name
placeholder), model step (no badges/eyebrow/picker, "Advanced
settings"), review (pill alive, new copy, no checklist).
- The retargeted role test asserts the hire payload carries `role:
"general"` and the typed name — it is the test that catches the silent
failure below.

## Risks

- **A latent silent failure, now closed.** `handleGiveHeartbeat` returns
early when `agentRole` is empty. With the picker removed and no default,
Connect would have hired nobody and shown no error. The default closes
it; the guard stays for any future path that clears the role.
- **Behavioral change:** every onboarding hire is filed as `general`
rather than a chosen role. The role remains editable in the app.
- **Behavioral change:** the model is no longer chosen during
onboarding. Every adapter offered here resolves its own default in
`buildAdapterConfig`, and the model is changeable later.
- **Assets:** the pill carries its own gradient fills and does not
follow the theme. That is deliberate — the agent looks like itself on
either ground.

## Model Used

Claude Opus 5 (`claude-opus-5`) via Claude Code, with tool use and code
execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links
- [x] My branch name describes the change and contains no internal
Paperclip ticket id
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [ ] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-26 21:39:55 -07:00
TonioandClaude Opus 5 88a0f885e8 Brand lockup, no idle env-check card, no Mission row (#12074)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - New customers meet the product through an onboarding arc that ends
in the tenant wizard
> - A staging walk of that arc found three rough edges: an ad-hoc icon
standing in for the brand, an environment-check card that narrates a
probe the flow already runs on its own, and a review checklist that
still lists a Mission the arc stopped asking for
> - Each one makes the product look less finished than it is at the
exact moment a customer decides what it is
> - This pull request renders the brand lockup, hides the idle
environment-check card while keeping the probe and its failure surface,
and drops the Mission row
> - The benefit is a first-session arc that reads as one product, with
no controls for questions nobody was asked

## Linked Issues or Issue Description

No public issue exists. The changes come from walking the sign-up arc on
a staging fleet.

**What happened:**
The model step shows an "Adapter environment check" card with a "Test
now" button even though pressing Connect runs the same probe and blocks
a failing hire. The review step lists "Mission" in its checklist
although onboarding no longer asks for one. The auth page renders a
sparkles icon beside the word "Paperclip" instead of the brand lockup.

**Expected behavior:**
The model step shows the check only when a probe has found something to
fix. The review checklist lists only what onboarding set up. The brand
renders as the lockup asset used across surfaces.

**Actual behavior:**
An idle card narrates a probe that runs regardless. A permanent
unchecked row marks a question nobody was asked. The brand is a generic
icon plus text.

**Steps to reproduce:**
1. Sign up on a staging fleet and enter the tenant wizard.
2. On "Create your first agent", choose a role and press Next: the model
step shows the "Adapter environment check" card before anything has been
probed.
3. Continue to Review: the checklist lists "Mission" as a permanently
unchecked row.

**Additional context:**
The Mission row outlived the removal of the mission step (#11935). The
environment probe itself still runs on Connect and blocks a failing
hire; only its idle card is at issue. The brand lockup lands across all
three surfaces in the same round — paperclip-cloud#270 and
paperclip-id#58 carry the other halves.

## What Changed

- `PaperclipLockup` renders the brand asset (mark + wordmark, one
geometry, `fill="currentColor"`); the auth page uses it in place of the
sparkles icon.
- The adapter environment check's idle card (explainer + "Test now") no
longer renders. The probe still runs on Connect and still blocks a
failing hire.
- The check's failure content still renders when a probe has found
something — the blocking error points the customer at "the reported
checks", so they stay visible.
- Connect retries a cached failed probe instead of reusing it. With
"Test now" gone, Connect is the only retry, and a stale fail would lock
out a machine the customer has since fixed.
- The review checklist drops its "Mission" row.

## Verification

Run the tenant suite:

```
cd ui && npx vitest run
```

- 4365 tests pass across 471 files; `npx tsc --noEmit` is clean.
- New test drives the wizard to the model step and asserts the
environment-check card is absent, anchored on "Connect a model" so an
unrendered step cannot pass as an absence.
- The review assertions anchor on the remaining rows ("Organization
name", "Agent created", "Model connected").

## Risks

- **Behavioral change:** a cached *failed* probe is re-run on Connect
instead of reused. Pass and warn results are still cached. This only
affects the retry path that "Test now" used to serve.
- **Hidden, not removed:** the environment check machinery is intact;
only the idle card is gone. A failing probe still blocks the hire and
still shows its checks.
- **Brand:** the wordmark now ships inside an SVG; its accessible name
carries the text. Screen readers announce "Paperclip" as before.

## Model Used

Claude Fable 5 (`claude-fable-5`) via Claude Code, with tool use and
code execution; earlier rounds on this branch's predecessor used Claude
Opus 5 (`claude-opus-5`).

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links
- [x] My branch name describes the change and contains no internal
Paperclip ticket id
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [ ] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-24 01:35:15 -07:00
TonioandClaude Opus 5 3ff636bc48 Drop the mission step from the wizard arc (#11935)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - New customers arrive through an onboarding arc that spans Paperclip
Cloud and the tenant app
> - Cloud's naming screen stopped asking for the company mission, but
the tenant wizard still decided its first step by asking whether the
company had one
> - Every Cloud-created company therefore looked mission-less on
arrival, so every walk detoured through a "Define your mission" screen
the design had already removed
> - This pull request removes that step from the arc, and makes step 1
create the company itself
> - The benefit is a shorter arc that matches the design, and a
three-step progress strip that now counts the steps that exist

## Linked Issues or Issue Description

No public issue exists. The problem was found by walking staging end to
end.

**What happened:**
A new customer signs in, names their organization, and waits for it to
build. The tenant wizard then asks "Define your mission" before it asks
for the first agent. Cloud no longer collects a mission, so this screen
appears for every new customer.

**Expected behavior:**
The wizard asks for the first agent, the model, and a review. The
progress strip counts three steps.

**Actual behavior:**
The wizard asks for the mission first. The progress strip counts five
segments, because the run does not enter on the agent arc.

**Additional context:**
Three merged pull requests built the mission-based step choice this
change removes: #11352, #11416 and #11429. The mission is now collected
later, inside the tenant app, so onboarding does not ask for it at all.

## What Changed

- `onboardingStepForCompany` always returns the agent step. The
`companyHasMission` parameter is removed, because it cannot change the
answer.
- `resolveRouteOnboardingOptions` no longer accepts `companyHasMission`.
- The dashboard no longer waits for the goal lookup before it opens the
wizard. That wait only chose a step, and the step is now fixed.
- Step 1 creates the company in a new `handleCreateCompany`. Company
creation used to sit at the end of `handleConfirmMission`.
- No company goal is written during onboarding.
- The three-step strip now shows on the agent, model and review steps,
because every Cloud-first run enters on the arc.
- The full-length bar drops its second segment. No run can fill it.
- The grow path keeps its step 2 questionnaire. Only the create path
skips ahead.
- Back from the agent step goes to the screen the run came from.
- Four end-to-end specs no longer drive the wizard through the mission
step.

## Verification

Run the tenant test suite:

```
cd ui && npx vitest run
```

- 4356 tests pass. 471 files pass.
- `npx tsc --noEmit` reports no errors.
- Fault injection: forcing `skipsMissionStep` to `true` fails the grow
questionnaire test. Removing the Back rule fails the Back test. Both
tests fail on the exact defect they guard.
- The three-step strip is asserted by an existing test. It checks `Step
1 of 3` and `aria-label="Create your first agent"`.

Manual check on staging after the paired Cloud change:

1. Open a new incognito window.
2. Sign in with a new account.
3. Name the organization.
4. Confirm the wizard shows "Create your first agent" and "Step 1 of 3".

## Risks

- **Behavioral change.** Onboarding no longer writes a company goal. An
agent hired during onboarding starts without a seeded mission. This is
intended. The mission moves to the tenant app.
- **Dead code.** `ONBOARDING_MISSION_STEP` and the mission screen stay
in the codebase, but nothing in the app opens them. They wait for the
surface that collects the mission later.
- **Grow path.** The grow path is unchanged, but it shares step 2 with
the removed screen. New tests cover it.
- **Superseded work.** #11352, #11416 and #11429 tuned the mission-based
step choice. This change removes the branch they tuned.

## Model Used

Claude Opus 5 (`claude-opus-5`), extended thinking, with tool use and
code execution through Claude Code.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links
- [x] My branch name describes the change and contains no internal
Paperclip ticket id
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [ ] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-22 02:28:45 -07:00
TonioandClaude Opus 5 24913064ff feat(commitperclip): surface the Co-Authored-By trailers a squash merge needs (#11498)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work, and it takes contributions from outside the core team
> - Those contributions arrive as PRs, and this repository squash-merges
every one of them
> - A squash collapses the whole branch into a single commit authored by
whoever pressed the button
> - So when a maintainer rebases and lands a contributor's stale PR, the
contributor's name survives only if the squash message carries a
`Co-Authored-By` trailer
> - Nothing prompts for that trailer, and the PR page keeps showing the
original author either way, so losing it is invisible at the moment it
happens
> - This pull request has commitperclip detect the situation and print
the exact trailers to paste
> - The benefit is that keeping an outside contributor's name is a
default rather than something a maintainer has to remember

## Linked Issues or Issue Description

No public issue exists. The problem follows, and it is not hypothetical.

**What happened?**

#11370, #11371 and #11379 landed @stubbi's work yesterday. Each of those
PRs carries a comment from me telling them their authorship would be
preserved. All three squash commits went in without a `Co-Authored-By`
trailer, so `git log` credits none of them:

| commit | landed from | credited |
| --- | --- | --- |
| `66515582e` | #9900 | Claude only |
| `bc0b5a164` | #9501 | Claude only |
| `35a9b9873` | #8982 | Claude only |
| `6542ad1f4` | #11259 | ✅ Jannes Stubbemann + Claude |

The last one has the trailer because that message was written by hand
with the contributor in mind. The only difference between the two
outcomes was memory. Master history cannot be rewritten, so those three
are now credited by comment on the original PRs — which is a worse
record than a commit trailer, and the reason to make this automatic.

**Expected behavior**

When a branch carries commits by someone other than the PR author, the
merger is told what trailers the squash needs.

**Paperclip version or commit**

`master` at `92047cac4`.

## What Changed

- `.github/scripts/check-pr-coauthors.mjs` — new gate.
- `.github/scripts/run-quality-gates.mjs` — fetches the PR's commits and
runs it.
- `.github/scripts/tests/check-pr-coauthors.test.mjs` — 12 cases.
- `.github/workflows/pr.yml` — runs `.github/scripts/tests/`.

### Informational, not a failure

The squash message does not exist while the PR is open. This can neither
be verified there nor fixed there, so failing a PR on it would block
work on something its author cannot satisfy. The gate notices that the
situation applies and prints the lines to paste.

Run against #11370's actual commits it produces exactly what was
missing:

```
This branch carries commits by stubbi. Squash-merging drops that authorship
unless the squash message carries their trailers, and nothing else will notice
if it does not. Add to the squash body when merging:

      Co-Authored-By: Jannes Stubbemann <stubbi@users.noreply.github.com>
```

### Edge cases it handles

Bots skipped; the PR author's own commits skipped; logins compared
case-insensitively (`PR_AUTHOR` does not always arrive in the same case
as the commit author login); each contributor listed once however many
commits they wrote; and a commit GitHub could not match to an account
falls back to its raw git author — that identity being the one most
likely to be lost, not least likely.

Paging stops at the API's own 250-commit ceiling rather than spinning on
full pages of nothing new.

### The test directory was not running

`.github/scripts/tests/` held ten test files covering the existing
gates, and no workflow ran any of them. Adding an eleventh would have
meant adding a test that never executes, so `pr.yml` now runs the
directory. All **149** pass, including the 137 that were already there
and previously unverified in CI.

## Verification

- 149 tests pass via `node --test '.github/scripts/tests/*.test.mjs'` —
the exact command CI now runs.
- The gate was run against the real commit shape from #11370 and
produces the missing trailer verbatim.

This PR is its own negative control: the branch carries only my commits,
so the new gate should stay silent on it. If commitperclip prints a
co-author note below, the gate is wrong.

## Risks

Low. Informational output only — it cannot fail a PR, and `allPassed` is
unchanged.

It adds one API call per gate run (`/pulls/{n}/commits`), fetched in the
same `Promise.all` as the existing PR and files calls.

Enabling the previously-unrun test directory could in principle surface
a pre-existing failure; all 149 pass locally, so it does not.

Revert the commit to restore.

## Model Used

Claude Opus 5 (`claude-opus-5`), through Claude Code. Extended thinking
enabled. Tool use enabled: file read and edit, shell for test runs.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links
- [x] My branch name describes the change and contains no internal
ticket id
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-21 16:44:26 -07:00
TonioandClaude Opus 5 3d366ba15f Rebuild the onboarding agent arc on the prototype's step design (#11905)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The onboarding wizard in `ui/` hires that first agent. It runs three
steps: create the agent, connect a model, and review
> - A standalone prototype holds the agreed design for these steps.
#10786 ported that prototype, but #11067 reverted it in full because the
port deleted `OnboardingWizard.tsx` while four pull requests were
editing that file
> - Those four pull requests have since merged. The revert said the port
can "re-land incrementally", and this is that re-land
> - This pull request takes the presentational layer from the prototype
only. It keeps master's wizard as the source of behaviour, so the eight
onboarding fixes merged since the revert stay in place
> - The benefit is that the three agent steps match the agreed design,
and no merged fix is lost to get there

## Linked Issues or Issue Description

Refs #10786 — the first attempt to land this design.
Refs #11067 — the revert that asked for it to re-land in smaller steps.

No public issue exists for the re-land. The problem is described below.

**Subsystem affected**

The `ui` package. The change touches the onboarding wizard, the agent
capsule,
and one Storybook story. It adds four small presentational components
under
`ui/src/components/onboarding/`.

**Current behavior**

The wizard's agent steps do not match the prototype. Each step shows a
small
heading beside an icon, above a form. The agent capsule sits below that
heading and does not animate. The agent gets a name but no role, so
every
first agent is created as `ceo`.

The wizard also shows a five-segment progress bar on these steps. A
walker who
enters on the agent step cannot reach the first two segments, so two of
the
five can never be filled.

**Proposed behavior**

The three steps use the prototype's card, its centred display heading,
and its
footer. One capsule sits above the heading and stays mounted across all
three
steps, so it reads as one object being built rather than three screens
that
each show their own.

A three-segment strip counts these steps for a walker who enters on
them. The
full-length bar stays for a walker who starts at step one, so that count
never
restarts partway.

The agent step gains a role. The options come from the agent role enum,
not
from the prototype's mock list.

**Reason and benefit**

The design is agreed and already built once. Re-landing it
presentation-first
keeps the behaviour that master gained after the revert.

Sourcing roles from the enum matters. The prototype offers "Coder",
which is
not a valid role — the enum uses `engineer` — so a walker who picked it
would
fail validation at hire time.

**Breaking changes**

None. The wizard keeps its routes, its draft format, and its hire call.
The
draft gains one optional field, `agentRole`. A draft saved before this
change
loads without it and falls back to the default.

## What Changed

- Add `ui/src/components/onboarding/`: `Stepper`, `OnboardingCard`,
`OnboardingHeading`, `FooterNav`, `AgentPreview`, and shared motion
constants
- Rebuild wizard steps 3–5 on those parts: one card, the capsule above a
  centred heading, and one footer
- Hold one `AgentCapsule` across the three steps. It springs in once,
then
  morphs from dashed slot to traced outline to filled
- Add `strokeDraw` to `AgentCapsule`. It traces the outline instead of
  cross-fading it. The dashed layer holds until the trace ends
- Add a role select to the agent step. Choosing a role fills the name,
unless
  the walker typed one
- Show one progress indicator per run, not two
- Label strip segments by destination, not by number
- Add `motion` to the `ui` package
- Add a Storybook story for the strip and the capsule states

## Verification

Run the tests:

```
pnpm --filter @paperclipai/ui exec vitest run
pnpm --filter @paperclipai/ui exec tsc -p tsconfig.json --noEmit
```

4235 tests pass. The typecheck is clean.

To see the steps, start the app and open `/<PREFIX>/onboarding` for a
company
that has a company-level goal. The wizard opens on the agent step. Step
three
requires a hire.

Three absence assertions were checked by fault injection. Each one fails
when
the old behaviour returns:

- put the step counter back, and the "shows no step counter" test fails
- default `strokeDraw` to true, and the cross-fade test fails
- restore the timer gate on the strip, and the indicator test fails

## Risks

Low to medium.

`motion` is one new dependency in `ui`. #11067 gave dependency weight as
one
of three reasons to revert #10786, so this branch carries the smallest
set
that works. `motion` drives the step transitions and the capsule
choreography,
and three files import it.

An earlier revision of this branch also added `three` and
`@types/three`. Both
are removed. They existed for the 3D backdrop, which belongs to the auth
and
welcome screens rather than to these three steps, so nothing on this
branch
imported them.

The role select changes what the wizard sends. Before this change every
first
agent was hired as `ceo`. Now the walker chooses. The values come from
the
enum, so the server accepts all of them.

Steps 1 and 2 keep the older design. They do not run on the Cloud-first
path,
where the company already exists.

## Model Used

Claude Opus 5 (`claude-opus-5`), with extended thinking, tool use, and
code
execution. Used for the code, the tests, and this description.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-21 16:41:53 -07:00
TonioandClaude Opus 5 55464204a6 test(ui): wait for conditions in the last four fixed-turn loops (#11523)
The remaining instances of the pattern #11499 and #11521 replaced in the routing
tests, found by grepping `attempt < N` across the suite. Budgets of 20, 25 and
30 turns rather than 3 and 5, which is why they surfaced far less often -
SkillStudio was among the failures seen while verifying the earlier PRs.

Three of the four were reimplementations of `vi.waitFor` down to rethrowing the
last error, differing only in bounding on turns rather than on time. The fourth
was the text variant. Two `flushReact` helpers became dead with the loops that
used them and are removed.

Verified on the mechanism, since budgets this large pass until the machine is
loaded and so prove nothing by passing: a throwaway probe drove a 30-turn loop
and `vi.waitFor` against a value landing at turn 60. The loop throws,
`vi.waitFor` reaches it. Not a lateral move between arbitrary bounds.

This closes one spelling of the pattern, not the class, and the sweep that found
these was too narrow. Two other shapes do the same thing and a grep for
`attempt < N` cannot see either: a fixed-cycle helper, `flushReact(cycles = 4)`
in AgentToolsTab.test.tsx, and fixed-duration sleeps in AgentToolsTab,
CompanyContext, AgentConfigForm.render, Artifacts, Search and
ImportFromVaultDialog. `AgentToolsTab > autosaves installed apps for the current
agent` failed one of three full-suite runs here, holds both shapes, and is
untouched by this change. Refs #11484.

ui typecheck clean; the four files pass; full suite passes two of three runs,
the third failing only on that pre-existing instance.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-17 00:58:16 -07:00
TonioandClaude Opus 5 65907aa41f test(ui): wait for the route, not three turns, in the cases-routing test (#11521)
The instance #11499 named but did not include. `App.cases-routing.test.tsx`
carried the identical fixed-turn loop that PR replaced in its sibling
`App.activity-routing.test.tsx` - three macrotasks instead of five, otherwise
the same helper - so it fails the same way when the suite runs many workers in
parallel and the container has not filled yet. It was one of the failures
observed while verifying #11499.

Same one-line replacement: `vi.waitFor` retries against a time budget, so a
loaded worker gets more turns rather than a failure. The two helpers are
identical again.

Verified on the mechanism rather than on a green run, because the old loop
passes in isolation too - that is what made this a flake and not a failure. A
throwaway probe drove both helpers against a container whose text lands after
ten macrotasks: the three-turn loop throws, `vi.waitFor` resolves. That is the
condition a loaded CI worker creates. The probe was deleted rather than
committed; it tests a test helper and had one question to answer.

This closes one named instance, not the class. The full ui suite passed three
consecutive times with no failure in any file, but the other instances seen
during #11499's verification - TaskChatComposer, RequestCollapsedSidebar -
simply did not recur, so they are rarer rather than fixed. Refs #11484.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-17 00:42:54 -07:00
TonioandClaude Opus 5 e07d605dfc test(ui): wait for conditions, not durations, in three flaky tests (#11499)
Three tests yielded a fixed number of macrotasks before asserting - five in one
case, one in another - which is ample on an idle machine and not when the suite
runs many workers in parallel. The container was still empty, or the state had
not landed, and the assertion failed on behaviour that works. `vi.waitFor`
retries against a time budget instead, so a loaded worker gets more turns
rather than a failure.

`DocumentAnnotationPopover` is a different race and is fixed differently. The
popover element is in the DOM as soon as React commits, while the effect that
registers the document-level keydown and pointerdown listeners runs afterwards.
A test dispatching in that gap loses the event outright, and a lost event
cannot be recovered by retrying an assertion - so the render is wrapped in
`act` to flush passive effects, and the waits only cover the smaller race that
remains.

Refs #11484.

Verified stable over six consecutive runs of the three files, and the full ui
suite passes. Other instances of the same class remain: the full suite still
shows an occasional failure in a different unrelated test on each run.
`App.cases-routing.test.tsx:104-108` is the clearest one - the identical
fixed-turn loop this PR replaced in its sibling `App.activity-routing.test.tsx`,
three turns instead of five - and takes the same one-line fix.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-17 00:29:42 -07:00
TonioandClaude Opus 5 870c305410 refactor(ui): drop the TZ pin now the fixtures are anchored (#11508)
#11480 pinned `TZ: "UTC"` in the vitest config because several suites asserted
local-time renders from UTC instants. That made the suite green everywhere, but
by suppressing the variable rather than fixing what depended on it: afterwards
no test could observe non-UTC behaviour, and a fixture quietly regaining a
local-time dependency would not be caught.

#11478 anchored those fixtures to the clock under test, which is the real fix,
so the pin now carries only its cost. Removed.

The order was load-bearing and is now satisfied. Measured on master before
#11478 landed, removing the pin failed four date-dependent tests at UTC+9 and
ten at UTC+12 - IssueProperties, IssueThreadInteractionCard, SummarySlotCard
and attention, all of which that PR anchors. Re-measured on master at
40e7add71 with the pin removed: 4117 pass at UTC+14, UTC+9 and UTC-11, against
a control of 4117 with the pin. The prerequisite is demonstrated rather than
assumed.

No test accompanies this, and the prefix says so. The change deletes
configuration, and what verifies it is the existing suite run at several
offsets - not expressible as a test case without a harness that re-runs vitest
under a different TZ.

A caution for anyone reading a failure here later: an earlier pass at this
misread the parallel-worker flakes from #11499 as timezone failures, because a
single run per zone showed a clean east-of-UTC pattern that was really noise.
Date-dependent failures are consistent across runs and name date-handling
tests; the flakes vary between runs and name unrelated pages.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-17 00:18:13 -07:00
TonioandClaude Opus 5 6d0adbfb5d refactor(ui): retire the shared-cache reasoning the account key made obsolete (#11507)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Three places in the UI turn the company list into an authorization
verdict: the invite landing page, the onboarding draft gate, and company
auto-selection
> - Each grew a defense when one `["companies"]` cache entry answered
for every account, and each documented that hazard at length
> - #11488 keyed the entry by account, so the hazard those comments
describe can no longer happen
> - The comments stayed, and a comment that describes a trap that no
longer exists is how the next reader removes a mechanism that is still
holding something up
> - This pull request replaces that reasoning with what the mechanisms
actually do now, and removes the one condition that genuinely went dead
> - The benefit is that the next person to simplify these gates has
accurate reasons to work from

## Linked Issues or Issue Description

No public issue exists. Follow-up to #11488, #11430 and #11417. The
problem follows.

**What happened?**

The three gates were written against a shared, account-less company
cache. #11488 keyed that entry by account, which made the documented
hazard impossible — but the documentation stayed. Each gate now carries
a long explanation of a cross-account leak that the key prevents, while
the mechanism it explains is in fact still required for a different and
unrelated reason.

That is a maintenance hazard in a specific direction: a reader who
checks the comment against the code concludes the mechanism is obsolete,
removes it, and reintroduces a failure the comment never mentioned.

**Expected behavior**

The reasoning next to each gate describes why the gate is there now.

**Steps to reproduce**

Read the comment above `ownershipDecidable` in `OnboardingWizard.tsx`
against `master`. It justifies `isSuccess` on the grounds that "after an
account switch the retained value is the previous account's list", which
the account-keyed entry makes impossible.

**Paperclip version or commit**

`master` at `0817fbad9`.

## What Changed

- `ui/src/pages/InviteLanding.tsx` — dropped the
`Boolean(sessionQuery.data)` conjunct from `membershipListIsCurrent`;
rewrote the comment.
- `ui/src/components/OnboardingWizard.tsx` — replaced the shared-cache
explanation above the ownership gate with the reason the gate still
exists.
- `ui/src/hooks/useSignOut.ts` — corrected the sweep's rationale, which
cited the company list as its example of data the next account could
read.

### The one dead condition

`membershipListIsCurrent` tested `Boolean(sessionQuery.data) &&
companiesQuery.isFetchedAfterMount`. The first term cannot be false when
the second is true: the query is `enabled` only while a session exists,
so the flag cannot be set without one. The lapsed-session case it looked
like it covered is covered by the keying instead — the observer re-keys
to the anonymous entry and holds no data to leak.

Tests pass with it removed, but that only shows no test distinguishes
it, which is why the reasoning above is recorded in the code rather than
left for the next reader to redo.

### What is deliberately kept

Each gate turned out to be load-bearing for a reason that has nothing to
do with accounts:

- **InviteLanding** still waits for a list fetched this mount. A pending
query reads as an empty list, which reads as "not a member", which
auto-accepts an invite the customer may already hold.
- **OnboardingWizard** still forces a fetch with `staleTime: 0`. A
cached list is the right account's but can be thirty seconds old, so a
company created moments ago in another tab is missing from it — and
missing reads as "you do not own this", which *deletes* the draft rather
than withholding it.
- **CompanyProvider** still clears the live selection on an account
change. That is component state and does not change key with the query.

Removing them as redundant is the mistake the stale comments invited;
this change is what makes that argument harder to make by accident.

## Verification

- `pnpm tsc -b` in `ui`: clean.
- `InviteLanding.test.tsx`, `OnboardingWizard.test.tsx`,
`CompanyContext.test.tsx`, `useSignOut.test.tsx`,
`companies-query.test.ts`: **67 passed**, run twice.

No behaviour change is claimed and none is intended: the only
non-comment edit is the removal of a condition that cannot alter the
expression's value.

**Not done:** no browser run. Nothing here is observable at runtime.

## Risks

Low. Comments, plus one condition shown to be unreachable-false.

The risk that remains is a documentation risk in the other direction: if
the keying is ever reverted or bypassed, these comments will understate
what the gates protect against. They name #11488 so that connection is
findable.

**This does not close the class.** Account-scoped entries other than the
list — `["companies", id]`, stats, and the rest — still survive an
account change that skips the sign-out button. That is unclaimed work,
and larger than this.

## Model Used

Claude Opus 5 (`claude-opus-5`), through Claude Code. Extended thinking
enabled. Tool use enabled: file read and edit, shell for typecheck and
test runs.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-16 22:34:55 -07:00
TonioandClaude Opus 5 40e7add71c test(ui): make the suite independent of the machine timezone (#11478)
Several suites asserted local-time renders from UTC instants, or built date
fixtures from the real clock, so they only held where local time happened to
match. CI runs in UTC and never reported it; a contributor anywhere else saw
failures on a clean checkout.

Seven files, anchored to the clock under test rather than to the machine's.
The set grew twice while being fixed: the two tests the issue named surfaced
three more at UTC+9, and those surfaced two more at UTC+14.

`StatusCards/format` is the interesting one. `rollupUpdatesToday` filters on the
UTC calendar day to match the server token cap, while the fixtures were built
on the local day. West of UTC that lands `iso(0)` in the previous UTC day for
the stretch between UTC midnight and local midnight - about seven hours a day
at UTC-7 - and east of UTC+12 "today at local noon" is already yesterday in UTC
outright. Either way the rows it means to count drop out. A run crossing
midnight UTC splits the same way.

Fixes #11476.

Deliberately left: IssueProperties.test.tsx:1515-1517 still pin the minute of
three timestamps against a UTC fixture. They pass at every offset tried,
including UTC+5:45, and the minute there is load-bearing - it distinguishes
Created from Started from Completed - so it wants more care than mechanical
anchoring.

Full ui suite 4113 pass. The TZ pin added by #11480 is still in place here and
is now redundant; #11508 removes it, stacked on this branch so it cannot land
without the anchoring it depends on.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-16 22:22:45 -07:00
TonioandClaude Opus 5 0817fbad92 fix(ui): scope invite membership checks to the signed-in account (#11417)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Which companies a person belongs to is an authorization fact the
server owns, and the UI caches the answer under a single `["companies"]`
key
> - That cache entry carries no account identity, and `main.tsx` sets
`staleTime: 30_000` for every query, so for thirty seconds after a
sign-in the previous account's list is served with no request at all
> - The invite landing page reads that list to decide whether the person
is already a member of the inviting company
> - A list that arrives with no loading state and no error therefore
looks authoritative while describing somebody else
> - This pull request makes the page trust only a list it fetched
itself, for the account signed in now
> - The benefit is that a membership decision stops depending on cache
freshness, which nothing in the app guarantees

## Linked Issues or Issue Description

No public issue exists. Refs #11380, #11382. The problem follows.

**What happened?**

`InviteLanding` read the shared `["companies"]` cache entry as proof of
membership in two places:

- The post-sign-in redirect called
`fetchQuery(companiesListQueryOptions)`, which returns the cached entry
without a request while it is inside the app-wide `staleTime`.
- An effect cleared the pending invite token whenever the cached list
contained the invited company.

Neither checked that the list belonged to the account signed in now. A
second account signing in on a warm tab, or a session that lapses
server-side, is enough to reach both.

**Expected behavior**

The page decides membership from a company list fetched for the current
session.

**Steps to reproduce**

1. On a self-hosted instance in `authenticated` mode, sign in as account
A, which belongs to company X.
2. Within thirty seconds, open an invite link for company X and sign in
as account B, which does not belong to it.
3. The page reads A's cached list, finds company X, and treats B as
already a member.

**Paperclip version or commit**

`master` at `2a4b4bc63`.

## What Changed

- `ui/src/pages/InviteLanding.tsx` — the membership query sets
`staleTime: 0` so it revalidates on mount, and the verdict is withheld
until that fetch lands, keyed on `isFetchedAfterMount`. The
token-clearing effect and the "already a member" branch both read
through that gate.
- `ui/src/pages/InviteLanding.tsx` — the post-sign-in path cancels
anything still in flight for the previous session, then forces a fetch
for the new one with `staleTime: 0`.
- `ui/src/pages/Auth.tsx` — sign-in resets the companies query instead
of invalidating it. Invalidation leaves the previous account's list
readable, and its fetch running, until the refetch returns.
- `ui/src/pages/InviteLanding.test.tsx` — coverage for the warm-cache
case, the token-clearing effect, and the `local_trusted` exemption.

### `local_trusted` is exempt

Those instances have no accounts, so the shared list is the only
identity there is. `membershipIsAccountScoped` is false there and the
gate stays open.

### Rebased onto the account-keyed cache

#11488 landed while this was open and keys the company list by account,
so the page can no longer reach another account's list at all. Two
things changed here as a result:

- The post-sign-in read now calls `fetchCompanyListForCurrentAccount`,
which replaces the `cancelQueries` plus forced `fetchQuery` this PR
originally carried. The helper is strictly stronger: it detaches the
in-flight `/companies` request inside the query function, and it
resolves the account identity past the session invalidation immediately
above rather than trusting the session entry still in the cache.
- The observer reads through `useCompanyListQuery`.

The mount-scoped `isFetchedAfterMount` gate is **kept**, not removed.
Its purpose has narrowed — cross-account leakage is now structurally
impossible, so what remains is holding the verdict until this page has a
list rather than acting on a pending one. It is still load-bearing:
disabling it fails two tests here. Removing a defense in the same change
that rebases onto a new foundation is the wrong order; that is a
follow-up once the keying has proven itself.

`Auth.tsx` can safely reset, because it navigates away on success and
`InviteLanding` mounts fresh afterward. Measurements of exactly when
that rewind does and does not bite are in
[#11380](https://github.com/paperclipai/paperclip/pull/11380#issuecomment-5300984911).

## Verification

- `InviteLanding.test.tsx`, `Auth.test.tsx`, `companies-query.test.ts`,
`CompanyContext.test.tsx` together: **51 passed**, run twice.
- `pnpm tsc -b`: clean.
- `InviteLanding.test.tsx` and `Auth.test.tsx` together: 21 passed.

Both failures are pre-existing and unrelated. Each reproduces on a tree
that does not contain this change, in files this change does not touch:

| Failure | Why it fails |
| --- | --- |
| `IssueProperties.test.tsx` | Timezone-dependent: expects `4:08 PM`,
gets `9:08 AM` |
| `StatusCards/format.test.ts` | Time-of-day dependent: "only counts
updates started today" breaks near midnight |

**Not done:** no manual two-account run in a browser. The path needs two
accounts on an `authenticated` instance, which a local dev instance
cannot exercise.

## Risks

Low. The failure direction is a membership verdict withheld for one
extra round trip, which resolves itself; the direction it removes is one
account's membership granted to another, which does not.

It adds one request per invite-page mount, on a query key the app
already uses.

**This does not close the class.** The shared list is still unscoped for
every other consumer. #11380 clears it on sign-out and #11382 handles
the onboarding draft gate; all three are needed, because an account can
change without passing through any one of those paths.

## Model Used

Claude Opus 5 (`claude-opus-5`), through Claude Code. Extended thinking
enabled. Tool use enabled: file read and edit, shell for typecheck and
test runs.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-16 21:55:50 -07:00
TonioandClaude Opus 5 327f59cac2 refactor(ui): key the company list by account (#11488)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Which companies a person belongs to is an authorization fact the
server owns, and the UI caches the answer for speed
> - It cached that answer under one `["companies"]` key with no account
attached, while `main.tsx` sets `staleTime: 30_000` for every query
> - So for thirty seconds after an account change, one person's list
answered questions asked about another, arriving with no loading state
and no error
> - Three separate consumers each grew their own defense against this,
and each was a place to forget one
> - This pull request keys the entry by account, so a list belonging to
someone else is not distrusted but unreachable
> - The benefit is that the protection stops depending on every future
consumer remembering to defend itself

## Linked Issues or Issue Description

No public issue exists. Refs #11380, #11382, #11417, #11430. The problem
follows.

**What happened?**

The company list lived in a single cache entry, `["companies"]`,
carrying no record of which account it was fetched for. Combined with
the app-wide 30s `staleTime`, any read within that window after an
account change returned the previous account's list — from cache, with
no request, no loading state and no error.

Every consumer that treats the list as an authorization fact had to know
this and defend itself:

- `InviteLanding` reads it to decide whether you already belong to the
inviting company (#11417).
- `OnboardingWizard` reads it to decide whether a saved draft belongs to
you (#11382, merged).
- `CompanyProvider` reads it to pick and persist your active company
(#11430).

All three defenses are correct. The problem is structural: the fourth
consumer has to invent a fourth one.

**Expected behavior**

A cached company list can only answer questions about the account it was
fetched for.

**Steps to reproduce**

1. On a self-hosted instance in `authenticated` mode, sign in as account
A, which belongs to company X.
2. Within thirty seconds, have account B become the session in that tab
— a second tab signing in, or A's session lapsing server-side.
3. Any consumer reading the company list receives A's list, and nothing
in the query result indicates it is not B's.

**Paperclip version or commit**

`master` at `ac91b7f3b`, which includes #11430.

## What Changed

- `ui/src/lib/queryKeys.ts` — `companies.list(userId)` replaces
`companies.all` as the list's entry. `companies.all` remains the prefix,
so it still matches for invalidation.
- `ui/src/api/companies-query.ts` — `companyListQueryOptions(userId)`
builds the keyed options; `useCompanyListQuery()` is the only observer
entry point and holds until the session settles, because the key cannot
be built before then; `fetchCompanyListForCurrentAccount(queryClient)`
covers imperative paths; `useAccountIdentity()` exposes the session
identity the key is built from.
- `ui/src/api/companies-query.ts` — the `/companies` detach moved into
the query function.
- `ui/src/context/CompanyContext.tsx` — drops the session-watching
refetch machinery the key now makes unnecessary (`removeQueries`, the
explicit replacement fetch, the awaiting gate). It still clears the live
selection on an account change, because that is component state and does
not change key with the query.
- `ui/src/pages/InviteLanding.tsx`,
`ui/src/components/OnboardingWizard.tsx` — read through the
account-aware API.
- Tests — the account-keyed guarantee, prefix invalidation still
reaching the list, the detach inside the query function, and updates
where suites seeded the old shared key.

### The existing defenses are deliberately left in place

The per-consumer gates in #11382, #11417 and #11430 are now belt and
braces. They are also what will catch this refactor if it is wrong
somewhere, so removing them in the same change that moves the foundation
would be the wrong order. Simplifying them is a follow-up, once this has
proven itself.

### Why `retry: 1` appears in CompanyProvider

An earlier measurement on #11430 found a retry on the replacement fetch
changed no outcome, because `removeQueries` made the observer rebind and
issue a second request for free. Keying by account removes that
mechanism and the free attempt with it. The retry now carries the
property the incidental refetch used to — a single blip during an
account change should not leave the customer with no companies until
they find "Try again". #11430's test for that property is unchanged and
still passes, which is how the gap was caught.

### A regression this went through, kept for the record

Gating the query on the session settling meant that while the account
was unknown the query was *disabled*, and a disabled query reports
`isLoading: false` with no data — which the provider defaults to an
empty list and reads as "asked, and owns nothing". That is the
destructive branch #11477 had just fixed, reached through a different
door: it would have cleared the customer's stored company on every cold
boot. #11477's test caught it during the rebase. `useCompanyListQuery`
now reports the wait for the account as part of the wait for the list.

### What this does not do

It does not scope the rest of the per-account cache. `["companies",
id]`, stats, and every other account-scoped entry still survive an
account change; that is the cache-lifetime work in #11380.

## Verification

- `pnpm vitest run` in `ui`: **4018 passed, 1 failed**.
- `pnpm tsc -b` in `ui`: clean.
- `companies-query.test.ts`: 6 passed. `CompanyContext.test.tsx`: 17
passed. `OnboardingWizard.test.tsx`: 13 passed.
`InviteLanding.test.tsx`: 13 passed.

The failure is the pre-existing timezone-dependent
`IssueProperties.test.tsx`, fixed by #11478.

Two behaviours are asserted rather than assumed, because the refactor is
only safe if they hold: that invalidating the `companies` prefix still
marks the account-keyed list stale (19 call sites depend on it), and
that the query function detaches the in-flight `/companies` request
before fetching.

**Not done:** no manual two-account run in a browser. The path needs two
accounts on an `authenticated` instance, which a local dev instance
cannot exercise.

## Risks

Moderate, and worth reading before approving.

**It touches `InviteLanding.tsx`, which #11417 also modifies**, so one
of the two will need a rebase — the conflict is mechanical (both change
how the same query is read).

This was #11481, which GitHub closed automatically when its base branch
(#11430's) was deleted on merge; reopening a pull request whose base
branch is gone is not permitted, so it continues here against `master`
with the same head and the same review already recorded on #11481.

**The list now waits for the session query.** The key cannot be built
before the account is known. In the app the session is already fetched
at boot by many components, so this is a dependency rather than an extra
request, but it does serialize: on a cold boot the list waits for the
session to land. Every test that renders a company-list consumer now
needs a session in the cache, which is why several suites gained a seed.

**A missing mock surfaces as a passing gate rather than an error.** The
detach inside the query function meant suites whose `companiesApi` mock
lacked `detachInflightList` had their query function throw, which read
as "decided" in the onboarding gate and mounted the wizard early. Fixed
in the affected suites; worth knowing as a failure mode.

## Model Used

Claude Opus 5 (`claude-opus-5`), through Claude Code. Extended thinking
enabled. Tool use enabled: file read and edit, shell for typecheck and
test runs.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-16 21:23:56 -07:00
TonioandClaude Opus 5 ac91b7f3b2 fix(ui): scope the company selection to the signed-in account (#11430)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Every company-scoped screen reads the active company from
`CompanyProvider`, which picks one from the `["companies"]` list and
remembers it in localStorage
> - That cache entry is shared app-wide and carries no account identity,
so it survives a change of account in the tab
> - The provider therefore auto-selects from whatever list is cached,
which can belong to the account that just went away
> - This pull request makes the provider watch the account and refuse to
derive a selection from a list fetched for a different one
> - The benefit is that the app stops pointing at a company the
signed-in account may not be able to see

## Linked Issues or Issue Description

No public issue exists. Refs #11380, #11382, #11417. The problem
follows.

**What happened?**

`CompanyProvider` auto-selects a company from the shared `["companies"]`
cache entry and writes that id to `localStorage`. Nothing ties that
entry to an account. When the account changes in the tab, the previous
account's list is still served, so the provider can select — and persist
— a company belonging to the account that just went away. Company-scoped
screens then render against a company the current account may not be
able to see.

Signing in through `Auth.tsx` invalidates the entry, so the in-app
sign-in path is covered. Two paths are not: a session that lapses
server-side, and a second account signing in on another tab. The
sign-out sweep in #11380 does not cover them either, because neither
presses the sign-out button.

**Expected behavior**

The company selection is derived only from a company list fetched for
the account that is signed in now.

**Steps to reproduce**

1. On a self-hosted instance in `authenticated` mode, sign in as account
A, which belongs to company X.
2. In a second tab, sign in as account B, which does not belong to
company X.
3. Return to the first tab. The session query refetches and reports
account B, while the company list is still account A's.
4. The provider keeps company X selected and leaves its id in
`localStorage`.

**Paperclip version or commit**

`master` at `6542ad1f4`.

## What Changed

- `ui/src/context/CompanyContext.tsx` — the provider observes
`queryKeys.auth.session`. On a change of session user it clears the live
selection, removes the shared company list, and holds auto-select until
a list fetched for the new account lands. The stored id is left alone on
purpose: `resolveBootstrapCompanySelection` re-validates it, so an
account signing back in keeps its company while an unrelated account
cannot inherit it.
- `ui/src/context/CompanyContext.tsx` — an errored list is treated as
undecided rather than as "no companies". With `retry: false` a single
network blip sticks, and the empty-list branch read it as proof the
account owns nothing and cleared the stored selection.
- `ui/src/api/client.ts` — new `detachInflightGet(path)`. GET coalescing
keys on the request path alone, so a `/companies` request issued under
the previous session could be joined by the replacement fetch and answer
it with the previous account's companies. Detaching leaves that request
to settle for its own callers and makes the next call issue a fresh one.
- `ui/src/api/companies.ts` — `companiesApi.detachInflightList()` wraps
that for the list path.
- `ui/src/context/CompanyContext.tsx` — `companyListUnavailable`
separates "no usable list because a request failed" from "this account
owns nothing", and `retryCompanies` gives a recovery action that
fetches. Both are derived from the query rather than tracked beside it;
a second copy of "did the last attempt succeed" drifted out of step
during review, reporting a failure over a later empty list that was
simply the truth.
- `ui/src/components/SidebarCompanyMenu.tsx` — renders "Couldn't load
companies" and a Try again item in place of "No companies", which is a
claim about the account that a failed request cannot support. This is
the menu `Sidebar` mounts, so it is the only place a customer can act on
the failure.
- `ui/src/components/CompanySwitcher.tsx` — the same treatment. The
application does not render this component (its only mount is a
Storybook story), so it is kept in step rather than relied on.
- `ui/src/context/CompanyContext.test.tsx`,
`ui/src/components/SidebarCompanyMenu.test.tsx`,
`ui/src/api/client.test.ts` — coverage for the account switch, a
same-account re-observation not churning, the detached GET, the failed
replacement and its recovery, a single blip self-healing, unavailability
not outliving the failure, and the sidebar rendering the recovery action
for a failure but plain "No companies" for an account that owns nothing.

### No `retry` override on the replacement fetch

The obvious fix for a failed replacement is a retry, and it is not
load-bearing here. A transient failure already gets a second attempt:
the observer rebinds to a fresh query on the render those state updates
schedule, and issues its own request — measured as two attempts with or
without the option. Retries would only add failed round trips before a
real outage is reported, and the outage is what needs a way out, which
is what `companyListUnavailable` and `retryCompanies` provide.

### Why `removeQueries` here, and why that does not generalise

Removal notifies no observer. What rebinds them at this call site is the
render the surrounding state updates schedule; every observer re-binds
to a fresh query on the next render. A caller without that guarantee
would leave mounted observers serving the previous account's value, so
this is not a pattern to lift elsewhere — the sign-out sweep in #11380
must use `resetQueries` instead, and its measurements are at
[#11380](https://github.com/paperclipai/paperclip/pull/11380#issuecomment-5300984911).

The inverse caveat holds for a local reset under an observer that stays
mounted, which is why #11417 and #11382 avoid `resetQueries`.

## Verification

- `pnpm vitest run` in `ui`: **4014 passed, 1 failed**.
- `pnpm tsc -b` in `ui`: clean.
- `CompanyContext.test.tsx`: 16 passed. `SidebarCompanyMenu.test.tsx`:
15 passed. `client.test.ts`: 9 passed.

The failure is pre-existing and unrelated: `IssueProperties.test.tsx`
expects `4:08 PM` and gets `9:08 AM`, a timezone-dependent assertion. It
reproduces on a tree without this change, and #11478 fixes it.

Each new test was confirmed to fail against the implementation it
covers, by reverting that change and re-running rather than by assuming.
The account-switch test fails without the fix (the selection stays on
the previous account's company and no refetch is issued); the
flag-clearing test fails without its clause (an empty list keeps reading
as "couldn't load").

**Not done:** no manual two-account run in a browser. The path needs two
accounts on an `authenticated` instance, which a local dev instance
cannot exercise.

## Risks

Low. The failure direction is a company selection withheld for one extra
round trip, which resolves when the list arrives. The direction it
removes is one account's company selected and persisted for another.

It adds one company-list request per account change, on a query key the
app already uses. It adds no request at boot: the session query it
observes is already fetched app-wide.

**This does not close the class.** Company-scoped entries other than the
list — `["companies", id]`, stats, and the rest of the per-account cache
— still survive an account change. That is the cache-lifetime work in
#11380, not this provider's.

## Model Used

Claude Opus 5 (`claude-opus-5`), through Claude Code. Extended thinking
enabled. Tool use enabled: file read and edit, shell for typecheck and
test runs, and a scratch vitest harness to measure `removeQueries` and
`resetQueries` notification behaviour against the installed
`@tanstack/query-core` 5.101.4.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-16 11:51:03 -07:00
TonioandClaude Opus 5 92047cac46 test(adapter-utils): make the next sandbox flake diagnosable (#11483)
`execution-target-sandbox` has failed twice in CI and not once in several
hundred local runs. This does not fix it. It makes the next occurrence carry
its own evidence, because a third unreproducible failure would teach nothing.

The observed signature was an empty stdout with exit code 0 - the child exited
cleanly having produced nothing, which is what a lost stdin frame looks like
from the test's side. Three mechanisms were checked and ruled out rather than
assumed: the helper resolving on `exit` rather than `close` (a 200-iteration
probe produced no truncations, and the failure was empty rather than partial);
the wrapper reporting exit before stdout drains (it already listens on
`close`); and frame writes racing (the stream wrapper's `writeEvent` is
synchronous and sequence-numbered).

Two candidates remain and the runtime tree separates them. A stdin queue frame
still present means the host wrote it and the wrapper never consumed it; a
drained queue with no output means it was consumed and the reply was lost on
the way back. The report prints that tree, both proxy streams, the exit code,
and the elapsed time - the last because the bridge and proxy run on 5s budgets
that are generous locally and tight on a runner sharing a box with 19 other
lanes.

Timeouts are deliberately unchanged. Raising them would probably make the
symptom go away, which is the reason not to do it blind.

The first revision capped the tree walk one level above the queue frames, so
"the queue is empty" and "the walk never looked" printed identically - the
distinction the report exists to make. Caught in review. Verifying that the
reporter printed something was not enough; it had to print the thing that
discriminates, which is now checked by planting a frame and forcing the
assertion.

adapter-utils typecheck clean; 44 pass, stable across repeated runs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-16 10:47:35 -07:00
TonioandClaude Opus 5 add65ba4a2 test(ui): make the suite pass outside UTC (#11480)
The ui suite was green in CI and red on a clean checkout in any other
timezone. GitHub's runners default to UTC, so nothing ever reported it. A
contributor elsewhere sees two failures on their first run, which reads as
"this project is broken" rather than "your clock differs from the runner's".

Two independent causes.

`IssueProperties` supplies a UTC instant and asserts on the local-time string
the UI renders from it - "2026-07-17T16:08:00.000Z" is expected to read
"Today, 4:08 PM". That holds only where local time is UTC. Pinned with
`env: { TZ: "UTC" }` rather than rewritten: those assertions are about what a
person sees, and "4:08 PM" is worth more to a reader than an expectation
computed from the same formatter the component uses, which would pass whatever
that formatter did.

`StatusCards/format` was wrong in two ways at once, and the pin hides only one,
so it is fixed directly. `rollupUpdatesToday` filters on the *UTC* day
boundary, while the test built fixtures from local noon on the real clock.
East of UTC+12, "today at local noon" is already yesterday in UTC and the rows
the test means to count are filtered out; and any run crossing midnight UTC
lands `iso(0)` and the function's default `now` on different days. The
fixtures now come from a fixed instant, passed as `now` - the parameter exists
for this, and the sibling test already used it.

Each fix was confirmed load-bearing by removing it under TZ=Pacific/Auckland.
Without the pin, IssueProperties fails; without the fixed instant, StatusCards
fails even with the pin removed, so neither rides on the other.

Full ui suite 4017 pass, 0 fail, in UTC, Pacific/Auckland and Asia/Kolkata.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-16 10:15:38 -07:00
TonioandClaude Opus 5 e384d0a2bd fix(ui): keep the stored company when the company request fails (#11477)
An error is not an answer, and this branch is destructive.

`companiesListQueryOptions` sets `retry: false`, so a request that fails before
ever succeeding leaves `data` undefined - which `CompanyProvider` defaults to
`{ companies: [], unauthorized: false }`. That is indistinguishable from "this
account was asked, and owns nothing", so `shouldClearStoredCompanySelection`
returned true and the effect removed the customer's stored company. With
`refetchOnWindowFocus: true`, a blip on focus during a cold load was enough,
and the next visit drops them onto whichever company sorts first.

The predicate now takes `errored` and refuses to clear on it. Required rather
than optional, so the compiler made both existing call sites state their
answer instead of inheriting a default.

Not clearing costs nothing: a stored id that no longer resolves is ignored by
`resolveBootstrapCompanySelection`, which checks it against the current list
before using it. Clearing wrongly costs the customer's selection, which cannot
be recovered.

Scoped deliberately. This file had been described as carrying the same defect
as the onboarding draft gate and the sign-out sweep, and that was overstated.
Those two *trusted* a stale list to answer "does this account own this
company?". This one validates membership against the current list and only
picks a default, so a stale list here self-corrects rather than leaking. The
failed-request branch is the part that is genuinely wrong, and it is the only
part changed. The transient re-decision during a background refetch is real,
self-correcting, and left alone.

Tested at both levels, because the predicate alone would not have caught it:
the provider is what defaults a failed request to an empty list, so the wiring
is where the decision goes wrong. Removing the guard fails both.

ui typecheck clean; full ui suite 4016 pass, with only the timezone-dependent
IssueProperties failure already present on master.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-16 09:58:46 -07:00
TonioandClaude Opus 5 0023c5c4a6 refactor(ui): record why sign-out resets rather than removes (#11473)
Comment only. Salvaged from `claude/sign-out-cache-note`, written while #11380
was in progress and never opened as a PR; that branch is deleted with this.

#11380 landed the reasoning for `resetQueries` but not the evidence, and not
the part that stops someone reversing it later.

The choice was measured rather than argued. Against query-core 5.101.4,
removal produced 0 notifications and left the observer holding the signed-out
account's session; reset produced 3 and null. The consequence of the former is
not only a stale read - CloudAccessGate's redirect fires on the session going
empty, so it never runs.

The opposite advice really does hold for a local reset of a key an observer is
still mounted against: reset rewinds the update counters `isFetchedAfterMount`
derives from while that observer keeps its bind-time baseline, so anything
gated on the flag withholds forever. Sign-out is not that case, because it
resets the session too and the consumer unmounts on the redirect.

Two sessions reached opposite recommendations on this API within a day, both
correct about different situations, which is the kind of thing a later reader
re-litigates without a note in the file.

The first revision of this note named the wrong consumers - InviteLanding and
the onboarding draft gate, neither of which reads `isFetchedAfterMount`.
`AppsConnect` is the only one on master and is what it names now.

ui typecheck clean; 13 pass across the sign-out suites.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-16 09:44:36 -07:00
TonioandClaude Opus 5 38752c4e5e fix(ui): clear account-scoped query caches on sign-out (#11380)
Sign-out invalidated two query keys - the auth session and health - and left
everything else in the cache. `invalidateQueries` also keeps serving the old
data while it refetches, so it was insufficient even for the two it did touch.
The company list was never touched at all, which is how one account's
companies could still be in hand when the next account signed in.

The self-hosted path now resets every account-scoped entry. Scoping is an
allowlist of instance-scoped roots rather than a list of things to clear, so a
query key added later is account-scoped unless someone deliberately says
otherwise - forgetting this file fails closed.

`health` is the only exemption. It carries deployment mode, bootstrap state
and Cloud metadata, nothing account-scoped, and `useCloudInstance` observes it
with `enabled: false` and leaves the fetch to CloudAccessGate. Dropping the
entry would strand every such observer on `null` until the gate happened to
refetch, flipping Cloud instances into their self-hosted rendering mid
sign-out. It is refreshed in place instead.

`resetQueries` rather than `removeQueries`: removal empties the cache without
notifying the observers already subscribed, so a mounted `useQuery` keeps
returning its last result until an unrelated re-render rebuilds it.
CompanyProvider sits above the router and stays mounted across the whole
sign-out and sign-in cycle, so that is the common case here rather than a
corner one. Reset notifies them, so the old data leaves the cache and
everything reading it.

The cloud path is untouched: it is a top-level navigation, and the document
reload builds a new QueryClient with nothing left to clear.

This is the root cause behind the onboarding draft-ownership gate added in
#11382. That gate stays, and its comment now says why: this fix covers the
sign-out button, not the question. An account can change without it - a
session lapsing server-side, a second account signing in on a warm tab, a
caller supplying the company context from somewhere else - so the gate stays
independent rather than deferring to this.

ui typecheck clean; full ui suite 4014 pass, with only the timezone-dependent
IssueProperties failure already present on master.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-15 07:55:14 -07:00
TonioandClaude Opus 5 b38d6ddb81 fix(onboarding): choose the launcher's step from the company's mission too (#11429)
The last of the three ways into onboarding for a company that already exists.
The route resolver and the dashboard both pick the step from whether the
company already has its mission; the "Add Agent" card on `/{prefix}/onboarding`
hardcoded the mission step. So the one entry point whose own copy reads "Add
another agent to X" was the one that stopped to ask X for the mission it
already had.

It now calls `onboardingStepForCompany`, like the other two. `matchedCompany`
moves above the early return because a hook cannot be called after it.

An unsettled or failed lookup still reads as "no mission" and costs the step,
which the customer can answer - the same fail-open rule the other callers
follow, and safe now that confirming the mission updates the company's existing
goal rather than adding a second one.

`OnboardingRoutePage` is exported so this can be driven directly. The
alternative was the whole `<App>` route table, which is a much heavier harness
for a question about one button's argument.

Four cases, and the first fails against the hardcoded step. The button lookup
asserts it matched something before clicking, because a lookup that silently
matches nothing turns the click into a no-op and the test into decoration.

This is the last piece of #11259 that had not landed.

ui typecheck clean; full ui suite 4010 pass. The one failure, in
IssueProperties, is timezone-dependent and reproduces on master.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-15 01:54:47 -07:00
6542ad1f4d fix(onboarding): carry an existing company's mission into the wizard (#11416)
Ports the substance of #11259, which predated the recent onboarding work and
was still open. One of the parts it solves is a regression #11352 introduced.

A company that already has its mission opens on the agent step, which is the
point of #11352. What that missed is that the mission field is filled only by
the step being skipped, and that the field is not decoration:
`composeCeoInstructions` seeds the lead agent's instructions from it, and the
Review checklist reads it. So every Cloud-seeded company hired its lead agent
with no Mission line at all, having been routed there precisely *because* it
had a mission. Before #11352 those companies dead-ended on the mission step; a
dead end became a quiet data loss, which is worse, because it completes.

`selectExistingCompanyMission` reads the company's own goal back into the shape
the mission field holds, and the wizard hydrates from it - only when the field
is empty, so a customer editing their mission is never overwritten by the
stored copy. The marker recording that hydration travels with the field it
describes, cleared wherever `companyGoal` is.

`isExistingCompanyMissionUnresolved` holds the hire while that read is
outstanding, counting an in-flight refetch over cached goals as unresolved.
That is the rule #11382 settled on a day earlier for a different consumer -
`isFetching`, not `isLoading`, because retained data is not an answer to the
question being asked now. #11259 had it first, on 11 August.

`canGoBackFromOnboardingStep` and `canJumpToOnboardingStep` bound how far back
a run can walk by the step it entered on. The Back button already applied that
rule inline; the progress bar applied only the "already completed" half, so a
run holding a company could still jump to step 1 - the step whose job is to
create one. The entry step is captured once, when the wizard opens, for the
same reason the step itself is.

`planMissionPersistence` came with them and turned out to be required rather
than tidying. Hydration sets `createdCompanyGoalId` from the company's existing
goal, and confirming the mission read that id as "already written" and skipped
the write, discarding the customer's edit. That skip was safe only while the id
could arrive one way - by writing. A goal in hand now means update it.

Each piece was checked by removing it and confirming a specific case fails.
The hydration case asserts on `saveInstructionsFile`'s content, the actual
consumer, rather than on the mission textarea, because the entry path never
renders that field and the navigation bound now prevents reaching it. One
caveat recorded rather than smoothed over: the reopen case fails only when
both marker-clears are removed, since `reset()` also clears the company id and
the next introduction routes through `clearCompanyScopedState`. They are kept
as one invariant rather than one guard plus a coincidence.

ui typecheck clean; full ui suite 4004 pass. Two failures remain, in
IssueProperties and StatusCards/format; both are date-dependent, both
reproduce on master with these changes stashed, and neither file is touched
here.

Co-Authored-By: Jannes Stubbemann <stubbi@users.noreply.github.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-15 01:42:03 -07:00
TonioandClaude Opus 5 484b1f626c fix(onboarding): verify draft ownership against a list fetched this session (#11382)
#11370 stopped onboarding restoring a saved draft when the company list had
*errored*. It still trusted the list when the list looked healthy — and the
wider door was exactly that. `main.tsx` sets `staleTime: 30_000` app-wide and
`Auth.tsx` invalidates rather than resets on sign-in, so `invalidateQueries`
keeps serving the previous account's companies with `isLoading` false and no
error at all. On a self-hosted instance, where sign-out does not reload the
page, signing in as a second account in the same warm tab could restore the
first account's draft. No request had to fail.

The wizard now judges ownership against a list it fetched for the current
session: its own `useQuery` on the shared key with `staleTime: 0`, gated so it
runs only when a parseable draft exists and adds no request otherwise.

Every clause of that gate earns its place, and each was verified by removing
it and watching a specific case fail:

- `isSuccess` ties the answer to this session. React Query retains the last
  good `data` when a refetch fails, so after an account switch the retained
  value is the previous account's list; a failed refetch flips status to error
  and this rejects it.
- The `unauthorized` check catches the opposite error. `companiesListQueryOptions`
  folds 401 and 403 into `{ companies: [], unauthorized: true }` rather than
  throwing, so an auth blip arrives as a *successful* empty list and would
  otherwise read as "this account owns nothing" and delete the draft.
- The mount gate keys on `isFetching`, not `isLoading`. `isLoading` is false
  whenever retained data exists, so a refetch over a warm cache mounted the
  wizard undecided — and with the wizard open, the persist effect overwrote
  the customer's own draft with defaults before the answer arrived. It still
  releases on failure, so the "Get Started" dead end stays fixed.
- An unreadable draft is judged, and cleared, before any of the above, and
  does not enable the query at all.

`isFetchedAfterMount` was in an earlier revision and is deliberately not here:
it is true after a failed refetch too, so it rejects nothing `isSuccess` has
not, and no test could distinguish it.

Worth recording how the first defect survived a check. I fault-injected it,
saw a test fail, and concluded the guard worked. It was failing for an
unrelated reason — the inner wizard mounted during the fetch and locked its
state initializers to defaults, so the draft could not appear whatever the
gate decided. Fixing the mount gate exposed the real behaviour. An injection
is only evidence if the failure it produces is the one being claimed.

This narrows onboarding only. The general fault is that a sign-out leaves
account-scoped caches in place, and account changes that skip the button —
a session lapsing server-side, a second account in a warm tab — reach the same
stale list. Tracked separately; this defence should not be removed as
redundant when that lands.

ui typecheck clean; full ui suite 3963 pass, with only the timezone-dependent
IssueProperties failure already present on master.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 22:59:41 -07:00
TonioandClaude Opus 5 35a9b98733 fix(ui): theme the onboarding wizard decorative panel instead of hardcoding dark (#11379)
The onboarding wizard's decorative right-hand panel, which holds the ASCII
paperclip illustration, hardcoded a near-black surface. `ThemeContext`
supports light and dark and follows `prefers-color-scheme`, so in light mode
a new customer met the product as a pale form beside a solid black rectangle
— on the one screen meant to introduce it. The glyphs inside already used
`text-muted-foreground`, so the panel was the only part ignoring the theme.

It now uses the paired `bg-muted` surface, which is defined in both themes
(`oklch(0.97 0 0)` light, `oklch(0.269 0 0)` dark), so the illustration reads
as ink on a surface either way and follows any future theme without another
fix.

The guard for it asserts the complete set of `bg-` classes on the panel is
`["bg-muted"]`, anchored to the `<AsciiArtAnimation />` wrapper rather than
scanning the file. Forbidding specific spellings is what failed here
originally: the first version checked `bg-[#rrggbb]` and silently stopped
guarding anything once master migrated the class to `bg-(--hex-1d1d1d)`.
Naming what is allowed cannot decay that way, and it catches named colours
like `bg-black` that no spelling list covered.

Lands the work from #8982 by @stubbi, whose two commits are included
unchanged with their authorship. The rebase and the guard are mine.

Tested: ui typecheck clean; both theme cases fail against four spellings of
the regression — `bg-(--hex-1d1d1d)`, `bg-[#1d1d1d]`, `bg-black`,
`bg-zinc-900` — where the original caught only one and my first widening
caught two. Full ui suite 3958 pass, with one timezone-dependent
IssueProperties failure present on master. All CI gates green; Greptile 5/5.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 12:06:20 -07:00
TonioandClaude Opus 5 66515582e4 fix(onboarding): do not restore wizard state for a company the user does not own (#11370)
Onboarding persists a draft to `localStorage`, including `createdCompanyId`.
That key is scoped per browser origin, not per account, so a browser that has
already run onboarding hands the stored company id to the next session on the
same origin — whoever is signed in. Every downstream call then targets a
company that account may not own: goals, agents and issues are created there,
and requests fail with authorization errors.

`restoreOnboardingState` now returns a saved draft only when the signed-in
account owns the company it names. Otherwise the draft is discarded and the
stale blob removed.

The wizard splits into a gate and an inner component because the inner one
has ~20 `useState(saved?.x ?? default)` initializers, and an initializer runs
only on the first render. Mounting before the restored draft is final locks
every field to its default with no way back, so the gate waits for the
company list while it is loading.

Ownership is judged only against a list that actually answered. Any company
query error makes it undecidable, whatever the list contains — the companies
cache is not account-scoped and survives sign-out, so a failed refetch after
an account switch can leave the previous account's companies in hand, and
trusting a non-empty list there would hand one account's draft to the next.
Nothing is restored and nothing is deleted in that state; the next successful
load decides.

Judging the draft and mounting the wizard are separate questions. The gate
withholds the wizard only while the list is *loading*, never on error: the
companies query sets `retry: false`, and with no companies the dashboard
offers a "Get Started" button wired to onboarding, so blocking there would
make that button do nothing at all. Mounting costs the draft nothing, because
the persist effect is itself gated on the wizard being open.

All four draft-storage call sites — read, write, cleanup, reset — go through
one guarded helper. Storage access throws outright where a browser denies it,
and each site sits in a render, an effect or a close handler, so an escaping
exception took down something the customer was using.

Lands the work from #9900 by @stubbi, whose two commits are included
unchanged with their authorship. The rebase, the error-path handling and the
storage guards are mine.

Follow-up filed separately: sign-out should remove account-scoped cached data
rather than invalidating two keys. This change is defensive and protects
onboarding only.

Tested: ui typecheck clean; 73 pass across the seven onboarding suites; full
ui suite 3952 pass, with one timezone-dependent IssueProperties failure
present on master in a file this does not touch. All CI gates green;
Greptile 5/5.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 11:08:48 -07:00
TonioandClaude Opus 5 bc0b5a1642 fix(ui): onboarding wizard keeps an invisible disabled adapter selected (#11371)
The wizard defaults `adapterType` to `claude_local`, and a saved draft can
name any adapter. The grid only renders adapters the server has enabled, so
on an instance where the held adapter is disabled nothing appears selected
while the wizard still holds it — and the first agent is hired on an adapter
the deployer turned off, which can never acquire a lease.

The selection now snaps to the first enabled, non-coming-soon adapter
whenever the held one is not visible, and adapter-specific model defaults
follow it.

The snap waits for the adapter registry to load. External adapter types are
registered into the UI registry only once the adapters query resolves, so
before that a saved external adapter is indistinguishable from a disabled
one — snapping on that transient list would replace the customer's choice
with a built-in and the persist effect would write it down. This gate fails
closed, unlike the fail-open gates in onboarding, because the directions of
harm are opposite: acting early silently rewrites a saved answer, while
waiting merely leaves the selection alone, which is the behaviour that
existed before the snap did.

The test file is named `OnboardingWizard.adapters.test.tsx` rather than
`OnboardingWizard.test.tsx`, which is the name #11370 uses for its restore
gate. Both merged cleanly onto master alone but collided with each other on
add/add, and nothing in either status showed it.

Lands the work from #9900's sibling, #9501, by @stubbi, whose commit is
included unchanged with their authorship. The rename and the registry gate
are mine.

Tested: ui typecheck clean; 51 pass across the adapter, hook, dialog,
config-form and wizard-step suites, including the other callers of the
adapter hook since that module changed. All CI gates green; Greptile 5/5.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 10:27:59 -07:00
TonioandClaude Opus 5 d95340b0b8 feat(ui): send a company with no agent into onboarding, at the right step (#11352)
A company with no agent cannot do anything: no runs, no tasks, nothing to
show. The dashboard says so in a banner with a link, which asks the customer
to notice a problem the product can fix for them. It is worse for a company
created by Paperclip Cloud: Cloud creates the company before the tenant
boots, so the companyless redirect never runs, and the customer arrives on an
empty dashboard straight out of a signup flow that already asked for a
mission.

The dashboard now opens onboarding when the agent list has loaded and is
empty, and onboarding opens on the agent step when the company already has
its mission — read from the company-level goal the seed writes, under the
query key the launch path already uses, so it shares a cache entry rather
than adding a fetch.

The step is decided once. `initialStep` is derived from the company list and
the goal list, so it changes on any retry, refetch or cache invalidation. An
effect that took it as a dependency called `setStep` on every one of those
and moved a customer who was already mid-flow. Gating the input only narrowed
that window; it could not close it. The step now belongs to the request that
opened the wizard: the effect reads it through a ref and is keyed on the
wizard opening or the company changing. `createdCompanyIdRef` beside it
already used this pattern for the same reason.

That exposed a path nothing had ever taken. A company reached the mission
step only by creating itself on step 1, so opening an existing company there
found code that had never run: `companyName` is only typed on step 1, and
both ways forward require it, so the step could not be completed at all; and
confirming advanced without writing anything, so the mission the customer
typed was discarded. Both fixed, and the write now reconciles against the
goal list rather than adding a second company goal, since the mission lookup
fails open and can send a company that has one back to that step.

Company-scoped state now stays with its company. `clearCompanyScopedState`
runs when the route replaces a company and when it withdraws one — the same
event, and clearing half of it left a goal id that made the next company skip
a mission it had never given. `stillTheSameCompany` guards all five async
writes, after the server work rather than before it, so a company switch
mid-flight cannot hand the new company the old one's goal, project, issue or
agent, and cannot leave a hired agent without its instructions file. The
keyboard path honours `loading` like every button already did.
`claimOnboardingOffer` makes onboarding an offer that stays declined for the
visit.

Route ownership is now recorded whenever the route names a company, including
one the wizard already holds. This changes a documented rule deliberately:
without it a self-created company was never withdrawn, so `/onboarding` would
show "create a company" while still holding the previous one and write the
customer's new mission into it.

Tested at the seam, because every defect on this branch lived between a value
and its consumer and the predicate tests passed at every stage.
`OnboardingWizard.step.test.tsx` renders the real wizard against the real
resolver and the real mission hook across 18 cases, and each was
fault-injected against the code it replaces rather than trusted on a green
run. That caught a case that passed against the broken code, and a race in
one of the guards.

ui typecheck clean; full ui suite 3923 pass, with one timezone-dependent
IssueProperties failure present on this branch's base in a file this change
does not touch. All CI gates green; Greptile 5/5.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 08:18:07 -07:00
aac6ce82e1 fix(ui): read the onboarding company prefix from the path, not the route match (#11351)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - New users meet the product through an onboarding wizard that creates
their company, their first agent, and a starter task
> - The wizard also serves an existing company, at
`/{PREFIX}/onboarding`, to add another agent to it
> - On that route the wizard ignores the company in the URL and opens at
"create a company" instead
> - It reads the prefix with `useParams()`, but it renders beside
`<Routes>` rather than inside it, so there is no route match to read
> - This pull request reads the prefix from the pathname, which is
available without a match
> - The benefit is that the URL a user follows decides what the wizard
asks them

## Linked Issues or Issue Description

No public issue exists for this. The problem follows.

**What happened?**

Open `/{PREFIX}/onboarding` for a company that already exists. The
wizard opens at step 1 and asks the user to create a company. The
company named in the URL is ignored.

**Expected behavior**

The wizard recognises the company in the URL and opens at step 2, so the
user adds an agent to that company instead of creating a second one.

**Steps to reproduce**

1. Create a company, so it has an issue prefix.
2. Go to `/{PREFIX}/onboarding`.
3. Read the first screen. It asks for a company name.

**Paperclip version or commit**

`master` at `5ca7b4c1f`.

**Deployment mode**

Any. This is client-side routing and does not depend on the server.

## What Changed

- `ui/src/lib/onboarding-route.ts` — adds
`companyPrefixFromOnboardingPath()`, which reads the prefix from the
pathname.
- `ui/src/components/OnboardingWizard.tsx` — uses that value when the
route match supplies none. One line, plus the import.
- `ui/src/lib/onboarding-route.test.ts` — six cases for the new
function.

`OnboardingWizard` renders beside `<Routes>` in `App.tsx`, so
`useParams()` returns nothing and `companyPrefix` was always
`undefined`. `resolveRouteOnboardingOptions` then took its no-prefix
branch every time. `useLocation()` needs only the router, not a match,
and the wizard already calls it.

The route match is still read first. If the wizard later moves inside
the route tree, this code does not need to change.

The new parser accepts the same shape as `isOnboardingPath()`: the
prefix is the first of exactly two segments. One test asserts the two
agree, because a disagreement would either open the wizard where no
company resolves, or resolve a company where onboarding is not served.

### Why the change is this small

Three pull requests are open against `OnboardingWizard.tsx` — #9900,
#9501 and #8982. A larger change there would collide with all three.
Almost all of this lands in `onboarding-route.ts`, a small file of pure
functions with existing tests.

## Verification

- `npx tsc --noEmit -p ui/tsconfig.json` — clean.
- `npx vitest run ui/src/lib/onboarding-route.test.ts` — 18 pass.
- `npx vitest run ui/src` — 3883 pass, 445 files.

One test shows the defect and the fix together. With `companyPrefix:
undefined`, which is what the wizard supplied before,
`resolveRouteOnboardingOptions` returns `{ initialStep: 1 }`. With the
parsed prefix it returns `{ initialStep: 2, companyId: "c1" }`.

**Pre-existing failures, unrelated:** `IssueProperties.test.tsx` and
`StatusCards/format.test.ts` fail on clean `origin/master` with these
changes stashed. Both look date-dependent.

**Not done:** no manual browser check. The behaviour is covered by unit
tests at the function boundary, and the wizard's own suite passes.

## Risks

Low. The route match is still preferred, so behaviour changes only where
`useParams()` gave nothing — which today is every render of this
component.

The parser returns a prefix only for a two-segment path ending in
`onboarding`, so no other route can start matching. An unknown prefix
already falls back to step 1 in `resolveRouteOnboardingOptions`, and
that path is unchanged.

To revert, remove the fallback in the wizard. The new function has no
other caller.

## Model Used

Claude Opus 5 (`claude-opus-5`), through Claude Code. Extended thinking
enabled. Tool use enabled: file read and edit, shell command execution
for typecheck and the test runs, and the GitHub CLI.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-08-13 21:45:39 -07:00
f0e6c0f549 feat(server): receive and apply the Paperclip Cloud onboarding seed (#11098)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Paperclip Cloud provisions a dedicated tenant stack for each
customer. During signup it asks for a mission, a name and role for the
first agent, and a first task.
> - Cloud pushes those answers into the new stack at activation, as
`POST /api/companies/:companyId/onboarding-seed`.
> - No route served that path. The tenant answered 404, so Cloud
recorded the push as unacknowledged and retried on every portfolio
fetch.
> - The failure was soft. The answers stayed durable in Cloud and the
stack still activated. But the stack opened on the empty first-run
wizard, and it asked the customer again for what they had already given.
> - This pull request adds the receiving endpoint. It validates the
seed, applies it, and acknowledges it.
> - The benefit is that a seeded stack opens with the mission, the agent
and the first task already in place.

## Linked Issues or Issue Description

No public GitHub issue covers this. The problem is described in-PR,
following the feature template.

**Subsystem affected**

server/ — Express REST API and orchestration services. Also
`packages/db` (one new table) and `packages/shared` (one new validator).

**Problem or motivation**

Paperclip Cloud collects onboarding answers at signup and pushes them to
the tenant stack at activation. The tenant had no route for that
request. It answered 404. Cloud treats a non-2xx as "not yet applied",
so it kept the answers and retried, but the stack itself stayed
unseeded. A customer who had already named their mission, their first
agent and their first task arrived at an empty first-run wizard that
asked for all three again.

**Proposed solution**

Serve `POST /api/companies/:companyId/onboarding-seed`. Validate the
body, apply it to the company, then acknowledge it.

The seed is customer free text, so it is bounded and validated in
`packages/shared` and read from the JSON body only. It is never read
from an `x-paperclip-cloud-*` header. That header set is the trusted
identity envelope: every member is derived server-side from the host
plus verified domain records, and that is exactly what makes it
trustworthy. Mixing user content into it would remove the property. A
test plants a mission on a cloud header and asserts that the body value
wins.

Application reuses the shapes the first-run wizard already produces, so
a seeded stack and a manually onboarded one look the same afterwards:

- The mission becomes the company-level goal. A multi-line mission
splits into a title and a description, as the wizard does.
- The agent becomes the company's first hire. Its free-text role ("Chief
of Staff") lands on `title`. The structural `role` stays `ceo`, which is
what the org chart and the default-instructions lookup read.
- The first task becomes an issue in the Onboarding project, assigned to
that agent.

Cloud retries until it gets a 2xx, and it reads any 2xx as "the tenant
holds this content". So the endpoint is idempotent per `revision`. A new
`company_onboarding_seeds` table records the applied revision together
with the goal, the agent and the issue it produced. A replay of a
revision that already matches is a successful no-op. A later revision —
the customer edited their answers — updates those three rows in place
instead of creating a second agent and a second task. The record is
written last, after every other write has landed, so a partial
application cannot present itself as acknowledged.

Everything is applied before the 200 is sent. This is an ordering
guarantee, not eventual consistency. The tests read the database
immediately after the response, with no waiting and no polling, so a
lazy receiver fails them on a fast machine as well as a slow one. That
matters because the redirect into the tenant dashboard is gated on this
acknowledgement.

**Alternatives considered**

Store the seed and let the tenant UI apply it on first load. Rejected:
the dashboard redirect is gated on the acknowledgement, so a background
apply would let the dashboard open before the agent and the task exist.
The whole point is that it must not.

Reuse `POST /companies/:companyId/agents` and `POST
/companies/:companyId/issues` over HTTP from Cloud. Rejected: it needs
three round trips with no shared idempotency key, and it moves the "did
all of it land?" decision to the caller.

**Roadmap alignment**

This completes an existing Cloud-to-tenant contract. It does not add a
new user-facing surface.

## What Changed

- Add `POST /api/companies/:companyId/onboarding-seed` in
`server/src/routes/onboarding-seed.ts`. It authenticates exactly as
`POST /api/companies/:companyId/logo` does, through
`assertCompanyAccess`.
- Add `server/src/services/onboarding-seed.ts`. It applies the mission,
the agent and the first task, and records the applied revision last.
- Add the `company_onboarding_seeds` table: schema, migration `0216`,
and journal entry. It holds the applied revision and the ids of the
goal, agent and issue the seed produced.
- Add `applyOnboardingSeedSchema` in `packages/shared`. It bounds
mission to 2000, agent name to 80, agent role to 120, task title to 200,
and task details to 2000 — the same limits Cloud enforces before it
sends.
- Mount the router in `server/src/app.ts` and register the path in the
OpenAPI document.
- Add `server/src/__tests__/onboarding-seed-route.test.ts` with 13
tests.
- The seeded agent is created on `claude_local`. This mirrors the
teams-catalog default for agents created server-side, where no human
runs an environment test first. `PAPERCLIP_ONBOARDING_SEED_ADAPTER_TYPE`
overrides it.

## Verification

```sh
pnpm typecheck                      # whole workspace, passes
npx vitest run \
  server/src/__tests__/onboarding-seed-route.test.ts \
  server/src/__tests__/openapi-routes.test.ts        # 15 passed
```

The suite runs against embedded Postgres with migrations applied, so
migration `0216` is exercised by every test.

The route tests cover:

- the happy path — mission, agent and task all applied, read immediately
after the 200
- replay of the same revision — no second agent, no second task, no
second goal, no second project
- a later revision — the goal, agent and task are updated in place
- a multi-line mission splitting into a goal title and description
- a revision-only seed
- the activity log entry written once, and not again on a replay
- a caller without access to the company — 403, and nothing written
- a body with no revision — 400
- each field bound past its limit — 400
- a mission planted on an `x-paperclip-cloud-*` header — ignored, body
wins
- an existing Onboarding project — reused, not duplicated

Not verified here: the full Cloud-to-tenant walk against a live stack.
That needs a deployed Cloud and a provisioned tenant together, which is
separate staging work.

## Risks

Migration `0216` creates one new table. It adds no column to an existing
table, rewrites nothing, and backfills nothing, so it is safe to apply
online. The migration safety check passes.

The endpoint writes to a company. Access is enforced by
`assertCompanyAccess`, the same gate the company logo write uses, and a
test covers the denial.

Behavioral note for stacks that already hold data. If a company already
has a non-built-in `ceo` agent, a first seed updates that agent's name
and title rather than creating a second lead. Likewise a seed adopts an
existing company-level goal rather than adding a parallel one. This is
deliberate: the seed is the customer's own stated answer from signup,
and two competing missions or two leads would be worse than one updated
in place. In the intended case — a stack that Cloud has just activated —
none of these exist yet.

The seeded agent is created on `claude_local` with an empty adapter
config. It is idle and needs the usual credential setup before it runs.
Seeding it does not start it.

## Update — rebased onto master + review hardening

Master moved on after this PR was cut, so it was **rebased onto
`master`** and
the seed migration was **renumbered from `0212` to `0216`** (the merged
#11101
took `0212_onboarding_first_task_unique`); the drizzle journal was
re-stitched
and `check:migrations` passes.

Two things landed on top of the original receiver:

- **Mission-only walk contract (PAP-67 r17.4).** The tenant now owns the
first
agent and the first task via #11101's server-owned onboarding path,
which
stamps `ONBOARDING_FIRST_TASK_ORIGIN_KIND` and races safely on the
partial
unique index `issues_onboarding_first_task_uq`. A comment in the apply
path
documents why this receiver leaves the first task to that path on the
cloud
walk, and a paperclip-cloud `node:test`
(`src/onboarding/walk-seed.test.ts`)
asserts the walk's seed carries no `agent`/`firstTask`. The receiver
retains
the agent/first-task code for its documented body contract, kept inert
on the
  cloud path by the mission-only seed.
- **Three Greptile P1 fixes** (`95622fa37`): concurrent application is
now
  serialized under a per-company `pg_advisory_xact_lock` (no duplicate
goal/agent/project/task on overlapping pushes); a revised first task
carries
its resolved `assigneeAgentId`/`goalId`; and the
`company.onboarding_seed_applied`
  audit write is best-effort so a logging failure can't leave the entry
  permanently absent. Two new regression tests cover the first two.

## Model Used

Claude Opus 5 (`claude-opus-5`), 1M context window, extended thinking,
with tool use and code execution. Used for the original codebase
investigation, the implementation, and the tests. The rebase, migration
renumber, mission-only contract, and the three P1 fixes were done with
Claude Opus 4.8 (`claude-opus-4-8`), extended thinking, with tool use
and code execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-12 22:54:07 -07:00
1e07d5b9aa fix(db): give the last two embedded-Postgres migration tests a timeout (#11313)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The `@paperclipai/db` package owns the database schema and its
migrations
> - Some migration tests start an embedded Postgres server and replay a
migration against it
> - An embedded Postgres server needs 7 to 12 seconds to start on a CI
runner
> - Vitest stops a test after 5 seconds unless the test sets its own
timeout
> - Two of these tests do not set a timeout, so they fail on CI before
they assert anything
> - This pull request gives both tests a 30 second timeout
> - The benefit is that unrelated pull requests stop failing on a test
they did not change

## Linked Issues or Issue Description

No public issue exists for this. The problem follows.

**What happened?**

The test `packages/db/src/company-secret-proposals-migration.test.ts`
fails on CI. The error is `Test timed out in 5000ms`. The test never
reaches its assertions. The suite reports `1 failed | 104 passed`.

The failure is not caused by the branch under test. It appeared on three
different branches in a few hours:

| Run | Head | Failing jobs |
| --- | --- | --- |
| 31630781317 | `95622fa3` | `General tests (workspaces-b)`, `verify`,
`e2e shard (2/3)`, `e2e` |
| 31652020976 | `feba90c9` | `General tests (workspaces-b)`, `verify`,
`e2e shard (3/3)`, `e2e` |
| 31651467721 | `f2115207` | `General tests (workspaces-a (1/2))`,
`verify` |

The `verify` job reads the result of the general tests. One timeout
therefore turns into two red checks. A reviewer sees two failures and
reads them as a regression.

**Expected behavior**

The test starts an embedded Postgres server, replays the migration, and
asserts the schema. It must pass on a normal CI runner.

**Steps to reproduce**

1. Open any pull request against `master`.
2. Wait for the job `General tests (workspaces-b)`.
3. Read the failure. The test times out after 5000 ms.

The failure needs a slow runner. A fast development machine starts
embedded Postgres in less than 5 seconds, so the test passes there.

**Paperclip version or commit**

`master` at `a09d7dcc0`.

**Deployment mode**

CI only. GitHub Actions, `ubuntu24` runner image.

## What Changed

- `packages/db/src/company-secret-proposals-migration.test.ts` — the
test now uses a 30 second timeout. The migration suites in this package
already use 20 to 60 seconds. 30 seconds is the most common value.
- `packages/db/src/status-card-migrations.test.ts` — the same change.
This test has the same defect. It does not fail yet because it replays
fewer statements. A fix to only one test moves the problem instead of
removing it.
- Both tests get a comment. The comment tells the next author why the 5
second default is too short.

These two tests were the only embedded-Postgres migration tests in the
package without a timeout.

## Verification

- Run `pnpm vitest run src/company-secret-proposals-migration.test.ts
src/status-card-migrations.test.ts` in `packages/db`. Both tests pass.
- These suites skip themselves when the Postgres binaries are absent. A
pass alone therefore proves nothing. Run the command with
`--reporter=verbose`. The output contains Postgres `NOTICE` messages,
for example `relation "status_cards" already exists, skipping`. These
messages prove the tests ran real SQL.
- Run the same command with `--testTimeout=1`. Both tests still pass.
This proves the per-test timeout overrides the global timeout. Before
this change, the same command fails immediately.
- All CI jobs on this pull request pass. The job `General tests
(workspaces-b)` passes. This job failed on the three runs listed above.

Not done: no attempt to reproduce the timeout on a development machine.
A fast machine starts embedded Postgres in less than 5 seconds, so the
failure does not occur there.

## Risks

Low risk. The change adds two timeout arguments to tests. It changes no
source code, no schema, and no dependency.

A longer timeout cannot hide a regression here. The tests assert the
same conditions as before. A migration that truly hangs now fails after
30 seconds. Before, it failed after 5 seconds with a message that
pointed at the wrong cause.

The `e2e` failures on the runs above have a different cause. The spec
`mcp-user-stories.spec.ts › US-9` fails with `502 — fetch failed` and
`fetch failed: bad port`. These errors come from MCP tool-connection
health checks. The failures hit different shards on different runs. This
pull request does not change that behavior. `e2e shard (2/3)` passes
here, which supports the view that those failures are unstable
infrastructure.

To revert, remove the two timeout arguments.

## Model Used

Claude Opus 5 (`claude-opus-5`), through Claude Code. Extended thinking
enabled. Tool use enabled: file read and edit, shell command execution
for the local test runs, and the GitHub CLI to read the failing CI logs.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-08-12 21:30:50 -07:00
TonioandPaperclip 384e5f6178 revert(ui): back out onboarding port (#10786) (#11067)
This reverts commit 11e56654f8.

#10786 ported the onboarding flow from a standalone prototype and
repointed `/onboarding` at the new `CloudOnboardingFlow`, deleting the
existing `OnboardingWizard.tsx` in the process. The ported flow is not
ready to be the shipping onboarding experience: it landed as a single
large port rather than an incremental migration, it pulled `motion`,
`three` and `@types/three` onto the UI dependency list for prototype
visuals, and it deleted the wizard that four in-flight pull requests
(#9900, #9501, #8982 and one more) were building on — those went
CONFLICTING the moment the file disappeared.

Rather than keep the half-migrated state on master while that is sorted
out, back the port out whole and re-land it incrementally. This restores
`OnboardingWizard.tsx` and the previous versions of the four e2e specs,
drops the `onboarding-preview.html` Vite entry, the DesignGuide
onboarding section and the `data-viz-misc` storybook story, and removes
the three prototype dependencies from `ui/package.json`.

This is an exact mechanical inverse of the squash commit — 41 files,
+2089/-3647, no hand edits. Reverting this commit restores all 41 files
byte for byte, so the port is recoverable in full when it is ready.

`pnpm-lock.yaml` is deliberately not touched. #10786 never updated it;
bot commit 4683f26c9 (#11036) added the `motion`/`three` entries
afterwards, so the lockfile is now ahead of the manifest. CI owns
lockfile updates (`.github/workflows/pr.yml`) and the policy job
regenerates it from the changed manifest.

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-08-07 14:55:34 -07:00
TonioandClaude Opus 5 11e56654f8 feat(ui): port onboarding flow from prototype; add cloud + local variants (#10786)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - First-run onboarding is the subsystem that turns a brand-new install
into a working company: it creates the company, its goal, a lead agent,
and that agent's first task
> - The existing `OnboardingWizard` carried all of that wiring
correctly, but its UI had drifted from the current design direction, and
a separate design prototype (`paperclip-onboard`) existed as a
standalone visual mock with no backend
> - Porting the prototype's *logic* would have thrown away working,
well-tested backend orchestration; leaving the two apart meant the
design never shipped
> - Separately, cloud and local (self-hosted) installs need meaningfully
different first runs — local has no sign-in and must let the user pick a
locally-installed CLI adapter — so a single linear wizard could not
serve both
> - This pull request rebuilds the presentational layer from the
prototype on top of the existing backend orchestration, and splits it
into two thin flow containers over a shared core
> - The benefit is that the shipped onboarding matches the intended
design, cloud and local can diverge without duplicating logic, and each
can later ship to a different app version while sharing one set of step
components

## Linked Issues or Issue Description

No existing issue — describing inline (feature request).

**What problem does this solve?**
Onboarding is the first thing a new user sees, and the shipped wizard
had drifted from the current design. In parallel, cloud and local
installs need different first-run paths: local has no hosted sign-in,
and its agent runs on a CLI adapter installed on the user's machine,
which the cloud path never has to ask about. There was no way to express
that difference without either forking the whole wizard or bolting
conditionals onto a single linear flow.

**Proposed solution**
Extract the onboarding step views and shell into a shared core, then
compose two thin flow containers (cloud and local) over it. Keep all
backend orchestration in the existing `useOnboardingFlow` hook so no
working logic is rewritten.

**Alternatives considered**
- *Single flow with a `variant` prop* — most DRY, but the two flows are
intended to ship on different app versions, and a shared file would have
to be split later anyway.
- *Two fully independent copies* — simplest per-flow, but every shared
refinement (spacing, motion, copy) would have to be made twice and would
drift.

## What Changed

- **Shared core** under `ui/src/components/onboarding/`:
`OnboardingScaffold` owns the full-screen shell and the single
`AnimatePresence` step crossfade, so both flows transition identically;
step views (Start / Company / Agent / Task), `FooterNav`, `AgentPreview`
and the motion constants are extracted for reuse.
- **`CloudOnboardingFlow`** — `start → company → agent → task`; mounted
in the real app via `OnboardingWizardVariant`. Behaviour matches the
retired wizard, including `previewMock` and the existing-company ("add
an agent") entry point.
- **`LocalOnboardingFlow`** — skips sign-in and adds an optional email
ask (with a privacy assurance), a local model/adapter step that hires
with `requireEnvProbe: true`, and a "star us on GitHub" interstitial
before completing. **Harness-only for now** — the real app still mounts
the cloud flow.
- **Deleted `OnboardingWizard.tsx`** (1,786 lines); updated its
Storybook stories and the `OnboardingWizardVariant` test to the new
components.
- **Orbiting 3D paperclip backdrop** behind the auth and welcome screens
(`three`), code-split so it only downloads on those screens; honours
`prefers-reduced-motion` and disposes its GL context on unmount.
- **`motion`** added for step transitions and the agent-capsule
choreography.
- Visual values routed through design tokens per `DESIGN.md`; `Stepper`
generalized to take a step total (backward compatible); `/design-guide`
page and the component index updated.
- **Standalone preview harness** (`ui/onboarding-preview.html`) with
`?flow=` and `?step=` for backend-free review, wired as a second Vite
rollup input.
- **Adapter env probe bound to the adapter it ran against.**
`hireLeadAgent` reused `adapterEnvResult` for any adapter, so when a
hire failed and the user picked a *different* local adapter and retried,
the previous adapter's verdict satisfied the `requireEnvProbe` guard
while the hire posted the new adapter's config — hiring it unprobed. The
cache is now keyed on the adapter type plus the exact config posted to
the test endpoint, the config is built once and shared by probe and
hire, a failed probe clears the cache, and `clearAdapterEnvResult()`
(called on adapter change) stops the step displaying a stale verdict.
Cloud is unaffected — it hires with `requireEnvProbe: false`. Reported
by Greptile.
- **E2E specs re-pointed at the new flow.** Four specs still drove the
deleted wizard (`onboarding`, `conference-room-typing-intro`,
`planning-mode-visual-verification`, `nux-phase4-screenshots`) and
failed with `element(s) not found` on `"Name your company"` /
`input[placeholder="Acme Corp"]`. Rather than repeat the new drive
sequence four times, `tests/e2e/onboarding-flow.ts` adds one driver per
step (`startCloudOnboarding`, `completeCompanyStep`,
`completeAgentStep`, `completeTaskStep`, `completeCloudOnboarding`) and
the specs import it, so the next flow change touches a single file. Two
now-dead `**/test-environment` route stubs went with it — the cloud flow
hires with `requireEnvProbe: false`, so that probe never fires.

## Verification

- `pnpm --filter @paperclipai/ui typecheck` — clean.
- `npx vitest run` over the onboarding suites
(`OnboardingWizardVariant`, `AgentCapsule`, `onboarding-launch`,
`onboarding-goal`, `onboarding-route`, `onboarding-adapter-config`) — 33
tests pass.
- `pnpm --filter @paperclipai/ui build` — succeeds; the three.js chunk
splits out separately (522 kB raw / 133 kB gzip) rather than entering
the main bundle.
- Both flows driven end-to-end in the preview harness in `previewMock`
(no database writes), plus the cloud flow rendered in the real
authenticated app at `/onboarding` to confirm the mount swap.
- The four re-pointed e2e specs pass locally against the new flow.
- New `ui/src/hooks/useOnboardingFlow.test.tsx` — 4 cases pinning the
adapter-probe cache (switch-adapter retry, cold path, explicit clear,
and the cloud flow's `requireEnvProbe: false`). Verified non-vacuous:
the switch-adapter case fails against the pre-fix code.
- Rebased onto current `master`; `pnpm-lock.yaml` is deliberately
**not** committed — `.github/workflows/pr.yml` regenerates it when a
manifest changes and shares it with downstream jobs as the `pr-lockfile`
artifact.

## Risks

- **Deleting `OnboardingWizard.tsx` is the one change that alters
existing app behaviour.** The cloud flow is intended to be
behaviour-equivalent, and its entry points are covered by the updated
`OnboardingWizardVariant` test, but this is the area to review most
closely.
- **Conflict risk with open PRs that touch the old wizard**: #9900,
#9501, #8982 and #6636 all modify
`ui/src/components/OnboardingWizard.tsx`, which this PR removes.
Whichever lands second will need its change re-applied to the new step
components. Flagging so ordering can be decided deliberately.
- **New dependencies**: `motion` and `three` (+ `@types/three`). `three`
is large, so it is lazily imported and code-split — it does not affect
the main bundle. Both are MIT.
- The **local flow is not reachable in the app** yet (harness/canary
only), so it carries no runtime risk today; wiring it up is a follow-up.
- The auth screens remain **presentational only** — they are not wired
to real auth, unchanged from before this PR.
- **Pre-existing, not introduced here:** `OnboardingWizardVariant`
renders outside `<Routes>` in `App.tsx`, so its `useParams()` never
resolves `:companyPrefix` and `/{prefix}/onboarding` opens the welcome
screen instead of jumping to the agent step. `master` has the identical
structure, so this PR faithfully ports existing behaviour; the working
"add an agent" entry is the launcher card behind the overlay, which is
what the screenshot spec drives. Worth a separate fix.

## Model Used

Claude Opus 5 (`claude-opus-5`) via Claude Code, with extended thinking
and tool use (repo search/edit, local test + build execution, and
browser-driven visual verification of the rendered flows). Portions of
the session also ran on `claude-opus-4-8` and `claude-fable-5`.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [ ] I will address all Greptile and reviewer comments before
requesting merge

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 23:35:05 -07:00
TonioandClaude Opus 5 492555aaf9 design(decisions): flatten decision cards to two task-borrowed types (#10474)
The Decisions queue ran five parallel colour/icon vocabularies chosen by
source kind, plus a separate severity badge, so two rows needing the same
response could look unrelated and none of it matched the task list.

Every row now resolves to one of two kinds, each borrowing the task status
it corresponds to: blocking renders as `blocked`, review as `in_review`,
both through StatusGlyph and the existing --status-task-icon-* tokens.
Source kinds keep their own wording; only colour and icon merge.

Card anatomy follows the design mock: no left accent rail, rounded cards
16px apart, a "/"-separated meta breadcrumb, a named See more / See less
control, and no separately tinted drawer when expanded. Verb order is
fixed across both states. Severity moves from chrome to a toolbar filter.

Four defects fixed along the way:
- blocked rows reported themselves as their own blocker (server-side)
- the task key was missing wherever the row's subject IS the task
- the task quicklook stuck open, because closing handed focus back to a
  trigger that opens on focus
- the card ring appeared on click, and only on cards with a toggle

Also: the standard task preview is aligned to its trigger's text and
scales out of it, the task eyebrow renders its project as a tile, and the
first motion tokens land alongside the disclosure and crossfade.

Supersedes #9574 and #9575.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 18:43:40 -07:00
TonioandPaperclip 3b9f36403e refactor(ui): vertically center task status circle in task list (#8376)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The web UI renders work items in a task/issue list and board
columns, each row leading with a status circle next to the title text
> - The status-circle wrapper `<span>` was not enforcing cross-axis
centering, so the circle aligned to the top of taller rows instead of
the title
> - This looked broken: the circle floated above the text rather than
sitting centered against it, across statuses
> - This pull request adds flex centering (`inline-flex` +
`items-center`) to the status-icon wrapper spans so the circle is
vertically centered with the title in every row
> - The benefit is a consistent, correctly aligned task list regardless
of task status or row height

## Linked Issues or Issue Description

No public GitHub issue exists; describing inline following the bug
report template.

### What happened?

In the task/issue list and board, the leading status circle was not
vertically centered with the task title text. On taller rows the circle
sat near the top instead of aligned with the title, so the list looked
misaligned across all statuses.

### Expected behavior

The status circle should be vertically centered with the title string in
every task row, regardless of status.

### Steps to reproduce

1. Open the task list (or board) view in the web UI.
2. Observe rows in different statuses, especially taller rows.
3. Notice the status circle is top-aligned rather than vertically
centered with the title text.

### Paperclip version or commit

`master` @ 411a9a51 (branch base for this PR).

### Deployment mode

Local dev (pnpm dev).

## What Changed

- `ui/src/components/IssueColumns.tsx`: added `items-center` to the
leading status-icon wrapper span (`inline-flex` container) in
`InboxIssueMetaLeading`.
- `ui/src/components/IssuesList.tsx`: added `inline-flex items-center`
to the two status-icon wrapper spans so the circle centers with the row
text (matching the `IssueColumns.tsx` change).

## Verification

- Manual: open the task list and board views; confirm the status circle
is vertically centered with the title across statuses and on taller
rows.
- Before/after matches the screenshots from the originating request
(broken: circle top-aligned; fixed: circle centered).
- `cd ui && npx tsc --noEmit` — type check passes for the touched
components.

## Risks

- Low risk. CSS-only change to wrapper spans; no logic, data, or API
changes. Limited to status-icon alignment in the task list/board rows.

## Model Used

- Claude Opus 4.6 (`claude-opus-4-6`), extended thinking + tool use, via
Claude Code.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [ ] I have added or updated tests where applicable (CSS-only alignment
change; no meaningful unit test — retitled `refactor:` per contribution
gate)
- [x] If this change affects the UI, I have included before/after
screenshots
- [x] I have updated relevant documentation to reflect my changes (none
required)
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-06-19 23:36:38 -07:00