mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-10 03:08:10 +02:00
codex/plugin-task-execution
2
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
11921075a4 |
Add first-task onboarding skill and Runner E2E coverage (#13517)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The first task helps a new user define and approve useful work. > - That workflow needs reusable instructions and tests against the production experience. > - Native Codex and Claude must load the assigned skill, including after resume. > - Maintainers need recorded conversations and precise failed checks to judge regressions. > - This pull request adds the first-task skill and a suite in the shared Runner E2E harness. > - It keeps behavior results separate from informational quality scores and incomplete recordings. ## Linked Issues or Issue Description **What existing behavior does this improve?** The first onboarding task and the Runner E2E report used to review it. **Current behavior** Onboarding embeds its policy in a hidden brief. Native Codex drops the skill-instructions setting at the Rust boundary. The shared E2E harness has no onboarding suite or full conversation view. **Proposed behavior** Assign and invoke `/first-task` for the onboarding task. Send selected Codex skills as structured protocol inputs. Run twelve scenarios across legacy Codex, legacy Claude, native Codex, and native ACPX Claude. Include all 48 cells in full campaigns. Show recorded chat, question and approval cards, exact checks, instructions, and billing in the shared dashboard. **Reason and benefit** Measure the real onboarding experience before changing prompts. Distinguish infrastructure failures, behavior failures, and unexercised journey steps. **Breaking changes** No database migration or production API change. First-task instructions now live in an assigned skill. The user-edited persona is preserved; the skill includes the maintainer-approved proposal-mode mapping and saved-plan requirement. Related: #11043 is earlier onboarding work. #13422 already fixes native Claude model pinning, context delivery, and read permissions on master; this branch includes those fixes through its base. The new Claude recovery test supplements them. ## What Changed - Extract and assign the first-task skill while retaining the production greeting and opening question. - Carry the Codex skill-instructions flag through thread start and resume. Resolve explicit task skill references only against assigned skills and send native skill inputs. - Invoke an unambiguously selected assigned skill through Claude ACPX’s native slash-command parser on initial and resumed turns, retaining the entire task/wake envelope as its argument. Do not carry that invocation into ordinary tasks. - Restore the saved single-task proposal modes: confirmation card, or saved plan with revision-targeted checkbox approval. Explicit plan requests also require a saved plan. - Add first-response and complete-journey cases with fixed user facts, acceptance checkpoints, durable outcome checks, and accounting for child runs. - Fail the eval when choice questions have fewer than two real options. Recognize planning documents without treating them as completed work. - Add optional, bounded quality judging as explicit post-processing. - Render full conversations and static interaction cards in the shared report. Conversations start folded. Show original and regraded results and incomplete journeys distinctly. - Keep credential-persistence scanning outside the first-task behavioral suite; retain public evidence redaction. - Refresh generated capability references after the API-reference edits. - Correct shared native question guidance and tool schemas: choices need at least two meaningful options; open-ended questions use canonical text fields with the required compatibility payload. Verify both formats through real tool-authority persistence. - Disable announcements automatically for every isolated Runner E2E process and label the gallery environment/provider/target explicitly. - Remove CI races in the GitHub connection browser test and native session recovery test by waiting for the actual async work before asserting its results. ## Verification - `pnpm exec vitest run server/src/services/onboarding-first-task-assets.test.ts server/src/__tests__/issue-onboarding-first-task-routes.test.ts`: 19 passed. - `pnpm --dir packages/paperclip-runner exec vitest run src/drivers/acpx/runtime-host.test.ts src/drivers/acpx/native-skill-prompt.test.ts src/cli/acpx-runtime-sidecar.test.ts`: 70 passed. Native command forwarding and the 1 MiB input boundary both failed before their fixes and passed afterward. Coverage includes changed skills on reopen, approval context, and an ordinary subsequent task. - Runner E2E unit suite: 306 passed. Harness typecheck passed. The 64 first-task fixture and grader tests also pass. - Full repository typecheck and build passed locally. Server typecheck and Runner build passed again after the native-command change. - Full GitHub Actions CI passed on `23e56447b`: all server/workspace/browser shards, Runner verification, typecheck/release registry, build, canary, policy, and Docker checks. Greptile reviewed this exact head at 5/5 with no unresolved threads. The earlier broad local run had database startup/timing failures that passed isolated retries; the complete remote suite is green. - Merge verification against current master: 312 harness tests and 13 native recovery tests passed. Regenerated semantic contracts and fixture hashes pass their consistency check. Full local typecheck and build also passed on the stacked queue branch. After merging the latest master and preserving the GitHub setup timing regression in the split browser suite, both focused GitHub browser tests passed. Three CI timing/startup flakes passed local verification and one remote retry; all latest-head checks are green. - Real pinned Claude SDK and Claude ACP JSON-RPC probes against a local mock API confirmed that `/skill-name` expands the assigned skill body before the model request and retains the task arguments. A prose mention does not. The probes made no paid model calls. The ACP probe used the current first-task skill body and retained the wake arguments. - [Full 48-case campaign and report](https://pages.paperclip.ing/runner-e2e-first-task-35053063880/): 44 passed after three interrupted Codex cases completed in targeted reruns. Original results, regrades, and all 51 executions remain in the report provenance. - [Claude campaign after the shared-question fix](https://pages.paperclip.ing/runner-e2e-first-task-claude-35099525201/): 10/12 passed with zero single-option failures. All 12 recorded the current assigned skill and corrected guidance. The failures exposed skipped skill invocation and a missing saved plan. This PR adds native command invocation and explicit saved-plan instructions; the subsequent report below still shows behavior failures. - [Fresh 12-case Claude report](https://pages.paperclip.ing/runner-e2e-first-task-claude-35102737804/) at `78452129e`: 10/12 pass after correcting two false proposal-matcher failures. The recordings said “Here is the task I will create and run/complete” in approval cards; the old matcher missed that word order. Regression tests failed before the fix and pass after it. Original results and offline regrade provenance remain linked. No agent rerun was needed. Zero single-option-question failures; two behavior failures remain: direct work before acceptance on a plain first message, and an explicit plan request without a saved plan. Neither check was relaxed. The follow-up `82087ac7e` fixes command-prefix size accounting; `94aefb1f3` fixes only that proposal matcher. - Report browser checks confirm folded conversations, rendered cards, explicit Local/Daytona labels, and no page errors. The published-object audit scanned 1,306 text files across 2,154 objects with no credential-format findings or prohibited files. Image pixels and unknown token formats are outside that scan. ## Risks - Model behavior is nondeterministic. One campaign is evidence, not a guarantee. The two remaining Claude behavior failures are visible in the report and require further product work; this PR does not claim all onboarding scenarios pass. - The suite checks persisted Paperclip effects. It cannot prove the absence of arbitrary external effects. - Historical recordings can miss later journey steps. These remain incomplete, never passes. - Native profiles switch runtime after the production onboarding wizard because it does not yet expose a native option. - Quality scores are informational and cannot override behavioral failures. ## Model Used OpenAI Codex, GPT-6, with reasoning, repository tools, and code execution. The exact deployed model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
5acf56658b |
feat(onboarding): first task opens as a chat with a chief of staff (#13068)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Onboarding ends by handing a new user to their first agent on a seeded first task > - Today the wizard asks for a mission up front, the UI composes what the agent is told, and the agent starts running before the user says anything > - New users get a cold, ticket-shaped start, and nobody can edit the agent's brief or persona without a code change > - This pull request makes the first task a short chat: a four-step wizard, a chief-of-staff persona, a greeting plus a two-option opening card, server-owned markdown texts, and no run until the user answers > - It also gives question cards one consistent action row (Cancel / Skip / Next), makes agent hires idempotent within a run, and turns the Paperclip Runner flag on by default for self-hosted instances > - The benefit is a first run the user steers, with texts a board operator can edit as markdown ## Linked Issues or Issue Description No public GitHub issue exists for this change. The feature request fields follow. Related PRs and issues: - Refs #11043 — an earlier draft of the first-task onboarding experience. This PR supersedes it. - Refs #11280 — a report about the onboarding first-task route test. This PR extends that test file. ### Subsystem affected Onboarding wizard, the seeded first task and its texts, task-chat question cards, agent hiring, and the instance experimental settings. ### Problem or motivation The onboarding wizard collects a mission through two extra steps and a questionnaire. The UI then composes the first agent's instructions and the first task description from those answers. The first task wakes the agent at once, so the agent runs and posts before the user types a word. Board operators cannot change the greeting, the brief, or the persona without editing TypeScript. Question cards in chat behave differently per adapter, and a single-select pick submits on click. A misread hire response could create a duplicate agent that the creating agent cannot remove. ### Proposed solution Reduce the wizard to four steps and stop the UI from authoring agent texts. Move the greeting, the brief, the chief-of-staff persona, and the opening question into markdown and JSON files that the server loads at runtime. Seed the persona onto the first agent through an explicit hire marker. Do not wake the first task until the user answers the opening card or types. Give every question card the same Cancel / Skip / Next actions. Add an experimental toggle that switches the single-task proposal between one confirmation card and a plan document with a checkbox card. Make agent hires idempotent within a run. ### Alternatives considered - Keep the mission questionnaire and feed it into the brief. Rejected: the agent asks better questions in chat, and the wizard gets shorter. - Keep the first task open-ended with a plain composer. Rejected: a two-option card gives the user a clear first move. - Derive the plan-document behaviour from the user's intent only. Rejected in favour of an explicit experimental toggle so operators can choose. - Key the "pick does not submit" behaviour off the presence of a submit label. Rejected: several adapters set a submit label on single-select cards, and their cards would change behaviour. ### Roadmap alignment `ROADMAP.md` lists no planned core work on onboarding or the first task. This change refines the existing flow and does not duplicate planned work. ## What Changed - Wizard: four steps (Name your organization, Create your first agent, Connect a model, Review). The front door and both mission steps are removed with their state and saved-progress keys. The UI no longer composes the first agent's instructions or the first task description. - Server-owned texts: the greeting, the brief with two proposal variants, the chief-of-staff persona, the opening question, and a README live in `server/src/onboarding-assets/first-task/` and load at runtime. The create route stores the assembled brief and ignores any client description. - Persona seed: an `onboardingFirstAgent` marker on the hire lets the server seed the chief-of-staff persona over the first agent's entry file. Board-authored hires only. The persona tells the agent the hire response shape and to list agents before it acts on an unclear result. - No auto-run: the first task does not queue an assignment wake. The stranded-assignment reconciler leaves it idle until a user comment or an answered card exists. - Opening card: the server seeds an `ask_user_questions` card right after the greeting with two options: "Interview me and propose a plan and an agent team to execute it." and "I have a task in mind" with free text. Answering wakes the agent. - Experimental toggle `enableFirstTaskPlanProposal` (default off): the single-task proposal is one confirmation card, or a plan document plus a checkbox card when on. - Question cards: every `ask_user_questions` card renders Cancel, Skip, and Next (the submit label on the last question). Skip hides on required questions. Picking an option no longer advances or submits by itself. - Wizard guards: the dashboard's agentless offer ignores a cached empty agent list while a refetch is in flight. The hire step adopts an agent that already carries the typed name instead of hiring "Name 2". - Agent hires are idempotent within a run: a retry of the identical request under the same run id returns the existing agent with `200` and `idempotent: true`. The fingerprint covers the whole validated request, so a corrected payload is a new hire. Lookup, create, and activity record run under one lock per company and run, so overlapping retries cannot both create. - The Paperclip Runner experimental flag defaults to on for self-hosted instances. Cloud keeps its declared default: a managed instance whose tenant row and managed overlay omit the flag resolves it to off. - Question cards: a send that finds an earlier required answer missing returns to that question with a message instead of failing silently. - The two onboarding e2e specs follow the new wizard: the front door and growth intake shots are gone, and the planning-mode spec dismisses the opening card before it reads the composer. - Docs: `docs/board-operator/editing-first-task-texts.md` explains how to edit the texts and the toggle. ## Verification Commands, run from the repo root: ``` pnpm -r --filter './packages/*' --filter '!@paperclipai/paperclip-runner' build pnpm --filter ./packages/shared typecheck pnpm --filter ./ui typecheck pnpm --filter ./server exec tsc --noEmit pnpm check:token-gates pnpm --filter ./ui exec vitest run OnboardingWizard onboarding QuestionForm InteractionCard ProtocolCard TaskChatComposer Dashboard feature PAPERCLIP_IN_WORKTREE=false pnpm --filter ./server exec vitest run onboarding-first-task heartbeat-process-recovery agent-hire-idempotency instance-settings agent-skills-routes issue-onboarding onboarding-greeting --testTimeout=90000 ``` Results on this branch: - Typecheck is clean for shared, ui, and server. - Token gates: 4 of 4 clean. - UI: 344 tests pass across 23 files. - Server: all suites pass. The first test in `agent-skills-routes` has its own 10 s cap and needs about 15 s on my laptop for the app cold start. It passes with a longer cap. This PR does not change that cap. Manual steps on a dev instance: 1. Open `/onboarding`. Confirm four steps: Name your organization, Create your first agent, Connect a model, Review. 2. Finish the wizard. Confirm the first task shows the chief-of-staff greeting and the opening card with two options. Confirm no run starts. 3. Pick "Interview me…". Confirm no run starts. Press Continue. Confirm a run starts and an interview card of 3–4 questions arrives. 4. On a fresh organization, pick "I have a task in mind", type a task, and press Continue. Confirm a proposal arrives as one confirmation card. 5. Turn on Settings → Experimental → "First task: propose with a plan document" and repeat step 4. Confirm a plan document and a checkbox card arrive. 6. Visit the dashboard after the hire. Confirm the wizard does not reopen and one agent exists. 7. Open any question card. Confirm Cancel returns the plain composer with the card still pending, Skip advances an optional question, and Next moves to the next question. Design reference with flow diagrams, chat mock-ups, and live captures: https://pages.paperclip.ing/first-task-flow/proposed/ ## Risks - `pnpm dev` now builds the runner daemon because the Paperclip Runner flag is on by default. Developers without a Rust toolchain must set `PAPERCLIP_RUNNER_BINARY` or turn the flag off. Self-hosted instances that never set the flag now let qualified agents use the runner. - The wizard drops the mission steps and their saved-progress keys. A user who is mid-wizard on an older build restarts at step 1 after an upgrade. Existing organizations are not touched. - The first task no longer runs on its own. A user who neither answers the card nor types sees no agent activity. This is intended. - The persona seed applies only to hires that carry the marker from the wizard. API hires are unchanged. - Hire idempotency is scoped to one run id and to the exact request. Retries across runs, or with a changed payload, still create a second agent. The lock is per server process, which matches how an instance serves its API. - Single-select question cards no longer submit on pick. Users of adapters that relied on that behaviour now press Next. - No database migrations. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Claude (Anthropic) through Claude Code. `claude-fable-5-1` with extended thinking, tool use, and code execution wrote most commits. `claude-opus-4-8` wrote the toggle, texts, wizard, and idempotency commits, as the `Co-Authored-By` trailers show. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |