mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-08 11:13:44 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Users create tasks with a title and a description. > - A required title adds work when the prompt already explains the request. > - An agent can name the task once it reads that request. > - This pull request accepts prompt-only tasks and starts them with a short prompt slice. > - A scoped title tool lets the assigned agent replace that slice early without changing execution state. > - A live browser eval checks the real agent call, saved title, audit entry, and preservation of user titles. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: task creation, shared contracts, database, server, runner tools, and board UI. **Problem or motivation** Users must currently write a title before they can submit a detailed task prompt. The agent has enough context to write a useful title itself. **Proposed solution** Make the title optional when a description is present. Save the first 120 characters of the normalized prompt as a provisional title. Ask the assigned agent to call `set_task_title` early. Use an atomic provisional-title guard to preserve titles supplied or edited by users. Keep explicit titles supported. Related: #14543 and #14556 concern empty-title submission. This change intentionally enables that submission when a prompt is present, instead of requiring a title. ## What Changed - Add the `titleNeedsGeneration` field with an idempotent migration. Keep existing titles unchanged. - Add `PUT /api/issues/:id/title` and the native and legacy `set_task_title` tool. Enforce company access, active-run ownership, shared, bounded retry receipts across native/HTTP calls, and transactional audit logging. Refresh external-object links after commit, with the same feature gate and plugin detectors as ordinary title edits. - Add early naming guidance in Standard, Ask, and Plan task context. Preserve the description, status, and assignment. - Allow prompt-only root and child task creation, plus draft restoration in the New Task dialog. Keep user titles supported. - Add an opt-in Product E2E suite for prompt-only Standard and Ask tasks, plus an explicit-title control. It checks actual provider calls within the first five tools, persisted state, audit attribution, and the reloaded UI. - Preserve a closed vocabulary of API key maintenance phrases in declared prose while rejecting opaque credential suffixes. Add one bounded naming retry after wording is rejected, without treating the rejected call as a saved title. - Repair the native cleanup receipt check exposed during full verification: accept matching input digests, retain legacy input checks, and reject conflicting receipts. ## Verification - Live Product E2E on `f43478473800e3a46b85c5ee79677efdb15108e7`: **3/3 passed** with native Codex `gpt-5.4-mini`, first attempts only, automatic retries disabled. Standard and Ask each saved “Rotate expired API key” on their first tool call, with matching persisted state and a single same-run audit entry. The explicit-title control retained its user title with zero title writes. All three verified the reloaded browser UI. - Campaign: `local-2026-09-30T21-30-11-021Z`. Earlier failed campaigns are retained separately; they exposed credential-prose handling and prompted the naming recovery fix. No failed result was regraded or deleted. - Reproduce with `pnpm test:e2e:runner -- --id task-titles.runner-codex-mini.local.prompt-title-standard --id task-titles.runner-codex-mini.local.prompt-title-ask --id task-titles.runner-codex-mini.local.preserve-explicit-title --max-automatic-retries 0` and an authorized provider key. - Full `pnpm -r typecheck` and `pnpm build` passed on the latest commit. The runner build used the configured external eval source tree. - Product E2E unit suite: **61 files, 818 tests passed**; E2E typecheck and UI token gates passed. - Title API/native regressions cover prompt-only and explicit child creation, user edits, ownership/company isolation, external reference refresh, cross-surface retry replay, and the 64-key limit without receipt eviction. All passed. Prompt-context coverage: **44 tests passed**. - Rust credential regressions: **35 tests passed**, including benign maintenance qualifiers and opaque credential rejection in every declared prose field. Catalog/report reconciliation: **28 tests passed**. Native recovery: **560 tests passed**. - Broad local `pnpm test:run`: **14,555 tests passed** in the general server group; two suites failed to initialize embedded PostgreSQL and the existing 40,000-file Git streaming stress test exceeded its 300-second macOS timeout. All three suites then passed in isolation (**5 tests passed**) without code or timeout changes. The original full local command exited nonzero and is not being represented as a clean full run. - Latest-head GitHub checks are green: **53 passed, 4 skipped, zero failed or pending**, including all test shards and the canary packaging dry run. Greptile reviewed the same commit at **5/5**, with zero unresolved review threads. ## Risks - The additive database field must reach the server and UI together. The migration uses `IF NOT EXISTS` and defaults existing tasks to a final title. - Title generation depends on the assigned agent running. Tasks without a run keep their provisional title. - Live qualification covers the native Codex path in Standard and Ask modes. API/legacy and Plan behavior have deterministic coverage. - The credential-prose exception validates the entire suffix against a closed maintenance vocabulary. Unknown suffixes, assignments, quoted values, credential prefixes, and diagnostics retain strict checks. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, tool use, and code execution. The exact deployment ID and context window are not exposed in this session. The live eval uses the native Codex `gpt-5.4-mini` profile. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
297 lines
17 KiB
Markdown
297 lines
17 KiB
Markdown
# Runner E2E fixture authoring
|
|
|
|
The fixture catalog is executable production-contract data. Keep it small,
|
|
typed, deterministic, and free of raw credentials.
|
|
|
|
## Suites and matrices
|
|
|
|
A `RunnerSuiteFixture` declares one durable testing purpose: stable ID, label,
|
|
description, profiles, environments, cases, expected size, and definition or
|
|
ranking metadata. Its execution IDs are globally prefixed as
|
|
`<suite>.<profile>.<environment>.<case>`. Add a new suite when the testing
|
|
purpose or desired cross-product differs; do not inflate an existing suite with
|
|
unrelated dimensions.
|
|
|
|
The suite definition fingerprint is historical comparison metadata. Any
|
|
profile, model qualification, environment, task, or ranking-snapshot change
|
|
must change that fingerprint automatically so the dashboard can annotate the
|
|
boundary instead of silently joining unlike totals.
|
|
|
|
## Agent profiles
|
|
|
|
Add `RunnerProfileFixture` entries in `catalog.ts`. A profile declares:
|
|
|
|
- a stable ID and searchable groups;
|
|
- legacy or native generation;
|
|
- adapter/provider and required credential;
|
|
- a model imported from its adapter constant or qualified runner profile;
|
|
- supported environment IDs;
|
|
- expected runtime metadata; and
|
|
- an agent payload factory.
|
|
|
|
Do not duplicate model IDs, qualification decisions, CLI versions, or runner
|
|
artifact rules. Codex profiles import `DEFAULT_CODEX_LOCAL_MODEL`, OpenCode
|
|
profiles import `QUALIFIED_OPENCODE_MODEL`, and ACPX profiles import
|
|
`QUALIFIED_ACPX_PROFILES`. Add or qualify models at their owning production
|
|
source first.
|
|
|
|
OpenRouter breadth profiles are generated from `openrouter-models.json`, not
|
|
written by hand. That reviewed snapshot must contain exactly five unique,
|
|
available, tool-capable models with rank, canonical ID, display name, supported
|
|
parameters, source URL, capture time, and verified content hash. Refresh it
|
|
manually with `pnpm test:e2e:runner:models:update`; nightly campaigns never
|
|
change fixture definitions.
|
|
|
|
Agent `adapterConfig.env` values must be `{type:"secret_ref", secretId,
|
|
version:"latest"}` objects supplied to the factory. A fixture source containing
|
|
a raw secret-looking value is rejected by catalog validation.
|
|
|
|
The manual Grok subscription profile uses `GROK_AUTH_JSON` as an explicit login
|
|
fixture. It does not put this credential in agent configuration or substitute an
|
|
API key. Setup seeds a new company-scoped Grok home inside the disposable instance
|
|
with mode 0700 and an exclusive mode-0600 auth file. Setup rejects redirected,
|
|
occupied, or nonisolated homes. Production runner discovery and refresh operate on
|
|
that company login; teardown destroys it after the remote environment is removed.
|
|
This fixture tests subscription execution, not the interactive browser login flow.
|
|
|
|
## Environments
|
|
|
|
An `EnvironmentFixture` declares driver/provider, credential requirements,
|
|
attempt deadline, lifecycle behavior, expected execution target, and a payload
|
|
factory validated by the shared environment schema.
|
|
|
|
The local environment is instance-managed: company creation ensures it exists,
|
|
and the public API intentionally rejects a second local environment. The setup
|
|
registry therefore discovers that row through the public environments API.
|
|
This still provides full isolation because every cell starts a new Paperclip
|
|
instance and database.
|
|
|
|
Daytona creates sandbox environments through the public API. The core fixture
|
|
keeps `reuseLease:false` and `runnerLifecycleMode:"per_turn"`. The dedicated
|
|
warm-continuity fixture uses `reuseLease:true` and
|
|
`runnerLifecycleMode:"warm"`; its distinct `configurationKey` is part of the
|
|
suite fingerprint even though both fixtures report `environmentId:"daytona"`.
|
|
Keep short provider cleanup backstops, a Daytona secret reference, and an
|
|
immutable image digest. Teardown
|
|
must delete the environment with reusable-lease destruction and must fail the
|
|
cell if cleanup cannot be confirmed. Keep CPU, memory, and disk explicit: lease
|
|
metadata and the per-test public-list-price runtime estimate depend on that
|
|
pinned billable resource shape. Changing it requires updating billing tests and
|
|
reviewing the versioned Daytona rates in `billing.ts`.
|
|
|
|
## Usage and billing data
|
|
|
|
Do not add fixture-authored token or dollar expectations. The live harness
|
|
reads usage from selected public heartbeat-run records and records coverage per
|
|
run. Provider-reported dollars remain distinct from runtime estimates. A zero
|
|
or missing native usage payload is `unavailable` unless a real token-bearing
|
|
receipt or provider cost proves otherwise. New execution environments must
|
|
provide lease/resource metadata for a runtime estimate or explicitly remain
|
|
`unavailable`; never infer that missing billing data means free execution.
|
|
|
|
Future providers (SSH, E2B, Modal, Cloudflare, Kubernetes, Novita, exe.dev)
|
|
should implement the same setup/probe/cleanup contract before being added to a
|
|
matrix. Unsupported profile/environment combinations belong in
|
|
`supportedEnvironments`, not in ad hoc test conditionals.
|
|
|
|
## Task cases and matchers
|
|
|
|
A `RunnerTaskFixture` owns a work mode, a typed flow, expected run count,
|
|
nonce-based title/prompt/marker factories, per-environment attempt deadlines,
|
|
deterministic matchers, and expected terminal state. Single-turn prompts should
|
|
make one bounded request with observable output and no nondeterministic judging.
|
|
The `plan_revision_acceptance` flow must also provide revision-request and Plan
|
|
marker factories. `question_resume_completion` must define the deterministic
|
|
browser answer and prove exactly two successful runs with no pending
|
|
interaction. `plan_approval_completion` must target the exact two-step
|
|
canonical Plan revision, capture its pending UI, approve in the browser, and
|
|
prove exactly two successful runs. `warm_three_turn` provides exactly two
|
|
browser follow-up messages, preserves one project/execution-workspace scope,
|
|
verifies host file contents after every turn, and finishes within three
|
|
ten-minute turn deadlines. The ordinary warm fixture uses managed instructions,
|
|
updates AGENT_HOME each turn, and verifies memory, an unchanged 8 MiB binary and
|
|
a deletion through public file APIs. Native turns 2 and 3 must copy/hash only the
|
|
changed memory file, with a saved receipt and the same provider PID. Journal and
|
|
Git stress fixtures retain fixed external bundles as controls. Keep the stable-PID
|
|
oracle strict; `instruction-persistence` also covers cold restarts and quota handling.
|
|
Native turns 1 and 2 include an actionable human review in the completion report's `attentionRequests`. Paperclip creates the review gate from that report. An explicit question-tool wait yields the turn and suppresses its final prose, so it is not interchangeable with this completion-review fixture. Turn 3 reports Done without another review.
|
|
|
|
Every selected case runs in its own isolated Paperclip process, and independent
|
|
cases may run concurrently. Follow-up turns inside one case retain their shared
|
|
task state. Each case creates and tears down its own company, secrets,
|
|
environment selection, agent, and browser-created task. The current plan case
|
|
proves three runs on the same issue: publish a two-step Plan,
|
|
request a three-step revision through the UI, and accept the exact new revision
|
|
through the UI before verifying implementation and Done.
|
|
|
|
The matcher union supports message exact/contains/regex/ordered checks, issue
|
|
and run state, runtime/environment metadata, files, artifacts, JSON paths, and
|
|
JSON Schema. The initial cases use normalized `message_contains` plus state,
|
|
runtime, and environment assertions; the plan flow additionally verifies
|
|
canonical document revision IDs, bodies, step counts, interaction targets, and
|
|
visible previews. Add matcher behavior and credential-free tests together.
|
|
|
|
Adding a task expands its suite's matrix. Update the suite's intentional size,
|
|
the complete-catalog size, and credential-free unit tests in the same change.
|
|
Paid tests never silently skip a missing credential or unsupported artifact.
|
|
|
|
## Prompt-only task title fixtures
|
|
|
|
`task-titles.ts` defines a bounded ordinary writing request and an independent
|
|
title oracle. Its `single_turn` cases leave the title field empty or supply an
|
|
explicit control title. The harness captures the exact browser creation response
|
|
instead of searching by a title that the agent may already have changed. It
|
|
never patches the title itself. Normal production instructions own the early
|
|
naming behavior; fixture prompts and agent instruction bundles contain no naming
|
|
hints. Existing company/secret/environment/agent registry dependencies are reused,
|
|
with 500-cent company and agent budgets and normal instance teardown.
|
|
|
|
Keep the call input, successful result, execution receipt, saved task, and
|
|
agent/run-attributed audit correlated. Missing evidence must fail. The first-five
|
|
tool-call bound counts calls in the initial provider run, including discovery.
|
|
The title must describe API-key rotation without requiring one exact wording.
|
|
The control must retain its title throughout, not merely restore it at the end.
|
|
The source digest versions the grader and request in catalog metadata. See
|
|
[Automatic task titles](README.md#automatic-task-titles) for live selectors,
|
|
coverage limits, evidence, and calibration.
|
|
|
|
## New Paperclip object fixtures
|
|
|
|
The explicit-only `lifecycle-baseline` suite reuses this registry and existing
|
|
continuation, chat and governed-action flows. Its narrative pairs require actual
|
|
agent/run-attributed comments or exact visible responses. See
|
|
[the live baseline contract](LIFECYCLE-BASELINE.md) for selectors and proof boundaries.
|
|
|
|
Register new objects in `live-fixtures.ts` with explicit dependencies in
|
|
`FixtureRegistry`. Setup must use a public API. Teardown runs in reverse order
|
|
and is invoked after partial setup failures. Direct database writes and private
|
|
test-only runner endpoints are prohibited.
|
|
|
|
The expected dependency shape is:
|
|
|
|
```text
|
|
company
|
|
└── encrypted secrets
|
|
└── environment
|
|
└── agent
|
|
└── browser-created task
|
|
```
|
|
|
|
Projects, goals, apps, and configuration fixtures can be inserted into that
|
|
graph without changing the launcher. Keep returned fixture state to IDs and
|
|
sanitized metadata; never retain raw secret values.
|
|
|
|
## Required checks
|
|
|
|
Run before a fixture change is reviewed:
|
|
|
|
```bash
|
|
pnpm test:e2e:runner:unit
|
|
pnpm test:e2e:runner:typecheck
|
|
pnpm test:e2e:runner -- --list
|
|
```
|
|
|
|
Then run the narrowest paid cell that exercises the fixture. A full matrix is a
|
|
manual or scheduled campaign, not a PR requirement.
|
|
|
|
|
|
## Persistent chat fixtures
|
|
|
|
`chat-cases.ts` defines the eight-case `agent-chat` suite; `chat-flow.ts` drives the
|
|
production composer, plan revision/approval controls, questions, reset command,
|
|
and project cards. Keep its 28 local cells intentional. `expectedRunCount`
|
|
counts provider turns, including cancelled and handed-off task runs, but excludes
|
|
synthetic `/new` runs. Assertions must inspect all company runs because ordinary
|
|
issue lists exclude the source conversation. `assertChatHandoff` rejects missing
|
|
projects/plans, chat children, wrong assignees, and execution before plan commit.
|
|
|
|
Retained `api-state.json`, `chat-handoff.json`, and plan-revision evidence
|
|
include persisted comments, session generations, run context and logs, project
|
|
workspaces, task documents, and ordering. They pass through the normal sanitizer.
|
|
Screenshots are allowlisted to the exact disposable agent chat. Cleanup cancels
|
|
all active runs in the isolated company, including handed-off work; usage from
|
|
failed and cancelled runs must not disappear from campaign totals.
|
|
|
|
|
|
Warm three-turn continuity grades the exact workspace file after each turn,
|
|
task completion, and sandbox/session identity. It also requires a visible
|
|
persisted final reply with each turn marker once and in order. It does not
|
|
grade exact final-reply wording; the hello
|
|
and continuation fixtures retain those exact-response checks. This separates
|
|
workspace persistence failures from model response-format variance.
|
|
|
|
`chat-hardening.ts` adds the explicit-only `agent-chat-hardening` journeys. Use
|
|
the ordinary public APIs to seed source documents and blockers. Keep the answer
|
|
out of the user's status/review request. Grade the exact source values, latest
|
|
blocker, preserved task identities, worker-authored output, and real executions.
|
|
The status request asks for JSON so the grader can distinguish the current
|
|
blocker from a historical mention and compare active-run count separately from
|
|
task status. The request must not reveal those expected values.
|
|
Capture the source after seeding and compare every field in the public issue
|
|
update contract, plus labels, dependencies, and dedicated-endpoint settings.
|
|
Derived inbound references may change when the chat legitimately cites a task.
|
|
The lost-acknowledgement probe may interrupt only the fixture browser's own
|
|
comment request after the real server has committed it. Retain its request ID
|
|
and replay that same request through the public API after restarting the server.
|
|
Never fabricate tool results or repair task state after a failed assertion.
|
|
|
|
`chat-stories.ts` seeds an ordinary file wait in the isolated agent's actual
|
|
home workspace; native Codex intentionally cannot see arbitrary host temp files.
|
|
The observed run workspace must match the fixture location. This is a deterministic interruption
|
|
boundary. The real provider command writes the readiness file and waits at most
|
|
two minutes. The harness must persist the next browser message while the same
|
|
run is active before supplying the brief. Always release the wait in `finally`.
|
|
Save boundary observations independently of the final outcome. The final answer
|
|
must recover a brief reference absent from both prompts; the revision oracle
|
|
also reads the actual conversation plan. Fixture setup never enables native API
|
|
tools for this suite. Do not describe its prepared-agent settings case as a
|
|
production onboarding qualification.
|
|
|
|
The `agent-chat-qualification` local fixtures use public APIs to seed two workers
|
|
and a task with a saved plan, or read-only tasks with contradictory historical
|
|
comments. Ordinary Node file waits in the isolated agent workspace establish
|
|
observable active execution; no provider output or database outcome is fabricated.
|
|
A worker-crash case sends SIGKILL only to a positively identified running native
|
|
worker PID, then uses the production Retry button. Each gate is released in a
|
|
finally block. Source facts and boundary state are retained with the attempt.
|
|
The lifecycle suite also includes two legacy disposition-repair probes. Their
|
|
first provider turn intentionally omits task disposition, and their second turn
|
|
must be an automatic, causally bound repair that records completion. They use
|
|
public task comments/status APIs and run-detail evidence; no private runtime
|
|
hooks or database mutations are used by the fixture.
|
|
|
|
The explicit-only `extended-harnesses` suite uses five bounded journeys for each
|
|
pending ACP candidate on local and Daytona. Candidate profile metadata includes
|
|
the exact authenticated discovery choice without promoting it to a product
|
|
default. Its file case anchors the task to a public project workspace, validates
|
|
the model's claimed result by reading the actual final bytes, and also exercises
|
|
remote copy-back. Keep candidate admission scoped to the selected model and the
|
|
isolated operator environment; ordinary agent configuration must not enable it.
|
|
|
|
## Persistent agent files
|
|
|
|
The `instruction_persistence` flow uses production managed storage and public file
|
|
APIs. The browser creates a supporting file, then a real agent edits its registered
|
|
AGENT_HOME with ordinary filesystem tools. Independent oracles verify instructions,
|
|
nested text, binary download bytes, and a stopped-run save receipt without new
|
|
revision history. The harness restarts the server and creates a fresh browser task
|
|
without disclosing the saved nonces. Its readback oracle downloads and verifies an
|
|
attachment's bytes and SHA-256, rather than accepting a filename or model claim.
|
|
A third task uploads a ready attachment and waits in an ordinary bounded shell
|
|
command while the board changes the current file through the public API. Stopped
|
|
cleanup must preserve the original candidate as a conflict. The browser reviews
|
|
current and incoming files and applies the run edits against the reviewed current
|
|
directory hash. All three tasks' runs count toward billing and teardown. The suite
|
|
is explicit-only. No private control-plane hooks or direct database writes are used.
|
|
|
|
## Direct blocker fixtures
|
|
|
|
`blocker-cases.ts`, `blocker-fixtures.ts`, `blocker-flow.ts`, and
|
|
`blocker-scoring.ts` define the explicit local legacy `blocker-guidance` suite.
|
|
Its fixture registry creates a manager through the public API and assigns the
|
|
production operational skill to worker and manager. Company-wide evidence and
|
|
cleanup include unexpected manager runs. The grader checks saved human input,
|
|
requester identity for scope questions, ownership history, no additional work or
|
|
hires, and the browser-answer continuation. See [Direct blocker guidance](README.md#direct-blocker-guidance)
|
|
for coverage boundaries and run commands.
|