mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 20:05:57 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Users create tasks with a title and a description. > - A required title adds work when the prompt already explains the request. > - An agent can name the task once it reads that request. > - This pull request accepts prompt-only tasks and starts them with a short prompt slice. > - A scoped title tool lets the assigned agent replace that slice early without changing execution state. > - A live browser eval checks the real agent call, saved title, audit entry, and preservation of user titles. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: task creation, shared contracts, database, server, runner tools, and board UI. **Problem or motivation** Users must currently write a title before they can submit a detailed task prompt. The agent has enough context to write a useful title itself. **Proposed solution** Make the title optional when a description is present. Save the first 120 characters of the normalized prompt as a provisional title. Ask the assigned agent to call `set_task_title` early. Use an atomic provisional-title guard to preserve titles supplied or edited by users. Keep explicit titles supported. Related: #14543 and #14556 concern empty-title submission. This change intentionally enables that submission when a prompt is present, instead of requiring a title. ## What Changed - Add the `titleNeedsGeneration` field with an idempotent migration. Keep existing titles unchanged. - Add `PUT /api/issues/:id/title` and the native and legacy `set_task_title` tool. Enforce company access, active-run ownership, shared, bounded retry receipts across native/HTTP calls, and transactional audit logging. Refresh external-object links after commit, with the same feature gate and plugin detectors as ordinary title edits. - Add early naming guidance in Standard, Ask, and Plan task context. Preserve the description, status, and assignment. - Allow prompt-only root and child task creation, plus draft restoration in the New Task dialog. Keep user titles supported. - Add an opt-in Product E2E suite for prompt-only Standard and Ask tasks, plus an explicit-title control. It checks actual provider calls within the first five tools, persisted state, audit attribution, and the reloaded UI. - Preserve a closed vocabulary of API key maintenance phrases in declared prose while rejecting opaque credential suffixes. Add one bounded naming retry after wording is rejected, without treating the rejected call as a saved title. - Repair the native cleanup receipt check exposed during full verification: accept matching input digests, retain legacy input checks, and reject conflicting receipts. ## Verification - Live Product E2E on `f43478473800e3a46b85c5ee79677efdb15108e7`: **3/3 passed** with native Codex `gpt-5.4-mini`, first attempts only, automatic retries disabled. Standard and Ask each saved “Rotate expired API key” on their first tool call, with matching persisted state and a single same-run audit entry. The explicit-title control retained its user title with zero title writes. All three verified the reloaded browser UI. - Campaign: `local-2026-09-30T21-30-11-021Z`. Earlier failed campaigns are retained separately; they exposed credential-prose handling and prompted the naming recovery fix. No failed result was regraded or deleted. - Reproduce with `pnpm test:e2e:runner -- --id task-titles.runner-codex-mini.local.prompt-title-standard --id task-titles.runner-codex-mini.local.prompt-title-ask --id task-titles.runner-codex-mini.local.preserve-explicit-title --max-automatic-retries 0` and an authorized provider key. - Full `pnpm -r typecheck` and `pnpm build` passed on the latest commit. The runner build used the configured external eval source tree. - Product E2E unit suite: **61 files, 818 tests passed**; E2E typecheck and UI token gates passed. - Title API/native regressions cover prompt-only and explicit child creation, user edits, ownership/company isolation, external reference refresh, cross-surface retry replay, and the 64-key limit without receipt eviction. All passed. Prompt-context coverage: **44 tests passed**. - Rust credential regressions: **35 tests passed**, including benign maintenance qualifiers and opaque credential rejection in every declared prose field. Catalog/report reconciliation: **28 tests passed**. Native recovery: **560 tests passed**. - Broad local `pnpm test:run`: **14,555 tests passed** in the general server group; two suites failed to initialize embedded PostgreSQL and the existing 40,000-file Git streaming stress test exceeded its 300-second macOS timeout. All three suites then passed in isolation (**5 tests passed**) without code or timeout changes. The original full local command exited nonzero and is not being represented as a clean full run. - Latest-head GitHub checks are green: **53 passed, 4 skipped, zero failed or pending**, including all test shards and the canary packaging dry run. Greptile reviewed the same commit at **5/5**, with zero unresolved review threads. ## Risks - The additive database field must reach the server and UI together. The migration uses `IF NOT EXISTS` and defaults existing tasks to a final title. - Title generation depends on the assigned agent running. Tasks without a run keep their provisional title. - Live qualification covers the native Codex path in Standard and Ask modes. API/legacy and Plan behavior have deterministic coverage. - The credential-prose exception validates the entire suffix against a closed maintenance vocabulary. Unknown suffixes, assignments, quoted values, credential prefixes, and diagnostics retain strict checks. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, tool use, and code execution. The exact deployment ID and context window are not exposed in this session. The live eval uses the native Codex `gpt-5.4-mini` profile. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
87 lines
6.6 KiB
TypeScript
87 lines
6.6 KiB
TypeScript
import { describe, expect, it } from "vitest";
|
|
import { runnerMatrix, runnerSuites } from "./catalog.js";
|
|
import { parseRunnerSelectors, selectRunnerExecutions } from "./selectors.js";
|
|
import { discoverReportCatalog } from "./report-catalog.js";
|
|
import { gradeTaskTitle, taskTitlePrompt, taskTitleTasks, TASK_TITLE_EARLY_CALL_LIMIT, type TaskTitleEvidenceInput } from "./task-titles.js";
|
|
|
|
const prompt = taskTitlePrompt("fixture");
|
|
const provisional = prompt.replace(/\s+/g, " ").slice(0, 120);
|
|
const title = "API key rotation checklist";
|
|
const event = (eventType: string, payload: unknown) => ({ eventType, payload: { prpEvent: { payload } } });
|
|
const start = (id: string, name: string, input: unknown = {}) => event("item.started", { item: { id, name, type: "tool_use", input } });
|
|
function evidence(): TaskTitleEvidenceInput {
|
|
return {
|
|
prompt, explicitTitle: "", agentId: "agent", runId: "run",
|
|
submitted: { description: prompt },
|
|
initial: { id: "task", title: provisional, titleNeedsGeneration: true, description: prompt },
|
|
final: { id: "task", title, titleNeedsGeneration: false, description: prompt, assigneeAgentId: "agent" },
|
|
events: [
|
|
start("context", "get_task_context"),
|
|
start("naming", "set_task_title", { title, onlyIfProvisional: true, idempotencyKey: "initial-title" }),
|
|
event("item.completed", { item: { id: "naming", tool_use_id: "naming", type: "tool_result", result: { id: "task", title, changed: true, titleNeedsGeneration: false } } }),
|
|
event("tool.execution.completed", { name: "set_task_title", executionId: "naming", status: "completed" }),
|
|
start("finish", "paperclip_finish"),
|
|
],
|
|
activity: [{ entityId: "task", entityType: "issue", action: "issue.updated", actorType: "agent", actorId: "agent", runId: "run", details: { title, previous: { title: provisional } } }],
|
|
};
|
|
}
|
|
|
|
describe("task title Product E2E", () => {
|
|
it("registers six explicit-only local cells, with production instructions and no naming hints in the request", () => {
|
|
const suite = runnerSuites.find(entry => entry.id === "task-titles")!;
|
|
const selected = selectRunnerExecutions(parseRunnerSelectors(["--suite", suite.id]));
|
|
expect(selected).toHaveLength(6);
|
|
expect(selected.every(entry => entry.environment.id === "local" && entry.profile.provider === "codex")).toBe(true);
|
|
expect(selectRunnerExecutions(parseRunnerSelectors(["--all"])).some(entry => entry.suite.id === suite.id)).toBe(false);
|
|
expect(suite.definitionMetadata).toMatchObject({ instructions: "production", providerRuns: 1, budgetMonthlyCents: 500 });
|
|
for (const task of taskTitleTasks) {
|
|
expect(task).toMatchObject({ expectedRunCount: 1, flow: "single_turn" });
|
|
expect(task.buildPrompt("fixture")).not.toMatch(/title|set_task|\/api\/|finish_task|paperclip_finish/i);
|
|
}
|
|
const discovered = discoverReportCatalog({ catalog: runnerMatrix, expected: selected.map(entry => entry.id), results: [] });
|
|
expect(discovered.filter(entry => entry.suite.id === suite.id)).toHaveLength(6);
|
|
});
|
|
|
|
it("accepts early, correlated tool evidence plus the same agent/run's persisted title mutation", () => {
|
|
expect(gradeTaskTitle(evidence())).toMatchObject({ passed: true, initialTitle: provisional, finalTitle: title });
|
|
});
|
|
|
|
it.each([
|
|
["missing evidence", (input: TaskTitleEvidenceInput) => { input.events = []; input.activity = []; }],
|
|
["a missing creation request", (input: TaskTitleEvidenceInput) => { input.submitted = undefined; }],
|
|
["a narrated call", (input: TaskTitleEvidenceInput) => { input.events = [event("message", { text: `set_task_title(${title}) succeeded` })]; }],
|
|
["a stale prompt slice", (input: TaskTitleEvidenceInput) => { (input.final as Record<string, unknown>).title = provisional; }],
|
|
["a generic new title", (input: TaskTitleEvidenceInput) => { (input.final as Record<string, unknown>).title = "Complete the requested task"; }],
|
|
["a title supplied by the fixture", (input: TaskTitleEvidenceInput) => { input.submitted = { title: provisional, description: prompt }; }],
|
|
["missing provisional marker", (input: TaskTitleEvidenceInput) => { (input.initial as Record<string, unknown>).titleNeedsGeneration = false; }],
|
|
["a late title call", (input: TaskTitleEvidenceInput) => { input.events = [...Array.from({ length: TASK_TITLE_EARLY_CALL_LIMIT }, (_, i) => start(`read-${i}`, "get_task_context")), ...input.events]; }],
|
|
["naming after five shell tools", (input: TaskTitleEvidenceInput) => { input.events = [...Array.from({ length: TASK_TITLE_EARLY_CALL_LIMIT }, (_, i) => event("tool.execution.started", { executionId: `shell-${i}`, name: "shell" })), ...input.events]; }],
|
|
["naming after completion", (input: TaskTitleEvidenceInput) => { input.events = [start("already-finished", "paperclip_finish"), ...input.events]; }],
|
|
["an unmatched result", (input: TaskTitleEvidenceInput) => { input.events = input.events.filter(row => row.eventType !== "item.completed"); }],
|
|
["no successful execution receipt", (input: TaskTitleEvidenceInput) => { input.events = input.events.filter(row => row.eventType !== "tool.execution.completed"); }],
|
|
["another run's mutation", (input: TaskTitleEvidenceInput) => { (input.activity[0] as Record<string, unknown>).runId = "other-run"; }],
|
|
["a board-authored mutation", (input: TaskTitleEvidenceInput) => { (input.activity[0] as Record<string, unknown>).actorType = "user"; }],
|
|
["a different task", (input: TaskTitleEvidenceInput) => { (input.final as Record<string, unknown>).id = "other-task"; }],
|
|
["a truncated prompt", (input: TaskTitleEvidenceInput) => { (input.final as Record<string, unknown>).description = provisional; }],
|
|
] as const)("rejects %s", (_label, corrupt) => {
|
|
const input = evidence(); corrupt(input);
|
|
expect(gradeTaskTitle(input).passed).toBe(false);
|
|
});
|
|
|
|
it("preserves an explicit title and rejects a rename even if the original is later restored", () => {
|
|
const input = evidence();
|
|
input.explicitTitle = "My chosen title";
|
|
input.submitted = { title: input.explicitTitle, description: prompt };
|
|
input.initial = { ...(input.initial as object), title: input.explicitTitle, titleNeedsGeneration: false };
|
|
input.final = { ...(input.final as object), title: input.explicitTitle };
|
|
input.events = [start("context", "get_task_context"), start("finish", "paperclip_finish")];
|
|
const forbiddenWrite = input.activity[0]; input.activity = [];
|
|
expect(gradeTaskTitle(input).passed).toBe(true);
|
|
const calls = input.events; input.events = [];
|
|
expect(gradeTaskTitle(input).passed).toBe(false);
|
|
input.events = calls;
|
|
input.activity = [forbiddenWrite];
|
|
expect(gradeTaskTitle(input).passed).toBe(false);
|
|
});
|
|
});
|