mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 21:05:21 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Users create tasks with a title and a description. > - A required title adds work when the prompt already explains the request. > - An agent can name the task once it reads that request. > - This pull request accepts prompt-only tasks and starts them with a short prompt slice. > - A scoped title tool lets the assigned agent replace that slice early without changing execution state. > - A live browser eval checks the real agent call, saved title, audit entry, and preservation of user titles. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: task creation, shared contracts, database, server, runner tools, and board UI. **Problem or motivation** Users must currently write a title before they can submit a detailed task prompt. The agent has enough context to write a useful title itself. **Proposed solution** Make the title optional when a description is present. Save the first 120 characters of the normalized prompt as a provisional title. Ask the assigned agent to call `set_task_title` early. Use an atomic provisional-title guard to preserve titles supplied or edited by users. Keep explicit titles supported. Related: #14543 and #14556 concern empty-title submission. This change intentionally enables that submission when a prompt is present, instead of requiring a title. ## What Changed - Add the `titleNeedsGeneration` field with an idempotent migration. Keep existing titles unchanged. - Add `PUT /api/issues/:id/title` and the native and legacy `set_task_title` tool. Enforce company access, active-run ownership, shared, bounded retry receipts across native/HTTP calls, and transactional audit logging. Refresh external-object links after commit, with the same feature gate and plugin detectors as ordinary title edits. - Add early naming guidance in Standard, Ask, and Plan task context. Preserve the description, status, and assignment. - Allow prompt-only root and child task creation, plus draft restoration in the New Task dialog. Keep user titles supported. - Add an opt-in Product E2E suite for prompt-only Standard and Ask tasks, plus an explicit-title control. It checks actual provider calls within the first five tools, persisted state, audit attribution, and the reloaded UI. - Preserve a closed vocabulary of API key maintenance phrases in declared prose while rejecting opaque credential suffixes. Add one bounded naming retry after wording is rejected, without treating the rejected call as a saved title. - Repair the native cleanup receipt check exposed during full verification: accept matching input digests, retain legacy input checks, and reject conflicting receipts. ## Verification - Live Product E2E on `f43478473800e3a46b85c5ee79677efdb15108e7`: **3/3 passed** with native Codex `gpt-5.4-mini`, first attempts only, automatic retries disabled. Standard and Ask each saved “Rotate expired API key” on their first tool call, with matching persisted state and a single same-run audit entry. The explicit-title control retained its user title with zero title writes. All three verified the reloaded browser UI. - Campaign: `local-2026-09-30T21-30-11-021Z`. Earlier failed campaigns are retained separately; they exposed credential-prose handling and prompted the naming recovery fix. No failed result was regraded or deleted. - Reproduce with `pnpm test:e2e:runner -- --id task-titles.runner-codex-mini.local.prompt-title-standard --id task-titles.runner-codex-mini.local.prompt-title-ask --id task-titles.runner-codex-mini.local.preserve-explicit-title --max-automatic-retries 0` and an authorized provider key. - Full `pnpm -r typecheck` and `pnpm build` passed on the latest commit. The runner build used the configured external eval source tree. - Product E2E unit suite: **61 files, 818 tests passed**; E2E typecheck and UI token gates passed. - Title API/native regressions cover prompt-only and explicit child creation, user edits, ownership/company isolation, external reference refresh, cross-surface retry replay, and the 64-key limit without receipt eviction. All passed. Prompt-context coverage: **44 tests passed**. - Rust credential regressions: **35 tests passed**, including benign maintenance qualifiers and opaque credential rejection in every declared prose field. Catalog/report reconciliation: **28 tests passed**. Native recovery: **560 tests passed**. - Broad local `pnpm test:run`: **14,555 tests passed** in the general server group; two suites failed to initialize embedded PostgreSQL and the existing 40,000-file Git streaming stress test exceeded its 300-second macOS timeout. All three suites then passed in isolation (**5 tests passed**) without code or timeout changes. The original full local command exited nonzero and is not being represented as a clean full run. - Latest-head GitHub checks are green: **53 passed, 4 skipped, zero failed or pending**, including all test shards and the canary packaging dry run. Greptile reviewed the same commit at **5/5**, with zero unresolved review threads. ## Risks - The additive database field must reach the server and UI together. The migration uses `IF NOT EXISTS` and defaults existing tasks to a final title. - Title generation depends on the assigned agent running. Tasks without a run keep their provisional title. - Live qualification covers the native Codex path in Standard and Ask modes. API/legacy and Plan behavior have deterministic coverage. - The credential-prose exception validates the entire suffix against a closed maintenance vocabulary. Unknown suffixes, assignments, quoted values, credential prefixes, and diagnostics retain strict checks. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, tool use, and code execution. The exact deployment ID and context window are not exposed in this session. The live eval uses the native Codex `gpt-5.4-mini` profile. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
127 lines
8.4 KiB
TypeScript
127 lines
8.4 KiB
TypeScript
import { createHash } from "node:crypto";
|
|
import { readFileSync } from "node:fs";
|
|
import type { RunnerTaskFixture } from "./types.js";
|
|
|
|
export const TASK_TITLE_BUDGET_CENTS = 500;
|
|
export const TASK_TITLE_EARLY_CALL_LIMIT = 5;
|
|
export const taskTitleDefinitionDigest = createHash("sha256").update(readFileSync(new URL(import.meta.url))).digest("hex");
|
|
|
|
// Deliberately no title/tool/API instructions: production guidance must cause naming.
|
|
export function taskTitlePrompt(nonce: string) {
|
|
return [
|
|
"I am preparing a short internal handover for a teammate who is taking over a routine maintenance job tomorrow. Keep the answer practical and brief.",
|
|
"Give me a three-item checklist for safely rotating an expired API key: create its replacement, update the consuming service, and verify access before revoking the old key.",
|
|
`Include receipt CHECKLIST-${nonce} in your answer. This is writing only; do not access a real account, change credentials, or create files.`,
|
|
].join("\n\n");
|
|
}
|
|
|
|
export const taskTitleTasks: readonly RunnerTaskFixture[] = ([
|
|
["prompt-title-standard", "Name a prompt-only task", "standard"],
|
|
["prompt-title-ask", "Name a prompt-only Ask task", "ask"],
|
|
["preserve-explicit-title", "Preserve the user's title", "standard"],
|
|
] as const).map(([id, label, workMode]) => ({
|
|
id, label, groups: [], workMode,
|
|
flow: "single_turn", expectedRunCount: 1,
|
|
attemptTimeoutMs: { local: 6 * 60_000, daytona: 6 * 60_000 },
|
|
expectedTerminalState: { issue: "done", run: "succeeded" },
|
|
buildTitle: nonce => id === "preserve-explicit-title" ? `Teammate handover ${nonce}` : "",
|
|
buildPrompt: taskTitlePrompt,
|
|
buildVisibleMarker: nonce => `CHECKLIST-${nonce}`,
|
|
buildMatchers: (nonce, execution) => [
|
|
{ kind: "message_contains", expected: `CHECKLIST-${nonce}` },
|
|
{ kind: "issue_status", expected: "done" },
|
|
{ kind: "run_status", expected: "succeeded" },
|
|
{ kind: "runtime_mode", expected: execution.profile.expectedRuntimeMode },
|
|
{ kind: "environment", expected: execution.environment.id },
|
|
],
|
|
}));
|
|
|
|
const record = (value: unknown): Record<string, unknown> =>
|
|
value !== null && typeof value === "object" && !Array.isArray(value) ? value as Record<string, unknown> : {};
|
|
|
|
export interface TaskTitleEvidenceInput {
|
|
initial: unknown;
|
|
final: unknown;
|
|
submitted: unknown;
|
|
prompt: string;
|
|
explicitTitle: string;
|
|
agentId: string;
|
|
runId: string;
|
|
events: readonly { eventType?: string; payload?: unknown }[];
|
|
activity: readonly unknown[];
|
|
}
|
|
|
|
/** Grade real call/result/receipt evidence and persisted state, never model prose. */
|
|
export function gradeTaskTitle(input: TaskTitleEvidenceInput) {
|
|
const initial = record(input.initial), final = record(input.final), submitted = record(input.submitted);
|
|
const provisional = input.prompt.trim().replace(/\s+/g, " ").slice(0, 120);
|
|
const checks: Array<{ id: string; passed: boolean; detail: string }> = [];
|
|
const check = (id: string, passed: boolean, detail: string) => checks.push({ id, passed, detail });
|
|
const starts = new Map<string, { name: string; input: Record<string, unknown>; position: number; index: number }>();
|
|
const results = new Map<string, Record<string, unknown>>();
|
|
const completions = new Map<string, number>();
|
|
let position = 0;
|
|
input.events.forEach((event, index) => {
|
|
const payload = record(record(record(event.payload).prpEvent).payload);
|
|
const item = record(payload.item);
|
|
if (event.eventType === "tool.execution.started" && typeof payload.executionId === "string" && !starts.has(payload.executionId)) {
|
|
starts.set(payload.executionId, { name: String(payload.name ?? ""), input: {}, position: ++position, index });
|
|
}
|
|
if (event.eventType === "item.started" && item.type === "tool_use" && typeof item.id === "string" && typeof item.name === "string") {
|
|
const prior = starts.get(item.id);
|
|
starts.set(item.id, { name: item.name, input: record(item.input), position: prior?.position ?? ++position, index: prior?.index ?? index });
|
|
}
|
|
if (event.eventType === "item.completed" && item.type === "tool_result" && typeof item.id === "string"
|
|
&& (item.tool_use_id === undefined || item.tool_use_id === item.id)) results.set(item.id, record(item.result));
|
|
if (event.eventType === "tool.execution.completed" && payload.name === "set_task_title" && payload.status === "completed" && typeof payload.executionId === "string") {
|
|
completions.set(payload.executionId, index);
|
|
}
|
|
});
|
|
const calls = [...starts].filter(([, call]) => call.name === "set_task_title");
|
|
const titleWrites = input.activity.map(record).filter(row => row.action === "issue.updated"
|
|
&& row.entityType === "issue" && row.entityId === initial.id && typeof record(row.details).title === "string");
|
|
check("creation-request", submitted.description === input.prompt
|
|
&& (input.explicitTitle ? submitted.title === input.explicitTitle : !Object.hasOwn(submitted, "title")), "Browser submitted the full prompt with the requested title, or omitted the title entirely.");
|
|
check("provider-tool-evidence", starts.size > 0 && Boolean(input.runId), "Provider tool evidence from the initial run is present; narration alone is insufficient.");
|
|
check("initial-title", initial.title === (input.explicitTitle || provisional)
|
|
&& initial.titleNeedsGeneration === !Boolean(input.explicitTitle), "Creation response proves the initial title and server-owned generation marker before the agent runs.");
|
|
check("same-task-and-description", typeof initial.id === "string" && final.id === initial.id
|
|
&& initial.description === input.prompt && final.description === input.prompt
|
|
&& final.assigneeAgentId === input.agentId, "The original task, assignment, and complete prompt survive naming.");
|
|
check("generation-settled", final.titleNeedsGeneration === false, "The stored task no longer needs a generated title.");
|
|
if (input.explicitTitle) {
|
|
check("explicit-title-preserved", final.title === input.explicitTitle && titleWrites.length === 0
|
|
&& calls.every(([id]) => results.get(id)?.changed === false), "No title-changing audit or tool result may overwrite a supplied title, even temporarily.");
|
|
} else {
|
|
const title = typeof final.title === "string" ? final.title : "";
|
|
check("descriptive-title", title.trim() === title && title.length >= 8 && title.length <= 120
|
|
&& title !== provisional && !input.prompt.startsWith(title) && !title.includes("CHECKLIST-")
|
|
&& /(rotat|replac|renew|expir)/i.test(title) && /\b(api|keys?|credentials?)\b/i.test(title), "A concise title describes the key-rotation request instead of copying the prompt prefix or answer receipt.");
|
|
const terminalIndex = input.events.findIndex(event => event.eventType === "run.terminal");
|
|
const finishIndex = Math.min(
|
|
[...starts.values()].find(call => ["paperclip_finish", "finish_task"].includes(call.name))?.index ?? input.events.length,
|
|
terminalIndex < 0 ? input.events.length : terminalIndex,
|
|
);
|
|
const proven = calls.filter(([id, call]) => {
|
|
const result = results.get(id);
|
|
const completedAt = completions.get(id);
|
|
return result !== undefined && call.position <= TASK_TITLE_EARLY_CALL_LIMIT && call.input.onlyIfProvisional === true
|
|
&& typeof call.input.idempotencyKey === "string" && call.input.idempotencyKey.trim().length > 0
|
|
&& call.input.title === title && result.id === initial.id && result.title === title
|
|
&& result.changed === true && result.titleNeedsGeneration === false
|
|
&& completedAt !== undefined && completedAt > call.index && completedAt < finishIndex;
|
|
});
|
|
check("early-title-tool", proven.length > 0, `A correlated set_task_title call must succeed among the first ${TASK_TITLE_EARLY_CALL_LIMIT} tools of the initial run, before completion.`);
|
|
check("agent-title-audit", titleWrites.length === 1 && titleWrites[0]?.actorType === "agent"
|
|
&& titleWrites[0]?.actorId === input.agentId && titleWrites[0]?.runId === input.runId
|
|
&& record(titleWrites[0]?.details).title === title
|
|
&& record(record(titleWrites[0]?.details).previous).title === provisional, "The same assigned agent and initial run must own the single persisted title mutation.");
|
|
}
|
|
return {
|
|
schema: "paperclip.task-title-eval.v1", passed: checks.every(entry => entry.passed), checks,
|
|
initialTitle: initial.title, finalTitle: final.title,
|
|
titleCalls: calls.map(([id, call]) => ({ id, position: call.position, completed: completions.has(id), result: results.get(id) ?? null })),
|
|
titleWrites,
|
|
};
|
|
}
|