mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-10 12:07:09 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Users create tasks with a title and a description. > - A required title adds work when the prompt already explains the request. > - An agent can name the task once it reads that request. > - This pull request accepts prompt-only tasks and starts them with a short prompt slice. > - A scoped title tool lets the assigned agent replace that slice early without changing execution state. > - A live browser eval checks the real agent call, saved title, audit entry, and preservation of user titles. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: task creation, shared contracts, database, server, runner tools, and board UI. **Problem or motivation** Users must currently write a title before they can submit a detailed task prompt. The agent has enough context to write a useful title itself. **Proposed solution** Make the title optional when a description is present. Save the first 120 characters of the normalized prompt as a provisional title. Ask the assigned agent to call `set_task_title` early. Use an atomic provisional-title guard to preserve titles supplied or edited by users. Keep explicit titles supported. Related: #14543 and #14556 concern empty-title submission. This change intentionally enables that submission when a prompt is present, instead of requiring a title. ## What Changed - Add the `titleNeedsGeneration` field with an idempotent migration. Keep existing titles unchanged. - Add `PUT /api/issues/:id/title` and the native and legacy `set_task_title` tool. Enforce company access, active-run ownership, shared, bounded retry receipts across native/HTTP calls, and transactional audit logging. Refresh external-object links after commit, with the same feature gate and plugin detectors as ordinary title edits. - Add early naming guidance in Standard, Ask, and Plan task context. Preserve the description, status, and assignment. - Allow prompt-only root and child task creation, plus draft restoration in the New Task dialog. Keep user titles supported. - Add an opt-in Product E2E suite for prompt-only Standard and Ask tasks, plus an explicit-title control. It checks actual provider calls within the first five tools, persisted state, audit attribution, and the reloaded UI. - Preserve a closed vocabulary of API key maintenance phrases in declared prose while rejecting opaque credential suffixes. Add one bounded naming retry after wording is rejected, without treating the rejected call as a saved title. - Repair the native cleanup receipt check exposed during full verification: accept matching input digests, retain legacy input checks, and reject conflicting receipts. ## Verification - Live Product E2E on `f43478473800e3a46b85c5ee79677efdb15108e7`: **3/3 passed** with native Codex `gpt-5.4-mini`, first attempts only, automatic retries disabled. Standard and Ask each saved “Rotate expired API key” on their first tool call, with matching persisted state and a single same-run audit entry. The explicit-title control retained its user title with zero title writes. All three verified the reloaded browser UI. - Campaign: `local-2026-09-30T21-30-11-021Z`. Earlier failed campaigns are retained separately; they exposed credential-prose handling and prompted the naming recovery fix. No failed result was regraded or deleted. - Reproduce with `pnpm test:e2e:runner -- --id task-titles.runner-codex-mini.local.prompt-title-standard --id task-titles.runner-codex-mini.local.prompt-title-ask --id task-titles.runner-codex-mini.local.preserve-explicit-title --max-automatic-retries 0` and an authorized provider key. - Full `pnpm -r typecheck` and `pnpm build` passed on the latest commit. The runner build used the configured external eval source tree. - Product E2E unit suite: **61 files, 818 tests passed**; E2E typecheck and UI token gates passed. - Title API/native regressions cover prompt-only and explicit child creation, user edits, ownership/company isolation, external reference refresh, cross-surface retry replay, and the 64-key limit without receipt eviction. All passed. Prompt-context coverage: **44 tests passed**. - Rust credential regressions: **35 tests passed**, including benign maintenance qualifiers and opaque credential rejection in every declared prose field. Catalog/report reconciliation: **28 tests passed**. Native recovery: **560 tests passed**. - Broad local `pnpm test:run`: **14,555 tests passed** in the general server group; two suites failed to initialize embedded PostgreSQL and the existing 40,000-file Git streaming stress test exceeded its 300-second macOS timeout. All three suites then passed in isolation (**5 tests passed**) without code or timeout changes. The original full local command exited nonzero and is not being represented as a clean full run. - Latest-head GitHub checks are green: **53 passed, 4 skipped, zero failed or pending**, including all test shards and the canary packaging dry run. Greptile reviewed the same commit at **5/5**, with zero unresolved review threads. ## Risks - The additive database field must reach the server and UI together. The migration uses `IF NOT EXISTS` and defaults existing tasks to a final title. - Title generation depends on the assigned agent running. Tasks without a run keep their provisional title. - Live qualification covers the native Codex path in Standard and Ask modes. API/legacy and Plan behavior have deterministic coverage. - The credential-prose exception validates the entire suffix against a closed maintenance vocabulary. Unknown suffixes, assignments, quoted values, credential prefixes, and diagnostics retain strict checks. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, tool use, and code execution. The exact deployment ID and context window are not exposed in this session. The live eval uses the native Codex `gpt-5.4-mini` profile. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
80 lines
5.5 KiB
TypeScript
80 lines
5.5 KiB
TypeScript
import { test, expect } from "@playwright/test";
|
|
import { readFile } from "node:fs/promises";
|
|
import path from "node:path";
|
|
import { seedWorkspaceBootstrap } from "./seed.mjs";
|
|
|
|
// Opt-in: requires the isolated test-drive launched with this suite's Git shim.
|
|
// Never attach these fault-injection tests to a regular development instance.
|
|
const base = process.env.WORKSPACE_BOOTSTRAP_TEST_URL;
|
|
test.skip(!base, "Launch the disposable workspace-bootstrap test-drive; see README.md");
|
|
|
|
for (const persistent of [false, true]) {
|
|
test(persistent ? "exhausts the shared retry budget with an actionable stop" : "recovers and completes through the task UI without manual Retry", async ({ page }, info) => {
|
|
const fixture = await seedWorkspaceBootstrap(base!, persistent);
|
|
const api = async (route: string) => {
|
|
const response = await page.request.get(`${base}/api${route}`);
|
|
expect(response.ok(), await response.text()).toBeTruthy();
|
|
return response.json();
|
|
};
|
|
const title = `${persistent ? "Bound persistent failure" : "Recover and preserve work"} ${Date.now()}`;
|
|
await page.goto(`${base}/${fixture.prefix}/dashboard`);
|
|
const announcement = page.getByRole("button", { name: "Dismiss announcement" });
|
|
if (await announcement.isVisible()) await announcement.click();
|
|
await page.getByRole("link", { name: "Tasks", exact: true }).click();
|
|
await page.getByRole("button", { name: "New Task", exact: true }).last().click();
|
|
await page.getByPlaceholder("Task title (optional)", { exact: true }).fill(title);
|
|
await page.getByRole("button", { name: "Assignee", exact: true }).click();
|
|
await page.getByRole("button", { name: fixture.agentName, exact: true }).click();
|
|
// Let the closing popover unmount before clicking another popover trigger.
|
|
await expect(page.getByRole("textbox", { name: "Search assignees...", includeHidden: true })).toHaveCount(0);
|
|
// The preceding popover's focus restoration can consume the first click.
|
|
await expect(async () => {
|
|
if (!await page.getByRole("textbox", { name: "Search projects..." }).isVisible()) {
|
|
await page.getByRole("button", { name: "Project", exact: true }).click();
|
|
}
|
|
await expect(page.getByRole("textbox", { name: "Search projects..." })).toBeVisible({ timeout: 1_000 });
|
|
}).toPass({ timeout: 10_000, intervals: [1_000] });
|
|
await page.getByRole("button", { name: fixture.projectName, exact: true }).click();
|
|
await expect(page.getByRole("textbox", { name: "Search projects...", includeHidden: true })).toHaveCount(0);
|
|
await page.getByRole("button", { name: "Create Task", exact: true }).click();
|
|
const taskLink = page.getByRole("complementary").getByRole("link", { name: title, exact: true });
|
|
await expect(taskLink).toBeVisible();
|
|
// Follow the UI's actual persisted link, including across task-tab state.
|
|
const href = await taskLink.getAttribute("href");
|
|
expect(href).toBeTruthy();
|
|
await page.goto(new URL(href!, base).href);
|
|
const tasks = await api(`/companies/${fixture.companyId}/issues`);
|
|
const task = tasks.find((row: { title: string }) => row.title === title);
|
|
const runs = async () => (await api(`/companies/${fixture.companyId}/heartbeat-runs`)).filter((row: { agentId: string }) => row.agentId === fixture.agentId);
|
|
await expect.poll(async () => (await runs()).some((row: { errorCode: string }) => row.errorCode === "workspace_git_scan_timeout"), { timeout: 30_000 }).toBe(true);
|
|
await expect(page.getByText(/Agent resumes in/)).toBeVisible();
|
|
await page.screenshot({ path: info.outputPath("scheduled-retry.png"), fullPage: true });
|
|
await expect.poll(async () => (await api(`/issues/${task.id}`)).status, { timeout: 150_000 }).toBe(persistent ? "blocked" : "done");
|
|
// The worker updates the task before its process exit is persisted.
|
|
await expect.poll(async () => (await runs()).filter((row: { status: string }) => ["running", "queued", "scheduled_retry"].includes(row.status)).length).toBe(0);
|
|
const history = await runs();
|
|
expect(history).toHaveLength(persistent ? 3 : 2);
|
|
expect(history.filter((row: { status: string }) => row.status === "succeeded")).toHaveLength(persistent ? 0 : 1);
|
|
if (persistent) {
|
|
await expect(page.getByText("Workspace scan timed out", { exact: true })).toBeVisible();
|
|
await expect(page.getByText("No live execution path", { exact: true })).toHaveCount(0);
|
|
await page.waitForTimeout(35_000);
|
|
expect(await runs()).toHaveLength(3);
|
|
} else {
|
|
await expect(page.getByText(/Workspace recovered automatically\./)).toBeVisible();
|
|
const project = await api(`/projects/${fixture.projectId}`);
|
|
const sourceCopy = project.workspaces.find((row: { name: string }) => row.name === "Source copy");
|
|
const completed = history.find((row: { status: string }) => row.status === "succeeded");
|
|
// The worker already verified the managed copy with its run-scoped API.
|
|
// Independently confirm the configured source was never overwritten.
|
|
expect(sourceCopy).toBeTruthy();
|
|
expect(completed.scheduledRetryAttempt).toBe(1);
|
|
expect(await readFile(path.join(fixture.source, "README.md"), "utf8")).toBe("Existing uncommitted work\n");
|
|
}
|
|
await page.reload();
|
|
await expect(page.getByRole("button", { name: new RegExp(`^Change status \\(current: ${persistent ? "Blocked" : "Done"}`) }).first()).toBeVisible();
|
|
await expect(page.getByText(/Agent resumes in/)).toHaveCount(0);
|
|
await page.screenshot({ path: info.outputPath("settled-task.png"), fullPage: true });
|
|
});
|
|
}
|