Files
DottaandPaperclip 018993140f feat: let agents name prompt-only tasks (#14761)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Users create tasks with a title and a description.
> - A required title adds work when the prompt already explains the
request.
> - An agent can name the task once it reads that request.
> - This pull request accepts prompt-only tasks and starts them with a
short prompt slice.
> - A scoped title tool lets the assigned agent replace that slice early
without changing execution state.
> - A live browser eval checks the real agent call, saved title, audit
entry, and preservation of user titles.

## Linked Issues or Issue Description

**Subsystem affected**

Cross-cutting: task creation, shared contracts, database, server, runner
tools, and board UI.

**Problem or motivation**

Users must currently write a title before they can submit a detailed
task prompt. The agent has enough context to write a useful title
itself.

**Proposed solution**

Make the title optional when a description is present. Save the first
120 characters of the normalized prompt as a provisional title. Ask the
assigned agent to call `set_task_title` early. Use an atomic
provisional-title guard to preserve titles supplied or edited by users.
Keep explicit titles supported.

Related: #14543 and #14556 concern empty-title submission. This change
intentionally enables that submission when a prompt is present, instead
of requiring a title.

## What Changed

- Add the `titleNeedsGeneration` field with an idempotent migration.
Keep existing titles unchanged.
- Add `PUT /api/issues/:id/title` and the native and legacy
`set_task_title` tool. Enforce company access, active-run ownership,
shared, bounded retry receipts across native/HTTP calls, and
transactional audit logging. Refresh external-object links after commit,
with the same feature gate and plugin detectors as ordinary title edits.
- Add early naming guidance in Standard, Ask, and Plan task context.
Preserve the description, status, and assignment.
- Allow prompt-only root and child task creation, plus draft restoration
in the New Task dialog. Keep user titles supported.
- Add an opt-in Product E2E suite for prompt-only Standard and Ask
tasks, plus an explicit-title control. It checks actual provider calls
within the first five tools, persisted state, audit attribution, and the
reloaded UI.
- Preserve a closed vocabulary of API key maintenance phrases in
declared prose while rejecting opaque credential suffixes. Add one
bounded naming retry after wording is rejected, without treating the
rejected call as a saved title.
- Repair the native cleanup receipt check exposed during full
verification: accept matching input digests, retain legacy input checks,
and reject conflicting receipts.

## Verification

- Live Product E2E on `f43478473800e3a46b85c5ee79677efdb15108e7`: **3/3
passed** with native Codex `gpt-5.4-mini`, first attempts only,
automatic retries disabled. Standard and Ask each saved “Rotate expired
API key” on their first tool call, with matching persisted state and a
single same-run audit entry. The explicit-title control retained its
user title with zero title writes. All three verified the reloaded
browser UI.
- Campaign: `local-2026-09-30T21-30-11-021Z`. Earlier failed campaigns
are retained separately; they exposed credential-prose handling and
prompted the naming recovery fix. No failed result was regraded or
deleted.
- Reproduce with `pnpm test:e2e:runner -- --id
task-titles.runner-codex-mini.local.prompt-title-standard --id
task-titles.runner-codex-mini.local.prompt-title-ask --id
task-titles.runner-codex-mini.local.preserve-explicit-title
--max-automatic-retries 0` and an authorized provider key.
- Full `pnpm -r typecheck` and `pnpm build` passed on the latest commit.
The runner build used the configured external eval source tree.
- Product E2E unit suite: **61 files, 818 tests passed**; E2E typecheck
and UI token gates passed.
- Title API/native regressions cover prompt-only and explicit child
creation, user edits, ownership/company isolation, external reference
refresh, cross-surface retry replay, and the 64-key limit without
receipt eviction. All passed. Prompt-context coverage: **44 tests
passed**.
- Rust credential regressions: **35 tests passed**, including benign
maintenance qualifiers and opaque credential rejection in every declared
prose field. Catalog/report reconciliation: **28 tests passed**. Native
recovery: **560 tests passed**.
- Broad local `pnpm test:run`: **14,555 tests passed** in the general
server group; two suites failed to initialize embedded PostgreSQL and
the existing 40,000-file Git streaming stress test exceeded its
300-second macOS timeout. All three suites then passed in isolation (**5
tests passed**) without code or timeout changes. The original full local
command exited nonzero and is not being represented as a clean full run.
- Latest-head GitHub checks are green: **53 passed, 4 skipped, zero
failed or pending**, including all test shards and the canary packaging
dry run. Greptile reviewed the same commit at **5/5**, with zero
unresolved review threads.

## Risks

- The additive database field must reach the server and UI together. The
migration uses `IF NOT EXISTS` and defaults existing tasks to a final
title.
- Title generation depends on the assigned agent running. Tasks without
a run keep their provisional title.
- Live qualification covers the native Codex path in Standard and Ask
modes. API/legacy and Plan behavior have deterministic coverage.
- The credential-prose exception validates the entire suffix against a
closed maintenance vocabulary. Unknown suffixes, assignments, quoted
values, credential prefixes, and diagnostics retain strict checks.

## Model Used

OpenAI Codex, based on GPT-6, with reasoning, tool use, and code
execution. The exact deployment ID and context window are not exposed in
this session. The live eval uses the native Codex `gpt-5.4-mini`
profile.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 16:48:24 -05:00
..

Workspace bootstrap recovery acceptance

This opt-in suite uses a disposable loopback instance, real Git repositories, the real Git scan scheduler, durable retries, and the task UI. A deterministic process adapter verifies the prepared files and completes its task through the run-scoped API. It makes no model calls. The test Git executable injects a real timeout only when the test explicitly arms a marker inside its repository.

From a development worktree, start the foreground instance:

BOOTSTRAP_FIXTURE_KEY=not-a-provider-key \
PATH="$PWD/tests/e2e/workspace-bootstrap/bin:$PATH" \
NODE_ENV=test \
node cli/node_modules/tsx/dist/cli.mjs cli/src/index.ts test-drive \
  --harness codex --api-key-env BOOTSTRAP_FIXTURE_KEY \
  --company-name 'Workspace Recovery QA' --no-browser

Set BOOTSTRAP_REAL_GIT to the absolute Git executable if it is not /usr/bin/git. Do not point this fixture at a normal development or production instance. The seeder checks the loopback host, trusted-local mode, and company name. It does not create tasks; each test creates its task through the UI.

Use the ready URL printed by test-drive:

WORKSPACE_BOOTSTRAP_TEST_URL=http://127.0.0.1:3100 \
pnpm exec playwright test --config tests/e2e/workspace-bootstrap/playwright.config.ts

The tests prove that one timeout schedules a retry and completes without manual Retry, that persistent timeouts stop after three total runs, and that the final state survives reload. Source edits remain intact and ignored private files are not copied. Screenshots and failure traces go to test-results/workspace-bootstrap. Stop test-drive with Ctrl-C. Its disposable data directory remains for inspection.