Files
PaperClipAI/tests/runner-e2e/everyday-cases.ts
T
DottaandPaperclip 4cfc0e3f6c fix(runner-e2e): provision complete worker prerequisites (#13674)
## Thinking Path

> - Paperclip manages work performed by AI agents.
> - Runner E2E tests verify that work in local and remote environments.
> - The full catalog exposed setup failures before agents could execute
tasks.
> - ACPX Codex missed sandbox provisioning, and new artifact cases
missed Docker preparation.
> - A host Claude version probe also stopped native Claude stories that
use a packaged provider.
> - This pull request repairs trusted worker setup and keeps its
selectors covered by catalog tests.
> - The resulting reruns can measure behavior instead of missing
prerequisites.

## Linked Issues or Issue Description

Follow-up to #13655. Full-catalog campaign:
https://github.com/paperclipai/paperclip/actions/runs/35417932353.

**What happened?**

ACPX Codex could not start its sandbox. New artifact stories missed
Docker preparation. Native Claude stories tried to spawn an unrelated
host CLI. The report job failed on trusted lockfile drift.

**What did you expect to happen?**

Prepare each worker's required capabilities before paid execution and
publish the retained results from trusted code.

**Steps to reproduce**

Run local ACPX Codex, everyday agent-review-handoff, or native Claude
everyday cells in the full-stack workflow at 43acbcc39.

## What Changed

- Include local ACPX Codex in exact-binary sandbox provisioning and
preflight.
- Prepare the pinned artifact oracle for agent-review-handoff. Do not
require it for skill creation.
- Remove host CLI version probes from native everyday stories.
- Retry GitHub actor lookup up to three times, with a deadline, while
retaining all authorization checks.
- Resolve reporter dependencies from the trusted checkout before its
frozen install.
- Test worker selections against the catalog and document the
default-branch requirement.

## Verification

- `pnpm test:e2e:runner:unit`: 383 tests pass.
- `pnpm test:e2e:runner:typecheck`: passes.
- PR checks: 54 passed, two intentionally skipped. Greptile: 5/5, no
inline findings.
- The paid workflow reads trusted setup from master. These setup changes
need to land before the affected Linux cells can verify them. Runtime
fixes and other affected reruns are on a separate branch.

## Risks

The setup selectors decide which workers receive sandbox policy and
Docker preparation. Catalog coverage checks their scope. Actor lookup
still fails closed. Reporter dependency resolution uses only the trusted
checkout; target branch code does not receive publication credentials.
No production prompt changes.

## Model Used

OpenAI GPT-6 via Codex, with repository inspection, code editing, and
test execution. The exact API model ID and context-window size are not
exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-19 09:27:15 -05:00

118 lines
6.1 KiB
TypeScript

import type { RunnerProfileFixture, RunnerTaskFixture } from "./types.js";
/** No fixture completion/API directions: the normal execution prompt owns that contract. */
export function productionStoryProfile(
profile: RunnerProfileFixture,
): RunnerProfileFixture {
return {
...profile,
buildAgent(input) {
const payload = profile.buildAgent(input);
return {
...payload,
name: `Studio Lead ${input.executionId}`,
role: "ceo",
title: "Studio Lead",
capabilities:
"Builds small projects, delegates implementation, reviews delivered work, and hires teammates when requested.",
instructionsBundle: {
entryFile: "AGENTS.md",
files: {
"AGENTS.md":
"You lead a small software studio. Help the user build useful small projects. Respect their requirements and verify the delivered work.",
},
},
};
},
};
}
export const SLUGIFY_REQUIREMENTS = `Build a small dependency-free Python command-line tool in slugify.py. It accepts one positional string and prints its slug: trim whitespace, lowercase, replace each sequence outside ASCII a-z and 0-9 with one hyphen, then trim hyphens. Expose slugify(text) for reuse. Include a README and automated tests. Deliver the source and tests as a downloadable ZIP. Work in the project workspace.`;
export const SLUGIFY_REVISION = `Add an optional --separator argument that accepts either - or _. Keep - as the default and preserve the earlier behavior. Deliver an updated ZIP, keeping the previous download available.`;
export const LATE_REQUIREMENT = `Also support --max-length as a positive integer. Truncate the final slug to that length, then remove any trailing separator. In particular, input " Queue--Ready!! " with --max-length 7 must print "queue-r". Preserve the default behavior.`;
const definitions = [
[
"build-revise",
"Build, download, and revise a project",
SLUGIFY_REQUIREMENTS,
2,
],
[
"delegate-feedback",
"Delegate implementation and preserve late feedback",
`Have Riley Builder implement the following as one child task. You review the downloaded result when it is ready. Keep implementation with Riley and post a progress note linking the child while work is underway. ${SLUGIFY_REQUIREMENTS}`,
4,
],
[
"agent-review-handoff",
"Delegate work through an agent review handoff",
`Have Riley Builder implement the following as one child task. The child must keep its original Riley assignee throughout. Before Riley finishes, require a native needs_review report with exactly one attention request: kind review, ownerClass agent, targetAgentId set to your exact lead agent id, and a summary naming you as the reviewer. The child must remain in_review while waiting. When the durable review wake arrives, inspect the child task context, approve the review through the native resolve_review tool with decision accept, then finish the parent task. Do not patch the child status, reassign the child, self-approve the child from the parent run, or bypass the review interaction. ${SLUGIFY_REQUIREMENTS}`,
3,
],
[
"hire-reuse",
"Hire one teammate, then reuse that agent",
`Hire exactly one agent named Morgan QA, reporting to you, using the same available AI connection and native runner configuration as you. Have Morgan implement the following in one child task, then review the result. ${SLUGIFY_REQUIREMENTS}`,
6,
],
[
"service-approve",
"Use a connection after approval",
`Use the connected page service to find recent pages and create a short Markdown briefing document on this task. Include the titles and verification code returned by the service.`,
2,
],
[
"service-decline",
"Respect a declined tool action",
`Use the connected page service to find recent pages and create a short Markdown briefing document on this task. Include the titles and verification code returned by the service. If I decline the action, do not try again or request another connection. Instead, give me a brief explanation that you could not retrieve the data. That explanation is the complete allowed fallback; no briefing is required after a decline.`,
2,
],
[
"connection-decline",
"Respect Not now on a new connection",
`Please connect Notion so you can read my recent pages and write a short briefing. If I choose Not now, do not try again or use another service. Instead, give me a brief explanation that you could not retrieve the pages. That explanation is the complete allowed fallback; no briefing is required after a decline.`,
2,
],
[
"recover-controller",
"Recover work after the server restarts",
SLUGIFY_REQUIREMENTS,
3,
],
[
"stop-redirect",
"Stop a task and send a new direction once",
SLUGIFY_REQUIREMENTS,
2,
],
[
"create-skill-studio",
"Create and edit a company skill",
"Create one company skill using this complete SKILL.md content:\n---\nname: release-readiness-checklist\ndescription: A bounded checklist for validating a release before handoff.\n---\n\n# Release Readiness Checklist\n\n1. Verify checks.\n2. Review evidence.\n3. Record the handoff.\n\nUse request key create-skill-e2e-001. After creating it, report the created skill and finish the task.",
1,
],
] as const;
export const everydayTasks: readonly RunnerTaskFixture[] = definitions.map(
([id, label, prompt, expectedRunCount]) => ({
id,
label,
groups: [],
workMode: "standard",
flow: "everyday_workflow",
expectedRunCount,
attemptTimeoutMs: { local: 12 * 60_000, daytona: 30 * 60_000 },
expectedTerminalState: { issue: "done", run: "succeeded" },
buildTitle: (nonce) => `${label} ${nonce}`,
buildPrompt: () => prompt,
buildVisibleMarker: (nonce) => `STUDIO_${nonce}`,
buildMatchers: () => [], // The workflow records independent artifact and lifecycle checks.
}),
);
/** Only these stories execute downloaded Python ZIPs in the pinned oracle. */
export function requiresEverydayArtifactOracle(caseId: string): boolean {
return ["build-revise", "delegate-feedback", "agent-review-handoff", "hire-reuse", "recover-controller", "stop-redirect"].includes(caseId);
}