mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-11 14:10:50 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The native runner must support project work, delegation, hiring, and service access. > - Browser tests exposed lost connection access, rejected helper events, and stalled recovery. > - Some eval failures also came from incorrect fixtures and decision controls. > - This pull request fixes those paths and adds eight everyday workflow stories. > - The tests retain observed failures and verify delivered files independently. > - The benefit is repeatable evidence for common user tasks and their remaining gaps. ## Linked Issues or Issue Description Related work: #13404 contains earlier workflow fixes. #13300 and #13470 changed the CI contracts used by the harness security tests. Merged companion: [paperclip-evals#22](https://github.com/paperclipai/paperclip-evals/pull/22). **What happened?** Native ACPX sessions did not receive the assigned connection gateway. Codex helper events could arrive before their spawn receipt and fail thread validation. A parent continuation could take a shared workspace before its child retried. A failed native continuation could leave the task status without a clear recovery blocker. The eval harness also confused tool approvals with new connection requests and could reject a valid delegated download. **Expected behavior** Keep assigned gateway access and its approval checks. Verify helper lineage before accepting helper progress. Let a waiting child proceed before automatic parent recovery. Preserve a failed task's recovery ownership. Grade the actual requested workflow and its delivered files. **Steps to reproduce** Run the everyday workflow suite with the native Codex and Claude profiles. Exercise service approval, connection refusal, delegated project work, and teammate reuse. The commands and case requirements are in `tests/runner-e2e/EVERYDAY-WORKFLOWS.md`. Use `pnpm test:runner-recovery` for controlled crash and replacement cases. ## What Changed - Pass the scoped connection gateway binding through the native ACPX host and sidecar. - Recognize Codex helper lineage from parent metadata and spawn receipts. Verify early helper events with `thread/read`. Keep helper events separate from root completion authority. - Guide agents to use persistent hiring, child tasks, dependency records, and a blocked handoff while waiting for a child. - Defer automatic parent recovery while a child has an active execution path in the same shared workspace. Allow parent recovery when the child needs review. - Record Blocked status and recovery evidence when a failed native continuation needs reconciliation, including existing active or escalated incidents. Preserve their owner and retry budget. - Add eight browser-driven workflow cases. Use real decision controls, explicit child feedback delivery, managed hiring credentials, and independent ZIP checks inside a bounded Docker sandbox. Verify sandbox availability before task creation. Record screenshot SHA-256 at capture. - Keep runner crash probes in controlled recovery tests. Preserve the original failure when cleanup also fails. - Display missing accounting and replay revisions as unavailable. Align harness security assertions with the approved CI changes. - Make the channel-rejection browser fixture bind its file after the send captures its payload. This prevents live refresh from removing the file before the simulated race. ## Verification - Full workspace `pnpm -r typecheck` passed after merging current master. - Runner E2E typecheck passed. Harness unit tests passed: 216/216. - Wake-queue database tests passed: 55/55. The two added existing-incident tests failed before the fix and pass after it. - Docker artifact calibration passed: 12/12. Host-file and host-loopback isolation tests failed before the fix and pass after it. Read-only delivery and output limits are also verified. - Full `pnpm build` passed. Targeted recovery tests passed: 83/83. - The channel-rejection browser test passed five consecutive runs after fixing the fixture race found in CI. - Local general-server (12,351 tests), UI (6,250), CLI (485), and workspace package groups passed. The monolithic run stopped at an unchanged lock-heartbeat fixture race; the isolated workspace group passed on rerun (shared: 747/747). A separate local serialized run passed 97 files before two socket errors in the unchanged issue-list route suite; that suite passed 15/15 on isolated rerun. These local full commands did not finish uninterrupted; the complete CI matrix below covers the remaining suites. - Final head `0fb293733fe307be7e6667ae8f1364077d0c6455`: **34 successful checks, 2 expected skips**, including every server/workspace shard, browser shard, native runner verification, build, and typecheck. [Final CI run](https://github.com/paperclipai/paperclip/actions/runs/34989136700). - Greptile reviewed this exact head at **5/5**; all review threads are resolved. Both Superagent security checks are successful. - ACPX credential-boundary tests passed: 118/118. Superagent accepted the runner/sidecar versus provider-environment trace and cleared its finding. - The latest paid local campaign on source `f6a2fdf7ac2af859826a2ae627ff4125a5478529` passed 22/24 cases: Sol 8/8, Claude 7/8, Mini 7/8. These results predate the merge with current master. - The two remaining failures are in `hire-reuse`: Claude exceeded the attempt deadline during final review; Mini made invalid deliverable tool calls and remained Blocked. - Six Daytona cases were not run because the matching immutable runner image was unavailable. This PR does not claim new remote model results. ## Risks The changes affect connection admission, helper identity, and recovery scheduling. Assigned gateway grants and user approval still govern service calls. The workspace admission gate still exists; the broader folder-sync design is separate work. Provider behavior can still cause the two recorded hiring failures. No database migration is required. Paid cases are opt-in and have bounded attempt deadlines. Project stories now require Docker and the documented pinned Python image on the harness host. ## Model Used OpenAI `gpt-6-astra` performed implementation, diagnosis, and substantive review. OpenAI `gpt-5.6-luna` assisted with verification, PR preparation, and review tracking. Both used repository tools and code execution. Context-window sizes were not recorded. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused checks and isolated reruns; full-run limitations are documented above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: OpenAI GPT-5.6 Luna <noreply@openai.com> Co-authored-by: Paperclip <noreply@paperclip.ing>
265 lines
8.6 KiB
TypeScript
265 lines
8.6 KiB
TypeScript
import { describe, it, expect } from "vitest";
|
|
import { runnerMatrix, runnerSuites } from "./catalog.js";
|
|
import { parseRunnerSelectors, selectRunnerExecutions } from "./selectors.js";
|
|
import { everydayTasks, productionStoryProfile } from "./everyday-cases.js";
|
|
import {
|
|
isStoryWorkspaceDeferral,
|
|
storyLifecycleChecks,
|
|
storyRepliesConsumed,
|
|
storyParentFinishedAfterChildren,
|
|
type StoryRun,
|
|
} from "./everyday-observations.js";
|
|
|
|
describe("manual everyday workflow catalog", () => {
|
|
it("is discoverable and explicitly selected without changing scheduled --all", () => {
|
|
const all = selectRunnerExecutions(parseRunnerSelectors(["--all"]));
|
|
expect(all.some((e) => e.suite.id === "everyday-workflows")).toBe(false);
|
|
const selected = selectRunnerExecutions(
|
|
parseRunnerSelectors(["--suite", "everyday-workflows"]),
|
|
);
|
|
expect(selected).toHaveLength(30);
|
|
expect(selected.every((e) => e.profile.generation === "native")).toBe(true);
|
|
expect(
|
|
selectRunnerExecutions(
|
|
parseRunnerSelectors(["--profile", "runner-codex"]),
|
|
).some((e) => e.suite.id === "everyday-workflows"),
|
|
).toBe(false);
|
|
const listed = selectRunnerExecutions(parseRunnerSelectors(["--list"]));
|
|
expect(listed.some((e) => e.suite.id === "everyday-workflows")).toBe(true);
|
|
});
|
|
it("does not schedule arbitrary runner-crash probes as model evals", () => {
|
|
const selected = selectRunnerExecutions(
|
|
parseRunnerSelectors(["--suite", "everyday-workflows"]),
|
|
);
|
|
expect(selected.some((e) => e.task.id.startsWith("recover-runner"))).toBe(false);
|
|
for (const id of ["recover-runner", "recover-runner-safe", "recover-runner-uncertain"])
|
|
expect(() => selectRunnerExecutions(parseRunnerSelectors([
|
|
"--id", `everyday-workflows.runner-codex.local.${id}`,
|
|
]))).toThrow();
|
|
expect(selected.some((e) => e.task.id === "recover-controller")).toBe(true);
|
|
expect(selected.some((e) => e.task.id === "stop-redirect")).toBe(true);
|
|
});
|
|
it("keeps ordinary prompts free of completion/API instructions", () => {
|
|
for (const task of everydayTasks)
|
|
expect(task.buildPrompt("sample")).not.toMatch(
|
|
/finish_task|paperclip_finish|PATCH|mark .*done|idempotencyKey/i,
|
|
);
|
|
const profile = runnerSuites.find((s) => s.id === "everyday-workflows")!
|
|
.profiles[0]!;
|
|
const value = productionStoryProfile(profile).buildAgent({
|
|
environmentId: "env",
|
|
environmentFixtureId: "local",
|
|
workspacePath: "/tmp/test",
|
|
secretRefs: {
|
|
OPENAI_API_KEY: {
|
|
type: "secret_ref",
|
|
secretId: "secret",
|
|
version: "latest",
|
|
},
|
|
},
|
|
executionId: "case",
|
|
});
|
|
expect(JSON.stringify(value.instructionsBundle)).not.toMatch(
|
|
/fixture|mark .*done|finish_task|api\/issues/i,
|
|
);
|
|
});
|
|
it("does not claim unsupported remote crash/hiring coverage", () => {
|
|
const remote = runnerMatrix.filter(
|
|
(e) =>
|
|
e.suite.id === "everyday-workflows" && e.environment.id === "daytona",
|
|
);
|
|
expect(remote).toHaveLength(6);
|
|
expect(new Set(remote.map((e) => e.task.id))).toEqual(
|
|
new Set(["build-revise", "delegate-feedback", "recover-controller"]),
|
|
);
|
|
});
|
|
});
|
|
|
|
describe("lifecycle oracle calibrated failures", () => {
|
|
const parent = {
|
|
id: "parent",
|
|
companyId: "company",
|
|
title: "Project",
|
|
status: "done",
|
|
assigneeAgentId: "lead",
|
|
};
|
|
const run: StoryRun = {
|
|
id: "run",
|
|
companyId: "company",
|
|
agentId: "lead",
|
|
status: "succeeded",
|
|
runtimeMode: "native",
|
|
runnerInstanceId: "runner",
|
|
contextSnapshot: { issueId: "parent" },
|
|
};
|
|
const score = (
|
|
runs: StoryRun[],
|
|
issues = [parent],
|
|
allowedInterruptedRuns: string[] = [],
|
|
) =>
|
|
storyLifecycleChecks({
|
|
issues,
|
|
runs,
|
|
parentId: parent.id,
|
|
leadId: "lead",
|
|
allowedInterruptedRuns,
|
|
});
|
|
it("accepts successful owned native work", () =>
|
|
expect(score([run]).every((c) => c.passed)).toBe(true));
|
|
it.each([
|
|
["legacy execution", { ...run, runtimeMode: "legacy" }, "native-runtime"],
|
|
[
|
|
"no runner identity",
|
|
{ ...run, runnerInstanceId: null },
|
|
"native-runtime",
|
|
],
|
|
["worker on parent", { ...run, agentId: "worker" }, "parent-owned-by-lead"],
|
|
["crash without recovery", { ...run, status: "failed" }, "successful-runs"],
|
|
["unsettled run", { ...run, status: "running" }, "settled"],
|
|
] as const)("rejects %s", (_label, bad, id) =>
|
|
expect(score([bad]).find((c) => c.id === id)?.passed).toBe(false),
|
|
);
|
|
it("does not exempt unrelated failures because another run was intentionally stopped", () => {
|
|
const checks = score(
|
|
[
|
|
{ ...run, status: "cancelled" },
|
|
{ ...run, id: "other", status: "failed" },
|
|
],
|
|
[parent],
|
|
["run"],
|
|
);
|
|
expect(checks.find((c) => c.id === "successful-runs")?.passed).toBe(false);
|
|
});
|
|
it("excludes a proven pre-dispatch workspace deferral without hiding executed failures", () => {
|
|
const deferred: StoryRun = {
|
|
id: "deferred",
|
|
companyId: "company",
|
|
agentId: "lead",
|
|
status: "cancelled",
|
|
errorCode: "workspace_busy",
|
|
resultJson: {
|
|
executionRecovery: {
|
|
kind: "workspace_wait",
|
|
providerWorkStarted: false,
|
|
},
|
|
},
|
|
};
|
|
expect(score([run, deferred]).every((c) => c.passed)).toBe(true);
|
|
expect(
|
|
score([run, { ...deferred, processPid: 123 }]).every((c) => c.passed),
|
|
).toBe(false);
|
|
expect(
|
|
score([run, { ...deferred, resultJson: {} }]).every((c) => c.passed),
|
|
).toBe(false);
|
|
});
|
|
it("rejects a finished answer left in review", () =>
|
|
expect(
|
|
score([run], [{ ...parent, status: "in_review" }]).find(
|
|
(c) => c.id === "tasks-done",
|
|
)?.passed,
|
|
).toBe(false));
|
|
});
|
|
|
|
describe("reply completion boundary", () => {
|
|
const run: StoryRun = {
|
|
id: "old",
|
|
companyId: "company",
|
|
agentId: "lead",
|
|
status: "succeeded",
|
|
runnerProfileJson: {
|
|
nativeExecutionInput: { task: { prompt: "Original task" } },
|
|
},
|
|
};
|
|
it("does not accept Done from the first run while a later request is still queued", () => {
|
|
expect(storyRepliesConsumed([run], ["later-comment"])).toBe(false);
|
|
const next = {
|
|
...run,
|
|
id: "new",
|
|
runnerProfileJson: {
|
|
nativeExecutionInput: { task: { prompt: "User reply later-comment" } },
|
|
},
|
|
};
|
|
expect(
|
|
storyRepliesConsumed(
|
|
[run, { ...next, status: "running" }],
|
|
["later-comment"],
|
|
),
|
|
).toBe(false);
|
|
expect(storyRepliesConsumed([run, next], ["later-comment"])).toBe(true);
|
|
});
|
|
});
|
|
|
|
describe("delegation completion order", () => {
|
|
it("rejects a parent closed before its worker finishes", () => {
|
|
const base: StoryRun = {
|
|
id: "lead-run",
|
|
companyId: "company",
|
|
agentId: "lead",
|
|
status: "succeeded",
|
|
nativeIssueId: "parent",
|
|
finishedAt: "2026-09-14T12:00:00Z",
|
|
};
|
|
const child = {
|
|
...base,
|
|
id: "worker-run",
|
|
agentId: "worker",
|
|
nativeIssueId: "child",
|
|
finishedAt: "2026-09-14T12:01:00Z",
|
|
};
|
|
expect(
|
|
storyParentFinishedAfterChildren([base, child], "parent", "lead", [
|
|
"child",
|
|
]),
|
|
).toBe(false);
|
|
expect(
|
|
storyParentFinishedAfterChildren(
|
|
[{ ...base, finishedAt: "2026-09-14T12:02:00Z" }, child],
|
|
"parent",
|
|
"lead",
|
|
["child"],
|
|
),
|
|
).toBe(true);
|
|
expect(
|
|
storyParentFinishedAfterChildren([base], "parent", "lead", ["child"]),
|
|
).toBe(false);
|
|
});
|
|
});
|
|
|
|
describe("cancelled queued workspace retry", () => {
|
|
const queued: StoryRun = {
|
|
id: "retry",
|
|
companyId: "company",
|
|
agentId: "agent",
|
|
status: "cancelled",
|
|
runtimeMode: "legacy",
|
|
runtimeModeResolvedAt: null,
|
|
startedAt: null,
|
|
retryOfRunId: "workspace-deferral",
|
|
scheduledRetryReason: "workspace_busy",
|
|
errorCode: "cancelled",
|
|
lastOutputSeq: 0,
|
|
resultJson: {
|
|
startupCancellation: {
|
|
requestedAt: "2026-09-14T19:10:23.057Z",
|
|
beforeNativeSelection: false,
|
|
},
|
|
},
|
|
};
|
|
it("does not mistake an unstarted cancelled retry for provider execution", () => {
|
|
expect(isStoryWorkspaceDeferral(queued)).toBe(true);
|
|
});
|
|
it.each([
|
|
{ startedAt: "2026-09-14T19:10:22Z" },
|
|
{ processPid: 42 },
|
|
{ runnerInstanceId: "runner" },
|
|
{ nativeSessionId: "session" },
|
|
{ lastOutputSeq: 1 },
|
|
{ usageJson: { outputTokens: 1 } },
|
|
{ runtimeModeResolvedAt: "2026-09-14T19:10:22Z" },
|
|
{ scheduledRetryReason: "other" },
|
|
{ resultJson: null },
|
|
])("retains a contradictory or unproven cancellation %j", (change) => {
|
|
expect(isStoryWorkspaceDeferral({ ...queued, ...change })).toBe(false);
|
|
});
|
|
});
|