mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-11 05:31:46 +02:00
test(runner): add everyday workflow evaluation harness (#13474)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The native runner must support project work, delegation, hiring, and service access. > - Browser tests exposed lost connection access, rejected helper events, and stalled recovery. > - Some eval failures also came from incorrect fixtures and decision controls. > - This pull request fixes those paths and adds eight everyday workflow stories. > - The tests retain observed failures and verify delivered files independently. > - The benefit is repeatable evidence for common user tasks and their remaining gaps. ## Linked Issues or Issue Description Related work: #13404 contains earlier workflow fixes. #13300 and #13470 changed the CI contracts used by the harness security tests. Merged companion: [paperclip-evals#22](https://github.com/paperclipai/paperclip-evals/pull/22). **What happened?** Native ACPX sessions did not receive the assigned connection gateway. Codex helper events could arrive before their spawn receipt and fail thread validation. A parent continuation could take a shared workspace before its child retried. A failed native continuation could leave the task status without a clear recovery blocker. The eval harness also confused tool approvals with new connection requests and could reject a valid delegated download. **Expected behavior** Keep assigned gateway access and its approval checks. Verify helper lineage before accepting helper progress. Let a waiting child proceed before automatic parent recovery. Preserve a failed task's recovery ownership. Grade the actual requested workflow and its delivered files. **Steps to reproduce** Run the everyday workflow suite with the native Codex and Claude profiles. Exercise service approval, connection refusal, delegated project work, and teammate reuse. The commands and case requirements are in `tests/runner-e2e/EVERYDAY-WORKFLOWS.md`. Use `pnpm test:runner-recovery` for controlled crash and replacement cases. ## What Changed - Pass the scoped connection gateway binding through the native ACPX host and sidecar. - Recognize Codex helper lineage from parent metadata and spawn receipts. Verify early helper events with `thread/read`. Keep helper events separate from root completion authority. - Guide agents to use persistent hiring, child tasks, dependency records, and a blocked handoff while waiting for a child. - Defer automatic parent recovery while a child has an active execution path in the same shared workspace. Allow parent recovery when the child needs review. - Record Blocked status and recovery evidence when a failed native continuation needs reconciliation, including existing active or escalated incidents. Preserve their owner and retry budget. - Add eight browser-driven workflow cases. Use real decision controls, explicit child feedback delivery, managed hiring credentials, and independent ZIP checks inside a bounded Docker sandbox. Verify sandbox availability before task creation. Record screenshot SHA-256 at capture. - Keep runner crash probes in controlled recovery tests. Preserve the original failure when cleanup also fails. - Display missing accounting and replay revisions as unavailable. Align harness security assertions with the approved CI changes. - Make the channel-rejection browser fixture bind its file after the send captures its payload. This prevents live refresh from removing the file before the simulated race. ## Verification - Full workspace `pnpm -r typecheck` passed after merging current master. - Runner E2E typecheck passed. Harness unit tests passed: 216/216. - Wake-queue database tests passed: 55/55. The two added existing-incident tests failed before the fix and pass after it. - Docker artifact calibration passed: 12/12. Host-file and host-loopback isolation tests failed before the fix and pass after it. Read-only delivery and output limits are also verified. - Full `pnpm build` passed. Targeted recovery tests passed: 83/83. - The channel-rejection browser test passed five consecutive runs after fixing the fixture race found in CI. - Local general-server (12,351 tests), UI (6,250), CLI (485), and workspace package groups passed. The monolithic run stopped at an unchanged lock-heartbeat fixture race; the isolated workspace group passed on rerun (shared: 747/747). A separate local serialized run passed 97 files before two socket errors in the unchanged issue-list route suite; that suite passed 15/15 on isolated rerun. These local full commands did not finish uninterrupted; the complete CI matrix below covers the remaining suites. - Final head `0fb293733fe307be7e6667ae8f1364077d0c6455`: **34 successful checks, 2 expected skips**, including every server/workspace shard, browser shard, native runner verification, build, and typecheck. [Final CI run](https://github.com/paperclipai/paperclip/actions/runs/34989136700). - Greptile reviewed this exact head at **5/5**; all review threads are resolved. Both Superagent security checks are successful. - ACPX credential-boundary tests passed: 118/118. Superagent accepted the runner/sidecar versus provider-environment trace and cleared its finding. - The latest paid local campaign on source `f6a2fdf7ac2af859826a2ae627ff4125a5478529` passed 22/24 cases: Sol 8/8, Claude 7/8, Mini 7/8. These results predate the merge with current master. - The two remaining failures are in `hire-reuse`: Claude exceeded the attempt deadline during final review; Mini made invalid deliverable tool calls and remained Blocked. - Six Daytona cases were not run because the matching immutable runner image was unavailable. This PR does not claim new remote model results. ## Risks The changes affect connection admission, helper identity, and recovery scheduling. Assigned gateway grants and user approval still govern service calls. The workspace admission gate still exists; the broader folder-sync design is separate work. Provider behavior can still cause the two recorded hiring failures. No database migration is required. Paid cases are opt-in and have bounded attempt deadlines. Project stories now require Docker and the documented pinned Python image on the harness host. ## Model Used OpenAI `gpt-6-astra` performed implementation, diagnosis, and substantive review. OpenAI `gpt-5.6-luna` assisted with verification, PR preparation, and review tracking. Both used repository tools and code execution. Context-window sizes were not recorded. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused checks and isolated reruns; full-run limitations are documented above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: OpenAI GPT-5.6 Luna <noreply@openai.com> Co-authored-by: Paperclip <noreply@paperclip.ing>
This commit is contained in:
58 files changed
+3600
-217
No files matched your search
@@ -4,6 +4,7 @@ import { eq } from "drizzle-orm";
|
||||
import { afterAll, afterEach, beforeAll, describe, expect, it } from "vitest";
|
||||
import type { Db } from "@paperclipai/db";
|
||||
import {
|
||||
activityLog,
|
||||
agentWakeupRequests,
|
||||
agents,
|
||||
companies,
|
||||
@@ -56,6 +57,7 @@ describeEmbeddedPostgres("wake-queue postgres adapter", () => {
|
||||
}, 20_000);
|
||||
|
||||
afterEach(async () => {
|
||||
await db.delete(activityLog);
|
||||
await db.delete(issueComments);
|
||||
// `heartbeat_runs.wakeup_request_id` references `agent_wakeup_requests.id`,
|
||||
// so the run row must go first.
|
||||
@@ -316,6 +318,62 @@ describeEmbeddedPostgres("wake-queue postgres adapter", () => {
|
||||
});
|
||||
}
|
||||
|
||||
it.each(["in_progress", "blocked"])("preserves recovery ownership and queued messages when a native task fails from %s", async (status) => {
|
||||
const companyId = await seedCompany();
|
||||
const agentId = await seedAgent({ companyId });
|
||||
const issueId = await seedIssue({ companyId, assigneeAgentId: agentId, status });
|
||||
const runId = await seedRun({ companyId, agentId, status: "failed", contextSnapshot: { issueId } });
|
||||
await db.update(heartbeatRuns).set({ runtimeMode: "native", errorCode: "thread_binding_mismatch" }).where(eq(heartbeatRuns.id, runId));
|
||||
const wakeId = await seedDeferredWake({ companyId, agentId, issueId });
|
||||
const adapter = createPostgresWakeQueueAdapter(db, stubDeps);
|
||||
await adapter.withIssueExecutionLock({ companyId, runId, now: new Date() }, async () => { throw new Error("must not replay an uncertain execution"); });
|
||||
const blockedIssue = (await db.select().from(issues).where(eq(issues.id, issueId)))[0];
|
||||
expect(blockedIssue.status).toBe("blocked");
|
||||
const entries = await db.select().from(activityLog).where(eq(activityLog.entityId, issueId));
|
||||
if (status === "in_progress") {
|
||||
expect(blockedIssue.blockedTransitionAt).not.toBeNull();
|
||||
expect(entries[0]).toMatchObject({ action: "issue.updated", details: { status: "blocked", previousStatus: "in_progress" } });
|
||||
} else expect(entries).toHaveLength(0);
|
||||
expect((await db.select().from(agentWakeupRequests).where(eq(agentWakeupRequests.id, wakeId)))[0].status).toBe("deferred_issue_execution");
|
||||
const action = (await db.select().from(issueRecoveryActions).where(eq(issueRecoveryActions.sourceIssueId, issueId)))[0];
|
||||
expect(action).toMatchObject({ ownerType: "board", cause: "native_continuation_requires_reconciliation" });
|
||||
if (status === "in_progress") expect(action.evidence).toMatchObject({ nativeFailureBlock: { runId, statusVersion: blockedIssue.statusVersion } });
|
||||
else expect(action.evidence.nativeFailureBlock).toBeUndefined();
|
||||
await adapter.withIssueExecutionLock({ companyId, runId, now: new Date() }, async () => { throw new Error("must not replay"); });
|
||||
expect(await db.select().from(issueRecoveryActions).where(eq(issueRecoveryActions.sourceIssueId, issueId))).toHaveLength(1);
|
||||
expect((await db.select().from(issues).where(eq(issues.id, issueId)))[0].statusVersion).toBe(blockedIssue.statusVersion);
|
||||
});
|
||||
|
||||
it.each(["active", "escalated"])("repairs a failed native task with an existing %s recovery action", async (status) => {
|
||||
const companyId = await seedCompany();
|
||||
const agentId = await seedAgent({ companyId });
|
||||
const issueId = await seedIssue({ companyId, assigneeAgentId: agentId, status: "in_progress" });
|
||||
const runId = await seedRun({ companyId, agentId, status: "failed", contextSnapshot: { issueId } });
|
||||
await db.update(heartbeatRuns).set({ runtimeMode: "native", errorCode: "runner_lost" }).where(eq(heartbeatRuns.id, runId));
|
||||
const wakeId = await seedDeferredWake({ companyId, agentId, issueId });
|
||||
const [existing] = await db.insert(issueRecoveryActions).values({
|
||||
companyId, sourceIssueId: issueId, status, kind: "active_run_watchdog",
|
||||
ownerType: "board", cause: "native_runner_restart_unverified", fingerprint: `restart:${runId}`,
|
||||
evidence: { runId, priorProof: "keep", automaticRecovery: { attempts: 2 } },
|
||||
nextAction: "Verify the previous execution stopped", attemptCount: 2, maxAttempts: 3,
|
||||
}).returning();
|
||||
const adapter = createPostgresWakeQueueAdapter(db, stubDeps);
|
||||
const release = () => adapter.withIssueExecutionLock({ companyId, runId, now: new Date() }, async () => { throw new Error("must not replay"); });
|
||||
await release();
|
||||
const [blocked] = await db.select().from(issues).where(eq(issues.id, issueId));
|
||||
expect(blocked.status).toBe("blocked");
|
||||
const actions = await db.select().from(issueRecoveryActions).where(eq(issueRecoveryActions.sourceIssueId, issueId));
|
||||
expect(actions).toHaveLength(1);
|
||||
expect(actions[0]).toMatchObject({ id: existing.id, status, cause: existing.cause,
|
||||
ownerType: "board", attemptCount: 2, maxAttempts: 3, nextAction: existing.nextAction,
|
||||
evidence: { runId, priorProof: "keep", automaticRecovery: { attempts: 2 },
|
||||
nativeFailureBlock: { runId, statusVersion: blocked.statusVersion } } });
|
||||
expect((await db.select().from(agentWakeupRequests).where(eq(agentWakeupRequests.id, wakeId)))[0].status).toBe("deferred_issue_execution");
|
||||
await release();
|
||||
expect((await db.select().from(issues).where(eq(issues.id, issueId)))[0].statusVersion).toBe(blocked.statusVersion);
|
||||
expect(await db.select().from(activityLog).where(eq(activityLog.entityId, issueId))).toHaveLength(1);
|
||||
});
|
||||
|
||||
it.each(["queued", "running", "scheduled_retry"])("does not promote another turn behind a %s successor without an execution lock", async (status) => {
|
||||
const companyId = await seedCompany();
|
||||
const agentId = await seedAgent({ companyId });
|
||||
|
||||
@@ -6,6 +6,7 @@ import { and, asc, eq, inArray, isNull, notInArray, or, sql } from "drizzle-orm"
|
||||
import type { Db } from "@paperclipai/db";
|
||||
import { extractIssueReferenceIdentifiers } from "@paperclipai/shared";
|
||||
import {
|
||||
activityLog,
|
||||
agentWakeupRequests,
|
||||
agents,
|
||||
chatActions,
|
||||
@@ -720,7 +721,7 @@ async function recordNativeTerminalRecoveryIfNeeded(tx: Db, run: HeartbeatRunRow
|
||||
if (!applies || isAcknowledgedNativeStop(run)) return false;
|
||||
|
||||
const existing = await tx
|
||||
.select({ id: issueRecoveryActions.id })
|
||||
.select({ id: issueRecoveryActions.id, evidence: issueRecoveryActions.evidence })
|
||||
.from(issueRecoveryActions)
|
||||
.where(
|
||||
and(
|
||||
@@ -733,6 +734,27 @@ async function recordNativeTerminalRecoveryIfNeeded(tx: Db, run: HeartbeatRunRow
|
||||
),
|
||||
)
|
||||
.limit(1);
|
||||
let nativeFailureBlock: { runId: string; statusVersion: number } | undefined;
|
||||
if (issue.status !== "blocked") {
|
||||
const projected = await issueService(tx).update(issue.id, { status: "blocked" }, tx);
|
||||
if (projected) {
|
||||
nativeFailureBlock = { runId: run.id, statusVersion: projected.statusVersion };
|
||||
await tx.insert(activityLog).values({
|
||||
companyId: issue.companyId, actorType: "system", actorId: "execution-recovery",
|
||||
action: "issue.updated", entityType: "issue", entityId: issue.id, runId: run.id,
|
||||
details: { status: "blocked", previousStatus: issue.status, reason: "native_continuation_requires_reconciliation" },
|
||||
});
|
||||
}
|
||||
}
|
||||
// Status projection is required even when restart/finalization created the
|
||||
// incident first. Preserve its owner, cause, retry budget, and prior evidence.
|
||||
if (nativeFailureBlock) {
|
||||
for (const action of existing) {
|
||||
await tx.update(issueRecoveryActions).set({
|
||||
evidence: { ...action.evidence, nativeFailureBlock }, updatedAt: now,
|
||||
}).where(and(eq(issueRecoveryActions.id, action.id), eq(issueRecoveryActions.companyId, issue.companyId)));
|
||||
}
|
||||
}
|
||||
if (!existing.length) {
|
||||
await tx
|
||||
.update(nativeRunFinalizations)
|
||||
@@ -760,7 +782,7 @@ async function recordNativeTerminalRecoveryIfNeeded(tx: Db, run: HeartbeatRunRow
|
||||
returnOwnerAgentId: run.agentId,
|
||||
cause: "native_continuation_requires_reconciliation",
|
||||
fingerprint: `native-continuation:${run.id}`,
|
||||
evidence: { runId: run.id, originalFailureCode: run.errorCode },
|
||||
evidence: { runId: run.id, originalFailureCode: run.errorCode, ...(nativeFailureBlock ? { nativeFailureBlock } : {}) },
|
||||
nextAction:
|
||||
"Inspect the original failure and reconcile the previous execution before continuing. Automatic recovery cannot start another incident.",
|
||||
maxAttempts: 3,
|
||||
|
||||
Reference in new issue
Block a user