mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-11 05:31:46 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The native runner must support project work, delegation, hiring, and service access. > - Browser tests exposed lost connection access, rejected helper events, and stalled recovery. > - Some eval failures also came from incorrect fixtures and decision controls. > - This pull request fixes those paths and adds eight everyday workflow stories. > - The tests retain observed failures and verify delivered files independently. > - The benefit is repeatable evidence for common user tasks and their remaining gaps. ## Linked Issues or Issue Description Related work: #13404 contains earlier workflow fixes. #13300 and #13470 changed the CI contracts used by the harness security tests. Merged companion: [paperclip-evals#22](https://github.com/paperclipai/paperclip-evals/pull/22). **What happened?** Native ACPX sessions did not receive the assigned connection gateway. Codex helper events could arrive before their spawn receipt and fail thread validation. A parent continuation could take a shared workspace before its child retried. A failed native continuation could leave the task status without a clear recovery blocker. The eval harness also confused tool approvals with new connection requests and could reject a valid delegated download. **Expected behavior** Keep assigned gateway access and its approval checks. Verify helper lineage before accepting helper progress. Let a waiting child proceed before automatic parent recovery. Preserve a failed task's recovery ownership. Grade the actual requested workflow and its delivered files. **Steps to reproduce** Run the everyday workflow suite with the native Codex and Claude profiles. Exercise service approval, connection refusal, delegated project work, and teammate reuse. The commands and case requirements are in `tests/runner-e2e/EVERYDAY-WORKFLOWS.md`. Use `pnpm test:runner-recovery` for controlled crash and replacement cases. ## What Changed - Pass the scoped connection gateway binding through the native ACPX host and sidecar. - Recognize Codex helper lineage from parent metadata and spawn receipts. Verify early helper events with `thread/read`. Keep helper events separate from root completion authority. - Guide agents to use persistent hiring, child tasks, dependency records, and a blocked handoff while waiting for a child. - Defer automatic parent recovery while a child has an active execution path in the same shared workspace. Allow parent recovery when the child needs review. - Record Blocked status and recovery evidence when a failed native continuation needs reconciliation, including existing active or escalated incidents. Preserve their owner and retry budget. - Add eight browser-driven workflow cases. Use real decision controls, explicit child feedback delivery, managed hiring credentials, and independent ZIP checks inside a bounded Docker sandbox. Verify sandbox availability before task creation. Record screenshot SHA-256 at capture. - Keep runner crash probes in controlled recovery tests. Preserve the original failure when cleanup also fails. - Display missing accounting and replay revisions as unavailable. Align harness security assertions with the approved CI changes. - Make the channel-rejection browser fixture bind its file after the send captures its payload. This prevents live refresh from removing the file before the simulated race. ## Verification - Full workspace `pnpm -r typecheck` passed after merging current master. - Runner E2E typecheck passed. Harness unit tests passed: 216/216. - Wake-queue database tests passed: 55/55. The two added existing-incident tests failed before the fix and pass after it. - Docker artifact calibration passed: 12/12. Host-file and host-loopback isolation tests failed before the fix and pass after it. Read-only delivery and output limits are also verified. - Full `pnpm build` passed. Targeted recovery tests passed: 83/83. - The channel-rejection browser test passed five consecutive runs after fixing the fixture race found in CI. - Local general-server (12,351 tests), UI (6,250), CLI (485), and workspace package groups passed. The monolithic run stopped at an unchanged lock-heartbeat fixture race; the isolated workspace group passed on rerun (shared: 747/747). A separate local serialized run passed 97 files before two socket errors in the unchanged issue-list route suite; that suite passed 15/15 on isolated rerun. These local full commands did not finish uninterrupted; the complete CI matrix below covers the remaining suites. - Final head `0fb293733fe307be7e6667ae8f1364077d0c6455`: **34 successful checks, 2 expected skips**, including every server/workspace shard, browser shard, native runner verification, build, and typecheck. [Final CI run](https://github.com/paperclipai/paperclip/actions/runs/34989136700). - Greptile reviewed this exact head at **5/5**; all review threads are resolved. Both Superagent security checks are successful. - ACPX credential-boundary tests passed: 118/118. Superagent accepted the runner/sidecar versus provider-environment trace and cleared its finding. - The latest paid local campaign on source `f6a2fdf7ac2af859826a2ae627ff4125a5478529` passed 22/24 cases: Sol 8/8, Claude 7/8, Mini 7/8. These results predate the merge with current master. - The two remaining failures are in `hire-reuse`: Claude exceeded the attempt deadline during final review; Mini made invalid deliverable tool calls and remained Blocked. - Six Daytona cases were not run because the matching immutable runner image was unavailable. This PR does not claim new remote model results. ## Risks The changes affect connection admission, helper identity, and recovery scheduling. Assigned gateway grants and user approval still govern service calls. The workspace admission gate still exists; the broader folder-sync design is separate work. Provider behavior can still cause the two recorded hiring failures. No database migration is required. Paid cases are opt-in and have bounded attempt deadlines. Project stories now require Docker and the documented pinned Python image on the harness host. ## Model Used OpenAI `gpt-6-astra` performed implementation, diagnosis, and substantive review. OpenAI `gpt-5.6-luna` assisted with verification, PR preparation, and review tracking. Both used repository tools and code execution. Context-window sizes were not recorded. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused checks and isolated reruns; full-run limitations are documented above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: OpenAI GPT-5.6 Luna <noreply@openai.com> Co-authored-by: Paperclip <noreply@paperclip.ing>
237 lines
7.5 KiB
TypeScript
237 lines
7.5 KiB
TypeScript
import { spawn } from "node:child_process";
|
|
import {
|
|
mkdir,
|
|
readFile,
|
|
readdir,
|
|
rm,
|
|
writeFile,
|
|
copyFile,
|
|
} from "node:fs/promises";
|
|
import path from "node:path";
|
|
import {
|
|
assertSecretFree,
|
|
findSecretLeak,
|
|
findSecretLeakInJsonValues,
|
|
redactText,
|
|
sanitizeJson,
|
|
} from "./redaction.js";
|
|
|
|
const TEXT_EXTENSIONS = new Set([
|
|
".json",
|
|
".xml",
|
|
".html",
|
|
".css",
|
|
".js",
|
|
".txt",
|
|
".log",
|
|
".md",
|
|
]);
|
|
const BINARY_EXTENSIONS = new Set([".png", ".webm"]);
|
|
const ACTIVE_CONTENT_EXTENSIONS = new Set([".svg"]);
|
|
const ALLOWED_DIRECTORIES = new Set([
|
|
"snapshots",
|
|
"playwright-output",
|
|
"blob-report",
|
|
"html-report",
|
|
]);
|
|
const ALLOWED_ROOT_FILES = new Set([
|
|
"result.json",
|
|
"final-state.png",
|
|
"failure.png",
|
|
"tool-review-pending.png",
|
|
"decision-pending.png",
|
|
"chat-plan-draft.png",
|
|
"chat-plan-revised.png",
|
|
"server.log",
|
|
"playwright.log",
|
|
"junit.xml",
|
|
]);
|
|
const REQUIRED_PASS_FILES = [
|
|
"final-state.png",
|
|
"result.json",
|
|
"junit.xml",
|
|
path.join("html-report", "index.html"),
|
|
path.join("snapshots", "fixtures.json"),
|
|
path.join("snapshots", "api-state.json"),
|
|
] as const;
|
|
|
|
export interface EvidencePackageResult {
|
|
files: string[];
|
|
leaks: Array<{ file: string; reason: string }>;
|
|
missing: string[];
|
|
}
|
|
|
|
async function walk(root: string, relative = ""): Promise<string[]> {
|
|
const directory = path.join(root, relative);
|
|
const entries = await readdir(directory, { withFileTypes: true }).catch(
|
|
() => [],
|
|
);
|
|
const files: string[] = [];
|
|
for (const entry of entries) {
|
|
const next = path.join(relative, entry.name);
|
|
if (entry.isDirectory()) files.push(...(await walk(root, next)));
|
|
else if (entry.isFile()) files.push(next);
|
|
}
|
|
return files;
|
|
}
|
|
|
|
function isAllowed(relative: string) {
|
|
const segments = relative.split(path.sep);
|
|
if (segments.length === 1)
|
|
return (
|
|
ALLOWED_ROOT_FILES.has(relative) ||
|
|
/^(?:plan|question)-[a-z0-9-]+\.png$/.test(relative)
|
|
);
|
|
return ALLOWED_DIRECTORIES.has(segments[0]);
|
|
}
|
|
|
|
async function inspectZip(source: string, secrets: readonly string[]) {
|
|
// Failure traces can exceed Node's child-process output buffer. Stream the
|
|
// expanded archive through the exact-value scanner with enough overlap to
|
|
// detect a credential split across stdout chunks, without retaining the
|
|
// archive in memory or weakening fail-closed evidence publication.
|
|
const overlap = Math.max(
|
|
256,
|
|
...secrets.map((secret) => Buffer.byteLength(secret, "utf8") + 16),
|
|
);
|
|
return new Promise<string | null>((resolve) => {
|
|
const unzip = spawn("unzip", ["-p", source], {
|
|
stdio: ["ignore", "pipe", "pipe"],
|
|
});
|
|
let carry = Buffer.alloc(0);
|
|
let leak: string | null = null;
|
|
let spawnError: Error | null = null;
|
|
|
|
unzip.stdout.on("data", (chunk: Buffer) => {
|
|
if (leak) return;
|
|
const data = Buffer.concat([carry, chunk]);
|
|
leak = findSecretLeak(data, secrets, { includeShapes: false });
|
|
carry = data.subarray(Math.max(0, data.length - overlap));
|
|
if (leak) unzip.kill();
|
|
});
|
|
// Drain diagnostics so a noisy unzip cannot block. Error text is withheld
|
|
// because archive paths and contents belong to private attempt evidence.
|
|
unzip.stderr.resume();
|
|
unzip.on("error", (error) => {
|
|
spawnError = error;
|
|
});
|
|
unzip.on("close", (code) => {
|
|
if (leak) return resolve(leak);
|
|
if (spawnError) {
|
|
return resolve(`zip could not be inspected: ${spawnError.message}`);
|
|
}
|
|
return resolve(
|
|
code === 0
|
|
? null
|
|
: `zip could not be inspected: unzip exited with code ${String(code)}`,
|
|
);
|
|
});
|
|
});
|
|
}
|
|
|
|
export async function packageEvidence(input: {
|
|
privateDir: string;
|
|
uploadDir: string;
|
|
secrets: readonly string[];
|
|
expectPassScreenshot: boolean;
|
|
}): Promise<EvidencePackageResult> {
|
|
await rm(input.uploadDir, { recursive: true, force: true });
|
|
await mkdir(input.uploadDir, { recursive: true });
|
|
const files: string[] = [];
|
|
const leaks: EvidencePackageResult["leaks"] = [];
|
|
const available = await walk(input.privateDir);
|
|
|
|
for (const relative of available.filter(isAllowed)) {
|
|
const source = path.join(input.privateDir, relative);
|
|
const extension = path.extname(relative).toLowerCase();
|
|
const destination = path.join(input.uploadDir, relative);
|
|
await mkdir(path.dirname(destination), { recursive: true });
|
|
if (ACTIVE_CONTENT_EXTENSIONS.has(extension)) {
|
|
// SVG can execute script when opened directly from an artifact. Keep the
|
|
// source in the disposable private attempt directory, but never admit it
|
|
// to the sanitized CI artifact.
|
|
continue;
|
|
} else if (TEXT_EXTENSIONS.has(extension)) {
|
|
const raw = await readFile(source, "utf8");
|
|
const parsed = extension === ".json" ? JSON.parse(raw) : null;
|
|
const leak =
|
|
extension === ".json"
|
|
? findSecretLeakInJsonValues(parsed, input.secrets)
|
|
: findSecretLeak(raw, input.secrets);
|
|
if (leak) leaks.push({ file: relative, reason: leak });
|
|
// Redacting an already serialized JSON string can change escape
|
|
// boundaries around shell commands. Sanitize parsed values instead so
|
|
// uploaded snapshots stay valid JSON.
|
|
const safe =
|
|
extension === ".json"
|
|
? `${JSON.stringify(sanitizeJson(parsed, input.secrets), null, 2)}\n`
|
|
: redactText(raw, input.secrets);
|
|
if (extension === ".json") {
|
|
const safeLeak = findSecretLeakInJsonValues(
|
|
JSON.parse(safe),
|
|
input.secrets,
|
|
);
|
|
if (safeLeak)
|
|
throw new Error(`Secret leak in ${relative}: ${safeLeak}`);
|
|
} else {
|
|
assertSecretFree(safe, input.secrets, relative);
|
|
}
|
|
await writeFile(destination, safe, "utf8");
|
|
files.push(relative);
|
|
} else if (extension === ".zip") {
|
|
const leak = await inspectZip(source, input.secrets);
|
|
if (leak) {
|
|
leaks.push({ file: relative, reason: leak });
|
|
continue;
|
|
}
|
|
await copyFile(source, destination);
|
|
files.push(relative);
|
|
} else if (BINARY_EXTENSIONS.has(extension)) {
|
|
// This raw-byte scan catches embedded plaintext credentials, but cannot
|
|
// inspect rendered pixels. Provider/UI raster and video files remain in
|
|
// the access-controlled CI artifact. The public publisher creates its
|
|
// own synthetic summary image from fixed labels and numeric/status data.
|
|
const raw = await readFile(source);
|
|
const leak = findSecretLeak(raw, input.secrets);
|
|
if (leak) {
|
|
leaks.push({ file: relative, reason: leak });
|
|
continue;
|
|
}
|
|
await copyFile(source, destination);
|
|
files.push(relative);
|
|
}
|
|
}
|
|
|
|
const missing = input.expectPassScreenshot
|
|
? [
|
|
...REQUIRED_PASS_FILES.filter((required) => !files.includes(required)),
|
|
...(files.some(
|
|
(file) =>
|
|
file.startsWith(`blob-report${path.sep}`) && file.endsWith(".zip"),
|
|
)
|
|
? []
|
|
: [path.join("blob-report", "*.zip")]),
|
|
]
|
|
: [];
|
|
const manifest = {
|
|
schema: "paperclip.runner-e2e.evidence/v1",
|
|
files: [...files].sort(),
|
|
leaks,
|
|
missing,
|
|
};
|
|
const manifestText = `${JSON.stringify(sanitizeJson(manifest, input.secrets), null, 2)}\n`;
|
|
const manifestLeak = findSecretLeakInJsonValues(
|
|
JSON.parse(manifestText),
|
|
input.secrets,
|
|
);
|
|
if (manifestLeak)
|
|
throw new Error(`Secret leak in evidence-manifest.json: ${manifestLeak}`);
|
|
await writeFile(
|
|
path.join(input.uploadDir, "evidence-manifest.json"),
|
|
manifestText,
|
|
"utf8",
|
|
);
|
|
files.push("evidence-manifest.json");
|
|
return { files, leaks, missing };
|
|
}
|