mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 21:05:21 +02:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - An agent needs personal files across tasks and sessions. > - AGENTS.md is one file in that directory. Supporting files need the same persistence. > - The Instructions Editor and agent runs must share one current directory. > - Concurrent runs should apply only the files they change. The last sync of the same file wins. > - This pull request uses existing file transport and removes temporary copies after sync. > - Old instruction-only sessions keep their restore contract. New saves do not create revision history. ## Linked Issues or Issue Description Refs #14325. This replaces its revision-oriented design with persistent agent files. Keep #14325 unmerged. Transport prerequisite #14416 merged first at `d172197117a14b80a1eb2d2835a0e7cce2679656`. This PR now targets master and remains below 100 changed files. Related work: #4513 and #8798 cover instruction tooling. This change handles run synchronization, cross-task personal files, browser editing, and old-session restoration. ## What Changed - Keep one current directory per company and agent. Point AGENT_HOME at a temporary working copy for each active run. Keep task files and provider HOME separate. - Restore text, binary files, and nested folders through workspace transport. Exclude remote agent files from task Git snapshots with a self-ignoring file inside the reserved runtime directory; never write through repository-controlled Git metadata. - Collect after the provider and child processes have stopped. Keep resumable conversation state. - Apply changed and deleted files under the agent lock. The last sync wins for the same file. Unrelated concurrent changes survive. - Remove temporary copies after successful sync, rejected sync, and staging failure. Register ownership before copying so restart recovery can remove interrupted preparation. Retry transient synchronization up to three times. Preserve the original remote lease reference until deletion succeeds; restart cleanup never acquires a replacement sandbox. Do not create captured directories or a conflict-review queue for new runs. - Keep browser editing, stale-draft protection, and streaming binary downloads. Keep the instruction entry and text editor limited to 1 MiB. - Keep historical agent-folder sync failures on their affected runs instead of repeating them above current saved instructions. Preserve legacy candidate review and current browser-save errors. Avoid duplicate quota warnings while retaining separate sync failures when they describe a different problem. - Require target-scoped caller grants for peer instruction access, while preserving self edits, responsible-user checks, and protected-change consent. - Treat full storage as a nonblocking run warning, never an agent pause or run-admission failure. Restore already-over-quota saved folders so ordinary agent cleanup can recover; warn on each run until cleanup. The run detail view shows the warning. - Allow 256 MiB per file, 2 GiB per directory, and 100,000 entries. Hash large files as streams. Check editor-save quotas with metadata instead of hashing unrelated files. - Preserve old native inputs, instruction-only copies, paths, digests, and pending legacy candidates. Adopt old revision heads once. New writes do not append history rows. - Add idempotent migration 0287 and verify upgrades from the preview tables and receipts. - Add nine interactive stories under **Agents / Persistent files**, including automatic incoming edits, stale browser drafts, and storage-limit diagnostics. ## Verification - Merge candidate: `4f5390107ec6ffd80a76d1d2e85530e66f21d079`, after merging current master and the landed transport prerequisite. Integration required no manual conflict resolution; the feature remains 99 changed files. Full workspace typecheck, production build, token gates, and 715 focused tests passed on this merge candidate. Fresh Greptile review is 5/5 with no unresolved findings. All 55 checks passed, with four conditional skips, including the build, typecheck, browser E2E, and canary dry run. A single retry recovered four jobs interrupted by runner shutdowns; no source changes were required. - Historical-warning UI fix: all 6,834 UI tests across 640 files passed, including regression coverage for three old failures, legacy preserved edits, and warnings scoped to the affected run. Full workspace typecheck, production build, Storybook build, and token gates passed. Browser-verified Storybook playtests passed for Historical Failures After Successful Save, Storage Limit, and Full Storage Run Warning. - Review follow-ups at `4e20c9fb2`: all 18 focused tests passed, including external Git directories, linked worktrees, symlinks, hardlinks, and distinct I/O failures alongside storage warnings. Server and UI typechecks, token gates, and the production build passed. - Storage warning regressions at `0724f3012`: all 33 directory tests and all five heartbeat-list tests passed, with no skips in their successful runs. They cover repeated runs while full, an already-over-quota saved folder, cleanup, warnings retained after unrelated save failures, and bounded warnings in large result JSON. Server typecheck passed after the final warning fixes. - Full workspace typecheck, production build, and token gates passed during this follow-up. Product E2E harness: 631 tests passed across 52 files; harness typecheck passed. Earlier native session/context and directory/legacy collection suites passed 537 tests; Runner unit/transport suites passed 329 tests. - **Real E2E at `0724f3012` (before this follow-up):** legacy local Codex and native Daytona Codex each passed six tasks, one server restart, seven independent assertions, and cleanup verification. Both prove browser-to-agent edits, agent-to-browser edits, nested/binary restoration, per-file last-sync-wins, a successful run after an oversized save rejection, and cleanup clearing the warning. - Native local Codex also passed the six-task quota flow before the final warning-retention fixes. That pass began at `918d1ed02` while the bounded-result warning fix was being edited, so it is not claimed as exact-final-head evidence. Its final-head rerun failed during embedded PostgreSQL bootstrap before any provider run: the macOS host had 87,365 of 87,381 SysV semaphores occupied. No unrelated services or kernel limits were changed. - The final-source report intentionally records **2/3 cells passed**, preserving the blocked native-local attempt: `tests/runner-e2e/results/agent-files-quota-final-20260928-report/`. Earlier failed attempts and provenance notes remain under `tests/runner-e2e/results/agent-files-quota-final-20260928-input/` and the original campaign directories. - Daytona used immutable image `ghcr.io/paperclipai/paperclip-daytona-runner@sha256:5643f0d801417cae3581833a1a3bc6715b325e028602738d2652c44cac5dc6bf` and its exact Linux runner binary. Controller source is `0724f3012`; image source is recorded separately. - Legacy-session compatibility and all three ACP Stop/resume browser regressions passed on the prior validated feature head `169fab46d5af21caa2269b4c1b29b69c933a6951`. They assert the same provider session is retained and interrupted writes are not replayed. Migration upgrade tests also passed earlier. - Nine interactive stories are under **Agents / Persistent files**, including **Full Storage Run Warning**. Its playtest and visual browser inspection passed; the warning states that runs continue and the editor remains available. - Prior-head checks on `4e20c9fb2`: 55 passed, two conditional jobs skipped, no failures or pending checks. All eight browser E2E shards and their aggregate passed. Fresh Greptile review is 5/5 with no findings; all review threads are resolved, the security scan passed, and GitHub reports no merge conflicts. - The broad local follow-up test run was interrupted after host semaphore exhaustion affected isolated PostgreSQL instances. It also encountered the existing macOS long-path fixture failure and two timeout failures. This is not a claim that the full local suite passed. Logs are retained; focused storage/warning tests passed. ## Risks - A later sync can overwrite an earlier edit to the same file, including a saved browser edit. There is no text merge or retained version. This is the intended last-sync-wins policy. - A save that exceeds a storage limit is rejected and its temporary copy is discarded. The run itself continues normally, and later runs restore the last saved files with a warning until cleanup. Transient sync failures get bounded retries. An I/O failure partway through a sync can leave some files updated; a failed receipt does not claim whole-folder success. - Larger folders increase copy time, network traffic, and temporary disk usage. Active runs still need working copies. Terminal runs do not accumulate archives. Operators must provision disk for agents and configured concurrency; these limits are not company-wide quotas. - A restored old native session remains instruction-only until a fresh session starts. Its original conflict fence and existing pending candidates remain compatible. - Provider processes close at the collection boundary. Conversation resume remains available, but warm process reuse is lost. - Backups must include the instance filesystem and database. External bundles keep their existing behavior until explicitly moved to managed storage. ## Model Used OpenAI Codex, GPT-6 family. The session does not expose a more specific model ID or context-window size. Reasoning, code execution, and browser tools assisted this change. Real provider E2E uses `gpt-5.6-sol`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Fry (Paperclip) <noreply@paperclip.ing>
233 lines
20 KiB
TypeScript
233 lines
20 KiB
TypeScript
import { createHash, randomBytes } from "node:crypto";
|
|
import { expect, type Page } from "@playwright/test";
|
|
import { pollUntil, type RunnerApi } from "./api.js";
|
|
import { captureFirstTaskAttachments } from "./first-task-attachments.js";
|
|
import { collectRunEvents } from "./run-observations.js";
|
|
import { createTaskThroughUi } from "./user-actions.js";
|
|
import type { LiveFixtureValues } from "./live-fixtures.js";
|
|
import type { MatrixExecution, RunnerTaskFixture } from "./types.js";
|
|
|
|
type Row = Record<string, any>;
|
|
export const instructionNonceLine = (nonce: string) => `Instruction persistence nonce: ${nonce}\n`;
|
|
export const instructionPersistenceTask: RunnerTaskFixture = {
|
|
id: "private-copy-persists", label: "Agent directory survives a fresh task",
|
|
groups: [], workMode: "standard", flow: "instruction_persistence",
|
|
expectedRunCount: 6, attemptTimeoutMs: { local: 20 * 60_000, daytona: 20 * 60_000 },
|
|
expectedTerminalState: { issue: "done", run: "succeeded" },
|
|
buildTitle: nonce => `Persist private instructions ${nonce}`,
|
|
buildVisibleMarker: () => "INSTRUCTIONS-VERIFIED",
|
|
buildPrompt: nonce => [
|
|
"Edit your own registered writable agent instruction entry with ordinary filesystem tools. The runtime guidance gives its exact private path.",
|
|
"Use Node.js built-in fs for these byte-preserving edits. Apply each append exactly once: inspect the existing suffix before retrying any command, because a warning does not imply that its writes failed.",
|
|
`Preserve its existing bytes and append exactly this UTF-8 suffix, represented as a JSON string: ${JSON.stringify(`\n${instructionNonceLine(nonce)}`)}`,
|
|
"Decode the JSON string once and append those bytes. Do not trim or normalize the existing file and do not add another blank line or separator.",
|
|
`In AGENT_HOME, create notes/retained.txt containing exactly ${JSON.stringify(`Personal file nonce: ${nonce}\n`)}. Create notes/bytes.bin with exactly the bytes [0,255,17,128,9]. Read notes/from-editor.txt and append exactly a newline followed by Edited by agent. and a final newline.`,
|
|
"Do not use update_agent_instructions, restore_agent_instructions, or an instructions API to save it. Do not edit repository AGENTS.md or the read-only loaded bundle.",
|
|
"After verifying the edits, upload a small text/plain attachment named agent-file-check.txt containing only 'Private file edits verified'. Use this attachment as your task completion evidence; the personal files themselves stay in AGENT_HOME.",
|
|
"Reply only Instruction copy edited without printing filesystem paths, then complete this task after the file edit. Paperclip will collect it after the provider stops; do not claim it has already persisted. Do not create further tasks.",
|
|
].join("\n"),
|
|
buildMatchers: () => [], // Independent current file and attachment oracle below.
|
|
};
|
|
|
|
export function gradeInstructionPersistence(input: { before: Row; after: Row; firstRunId: string; expectedContent: string; proof: Row | undefined; expectedProof: string; saveEvent?: Row; fileProof?: boolean }) {
|
|
return [
|
|
{ id: "directory-save-receipt", passed: input.saveEvent?.runId === input.firstRunId && input.saveEvent?.state === "saved" },
|
|
{ id: "nested-and-binary-files", passed: input.fileProof === true },
|
|
{ id: "exact-canonical-bytes", passed: input.after.content === input.expectedContent && input.after.contentHash === createHash("sha256").update(input.expectedContent).digest("hex") },
|
|
{ id: "fresh-task-downloaded-proof", passed: input.proof?.contentVerified === true && input.proof.body === input.expectedProof },
|
|
].map(check => ({ ...check, detail: check.passed ? `${check.id} verified independently` : `${check.id} missing or incorrect` }));
|
|
}
|
|
|
|
export async function runInstructionPersistenceFlow(input: {
|
|
page: Page; api: RunnerApi; fixtures: LiveFixtureValues; execution: MatrixExecution; nonce: string;
|
|
secrets: readonly string[]; deadlineAt: number;
|
|
restart(): Promise<void>;
|
|
observe(issue: Row, runs: Row[]): void;
|
|
capture(id: string, label: string, file: string): Promise<void>;
|
|
evidence(name: string, data: unknown): Promise<void>;
|
|
}) {
|
|
const { page, api, fixtures, execution, nonce } = input;
|
|
// Fixture names contain the campaign nonce. Use an unrelated value that the
|
|
// fresh task can obtain only from the saved entry (or forbidden task history).
|
|
const persistedNonce = randomBytes(16).toString("hex");
|
|
const filePath = `/api/agents/${fixtures.agent.id}/instructions-bundle/file?path=AGENTS.md`;
|
|
const before = await api.get<Row>(filePath);
|
|
if (typeof before.content !== "string" || !before.contentHash) throw new Error("Managed instructions must expose a current file hash");
|
|
const expectedContent = `${before.content}\n${instructionNonceLine(persistedNonce)}`;
|
|
const instructionsUrl = `/${fixtures.company.issuePrefix}/agents/${fixtures.agent.id}/instructions`;
|
|
const editorText = `Editor nonce: ${randomBytes(16).toString("hex")}`;
|
|
await page.goto(instructionsUrl);
|
|
await page.getByRole("button", { name: "Add agent file", exact: true }).click();
|
|
await page.getByPlaceholder("TOOLS.md").fill("notes/from-editor.txt");
|
|
await page.getByRole("button", { name: "Create", exact: true }).click();
|
|
await page.getByRole("group", { name: "Instruction file view" }).getByRole("button", { name: "edit", exact: true }).click();
|
|
await page.getByRole("textbox", { name: "Instruction file editor" }).fill(editorText);
|
|
await page.getByRole("button", { name: "Save changes", exact: true }).click();
|
|
await expect(page.getByRole("button", { name: "Save changes", exact: true })).toBeDisabled();
|
|
await expect(page.getByRole("button", { name: "History", exact: true })).toHaveCount(0);
|
|
let issue: Row = {};
|
|
let runs: Row[] = [];
|
|
async function create(title: string, prompt: string) {
|
|
await createTaskThroughUi({ page, issuePrefix: fixtures.company.issuePrefix!, agentName: fixtures.agent.name, title, prompt, workMode: "standard", projectName: fixtures.project?.name });
|
|
const found = await pollUntil({ label: `instruction task ${title}`, deadlineAt: input.deadlineAt,
|
|
load: async () => (await api.get<Row[]>(`/api/companies/${fixtures.company.id}/issues?limit=100`)).find(row => row.title === title), accept: row => Boolean(row) });
|
|
if (!found) throw new Error("Browser-created instruction task missing");
|
|
issue = found;
|
|
input.observe(issue, runs);
|
|
await page.goto(`/${fixtures.company.issuePrefix}/issues/${issue.identifier ?? issue.id}`);
|
|
}
|
|
async function settle(count: number) {
|
|
await pollUntil({ label: `instruction run ${count} completed`, deadlineAt: input.deadlineAt,
|
|
load: async () => {
|
|
issue = await api.get<Row>(`/api/issues/${issue.id}`);
|
|
const listed = await api.get<Row[]>(`/api/companies/${fixtures.company.id}/heartbeat-runs?limit=100`);
|
|
runs = await Promise.all(listed.map(row => api.get<Row>(`/api/heartbeat-runs/${row.id}`)));
|
|
runs.sort((a, b) => String(a.createdAt).localeCompare(String(b.createdAt)));
|
|
input.observe(issue, runs);
|
|
return { issue, runs };
|
|
},
|
|
accept: state => state.issue.status === "done" && state.runs.length === count && state.runs.every(row => row.status === "succeeded"),
|
|
reject: state => state.runs.some(row => ["failed", "cancelled", "timed_out"].includes(row.status)) ? "Instruction task provider run failed" : state.runs.length > count ? "Instruction task dispatched an extra run" : state.issue.status === "blocked" && state.runs.length === count && state.runs.every(row => row.status === "succeeded") ? `Instruction task reported a terminal blocker: ${JSON.stringify(state.issue.unblockDescriptor ?? {})}` : undefined,
|
|
});
|
|
expect(runs.every(row => row.runtimeMode === execution.profile.expectedRuntimeMode)).toBe(true);
|
|
await page.reload();
|
|
await expect(page.getByTestId("issue-detail-header").getByRole("button", { name: "Change status (current: Done)", exact: true })).toBeVisible();
|
|
}
|
|
await create(execution.task.buildTitle(nonce), execution.task.buildPrompt(persistedNonce));
|
|
await settle(1);
|
|
const firstRunId = runs[0]!.id;
|
|
const after = await pollUntil({ label: "stopped agent directory save", deadlineAt: Math.min(input.deadlineAt, Date.now() + 30_000),
|
|
load: () => api.get<Row>(filePath), accept: row => row.content === expectedContent });
|
|
const events = await collectRunEvents<Row>((afterSeq, limit) => api.get(`/api/heartbeat-runs/${firstRunId}/events?afterSeq=${afterSeq}&limit=${limit}`));
|
|
const saveEvent = events.find(row => row.eventType === "instruction_save" && row.payload?.state === "saved");
|
|
expect(saveEvent).toBeTruthy();
|
|
const readPersonal = (name: string) => api.get<Row>(`/api/agents/${fixtures.agent.id}/instructions-bundle/file?path=${encodeURIComponent(name)}`);
|
|
const note = await readPersonal("notes/retained.txt");
|
|
const fromEditor = await readPersonal("notes/from-editor.txt");
|
|
const binaryResponse = await api.request.get(`/api/agents/${fixtures.agent.id}/instructions-bundle/file?path=notes%2Fbytes.bin&download=true`);
|
|
expect(binaryResponse.ok()).toBe(true);
|
|
const binary = await binaryResponse.body();
|
|
const fileProof = note.content === `Personal file nonce: ${persistedNonce}\n` && fromEditor.content === `${editorText}\nEdited by agent.\n` && binary.equals(Buffer.from([0,255,17,128,9]));
|
|
expect(fileProof).toBe(true);
|
|
const history = await api.get<Row>(`/api/agents/${fixtures.agent.id}/instructions-bundle/history?path=AGENTS.md`);
|
|
expect(history.revisions).toHaveLength(0);
|
|
await input.evidence("instruction-first-save.json", { before, after, run: runs[0], events });
|
|
await input.capture("instruction-edited", "Private instructions saved after provider stop", "instruction-edited.png");
|
|
// A new server and a new issue cannot pass by retaining model conversation.
|
|
await input.restart();
|
|
expect((await api.get<Row>(filePath)).content).toBe(expectedContent);
|
|
await create("Read persisted instructions", [
|
|
"Read your own loaded agent instruction entry (or its current registered private copy) using ordinary filesystem tools.",
|
|
"Find the line beginning 'Instruction persistence nonce: '. Copy that entire line plus one final newline into instruction-proof.txt. Then append the exact bytes of notes/retained.txt from AGENT_HOME. Verify notes/bytes.bin contains the bytes [0,255,17,128,9]. Do not infer the value from this task title or other task history. Do not change your instructions.",
|
|
"Upload instruction-proof.txt as a text/plain task attachment named instruction-proof.txt using the normal artifact workflow. A local file alone is insufficient.",
|
|
`Reply with exactly ${execution.task.buildVisibleMarker(nonce)} and complete the task.`,
|
|
].join("\n"));
|
|
await settle(2);
|
|
const attachments = await captureFirstTaskAttachments(api, [{ ...issue, id: String(issue.id) }], input.secrets);
|
|
const proof = attachments.find(row => row.originalFilename === "instruction-proof.txt" || row.name === "instruction-proof.txt");
|
|
const final = await api.get<Row>(filePath);
|
|
expect(final.revision.id).toBe(after.revision.id);
|
|
const checks = gradeInstructionPersistence({ before, after, firstRunId, expectedContent, proof, expectedProof: `${instructionNonceLine(persistedNonce)}Personal file nonce: ${persistedNonce}\n`, saveEvent: { runId: firstRunId, state: saveEvent?.payload?.state }, fileProof });
|
|
await expect(page.getByTestId("task-chat-agent-bubble").filter({ hasText: execution.task.buildVisibleMarker(nonce) }).last()).toBeVisible();
|
|
await input.capture("final-state", "Fresh task downloaded the persisted instruction nonce", "final-state.png");
|
|
expect(checks.filter(check => !check.passed), "Independent instruction persistence checks").toEqual([]);
|
|
|
|
checks.push({ id: "editor-round-trip-no-history", passed: true, detail: "A browser-created supporting file was edited by the agent; all current bytes persisted without revision rows" });
|
|
const restored = final;
|
|
|
|
// The provider publishes an ordinary attachment before a bounded command wait.
|
|
// A board edit during that run is superseded only for the same changed file.
|
|
const conflictSuffix = `\nLast completed synchronization: ${nonce}\n`;
|
|
const expectedCandidate = `${restored.content}${conflictSuffix}`;
|
|
await create("Sync a concurrent instruction edit", [
|
|
"Use Node.js built-in fs. Apply the append exactly once, checking existing bytes before any retry.",
|
|
`Append exactly this UTF-8 suffix to your current registered writable instruction entry, represented as a JSON string: ${JSON.stringify(conflictSuffix)}`,
|
|
"Decode the JSON string once. Preserve all existing bytes. Do not use an instruction revision tool or instructions API.",
|
|
"After the file edit, upload a text/plain task attachment named instruction-candidate-ready.txt with the text ready. Use the ordinary artifact workflow.",
|
|
"Then execute the ordinary shell command sleep 45 and wait for it to finish. This gives the board time to edit the canonical instructions concurrently. Do not complete the task before that command finishes.",
|
|
"After the wait completes, reply Candidate edit ready and complete the task. Do not change the instructions again or claim that they saved.",
|
|
].join("\n"));
|
|
await pollUntil({ label: "provider staged concurrent instruction edit", deadlineAt: input.deadlineAt,
|
|
load: () => api.get<Row[]>(`/api/issues/${issue.id}/attachments`),
|
|
accept: rows => rows.some(row => row.originalFilename === "instruction-candidate-ready.txt" || row.name === "instruction-candidate-ready.txt") });
|
|
const active = await api.get<Row[]>(`/api/issues/${issue.id}/runs`);
|
|
expect(active.some(row => row.status === "running")).toBe(true);
|
|
const boardMarker = "Concurrent board instruction edit.";
|
|
await page.goto(instructionsUrl);
|
|
await page.getByText("AGENTS.md", { exact: true }).first().click();
|
|
await page.getByRole("group", { name: "Instruction file view" }).getByRole("button", { name: "edit", exact: true }).click();
|
|
const entryEditor = page.getByRole("textbox", { name: "editable markdown" });
|
|
await entryEditor.click();
|
|
await entryEditor.press("ControlOrMeta+End");
|
|
await entryEditor.press("Enter");
|
|
await entryEditor.pressSequentially(boardMarker);
|
|
await page.getByRole("button", { name: "Save changes", exact: true }).click();
|
|
await expect(page.getByRole("button", { name: "Save changes", exact: true })).toBeDisabled();
|
|
const board = await api.get<Row>(filePath);
|
|
expect(board.content).toContain(boardMarker);
|
|
expect(board.contentHash).not.toBe(restored.contentHash);
|
|
const unrelatedContent = `Concurrent independent file: ${nonce}`;
|
|
const unrelated = await api.request.put(`/api/agents/${fixtures.agent.id}/instructions-bundle/file`, {
|
|
data: { path: "notes/concurrent-editor.txt", content: unrelatedContent, baseHash: null },
|
|
});
|
|
expect(unrelated.ok()).toBe(true);
|
|
await page.goto(`/${fixtures.company.issuePrefix}/issues/${issue.identifier ?? issue.id}`);
|
|
await settle(3);
|
|
const syncRunId = runs[2]!.id;
|
|
const resolved = await pollUntil({ label: "last completed sync wins", deadlineAt: Math.min(input.deadlineAt, Date.now() + 30_000),
|
|
load: () => api.get<Row>(filePath), accept: row => row.content === expectedCandidate });
|
|
const candidates = await api.get<Row[]>(`/api/agents/${fixtures.agent.id}/instructions-bundle/candidates`);
|
|
expect(candidates.some(row => row.runId === syncRunId)).toBe(false);
|
|
expect((await readPersonal("notes/concurrent-editor.txt")).content).toBe(unrelatedContent);
|
|
const syncEvents = await collectRunEvents<Row>((afterSeq, limit) => api.get(`/api/heartbeat-runs/${syncRunId}/events?afterSeq=${afterSeq}&limit=${limit}`));
|
|
expect(syncEvents.some(row => row.eventType === "instruction_save" && row.payload?.state === "saved")).toBe(true);
|
|
await page.goto(instructionsUrl);
|
|
await expect(page.getByRole("button", { name: "Review preserved files", exact: true })).toHaveCount(0);
|
|
checks.push({ id: "per-file-last-sync-wins", passed: true, detail: "The later agent sync replaced the concurrent browser edit to its changed entry, preserved an unrelated new file, and created no conflict candidate" });
|
|
await page.goto(`/${fixtures.company.issuePrefix}/issues/${issue.identifier ?? issue.id}`);
|
|
await input.capture("last-sync-wins", "Concurrent changes synchronized per file without a conflict-review step", "last-sync-wins.png");
|
|
|
|
const quotaTask = (action: string, receipt: string) => [
|
|
"This is a controlled persistent-storage quota check. Use ordinary Node.js filesystem tools in your registered AGENT_HOME. Do not edit AGENTS.md.",
|
|
action,
|
|
`Upload a small text/plain task attachment named ${receipt}.txt containing the observed file size or cleanup result. This attachment is the primary task deliverable.`,
|
|
"Complete this task normally after uploading the receipt. A persistent-file storage warning is expected and must not prevent completion. Do not perform additional cleanup or change other personal files.",
|
|
].join("\n");
|
|
await create("Reach the agent file storage limit", quotaTask(
|
|
"Create quota-cache.bin using fs.openSync with flag w, fs.ftruncateSync(fd, 268435456), and fs.closeSync. This is a sparse fixture file, not a download. Verify its size using fs.statSync without reading the large contents.", "quota-full"));
|
|
await settle(4);
|
|
const fullRun = runs[3]!;
|
|
expect(fullRun.resultJson?.instructionSave).toMatchObject({ state: "saved", storageWarning: expect.stringContaining("Agent storage is full") });
|
|
await page.goto(`/${fixtures.company.issuePrefix}/agents/${fixtures.agent.id}/runs/${fullRun.id}`);
|
|
await expect(page.getByRole("note").filter({ hasText: "Agent storage warning" })).toContainText("Runs can continue");
|
|
// Public campaign screenshots are limited to sanitized task routes. Verify
|
|
// the warning in the real run UI, then capture its completed task outcome.
|
|
await page.goto(`/${fixtures.company.issuePrefix}/issues/${issue.identifier ?? issue.id}`);
|
|
await input.capture("storage-warning", "Task succeeded while agent storage reached its limit", "storage-warning.png");
|
|
|
|
await create("Keep running while agent storage is full", quotaTask(
|
|
"Verify quota-cache.bin already exists and its size is exactly 268435456. Grow only this file to 268435457 bytes with fs.truncateSync, then verify the new size. Leave it above the limit for this run's sync check.", "quota-exceeded"));
|
|
await settle(5);
|
|
const exceededRun = runs[4]!;
|
|
expect(exceededRun.resultJson?.instructionSave).toMatchObject({ state: "unavailable", errorCode: "AGENT_FILES_LIMIT_EXCEEDED", storageWarning: expect.stringContaining("Runs can continue") });
|
|
const fullEvents = await collectRunEvents<Row>((afterSeq, limit) => api.get(`/api/heartbeat-runs/${exceededRun.id}/events?afterSeq=${afterSeq}&limit=${limit}`));
|
|
expect(fullEvents.some(row => row.eventType === "instruction_save" && row.level === "warn" && row.payload?.state === "prepared" && row.payload?.storageWarning)).toBe(true);
|
|
|
|
await create("Clean up agent storage during a normal task", quotaTask(
|
|
"Verify restored quota-cache.bin has size 268435456: the rejected oversized edit must not have replaced its saved bytes. Delete quota-cache.bin with fs.unlinkSync, then write notes/after-quota.txt containing exactly 'Runs still work after quota cleanup'.", "quota-cleaned"));
|
|
await settle(6);
|
|
const cleanedRun = runs[5]!;
|
|
expect(cleanedRun.resultJson?.instructionSave).toMatchObject({ state: "saved", storageWarning: null });
|
|
expect((await readPersonal("notes/after-quota.txt")).content).toBe("Runs still work after quota cleanup");
|
|
const bundle = await api.get<Row>(`/api/agents/${fixtures.agent.id}/instructions-bundle`);
|
|
expect(bundle.files.some((file: Row) => file.path === "quota-cache.bin")).toBe(false);
|
|
await page.goto(`/${fixtures.company.issuePrefix}/agents/${fixtures.agent.id}/runs/${cleanedRun.id}`);
|
|
await expect(page.getByRole("note").filter({ hasText: "Agent storage warning" })).toHaveCount(0);
|
|
checks.push({ id: "storage-full-does-not-block-runs", passed: true, detail: "A run saved a file at quota, a subsequent run succeeded despite an oversized save rejection, and the next run removed the full file and cleared its warning; all three tasks completed" });
|
|
await input.evidence("api-state.json", { issue, runs, checks, canonicalInstructions: resolved, attachments });
|
|
await input.evidence("instruction-persistence.json", { checks, before, after, final, restored, board, candidates, resolved, syncEvents, runs, attachments, fullEvents });
|
|
await page.goto(`/${fixtures.company.issuePrefix}/issues/${issue.identifier ?? issue.id}`);
|
|
await input.capture("storage-recovered", "Agent completed a normal task and cleared storage warning after cleanup", "storage-recovered.png");
|
|
return { issue, runs, checks };
|
|
}
|