Clarify legacy issue document delivery in the operational skill

Preserve the tiny hire manual and original assigned-skill case. Add a focused public document/revision/link oracle and explicit presence/absence provenance for a matched skill-only comparison.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
This commit is contained in:
DottaandPaperclip committed 2026-10-02 15:24:01 -05:00
1 parent 5837aa4442
commit abd0b628ca
16 files changed
+247 -37

No files matched your search

@@ -0,0 +1,18 @@
# Legacy document skill repair qualification — 2026-10-02
**TL;DR: no repair performance result yet.** All four paired cases are pending. The original prompt-removal comparison observed two new classic Claude/OpenCode document-delivery failures; its retained verdicts remain unchanged. This focused comparison measures a small operational-skill repair while holding the tiny manual/shared prompts fixed, rather than remeasuring the combined instruction removal.
The approved repair adds early API-runtime document PUT/receipt/link guidance and a short generic issue-document reference. A valid saved-revision write receipt is sufficient; GET is used for existing-document updates or unclear/conflicting receipts. Explicit destinations, downloadable files, and native document-tool boundaries remain intact. The default manual stays eight words.
| Classic profile | Case | Pre-fix skill | Repaired skill |
| --- | --- | --- | --- |
| Claude | Original assigned skill | Pending | Pending |
| Claude | Explicit Paperclip document | Pending | Pending |
| OpenCode | Original assigned skill | Pending | Pending |
| OpenCode | Explicit Paperclip document | Pending | Pending |
The two variants select these four cells only, eight expected provider turns total. All model, effort, auth, tool, permission, budget, fixture and behavioral-grader sources match; only `skills/paperclip/SKILL.md` and the presence of `skills/paperclip/references/issue-documents.md` differ. The absent historical reference is explicitly fingerprinted as absent; its new content is not copied into the baseline.
The original request and pinned output procedure are preserved verbatim. The added explicit case independently reads the public task document and its saved revision/content and requires an agent comment linking the exact same-app document. Correct relative and same-app absolute URLs pass. Workspace-only output, missing revisions, wrong content, wrong document keys, other-origin URLs and user-only links fail calibration.
Local preparation: 909 E2E support tests and E2E typecheck pass; the narrow selector discovers exactly these four cells. Initial sandbox listener failures and two obsolete catalog-count assertions are retained separately; the unchanged behavioral oracles were not weakened. Exact-head prerequisites, immutable source freeze and protected GitHub dispatch are next. Actual costs and outcome deltas remain unknown until retained results are inspected. Existing Claude chat-memory and legacy ACP credential/receipt failures remain separate unresolved findings.
@@ -1,5 +1,7 @@
# Stock-harness live comparison — 2026-10-02
**TL;DR:** Two newly failing paired cases: classic Claude/OpenCode document delivery. Two newly passing cases: classic and native OpenCode ordered continuation. Seven unchanged failures and 13 unchanged passes. The extra ACP Claude storage symptom occurs beneath an existing credential-guard failure. All 48 results are available; no general performance equivalence is established.
The reduced instructions are **not yet qualified for merge**. Classic Claude and classic OpenCode pass the historical skill-output case but fail with the reduced instructions: they finish without saving a Paperclip task document. Native Codex and native Claude pass all three journeys in both variants. Single trials and ambiguous fixture storage wording limit causal attribution; the failed oracle remains unchanged.
The reductions and evaluation setup are in draft [PR #14948](https://github.com/paperclipai/paperclip/pull/14948). This report and its [safe evidence projection](2026-10-02-stock-harness-live-comparison.json) record the measured revisions rather than claiming the final documentation head was run through the full matrix.
@@ -96,3 +98,13 @@ At `1eb5ba420`, all 557 credential-free prerequisites pass (556 TypeScript plus
Repository typecheck and build pass. The complete local Vitest run has 14,870 passed, 83 skipped and three unrelated timing failures; the three affected files pass unchanged narrow reruns (25 auth, 16 reviewed-chat binding, four webhook tests). Original failures are retained rather than calling the first full run green.
PR #14948 stays draft while document delivery is unresolved. Unrepresented harnesses, legacy ACP credential safety, public-receipt clipping, the independent ACP base-replacement issue, existing Claude chat memory behavior and saved-session migration are not qualified away by these results. This small matrix measures skill/context/chat behavior, not broad coding quality or statistical equivalence.
## Retained delivery diagnosis and approved repair
Both failed classic agents loaded the operational Paperclip skill. This is not an observed skill-discovery failure. In the reduced Claude run, the agent wrote a workspace Markdown file before loading Paperclip and then completed the issue without a document API write. The reduced OpenCode run loaded Paperclip, accessed the assigned skill through the public skills API, and wrote a workspace task-document file; it made no document, attachment, or work-product delivery write. Both corresponding historical runs saved a Paperclip issue document.
Claude received the full operational skill, including its existing no-local-only delivery guidance. OpenCode's skill result was explicitly truncated in both variants: the early artifact rule survived, but the later generic document endpoint and plan-only write example were omitted. That truncation is existing behavior, not a newly measured regression. Removing always-on delivery reminders may interact with it, but the combined manual/shared reduction and ambiguous assigned-skill wording prevent a causal attribution to one layer.
Dotta approved a small early API-runtime skill recipe and a generic issue-document reference. The tiny manual remains eight words. The recipe checks the successful write's returned saved revision/content and links the returned key; it does not require another GET after a clear valid receipt. It preserves explicit destinations, downloadable-file delivery, and native document-tool boundaries.
The focused repair comparison will hold the tiny manual/shared prompts fixed and vary only `skills/paperclip/SKILL.md` plus the new `references/issue-documents.md`. Classic Claude and OpenCode each run the preserved original case and an added explicitly Paperclip-storage case: four cells per variant, eight expected provider turns in total. This measures the skill repair separately from the original reduction. The explicit case independently checks public saved body/revision and an agent comment linking the exact same-app document; local-only, missing-revision, wrong-document, local-path, and other-origin outcomes cannot pass. Current state: fixture preparation and source freeze, no repair provider run yet.
@@ -421,3 +421,8 @@ results, remaining exceptions, and follow-ups here before checking it off.
| 2026-10-02 | Matched default-manual/shared-prompt campaigns measured failures, with #14920 constant. | Candidate `f02d8d0df`: 15/24 pass. Historical `12c5433c6`: 15/24 pass after the one runner-shutdown recovery also timed out. Classic Claude/OpenCode Paperclip document delivery regressed in observed trials; keep PR #14948 draft. [Live report](2026-10-02-stock-harness-live-comparison.md). |
| 2026-10-02 | Cold prerequisite and packaging faults were repaired without weakening admission or behavioral graders. | Three setup attempts stopped before providers. Current `1eb5ba420` pilot passes 557 prerequisite checks and protected report publication; source/hash/cost evidence retained. Candidate cancellation and historical runner shutdown recovered only for missing cells. |
| 2026-10-02 | Final bounded recovery completed; no further model reruns. | Both matched cohorts have all 24 results, 15 pass and nine fail. Two classic skill deliveries regress; two OpenCode ordered cases pass only with reduced instructions. Overall parity is not behavioral equivalence. All 48 retained result/receipt projections are hashed and sanitized; original interruptions and partial unknown spend remain recorded. |
| 2026-10-02 | Dotta approved the measured legacy delivery repair and narrow follow-up qualification. | Early operational skill PUT/receipt/link guidance plus generic issue-document reference; eight-word manual retained. Focused original + explicit Paperclip-storage cases on classic Claude/OpenCode, with only the two skill sources varied. Source freeze/admission pending; no full matrix rerun. Claude chat-memory (F15) and ACP credential/receipt failures (F16) remain separate and unresolved. |
| 2026-10-02 | Native tool-description comparison completed all six paired cases. | Zero newly failing cases, three unchanged completion passes, three unchanged blocker failures. Claude/OpenCode blocker API matchers pass in both variants; UI matcher wrongly demanded marker-only replies. Codex's exact-action punctuation failure is unchanged. Original failures retained; provider-free grader correction/calibration in progress, no native fixed-prompt removal measured. |
- F18: Native blocker browser assertion required a marker-only reply despite asking for owner/action/reason. Correct marker-plus-explanation checks symmetrically, calibrate contradictory and future-condition replies, retain original verdicts.
- F19: Both failed classic agents loaded Paperclip; existing OpenCode skill truncation omits the late document recipe. Early skill delivery recipe approved; discovery pointers and restoring the manual are unnecessary for this observed issue.
+2
View File
@@ -15,6 +15,8 @@ You run in **heartbeats** — short execution windows triggered by Paperclip. Ea
In Paperclip, **task** and **issue** refer to the same work item. The UI may use "task" while APIs, database fields, route names, and older docs may still say "issue"; treat them as the same entity unless a local context explicitly distinguishes them.
**Task documents (API runtimes).** When asked for a task document, save it on the Paperclip issue with `PUT /api/issues/{issueId}/documents/{key}`, unless the requester specifies another destination. Confirm the returned document's saved revision and link it before reporting completion; read [references/issue-documents.md](references/issue-documents.md) for the payload and revision-safe updates. Downloadable files follow [Generated Artifacts and Work Products](#generated-artifacts-and-work-products).
## Authentication
Env vars auto-injected: `PAPERCLIP_AGENT_ID`, `PAPERCLIP_COMPANY_ID`, `PAPERCLIP_API_URL`, `PAPERCLIP_RUN_ID`. Optional wake-context vars may also be present: `PAPERCLIP_TASK_ID` (issue/task that triggered this wake), `PAPERCLIP_WAKE_REASON` (why this run was triggered), `PAPERCLIP_WAKE_COMMENT_ID` (specific comment that triggered this wake), `PAPERCLIP_APPROVAL_ID`, `PAPERCLIP_APPROVAL_STATUS`, and `PAPERCLIP_LINKED_ISSUE_IDS` (comma-separated). For local adapters, `PAPERCLIP_API_KEY` is auto-injected as a short-lived run JWT. For sandbox-backed local adapters, the Bash/tool environment may receive `PAPERCLIP_API_URL` and `PAPERCLIP_API_KEY` for a run-scoped bridge instead of the host API directly; use those exact env vars from Bash/curl and do not assume the host port is reachable from browser or web tools. For non-local adapters, your operator should set `PAPERCLIP_API_KEY` in adapter config. All requests use `Authorization: Bearer $PAPERCLIP_API_KEY`. All endpoints are under `/api`. Use JSON except for multipart attachment uploads and binary content downloads. Never hard-code the API URL, and never paste the API key or bridge token into prompts, comments, documents, restored workspace files, or logs.
@@ -0,0 +1,49 @@
# Issue documents through the API
Use this procedure for requested Paperclip task documents in API-based runtimes.
Honor a destination explicitly chosen by the requester. Native Runner runtimes
use their document tool and its contract. Downloadable files follow
[artifacts.md](artifacts.md).
## Create a document
Use the injected `PAPERCLIP_API_URL` and bearer `PAPERCLIP_API_KEY`. Include
`X-Paperclip-Run-Id: $PAPERCLIP_RUN_ID` on the write and send JSON with
`Content-Type: application/json`.
Choose a descriptive document key such as `report`; keys use lowercase letters,
numbers, underscores, or hyphens and are at most 64 characters. Plans use `plan`.
```text
PUT /api/issues/{issueId}/documents/{key}
```
```json
{
"title": "Task report",
"format": "markdown",
"body": "# Task report\n\nThe requested findings.",
"baseRevisionId": null
}
```
A successful write returns HTTP `201` for creation or `200` for an update,
with the saved document JSON. Check its `key`, `body`, and `latestRevisionId`
before claiming delivery. This receipt is sufficient; another GET is unnecessary
when it confirms the intended content and a saved revision.
Link the returned document key in the completion comment or response:
`/<prefix>/issues/<issue-identifier>#document-<returned-key>`.
Use the returned key because a locked document can redirect an agent's write
to a new document.
## Update or resolve an unclear write
For an existing document, first `GET /api/issues/{issueId}/documents/{key}`
and read its body and `latestRevisionId`. Send the updated body with
`baseRevisionId` set to that revision. On `409`, fetch the latest document and
reconcile changes before retrying; preserve the skill's bounded-write retry rule.
If a write's status or receipt is unclear, GET the document to check the saved
content and revision before retrying. If delivery cannot be confirmed, report
that limitation instead of claiming the requested document was saved.
+6
View File
@@ -227,3 +227,9 @@ publication failures remain retained; local reconstructed copies move folders
without editing results, graders or usage. Never dispatch a second development
campaign for an active target branch: workflow concurrency supersedes the older
run. Use a separate frozen-source branch when independent campaigns must overlap.
## Focused legacy delivery repair
The original 24 cells and original assigned-skill request/procedure remain unchanged. Two added explicit Paperclip-document cells apply only to classic Claude and OpenCode, making 26 catalog cells (50 expected turns if every cell is selected). The repair campaign selects only the original skill case and new document case for those two profiles: four cells per variant, eight expected turns total. No full matrix rerun is planned.
Both skill sources are recorded in definition/admission digests. The pre-fix baseline restores the old SKILL.md and records the new reference as absent, without copying the new recipe into that baseline. It holds the tiny manual/shared prompts and all fixture/model/auth/effort inputs fixed. The explicit document oracle checks actual public content/revision and an exact same-app document link, with plausible-negative calibrations.
+3 -3
View File
@@ -136,10 +136,10 @@ describe("runner E2E catalog", () => {
expect(localIntegrityTasks).toHaveLength(2);
expect(openRouterBreadthTasks).toHaveLength(3);
expect(runnerSuites.map((suite) => suite.expectedMatrixSize)).toEqual([
6, 30, 3, 16, 16, 2, 6, 8, 46, 23, 50, 20, 24, 52, 28, 18, 6, 6, 12, 10, 48, 16, 10, 2, 1, 1,
6, 30, 3, 16, 16, 2, 6, 8, 46, 23, 50, 20, 26, 52, 28, 18, 6, 6, 12, 10, 48, 16, 10, 2, 1, 1,
]);
expect(validateRunnerCatalog()).toHaveLength(460);
expect(new Set(runnerMatrix.map((entry) => entry.id)).size).toBe(460);
expect(validateRunnerCatalog()).toHaveLength(462);
expect(new Set(runnerMatrix.map((entry) => entry.id)).size).toBe(462);
expect(
runnerMatrix.filter((entry) => entry.suite.id === "core-compatibility"),
).toHaveLength(48);
+8 -5
View File
@@ -5,8 +5,8 @@ import { taskTitleTasks, taskTitleDefinitionDigest, TASK_TITLE_BUDGET_CENTS } fr
import { blockerTasks, blockerProfile } from "./blocker-cases.js";
import { accountingTasks } from "./accounting-cases.js";
import { continuationTasks } from "./continuation-cases.js";
import { contextIntegrityTasks } from "./context-integrity-cases.js";
import { productionDefaultHireProfile, stockHarnessSourceDigest } from "./stock-harness.js";
import { contextIntegrityTasks, paperclipDocumentTask } from "./context-integrity-cases.js";
import { productionDefaultHireProfile, stockHarnessSourceDigest, stockHarnessSkillSources } from "./stock-harness.js";
import { lifecycleLiveTasks, lifecycleLiveDefinitionDigest } from "./lifecycle-live-cases.js";
import { everydayTasks, productionStoryProfile } from "./everyday-cases.js";
@@ -1222,11 +1222,14 @@ export const runnerSuites: readonly RunnerSuiteFixture[] = [
profiles: [...contextIntegrityProfiles.filter(profile => !pendingContextIntegrityProfiles.includes(profile)),
runnerProfiles.find(profile => profile.id === "legacy-opencode")!].map(productionDefaultHireProfile),
environments: [localEnvironment],
tasks: [...contextIntegrityTasks, chatTasks.find(task => task.id === "continuity-restart")!],
expectedMatrixSize: 24,
tasks: [...contextIntegrityTasks, chatTasks.find(task => task.id === "continuity-restart")!, paperclipDocumentTask],
expectedMatrixSize: 26,
excludedExecutionIds: ["legacy-codex", "legacy-acp-codex", "legacy-acp-claude", "runner-codex", "runner-acpx-claude", "runner-opencode"]
.map(profile => `stock-harness.${profile}.local.${paperclipDocumentTask.id}`),
definitionMetadata: {
version: 1, instructions: "production-default-hire", scheduling: "explicit-only",
version: 2, instructions: "production-default-hire", scheduling: "explicit-only",
sourceDigest: stockHarnessSourceDigest(),
operationalSkillSources: stockHarnessSkillSources(),
grading: "public-default-bundle-and-delivered-prompts-plus-independent-lifecycle-oracles",
vendorBaseEvidence: "required deterministic Codex driver/runnerd/Rust gate; task success is not vendor-base proof",
paidCalls: "one skill turn, two ordered-comment turns, three chat turns per profile",
+17 -7
View File
@@ -5,10 +5,14 @@ export const CONTEXT_INTEGRITY_CASES = [
"assigned-skill-explicit-invocation",
] as const;
export type ContextIntegrityCase = (typeof CONTEXT_INTEGRITY_CASES)[number];
export const PAPERCLIP_DOCUMENT_CASE = "assigned-skill-paperclip-document" as const;
export type ContextIntegrityCase = (typeof CONTEXT_INTEGRITY_CASES)[number] | typeof PAPERCLIP_DOCUMENT_CASE;
export function isAssignedSkillContext(id: string): boolean {
return id === "assigned-skill-explicit-invocation" || id === PAPERCLIP_DOCUMENT_CASE;
}
export function contextIntegrityScenario(id: string, nonce: string) {
if (!CONTEXT_INTEGRITY_CASES.includes(id as ContextIntegrityCase))
if (!CONTEXT_INTEGRITY_CASES.includes(id as (typeof CONTEXT_INTEGRITY_CASES)[number]) && id !== PAPERCLIP_DOCUMENT_CASE)
throw new Error(`Unknown context-integrity case: ${id}`);
const marker = `CONTEXT_OK_${nonce.replace(/[^a-z0-9]/gi, "")}`;
const repeated = "Append an entry for every request, even identical wording.";
@@ -23,15 +27,17 @@ export function contextIntegrityScenario(id: string, nonce: string) {
prompt:
id === "ordered-comment-continuation"
? `Maintain a packing list report document. Start with the initial scope: passport and charger. Save the initial report and leave this task waiting for follow-up comments. When follow-up comments arrive, copy each user's exact wording verbatim into one ordered ledger, including identical wording repeated more than once. After the follow-up batch, save one report containing both initial items and every verbatim request in arrival order. Add a separate section headed "## Final scope" outside the quoted ledger. Put only the final requested scope name as plain text in that section. Then finish. Do not invent extra comments, child tasks, or deliverables.`
: "Use the assigned Context integrity output skill for this task, then follow its instructions and finish the task. Do not create child tasks or unrelated deliverables.",
: id === PAPERCLIP_DOCUMENT_CASE
? "Use the assigned Context integrity output skill. Save its requested document as a Paperclip document on this task, check the saved revision and content, and post a completion comment linking the saved document. A workspace file alone does not deliver this requested Paperclip document. Finish after delivery. Do not create child tasks or unrelated deliverables."
: "Use the assigned Context integrity output skill for this task, then follow its instructions and finish the task. Do not create child tasks or unrelated deliverables.",
comments: [repeated, repeated, changed] as const,
};
}
export const contextIntegrityTasks: readonly RunnerTaskFixture[] =
CONTEXT_INTEGRITY_CASES.map((id) => ({
function contextIntegrityTask(id: ContextIntegrityCase): RunnerTaskFixture {
return {
id,
label: id === "ordered-comment-continuation" ? "Ordered comment continuation" : "Assigned skill invocation",
label: id === "ordered-comment-continuation" ? "Ordered comment continuation" : id === PAPERCLIP_DOCUMENT_CASE ? "Explicit Paperclip document delivery" : "Assigned skill invocation",
groups: ["context-integrity"],
workMode: "standard",
flow: "context_integrity",
@@ -42,4 +48,8 @@ export const contextIntegrityTasks: readonly RunnerTaskFixture[] =
buildPrompt: (nonce) => contextIntegrityScenario(id, nonce).prompt,
buildVisibleMarker: (nonce) => contextIntegrityScenario(id, nonce).marker,
buildMatchers: () => [],
}));
};
}
export const contextIntegrityTasks: readonly RunnerTaskFixture[] = CONTEXT_INTEGRITY_CASES.map(contextIntegrityTask);
export const paperclipDocumentTask = contextIntegrityTask(PAPERCLIP_DOCUMENT_CASE);
+8 -8
View File
@@ -1,6 +1,6 @@
import type { Page } from "@playwright/test";
import { randomUUID } from "node:crypto";
import { contextIntegrityScenario } from "./context-integrity-cases.js";
import { contextIntegrityScenario, isAssignedSkillContext } from "./context-integrity-cases.js";
import { gradeContextIntegrity, type ContextIntegrityCheckpoint } from "./context-integrity-scoring.js";
import { pollUntil, type RunnerApi } from "./api.js";
import { createTaskThroughUi } from "./user-actions.js";
@@ -102,7 +102,7 @@ export async function runContextIntegrityFlow(input: {
let assignedVersionId: string | null = null;
const checkpoints: ContextIntegrityCheckpoint[] = [];
if (scenario.id === "assigned-skill-explicit-invocation") {
if (isAssignedSkillContext(scenario.id)) {
await api.patch("/api/instance/settings/experimental", { enableBetaSkills: true });
const markdown = `---\nname: ${scenario.skillKey}\ndescription: Context integrity output procedure.\n---\n\n# Context integrity output procedure\n\nWrite exactly one task document whose body contains the marker ${scenario.marker}. Finish the task after saving that document.`;
if (!markdown.startsWith(`---\nname: ${scenario.skillKey}\ndescription:`) || markdown.includes("\\n")) {
@@ -141,18 +141,18 @@ export async function runContextIntegrityFlow(input: {
api.get<Row[]>(`/api/issues/${issue!.id}/documents`),
api.get<Row[]>(`${companyPath}/skills`),
api.get<Row>(`/api/issues/${issue!.id}/queued-comments`),
scenario.id === "assigned-skill-explicit-invocation" ? api.get<Row>(`/api/agents/${fixtures.agent.id}/skills?companyId=${fixtures.company.id}`) : Promise.resolve({} as Row),
isAssignedSkillContext(scenario.id) ? api.get<Row>(`/api/agents/${fixtures.agent.id}/skills?companyId=${fixtures.company.id}`) : Promise.resolve({} as Row),
]);
const detailedDocuments = await Promise.all(documents.map((document) => api.get<Row>(`/api/issues/${issue!.id}/documents/${encodeURIComponent(String(document.key))}`)));
const assignedSkill = scenario.id === "assigned-skill-explicit-invocation"
const assignedSkill = isAssignedSkillContext(scenario.id)
? skillRows.find((skill) => skill.key === scenario.assignedSkill?.key || skill.slug === scenario.assignedSkill?.key)
: undefined;
const desiredSkill = (assignedState.desiredSkillEntries as Array<Row> | undefined)?.find((entry) => entry.key === scenario.assignedSkill?.key);
const runEvents = await Promise.all(runs.map((run) => api.get<Row[]>(`/api/heartbeat-runs/${run.id}/events?limit=1000`)));
const runLogs = await Promise.all(runs.map((run) => readContextIntegrityRunLog(api, run.id)));
const skillInvocationEvidence = scenario.id === "assigned-skill-explicit-invocation" && runEvents.some((events) => events.some((event) => containsExplicitSkillInput(event, String(assignedSkill?.slug ?? scenario.skillKey))));
skillRequestText = scenario.id === "assigned-skill-explicit-invocation" ? String(issue?.description ?? "") : "";
checkpoints.push({ phase, issue: { id: issue!.id, status: String(issue!.status) }, comments, queuedComments, documents: detailedDocuments as Array<{ key: string; body?: string | null }>, runs, runEvents: runEvents.flat(), runLogs, assignedSkill: assignedSkill ? { key: String(desiredSkill?.key ?? assignedSkill.key ?? assignedSkill.slug), runtimeName: String(assignedSkill.slug ?? ""), versionId: desiredSkill?.versionId === assignedVersionId ? assignedVersionId : null, markdown: String(assignedSkill.markdown ?? scenario.assignedSkill?.markdown ?? "") } : undefined, skillRequestText: scenario.id === "assigned-skill-explicit-invocation" ? skillRequestText : undefined, skillInvocationEvidence });
const skillInvocationEvidence = isAssignedSkillContext(scenario.id) && runEvents.some((events) => events.some((event) => containsExplicitSkillInput(event, String(assignedSkill?.slug ?? scenario.skillKey))));
skillRequestText = isAssignedSkillContext(scenario.id) ? String(issue?.description ?? "") : "";
checkpoints.push({ phase, issue: { id: issue!.id, status: String(issue!.status), identifier: String(issue!.identifier ?? ""), issuePrefix: String(fixtures.company.issuePrefix ?? ""), appOrigin: new URL(page.url()).origin }, comments, queuedComments, documents: detailedDocuments as Array<{ key: string; body?: string | null }>, runs, runEvents: runEvents.flat(), runLogs, assignedSkill: assignedSkill ? { key: String(desiredSkill?.key ?? assignedSkill.key ?? assignedSkill.slug), runtimeName: String(assignedSkill.slug ?? ""), versionId: desiredSkill?.versionId === assignedVersionId ? assignedVersionId : null, markdown: String(assignedSkill.markdown ?? scenario.assignedSkill?.markdown ?? "") } : undefined, skillRequestText: isAssignedSkillContext(scenario.id) ? skillRequestText : undefined, skillInvocationEvidence });
const checks = gradeContextIntegrity({ id: scenario.id, marker: scenario.marker, comments: scenario.comments, checkpoints });
input.observe(issue!, runs, checks);
await input.evidence("context-integrity.json", { schema: "paperclip.context-integrity.v1", scenario, budgetGuard, checkpoints, checks });
@@ -175,7 +175,7 @@ export async function runContextIntegrityFlow(input: {
await api.patch(`/api/agents/${fixtures.agent.id}/budgets`, {
budgetMonthlyCents: budgetGuard.agentMonthlyCents,
});
const taskPrompt = scenario.id === "assigned-skill-explicit-invocation"
const taskPrompt = isAssignedSkillContext(scenario.id)
? `${scenario.prompt}\n\nUse /${scenario.skillKey} for this request.`
: scenario.prompt;
skillRequestText = taskPrompt;
+25 -6
View File
@@ -1,11 +1,11 @@
import type { ContextIntegrityCase } from "./context-integrity-cases.js";
import { isAssignedSkillContext, PAPERCLIP_DOCUMENT_CASE, type ContextIntegrityCase } from "./context-integrity-cases.js";
export interface ContextIntegrityCheckpoint {
phase: "initial" | "comment-1" | "comment-2" | "comment-3" | "final";
issue: { id: string; status: string };
issue: { id: string; status: string; identifier?: string; issuePrefix?: string; appOrigin?: string };
comments: Array<Record<string, unknown>>;
queuedComments?: Record<string, unknown>;
documents: Array<{ key: string; body?: string | null }>;
documents: Array<{ key: string; body?: string | null; latestRevisionId?: string | null; latestRevisionNumber?: number | null }>;
runs: Array<Record<string, unknown>>;
assignedSkill?: { key: string; runtimeName?: string; versionId?: string | null; markdown?: string };
skillRequestText?: string;
@@ -117,7 +117,7 @@ export function gradeContextIntegrity(input: {
const expectedCommentIds = humanRows.slice(0, input.comments.length).map((comment) => String(comment.id ?? ""));
check("continuation-wake-comment-ids", Boolean(continuation) && JSON.stringify(wakeCommentIds(continuation)) === JSON.stringify(expectedCommentIds), "The deferred continuation wake must carry all public comment IDs in arrival order.");
}
if (input.id === "assigned-skill-explicit-invocation") {
if (isAssignedSkillContext(input.id)) {
const runtimeName = String(initial?.assignedSkill?.runtimeName ?? initial?.assignedSkill?.key ?? "");
const requestText = String(initial?.skillRequestText ?? "");
const references = requestText.split(/\s+/).map((word) => word.replace(/[.,]$/, ""));
@@ -126,8 +126,27 @@ export function gradeContextIntegrity(input: {
check("skill-source-marker", Boolean(initial?.assignedSkill?.markdown?.includes(input.marker)), "The assigned pinned skill source must contain the output marker.");
check("marker-not-in-request", !requestText.includes(input.marker) && !finalComments.some((comment) => comment.includes(input.marker)), "The output marker must originate from the assigned skill, not the task request or comments.");
}
const output = final ? oneOutput(final, input.marker, input.id === "assigned-skill-explicit-invocation") : undefined;
check("single-durable-output", Boolean(output) && final!.documents.length === 1, input.id === "assigned-skill-explicit-invocation" ? "Exactly one durable task document must contain the skill's marker." : "Exactly one durable packing report document must be saved.");
const output = final ? oneOutput(final, input.marker, isAssignedSkillContext(input.id)) : undefined;
check("single-durable-output", Boolean(output) && final!.documents.length === 1, isAssignedSkillContext(input.id) ? "Exactly one durable task document must contain the skill's marker." : "Exactly one durable packing report document must be saved.");
if (input.id === PAPERCLIP_DOCUMENT_CASE) {
check("saved-document-revision", Boolean(output?.latestRevisionId) && Number.isInteger(output?.latestRevisionNumber) && Number(output?.latestRevisionNumber) > 0,
"The public saved document must have a persisted revision and the requested content.");
const route = output && final?.issue.identifier && final.issue.issuePrefix
? `/${final.issue.issuePrefix}/issues/${final.issue.identifier}#document-${output.key}` : undefined;
const agentComments = (final?.comments ?? []).filter(comment => comment.authorType === "agent" || Boolean(comment.authorAgentId));
const origin = final?.issue.appOrigin;
const usableLink = Boolean(route && origin && agentComments.some(comment => {
const links = String(comment.body ?? "").matchAll(/\[[^\]]*\]\((?:<([^>]+)>|([^\s)]+))(?:\s+["'][^"']*["'])?\)/g);
return [...links].some(link => {
try {
const url = new URL(link[1] ?? link[2]!, origin);
return url.origin === origin && `${url.pathname}${url.hash}` === route && !url.search;
} catch { return false; }
});
}));
check("saved-document-link", usableLink,
"An agent completion comment must link the exact saved document on this task; a local path or claimed URL does not suffice.");
}
check("completed-task", final?.issue.status === "done", `Final task status: ${final?.issue.status ?? "missing"}.`);
check("successful-runs", Boolean(final?.runs.length) && final!.runs.every((run) => run.status === "succeeded"), "All recorded context-integrity runs must succeed.");
return checks;
@@ -0,0 +1,52 @@
import { describe, expect, it } from "vitest";
import { contextIntegrityScenario, PAPERCLIP_DOCUMENT_CASE } from "./context-integrity-cases.js";
import { gradeContextIntegrity, type ContextIntegrityCheckpoint } from "./context-integrity-scoring.js";
function recording() {
const scenario = contextIntegrityScenario(PAPERCLIP_DOCUMENT_CASE, "probe");
const initial: ContextIntegrityCheckpoint = {
phase: "initial", issue: { id: "issue", identifier: "DOC-1", issuePrefix: "DOC", appOrigin: "https://paperclip.example", status: "in_progress" },
documents: [], comments: [], runs: [{ id: "run", status: "running" }],
assignedSkill: { key: scenario.skillKey, runtimeName: scenario.skillKey, versionId: "version", markdown: `Write ${scenario.marker}` },
skillRequestText: `${scenario.prompt}\nUse /${scenario.skillKey}`,
};
const final: ContextIntegrityCheckpoint = {
...initial, phase: "final", issue: { ...initial.issue, status: "done" },
documents: [{ key: "report", body: `Requested ${scenario.marker}`, latestRevisionId: "revision", latestRevisionNumber: 1 }],
comments: [{ authorAgentId: "agent", body: "Saved [report](/DOC/issues/DOC-1#document-report)." }],
runs: [{ id: "run", status: "succeeded" }],
};
return { scenario, initial, final };
}
describe("explicit Paperclip document delivery", () => {
it("checks the actual saved content, revision and exact document link", () => {
const { scenario, initial, final } = recording();
expect(gradeContextIntegrity({ ...scenario, checkpoints: [initial, final] }).every(row => row.passed)).toBe(true);
expect(scenario.prompt).toContain("Paperclip document");
expect(scenario.prompt).not.toContain(scenario.marker);
});
it.each(["https://paperclip.example/DOC/issues/DOC-1#document-report", "<https://paperclip.example/DOC/issues/DOC-1#document-report>"])("accepts the same-app document link %s", target => {
const { scenario, initial, final } = recording();
final.comments[0]!.body = `Saved [report](${target}).`;
expect(gradeContextIntegrity({ ...scenario, checkpoints: [initial, final] }).every(row => row.passed)).toBe(true);
});
it.each(["local-only", "missing-revision", "wrong-content", "local-link", "wrong-document-link", "other-app-link", "user-link-only"])(
"rejects plausible %s delivery", variant => {
const { scenario, initial, final } = recording();
if (variant === "local-only") { final.documents = []; final.comments[0]!.body = "Saved report.md in the workspace."; }
if (variant === "missing-revision") final.documents[0]!.latestRevisionId = null;
if (variant === "wrong-content") final.documents[0]!.body = "A different report";
if (variant === "local-link") final.comments[0]!.body = "Saved [report](./report.md).";
if (variant === "wrong-document-link") final.comments[0]!.body = "Saved [report](/DOC/issues/DOC-1#document-other).";
if (variant === "other-app-link") final.comments[0]!.body = "Saved [report](https://other.example/DOC/issues/DOC-1#document-report).";
if (variant === "user-link-only") { delete final.comments[0]!.authorAgentId; final.comments[0]!.authorUserId = "user"; }
expect(gradeContextIntegrity({ ...scenario, checkpoints: [initial, final] }).some(row => !row.passed)).toBe(true);
},
);
it("preserves the original ambiguous request and assigned output procedure", () => {
expect(contextIntegrityScenario("assigned-skill-explicit-invocation", "probe").prompt).toBe(
"Use the assigned Context integrity output skill for this task, then follow its instructions and finish the task. Do not create child tasks or unrelated deliverables.",
);
});
});
+9 -3
View File
@@ -40,7 +40,7 @@ export const stockHarnessGates = [
required: ["renders standard assignment wake with task authority", "renders scoped planning wake authority"] },
{ id: "SH-eval", name: "Independent oracle and qualification admission", cwd: ".",
config: "tests/runner-e2e/vitest.config.ts",
files: ["tests/runner-e2e/stock-harness.test.ts", "tests/runner-e2e/stock-harness-checks.test.mjs",
files: ["tests/runner-e2e/paperclip-document.test.ts", "tests/runner-e2e/stock-harness.test.ts", "tests/runner-e2e/stock-harness-checks.test.mjs",
"tests/runner-e2e/stock-harness-admission.test.ts", "tests/runner-e2e/stock-harness-digest.test.ts",
"tests/runner-e2e/select-rerun-artifacts.test.ts"],
required: ["rejects old SHA before providers", "allows toolchain paths and excludes every present or future credential",
@@ -70,6 +70,7 @@ export function sourceFingerprint() {
...stockHarnessGates.flatMap(gate => gate.files.map(file => join(gate.cwd, file))),
"tests/runner-e2e/stock-harness.ts", "tests/runner-e2e/stock-harness-checks.mjs", "tests/runner-e2e/catalog.ts",
"tests/runner-e2e/stock-harness-admission.ts", "tests/runner-e2e/launch.ts", "tests/runner-e2e/runner.spec.ts",
"tests/runner-e2e/context-integrity-cases.ts", "tests/runner-e2e/context-integrity-flow.ts", "tests/runner-e2e/context-integrity-scoring.ts",
"packages/adapter-utils/src/server-utils.ts", "packages/shared/src/connection-intent-guidance.ts",
"server/src/onboarding-assets/default/AGENTS.md", "server/src/routes/agents.ts", "scripts/ensure-plugin-build-deps.mjs",
"packages/paperclip-runner/src/drivers/codex/codex-app-server-driver-impl.ts",
@@ -82,10 +83,15 @@ export function sourceFingerprint() {
"packages/paperclip-runner/runner/crates/runner-core/tests/codex_provider.rs",
]);
const sourceErrors = [];
sources.add("skills/paperclip/SKILL.md");
sources.add("skills/paperclip/references/issue-documents.md");
for (const source of [...sources].sort()) {
hash.update(source);
try { hash.update(readFileSync(join(root, source))); }
catch { sourceErrors.push(source); }
try { hash.update("present\0").update(readFileSync(join(root, source))); }
catch (error) {
if (source === "skills/paperclip/references/issue-documents.md" && error.code === "ENOENT") hash.update("absent\0");
else sourceErrors.push(source);
}
}
return { fingerprint: hash.digest("hex"), sourceErrors };
}
+14 -2
View File
@@ -1,15 +1,27 @@
import { describe, expect, it, vi } from "vitest";
import { readFileSync } from "node:fs";
import { stockHarnessSourceDigest } from "./stock-harness.js";
import { stockHarnessSourceDigest, stockHarnessSkillSources } from "./stock-harness.js";
vi.mock("node:fs", async importOriginal => ({ ...await importOriginal<typeof import("node:fs")>(), readFileSync: vi.fn() }));
describe("stock harness instruction revision", () => {
it.each(["server/src/onboarding-assets/default/AGENTS.md", "packages/adapter-utils/src/server-utils.ts", "packages/shared/src/connection-intent-guidance.ts"])(
it.each(["server/src/onboarding-assets/default/AGENTS.md", "packages/adapter-utils/src/server-utils.ts", "packages/shared/src/connection-intent-guidance.ts", "skills/paperclip/SKILL.md", "skills/paperclip/references/issue-documents.md"])(
"changes when the evaluated %s changes", source => {
vi.mocked(readFileSync).mockImplementation(() => Buffer.from("unchanged"));
const original = stockHarnessSourceDigest();
vi.mocked(readFileSync).mockImplementation(file => Buffer.from(String(file).endsWith(source) ? "changed instructions" : "unchanged"));
expect(stockHarnessSourceDigest()).not.toBe(original);
});
it("records an absent historical recipe without introducing its content or hiding other read errors", () => {
vi.mocked(readFileSync).mockImplementation(() => Buffer.from("unchanged"));
const present = stockHarnessSourceDigest();
vi.mocked(readFileSync).mockImplementation(file => {
if (String(file).endsWith("references/issue-documents.md")) throw Object.assign(new Error("absent"), { code: "ENOENT" });
return Buffer.from("unchanged");
});
expect(stockHarnessSkillSources()[1]).toEqual({ path: "skills/paperclip/references/issue-documents.md", present: false, sha256: null });
expect(stockHarnessSourceDigest()).not.toBe(present);
vi.mocked(readFileSync).mockImplementation(() => { throw Object.assign(new Error("unreadable"), { code: "EACCES" }); });
expect(stockHarnessSourceDigest).toThrow("unreadable");
});
});
+5 -3
View File
@@ -16,18 +16,20 @@ function recording(generation: "legacy" | "native" = "legacy"): StockHarnessEvid
}
describe("stock harness Product E2E", () => {
it("registers 24 explicit local cells with independent skill, continuation and chat oracles", () => {
it("registers 26 explicit local cells with the original cases and two focused delivery repair cells", () => {
const cells = runnerMatrix.filter(row => row.suite.id === "stock-harness");
expect(cells).toHaveLength(24);
expect(cells).toHaveLength(26);
expect(new Set(cells.map(row => row.profile.id))).toEqual(new Set([
"legacy-codex", "legacy-claude", "legacy-opencode", "legacy-acp-codex", "legacy-acp-claude",
"runner-codex", "runner-acpx-claude", "runner-opencode",
]));
expect(new Set(cells.map(row => row.task.id))).toEqual(new Set([
"ordered-comment-continuation", "assigned-skill-explicit-invocation", "continuity-restart",
"assigned-skill-paperclip-document",
]));
expect(cells.every(row => row.suite.manualOnly && row.environment.id === "local")).toBe(true);
expect(cells.reduce((turns, row) => turns + row.task.expectedRunCount, 0)).toBe(48);
expect(cells.reduce((turns, row) => turns + row.task.expectedRunCount, 0)).toBe(50);
expect(cells.filter(row => row.task.id === "assigned-skill-paperclip-document").map(row => row.profile.id).sort()).toEqual(["legacy-claude", "legacy-opencode"]);
expect(cells[0]!.suite.definitionMetadata?.sourceDigest).toMatch(/^[a-f0-9]{64}$/);
});
+14
View File
@@ -6,6 +6,19 @@ import type { RunnerProfileFixture } from "./types.js";
// constants. Otherwise a larger shipped manual could silently update the oracle.
export const STOCK_HIRE_IDENTITY = "You are an agent in a Paperclip company.\n";
export const STOCK_TEMPLATE_IDENTITY = "You are agent";
const SKILL_SOURCES = ["../../skills/paperclip/SKILL.md", "../../skills/paperclip/references/issue-documents.md"];
export function stockHarnessSkillSources() {
return SKILL_SOURCES.map(source => {
try {
return { path: source.replace(/^\.\.\/\.\.\//, ""), present: true,
sha256: createHash("sha256").update(readFileSync(new URL(source, import.meta.url))).digest("hex") };
} catch (error) {
if (source.endsWith("/references/issue-documents.md") && (error as NodeJS.ErrnoException).code === "ENOENT")
return { path: source.replace(/^\.\.\/\.\.\//, ""), present: false, sha256: null };
throw error;
}
});
}
const REMOVED_PROCEDURES = [
"Execution contract:",
"Start actionable work in this heartbeat",
@@ -115,5 +128,6 @@ export function stockHarnessSourceDigest() {
"../../packages/adapter-utils/src/server-utils.ts",
"../../packages/shared/src/connection-intent-guidance.ts",
]) hash.update(source).update(readFileSync(new URL(source, import.meta.url)));
hash.update(JSON.stringify(stockHarnessSkillSources()));
return hash.digest("hex");
}