Files
PaperClipAI/tests/runner-e2e/redaction.ts
T
Dotta af3023f1e3 fix(runner): repair paid provider startup paths (#12769)
## Thinking Path

> - Paperclip manages AI agents that perform work.
> - Paperclip Runner connects durable task runs to local provider
processes.
> - The full-stack paid matrix exposed failures after the runner
integrity repair.
> - Verified JavaScript entrypoints lost their relative module graph
when Linux executed them through descriptor paths.
> - Returned provider startup errors also remained pending and became
indeterminate after recovery.
> - Sparse Codex tool lifecycle events lost the `write_document`
identity before task transcript projection.
> - This pull request repairs those three boundaries and makes the
structured-question fixture deterministic.
> - The benefit is repeatable provider startup, exact failure replay,
and correct inline Plan placement.

## Linked Issues or Issue Description

Refs #12721 and #12700.

**What happened?**

The paid runner matrix failed ACPX and OpenCode startup before provider
session creation. The runner journal then replaced the original startup
error with an indeterminate recovery result. Native Codex saved a Plan
but rendered it only as a fallback card. A legacy Claude waiting reply
could also echo the reserved terminal marker before the answer arrived.

**Expected behavior**

Verified JavaScript providers must start from immutable
descriptor-backed artifacts. Returned startup failures must persist as
terminal failed command results. Native tool lifecycle updates must
preserve the `write_document` boundary. Pre-answer fixture output must
not contain the reserved terminal marker.

**Steps to reproduce**

1. Run the local provider cells in the Runner Full-Stack E2E workflow.
2. Observe ACPX and OpenCode fail during `session.open` before provider
execution.
3. Observe recovery report `execution_indeterminate` instead of the
original startup error.
4. Run the native Codex Plan cell and observe the fallback Plan card
after the tool activity row.
5. Run the legacy Claude structured-question resume cell and observe an
early marker echo in waiting prose.

**Paperclip version or commit**

`0f9452101740835ce0b1488a204bf48acd5bafc3`

**Deployment mode**

Local development with the paid GitHub Actions acceptance workflow.

## What Changed

- Bundle the ACPX sidecar and OpenCode proxy as self-contained Node ESM
entrypoints before hashing and verified descriptor launch.
- Anchor ACPX dynamic provider package resolution at a
controller-derived provider-pack root and keep that root out of the
provider child environment.
- Persist executor-returned startup errors as redacted durable failed
command results while retaining indeterminate recovery for true process
death.
- Coalesce sparse native tool items by stable ID so a late
`write_document` name, input, and result reach the transcript boundary
once.
- Forbid the structured-question fixture from spelling or announcing its
reserved terminal marker before the user answers.

## Verification

- Rust and TypeScript regression tests cover durable failed replay, true
crash ambiguity, bundle closure, package-root derivation, environment
filtering, exact Codex tool lifecycle coalescing, and prompt
determinism.
- Local execution is intentionally limited to formatters and static diff
checks. GitHub Actions will run tests, type checks, builds, and security
checks.
- After ordinary CI is green, scoped paid cells will validate one ACPX
launch, one OpenCode launch, native Codex Plan projection, and legacy
Claude structured resume before a complete matrix rerun.
- Prior failing matrix:
https://github.com/paperclipai/paperclip/actions/runs/33682434315

## Risks

- Bundling changes the bytes covered by provider launch hashes.
Provider-pack generation already hashes the final built files.
- ACPX still loads qualified provider packages dynamically. The
controller supplies a normalized package root, while existing version,
digest, path, and descriptor checks remain active.
- Durable `failed` is terminal. Replays return the same redacted result
and do not execute the provider effect twice.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

OpenAI Codex based on GPT-5 with agentic reasoning, repository
inspection, code editing, Git, parallel subagents, and GitHub Actions
coordination. The exact deployed snapshot and context-window size are
not exposed to this task.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either linked related public work or described the bug in
this PR
- [x] I have not referenced internal or instance-local Paperclip issues
or links
- [x] My branch name describes the change and contains no internal
ticket id
- [ ] I have run tests locally and they pass (intentionally deferred to
GitHub Actions)
- [x] I have added or updated tests where applicable
- [x] No documentation change is required for this runtime repair
- [x] I have considered and documented the risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-04 07:58:44 -05:00

192 lines
5.9 KiB
TypeScript

import { createReadStream } from "node:fs";
import { readdir } from "node:fs/promises";
import path from "node:path";
import { redactDiagnosticText } from "../../packages/adapter-utils/src/command-redaction.js";
const SECRET_SHAPES = [
/\bsk-ant-[A-Za-z0-9_-]{16,}\b/g,
/\bsk-(?:proj-)?[A-Za-z0-9_-]{16,}\b/g,
/\b(?:openrouter|daytona)[-_]?(?:api)?[-_]?key["'=:\s]+[A-Za-z0-9._-]{12,}\b/gi,
] as const;
export function normalizedSecrets(values: readonly (string | undefined)[]) {
return [
...new Set(
values
.map((value) => value?.trim())
.filter((value): value is string => Boolean(value)),
),
].sort((left, right) => right.length - left.length);
}
export function isEphemeralCodexRuntimeAuthFile(
paperclipHome: string,
file: string,
) {
const relative = path.relative(paperclipHome, file).split(path.sep).join("/");
return (
/^instances\/[^/]+\/companies\/[^/]+\/agents\/[^/]+\/codex-home\/auth\.json$/.test(
relative,
) ||
/^instances\/[^/]+\/runtime\/paperclip-runner\/durable-sessions\/[^/]+\/codex-home\/auth\.json$/.test(
relative,
) ||
/^instances\/[^/]+\/runtime\/paperclip-runner\/acpx\/acpx\/[^/]+\/codex-home\/auth\.json$/.test(
relative,
)
);
}
export function redactText(value: string, secrets: readonly string[]) {
let redacted = redactDiagnosticText(value, "[REDACTED]");
return redactKnownSecretsAndShapes(redacted, secrets);
}
function redactKnownSecretsAndShapes(
value: string,
secrets: readonly string[],
) {
let redacted = value;
for (const secret of normalizedSecrets(secrets))
redacted = redacted.split(secret).join("[REDACTED]");
for (const pattern of SECRET_SHAPES)
redacted = redacted.replace(pattern, "[REDACTED_SECRET_SHAPE]");
return redacted;
}
function redactStructuredText(value: string, secrets: readonly string[]) {
const knownSafe = redactKnownSecretsAndShapes(value, secrets);
if (
!/(?:Authorization\s*:\s*Bearer|(?:api[-_]?key|token|secret|password)\s*=)/i.test(
knownSafe,
)
) {
return knownSafe;
}
return redactKnownSecretsAndShapes(
redactDiagnosticText(knownSafe, "[REDACTED]"),
secrets,
);
}
export function findSecretLeak(
value: string | Buffer,
secrets: readonly string[],
options: { includeShapes?: boolean } = {},
): string | null {
const text = Buffer.isBuffer(value) ? value.toString("utf8") : value;
for (const secret of normalizedSecrets(secrets)) {
if (text.includes(secret)) return "exact secret value";
}
if (options.includeShapes !== false) {
for (const pattern of SECRET_SHAPES) {
pattern.lastIndex = 0;
if (pattern.test(text)) return "secret-shaped value";
}
}
return null;
}
/**
* Scan structured JSON one key/value string at a time. Scanning the serialized
* document for provider key shapes can join an object key, punctuation, and an
* unrelated value into a false positive that never existed in the payload.
*/
export function findSecretLeakInJsonValues(
value: unknown,
secrets: readonly string[],
options: { includeShapes?: boolean } = {},
): string | null {
if (typeof value === "string") return findSecretLeak(value, secrets, options);
if (Array.isArray(value)) {
for (const entry of value) {
const leak = findSecretLeakInJsonValues(entry, secrets, options);
if (leak) return leak;
}
return null;
}
if (value && typeof value === "object") {
for (const [key, entry] of Object.entries(value)) {
const keyLeak = findSecretLeak(key, secrets, options);
if (keyLeak) return keyLeak;
const valueLeak = findSecretLeakInJsonValues(entry, secrets, options);
if (valueLeak) return valueLeak;
}
}
return null;
}
export function sanitizeJson(
value: unknown,
secrets: readonly string[],
): unknown {
// Structured values are not shell diagnostics. Applying the command/JWT
// heuristic here would redact stable dotted identifiers such as our schema
// name. Exact loaded secrets and well-known provider key shapes are enough.
if (typeof value === "string") return redactStructuredText(value, secrets);
if (Array.isArray(value))
return value.map((entry) => sanitizeJson(entry, secrets));
if (value && typeof value === "object") {
return Object.fromEntries(
Object.entries(value).map(([key, entry]) => [
key,
sanitizeJson(entry, secrets),
]),
);
}
return value;
}
export function assertSecretFree(
value: string | Buffer,
secrets: readonly string[],
label: string,
options: { includeShapes?: boolean } = {},
) {
const leak = findSecretLeak(value, secrets, options);
if (leak) throw new Error(`Secret leak in ${label}: ${leak}`);
}
export async function findSecretLeakInDirectory(
root: string,
secrets: readonly string[],
options: {
includeShapes?: boolean;
ignoreFile?: (file: string) => boolean;
} = {},
): Promise<{ file: string; reason: string } | null> {
const overlap = Math.max(
256,
...normalizedSecrets(secrets).map((secret) => secret.length + 16),
);
const scan = async (
directory: string,
): Promise<{
file: string;
reason: string;
} | null> => {
const entries = await readdir(directory, { withFileTypes: true });
for (const entry of entries) {
const file = path.join(directory, entry.name);
if (entry.isDirectory()) {
const leak = await scan(file);
if (leak) return leak;
} else if (entry.isFile()) {
if (options.ignoreFile?.(file)) continue;
let carry = Buffer.alloc(0);
for await (const chunk of createReadStream(file)) {
const data = Buffer.concat([
carry,
Buffer.isBuffer(chunk) ? chunk : Buffer.from(chunk),
]);
const reason = findSecretLeak(data, secrets, options);
if (reason) return { file, reason };
carry = data.subarray(Math.max(0, data.length - overlap));
}
}
}
return null;
};
return scan(root);
}