mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-06 21:05:21 +02:00
## Thinking Path > - Paperclip manages AI agents and governs their work. > - Its native runner uses structured provider protocols for sessions and tools. > - Grok Build supports ACP over stdio, but the runner did not expose it. > - Native execution requires company-scoped credentials, verified identities, and permission gates. > - This change adds Grok through ACPX for local and Daytona execution. > - Subscription login and explicit API-key execution have separate credential paths. > - Qualification grades real tool outcomes, durable state, and browser workflows. ## Linked Issues or Issue Description Refs #13845, #13847, #13850, #13878, #13901, #13973, #13977, #13979. Add **Grok Build** to `paperclip_runner` with `provider: "acpx"`, `acpxAgent: "grok"`, and model `grok-4.7`. Existing legacy Grok agents keep their adapter. Merge the three companion fixes (#13973, #13977, #13979) before treating the integrated Product qualification as deployed behavior. ## What Changed - Synchronize shared, TypeScript, Rust, server, validation, and UI provider contracts. - Run Grok native ACP stdio through ACPX and the authenticated Paperclip MCP bridge. Verify the pinned executable and exact ACP model identity. - Prefer company subscription login. Support an explicit company-secret API key without automatic paid fallback. Fence refresh and copyback to the same account and remove private runtime credentials after containment. - Preserve selected permissions, cancellation, durable session identity, resume, and restart recovery. Keep unsupported steering and goals unavailable. Preserve missing usage and cost as unknown. - Package checksum-verified Grok Build 1.0.13 for Daytona with an immutable, signed image built on EC2. - Add deterministic admission, protocol, permissions, identity, credential, failure, and cleanup checks. Add the maintained 39-case protocol roster and separate subscription/API Product profiles. - Fix live-test findings in reasoning events, reloads, idle-owner retirement, credential-home cleanup, expired-login model discovery, launcher pinning, and rerun evidence selection. - Align control-plane state readers with the transport's 64 MiB bound while retaining identity, ownership, lifecycle, and size rejection checks. - Stabilize two asynchronous CI assertions while retaining actual outcome and filesystem-evidence checks. ## Verification Current integration head `f114948376056fe0b6b34c1496ae8667b59daa63` includes master `3447609d2247e75e55d91493dda91a608364f672` (2026-09-28). Two master advances during verification overlapped the eval catalog; the final merge preserves Grok qualification, completion updates, and bounded API-response reading in all 348 cells. All 77 focused catalog/eval/workflow tests pass. Both native stack layers (#14397) are mergeable, and both exact-head Greptile reviews are 5/5 with successful security scans and no unresolved review threads. All current-head CI is green: 56 successful checks/statuses and four intentional skips ([CI run](https://github.com/paperclipai/paperclip/actions/runs/36447097232)). Trunk code-owner requirements remain enforced. The review summary’s non-blocking saved-asset offset classification note concerns code already merged in #14301; those runtime files are identical to master and outside this stack’s diff. Historical live evidence below retains its original source revisions. Earlier integration checkpoint: `24fc9b94ca0afb21ccdc8d26dbb2e4b258ad72cb`. Refreshed against master `0f14d2612`, preserving Grok qualification alongside the new accounting and lifecycle suites. All 124 focused catalog, evidence, and service-worker checks pass. The current base workflow includes the explicitly selected public-install verification lane; follow-up #14024 supplies its verifier script. CI at that earlier checkpoint was green (56 successful checks/statuses, four intentional skips), and the review is 5/5 with no unresolved findings. Prior feature CI at `fd73f0a9b1ecdf4094685054028df71739ddc3e1` passed ([run 36148259902](https://github.com/paperclipai/paperclip/actions/runs/36148259902)); that is historical evidence, not a current-head result. Paid Product measurements use frozen integrated source `2d939a92b21dcaf5c77c88b54d96784d2ddd0699`, which combines the feature with #13973, #13977, and #13979. That source passed all 52 CI checks and clean 5/5 review. Later master syncs incorporate upstream changes. Their checks remain separate from these pinned live measurements. | Check | Result and source-pinned report | | --- | --- | | Subscription protocol roster | [39/39 first attempts; 206 assertions](https://d1p6rlowie26tp.cloudfront.net/runner-protocol-evals/campaigns/gha-36046839612-1/index.html), runtime `bc6833f7`, evals `92bb4b8c` | | API protocol roster | [39/39 first attempts; 206 assertions](https://d1p6rlowie26tp.cloudfront.net/runner-protocol-evals/campaigns/gha-35926577007-1/index.html), runtime `4a1061c8`, evals `3213dbec` | | Subscription full Product matrix | [16/16 first attempts; 144 assertions; cleanup passed](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36096908572-1/index.html), source `2d939a92` | | Subscription core repetitions | 18/18: tool use, planning approval, and Stop/resume each passed three times in local and Daytona profiles. The full matrix contains repetition one; [repeat two](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36104551060-1/index.html) and [repeat three](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36108100404-1/index.html) each passed 6/6. Total: 28 unique subscription attempts at `2d939a92`. | | API smoke and question continuation | [4/4 first attempts; cleanup passed](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36147315401-1/index.html), both environments at `2d939a92` | | Historical API Product coverage | [16/16 full matrix](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35875144860-1/index.html) and 18/18 core repetitions at `4a1061c8`; retained as measurements of that revision | | Native Daytona proof | Three subscription and three API MCP/permissions/resume runs passed at `bc6833f7`. Three expired-login admission and fenced refresh checks passed without inference. All test sandboxes were removed. | | Inspectable artifacts and UI | Current-source screenshots verify planning approval, direct Ask completion, question continuation after controller restart, and two downloadable project revisions. The project downloads pass 12 and 18 tests; all 40 independent artifact oracle checks pass. | | Provider-free checks | 116 eval-validator tests, 39 Grok definitions, and 359 enabled/external campaign cells pass. Continuation regressions above 2 MiB and 16 MiB failed before their fixes; 32 focused recovery/ownership/size checks pass. | The 32 unique current-source Product attempts have no failures, retries, or skipped cells, and all cleanup checks pass. Whole-workflow timing, model identity, image and provider-pack provenance, attempts, and accounting coverage are retained in the canonical reports. The report publisher's conservative `complete=false` flag is preserved; independent audits verify the exact selected source catalog and immutable result rows. Pins: Grok Build `1.0.13 (5e9a58528b76)`, ACPX `0.13.1`, ACP model `grok-4.7`. Linux binary SHA-256: `edf79521581bb5e6b95abef848491a6a742e860da3e237ebe86a280d30dce4c1`. Launcher SHA-256: `f0b698395a3704ed2ffaf84ea19bdb20c36c8a0a70b7c629c7b6ffe144e59e55`. Image: `ghcr.io/paperclipai/paperclip-daytona-runner@sha256:76b24edfd850219e949418b19e4ceba690e84d51d199ade426e484953329b5e9`. Image build source is `4196a4cd`, recorded separately from application source `2d939a92`; each campaign verifies the image signature and provider pack. Original failed campaigns remain available: [continuation bound](https://github.com/paperclipai/paperclip/actions/runs/36057718059), [scheduler/event capture](https://github.com/paperclipai/paperclip/actions/runs/36071063537), and [startup cleanup plus EC2 interruption](https://github.com/paperclipai/paperclip/actions/runs/36080870743). They retain their original grades. No Docker or Rust builds ran on the developer laptop for these follow-ups. ## Risks Merge packaging follow-up #14024 with this base before public release. The follow-up replaces the private Grok bridge package with a built-in launcher and makes the native binary an explicit sandbox prerequisite. Three separate, reviewed fixes are part of the tested integrated behavior: #13973 serializes task-run admission; #13977 captures complete event evidence; #13979 durably reconciles failed Daytona creation. Each has green CI and clean 5/5 review. Failed-create recovery has 277 plugin tests, 92 SDK tests, host-runtime recovery tests, and a real Daytona lost-deletion-receipt proof. The live proof uses a private file for journal persistence; database durability is covered by host tests. Worker death before delivery of a failure envelope remains outside that recovery mechanism. Subscription fixtures stage an authorized company login; interactive browser sign-in is not qualified. Local Product profiles ran on EC2 Linux. The temporary subscription credential was removed from the protected GitHub environment after all subscription audits, with absence verified. Runtime homes and refresh copyback remain ownership-fenced. Protocol results remain pinned to their original revisions; they are not relabeled as tests of the latest feature commit. New binary/model versions require qualification. Missing token usage and model cost remain unknown; runtime estimates do not establish a full bill. Automatic paid Grok scheduling remains disabled pending separate reviewed enablement. The 64 MiB bound can increase memory use for verbose sessions, and larger files still fail closed. No automatic legacy-agent migration occurs. ## Model Used OpenAI GPT-6 through Codex, with tool use and code execution. The exact serving model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
114 lines
5.1 KiB
JavaScript
114 lines
5.1 KiB
JavaScript
// The build id is stamped into this file at production build time (see
|
|
// stampServiceWorkerBuildId in vite.config.ts), so a deploy that changes only
|
|
// the app bundle still changes sw.js byte-for-byte. That is what makes the
|
|
// browser install a new worker, which — via skipWaiting + controllerchange —
|
|
// reloads parked tabs onto the fresh bundle. Left as the literal placeholder in
|
|
// dev, where HMR (not the worker) drives refreshes.
|
|
const BUILD_ID = "__PAPERCLIP_BUILD_ID__";
|
|
// Separate this allowlisted cache from older workers that cached arbitrary URLs.
|
|
const CACHE_NAME = `paperclip-public-assets-${BUILD_ID}`;
|
|
const privateRequests = new Set();
|
|
const privateCacheControl = /(?:^|,)\s*(?:no-store|private)(?:\s*(?:,|=)|\s*$)/i;
|
|
|
|
// Static recovery only: never cache or embed authenticated page content here.
|
|
function offlineNavigationResponse() {
|
|
return new Response(`<!doctype html>
|
|
<html lang="en">
|
|
<head><meta charset="utf-8"><meta name="viewport" content="width=device-width, initial-scale=1">
|
|
<meta name="color-scheme" content="light dark"><title>Paperclip is offline</title></head>
|
|
<body><main><h1>Paperclip is offline</h1>
|
|
<p>Check your connection, then reload this page to try again.</p>
|
|
<button type="button" onclick="window.location.reload()">Reload page</button>
|
|
</main></body></html>`, {
|
|
status: 503,
|
|
headers: { "Content-Type": "text/html; charset=utf-8", "Cache-Control": "no-store" },
|
|
});
|
|
}
|
|
|
|
async function evictRequest(request) {
|
|
await Promise.all((await caches.keys()).map(async (key) => {
|
|
const cache = await caches.open(key);
|
|
await cache.delete(request, { ignoreVary: true });
|
|
}));
|
|
}
|
|
|
|
self.addEventListener("install", () => {
|
|
self.skipWaiting();
|
|
});
|
|
|
|
self.addEventListener("activate", (event) => {
|
|
event.waitUntil(
|
|
caches.keys().then((keys) =>
|
|
Promise.all(keys.map((key) => caches.delete(key)))
|
|
)
|
|
);
|
|
self.clients.claim();
|
|
});
|
|
|
|
self.addEventListener("fetch", (event) => {
|
|
// Vite owns development module revalidation and HMR. Passing that graph
|
|
// through an offline worker can forward bodyless 304 responses on reload.
|
|
// Only a stamped production build has an offline-cache contract.
|
|
if (BUILD_ID.startsWith("__")) return;
|
|
const { request } = event;
|
|
const url = new URL(request.url);
|
|
// Only immutable Vite build assets have a public offline-cache contract.
|
|
// Never infer that application/extension responses are public from absent
|
|
// headers, or from an in-memory classification lost when this worker restarts.
|
|
const publicAsset = url.origin === self.location.origin && !url.search &&
|
|
/^\/assets\/[^/]+-[a-zA-Z0-9_-]{8,}\.[a-zA-Z0-9.]+$/.test(url.pathname);
|
|
|
|
// Explicitly private requests must bypass BOTH cache writes and offline
|
|
// fallback, including extension endpoints outside the host /api namespace.
|
|
if (request.method !== "GET" || url.pathname.startsWith("/api")) {
|
|
return;
|
|
}
|
|
if (request.cache === "no-store") {
|
|
privateRequests.add(request.url);
|
|
event.waitUntil(evictRequest(request).catch(() => {}));
|
|
return;
|
|
}
|
|
|
|
// Vite development modules use the browser's conditional-response cache.
|
|
// Passing their requests through fetch/respondWith can return a bodyless 304
|
|
// to the module loader on reload and leave the app root empty. They are not
|
|
// build assets and have no offline contract, so leave them to the browser.
|
|
if (url.origin === self.location.origin &&
|
|
/^\/(?:@fs|@vite|@id|src|node_modules)\//.test(url.pathname)) {
|
|
return;
|
|
}
|
|
|
|
// Network-first; only public build assets can use an offline fallback.
|
|
event.respondWith(
|
|
fetch(request)
|
|
.then(async (response) => {
|
|
const cacheControl = response.headers.get("cache-control") ?? "";
|
|
if (privateCacheControl.test(cacheControl)) {
|
|
// Revoke earlier cacheable responses too. Keep an in-memory denylist
|
|
// if storage is unavailable so offline fallback still fails closed.
|
|
privateRequests.add(request.url);
|
|
await evictRequest(request).catch(() => {});
|
|
} else if (response.ok && publicAsset && !privateRequests.has(request.url)) {
|
|
const clone = response.clone();
|
|
await caches.open(CACHE_NAME).then(async (cache) => {
|
|
await cache.put(request, clone);
|
|
// A concurrent response may have revoked this URL during put().
|
|
if (privateRequests.has(request.url)) await cache.delete(request, { ignoreVary: true });
|
|
}).catch(() => {});
|
|
}
|
|
return response;
|
|
})
|
|
.catch(async () => {
|
|
if (privateRequests.has(request.url)) return Response.error();
|
|
if (!publicAsset) return request.mode === "navigate" ? offlineNavigationResponse() : Response.error();
|
|
// Restrict lookup to this policy's cache; old arbitrary-response caches
|
|
// must not become fallback candidates if activation cleanup fails.
|
|
try {
|
|
const cached = await (await caches.open(CACHE_NAME)).match(request);
|
|
if (cached && !privateCacheControl.test(cached.headers.get("cache-control") ?? "")) return cached;
|
|
} catch { /* Unavailable cache storage is an offline miss. */ }
|
|
return Response.error();
|
|
})
|
|
);
|
|
});
|