Serialize native workspace finalization, validate streamed archives within bounded limits, and quietly recover unsafe exports from saved results. Preserve exact allocations for exhausted transient failures and provide export-only retry without rerunning the provider. Consolidates #14314, #14315, #14329, and #14334 while preserving the already-merged finalization label changes. Co-Authored-By: Paperclip <noreply@paperclip.ing>
14 KiB
Native workspace finalization ownership and recovery
Native workspace export and merge acquire a PostgreSQL advisory lock scoped to company and run before the first physical copyback. The live heartbeat and the reconciler share that lock. Recovery skips a busy owner without recording another workspace operation or spending a retry. Recovery also rechecks the coordinator's terminal state and retry time under ownership: an earlier sweep snapshot cannot start another export after live finalization publishes a permanent repair or delay. A completed workspace barrier is reread under ownership before export, and a committed coordinator cannot be overwritten by a late failure receipt.
The lock transaction holds no row locks. Ordinary progress and finalization
receipts remain visible through the normal database pool. A dedicated connection outside the application pool is
reserved for the duration of copyback and closed when it settles; even a one-connection application pool remains available for progress writes. Its run profile carries
nativeWorkspaceFinalizationOwner, an exact token, host, PID, and process-start
receipt. Losing the lock connection does not prove physical copyback stopped:
a contender still refuses a receipt whose controller is alive. The original
callback joins before its token is released, and publication checks the lock
connection and token. Graceful completion clears the receipt. If that cleanup write fails after the
callback joins, only the same exact controller boot retains positive in-process
join evidence and may resume after reconnecting; an unknown token or a new boot
does not inherit that authority. A controller that
has exited on the same host can be recovered automatically only when a durable
successful workspace barrier proves its copyback finished. A dead parent can
leave tar/Git children alive, so incomplete copyback requires operator stop
verification even on the same host. PID reuse is checked against its recorded
start time rather than trusted by PID alone.
Unverified copyback after controller replacement
The controller cannot verify a process on a foreign or unknown host, or orphaned
copyback children after an abrupt parent death before the success barrier. It surfaces
native_workspace_finalization_owner_unverified as board-owned recovery, with no
automatic provider wake. This is an intentional limit: elapsed time or a missing
database connection never proves the old copyback process stopped.
An instance operator must first verify through the deployment platform that the exact prior controller and its copyback subprocesses have stopped. Retain the sandbox, accepted native result, and workspace descriptor. Do not run a new provider turn, delete workspace contents, or relax archive confinement.
After that platform verification, use a database maintenance transaction to release only the exact receipt shown by the recovery action. Replace the four placeholders with the action's company, run, token, and source issue. The advisory lock prevents concurrent acquisition during this change; the token comparison prevents clearing a newer owner. A zero-row update means ownership changed and requires fresh inspection. This maintenance operation is for a full-control instance operator, not an agent tool.
BEGIN;
SELECT pg_advisory_xact_lock(hashtextextended(
'native-workspace-finalization:<company-id>:<run-id>', 0));
UPDATE heartbeat_runs
SET runner_profile_json = runner_profile_json - 'nativeWorkspaceFinalizationOwner'
WHERE company_id = '<company-id>'::uuid
AND id = '<run-id>'::uuid
AND native_issue_id = '<issue-id>'::uuid
AND runtime_mode = 'native'
AND runner_profile_json->'nativeWorkspaceFinalizationOwner'->>'token' = '<owner-token>'
RETURNING id;
COMMIT;
Resolve the existing board recovery action with a note containing the platform
stop evidence. The ordinary reconciliation sweep then resumes workspace
finalization from the accepted result. Confirm the run's native phase and
resultJson.finalizationPhase are committed, there is no nextAttemptAt, and
no workspace operation is still running. The accepted provider result is reused.
Automatic unsafe archive recovery
An unsafe link in a native workspace must not fail a completed task or require an operator to repair a sandbox. The accepted agent result remains authoritative.
Daytona validates every exported archive before extraction. If validation rejects an archive, export once more using files, directories, and relative symlinks whose resolved targets stay inside the workspace. The fallback does not follow or delete symlinks. It omits unsafe or unresolvable links, stores hard-linked files as ordinary bytes, and preserves directory exclusions and empty directories. The second archive passes the same confinement checks. A fixed informational message records the fallback in the provider log.
If native copyback still rejects the archive or its source confinement check,
Paperclip discards that export, records workspace_export_omitted at info level
in the local run log, and completes finalization using the original accepted
result. It does not create a task warning, recovery card, repair request, or new
provider turn. The normal completion policy still enforces ownership and explicit
workflow constraints. Lost remote files are an accepted tradeoff.
The live path and restart finalizer use the same policy under workspace finalization ownership. Ownership loss, transport failures, missing sandboxes, and other unrelated errors retain their existing handling. No unsafe archive is extracted. No host path or link target is copied into the informational run event.
Repairing a failed workspace export without rerunning the agent
Transient export failures retry three times, then produce
native_workspace_sync_out_retry_exhausted. The accepted result stays saved,
and the exact sandbox is stopped and retained for export-only recovery. The
reason identifies a transport/copyback failure, not an inferred disk-space cause.
New unsafe archives use the automatic policy above and do not create this hold.
Historical unsafe-archive holds recover automatically as described below; they do not use this operator flow.
A board member with runtime management access can complete the saved result:
- Read the recovery action, run, and environment lease. Verify the company, run, provider sandbox ID, and stopped-provider receipt. Preserve that exact sandbox; do not acquire a replacement or seed the host workspace over it.
- Through the provider console or official SDK, resume that sandbox and inspect the export failure while preserving all user files. Restore transport access or other provider prerequisites before retrying. Extending the sandbox retention window can be necessary during this repair.
- In the task's Workspace export needs repair notice, describe the repair and
choose Retry workspace export. The equivalent board API is
POST /api/issues/:id/recovery-actions/retry-workspace-exportwith{ "actionId": "<recovery-action-id>", "runId": "<run-id>", "repairNote": "<repair and preservation evidence>" }. - Confirm the same run commits, the saved result determines the task outcome, and the recovery action resolves. No new provider turn or wake is created. A reusable repair sandbox is stopped and retained again. An ephemeral sandbox follows its original destroy-after-turn policy only after the exact result commits and the workspace reference contains a finalized host-copy receipt.
Admission requires the same accepted result, task owner, descriptor, and lease;
confirmed prior provider stop; an available repaired sandbox; and no newer task
execution or competing current lease. It shares finalization ownership and
resumes only the recorded provider lease through its verified lifecycle method.
That method drains old activity and verifies the saved workspace identity before
reopening the provider's controller admission gate; an external console restart
alone does not reopen that gate. Admission rejects a missing or replacement
sandbox and revalidates after probing the exact workspace. A changed binding or unavailable sandbox
returns 409 without reopening work. A duplicate queued request is idempotent.
Unsafe archives use automatic salvage or omission during retry. Generic
ordinary recovery resolution cannot retry, mark done, or send the task for review
in place of export-only retry. Explicit board false-positive/cancellation
dispositions remain deliberate overrides; they do not claim successful copyback.
Before publishing terminal export failure, the finalizer atomically records a stop-only intent with the failed run. This includes ephemeral allocations: ordinary release policy cannot delete their unexported files. A crash before cleanup leaves enough exact lease, result, and plugin authority for restart recovery. The same intent remains while export-only retry is active.
Before resuming the retained sandbox, the controller durably records a stop-only
cleanup intent on that exact lease and removes the old stopped receipt. If the
resume reply is lost, probing fails, or admission changes, it stops and retains
the sandbox. A failed stop remains pending_cleanup; a restarted controller's
bounded cleanup sweep retries the verified stop without destroying saved files.
An in-flight request holds a 15-minute cleanup claim, so a controller crash may
delay that sweep until the claim expires. Neither an unconfirmed stop nor a
stale receipt grants admission, and a changed lease or competing sandbox owner
prevents cleanup from taking ownership. Retry export after the lease has a new
confirmed stopped receipt. No provider turn is created by this recovery.
Cleanup and its readiness probe use the plugin ID recorded on that lease. Another
plugin with the same provider name cannot take over; an unavailable original
plugin defers cleanup without consuming an attempt. Stop-only cleanup requires the exact plugin worker to advertise environmentStopLease. Older plugins without that hook receive neither release nor destroy; resume admission also defers until safe compensation is available.
New intents use schema v2 with an explicit plugin pin. A v1 intent created before
that field existed remains recoverable using only the plugin ID already recorded
on its exact lease; an explicit mismatched pin is still rejected.
The opt-in native-workspace-export-resume.live.test.ts is a provider-boundary
fault integration, separate from the browser Product E2E. After building the
Daytona plugin, run that exact Vitest file with PAPERCLIP_LIVE_EXPORT_RESUME=1,
DAYTONA_API_KEY, and PAPERCLIP_LIVE_EXPORT_RESUME_IMAGE set to an immutable
image digest. It creates one disposable ephemeral sandbox and database, injects probe and stop transport failures, and verifies a fresh runtime can stop the sandbox while preserving exact nonce bytes. It also injects three transient failures at the production finalizer boundary and verifies retained work plus export-only retry admission. This boundary test does not claim physical copyback or result commitment; historical browser Product E2E supplied that proof before automatic unsafe-export recovery. It deletes only that owned fixture after proof.
Historical unsafe exports
The reconciliation sweep automatically recovers accepted results that an older
controller terminalized with native_workspace_sync_out_unsafe_archive. It
omits the already-rejected export without accessing the old provider, so a
stopped, unavailable, or deleted sandbox cannot require user repair. The
omission barrier and result re-admission are atomic under finalization ownership.
The normal status arbiter still enforces current ownership, contracts, and gates.
This works with active, incorrectly resolved, or missing old repair actions. It clears the matching stale unsafe notices instead of restoring them. Newer runs, changed task ownership or contracts, another active recovery action, and explicit operator dispositions prevent replay of old results. No provider turn is created. Diagnostics appear only in the local run log; no unsafe-export activity notice, repair card, or recovery chip is created. Existing stop-only allocation intents retain their exact ownership and cleanup protections.
Daytona stop-only preservation disables provider auto-delete and refreshes the provider record to confirm the disabled policy before stopping. An unavailable or unconfirmed policy leaves cleanup pending; it never falls back to stop or delete. The original ephemeral destroy policy applies only after exact accepted-result copyback and commitment.
New ephemeral native allocations receive an acquisition-time workspace sentinel bound to their run and provider allocation. Export-only resume requires that proof and the unchanged durable intent. A legacy ephemeral allocation without this proof remains blocked; recovery never creates a sentinel or accepts a replacement during resume. The reusable workspace identity hash is unchanged.
Disabling auto-delete retains provider storage while repair waits. No automatic lease-age sweep deletes these released, confirmed-stopped allocations. Storage costs can continue until committed copyback applies the original cleanup policy or an operator explicitly deletes the allocation.
Generic stop-only cleanup
An explicit sandbox stop_and_retain also requires the dedicated provider hook
when no export repair is involved, including successful per-turn reusable runs
and startup cancellation. Before dispatch, the controller stores a separate
sandboxStopAndRetain intent bound to the company, run, lease, provider allocation,
and original plugin. The initial cleanup and restart sweep both require that
plugin to advertise environmentStopLease. Missing capability or failed stop
leaves pending cleanup; ordinary release and destroy are never fallback methods.
A matching stopped receipt clears the pending intent, leaves the lease released
and resumable, and saves sandboxStopAndRetainReceipt. Its request identity,
original plugin and method: "environmentStopLease" attest this dispatch route.
The accepted-result export intent and its commitment requirements are unchanged.
The built-in fake provider uses its explicit stopLease operation rather than a
plugin RPC. Its intent pins the registered built-in provider and its receipt says
builtin.stopLease. This provider owns no real process or filesystem; the receipt
models its lifecycle. Missing built-in stop support or an unconfirmed receipt
also remains pending through restart, without calling release or destroy.