Files
PaperClipAI/packages/adapter-utils/src/runtime-progress.ts
T
Devin FoleyandPaperclip c48feee190 Improve live agent feedback during sandboxed runs (#8915)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - A core part of that experience is watching active agent runs without
dropping into raw logs first
> - Local and sandbox-backed adapters already record useful run output,
progress, and tool activity
> - But active issue threads could sit visually stale while the agent
was syncing workspaces, tailing sandbox output, or emitting incremental
tool-call updates
> - Operators need timely, human-readable progress while preserving the
raw transcript underneath
> - This pull request streams sandbox run-log progress into runtime
status, keeps visible issue threads refreshed, and folds repeated ACPX
tool updates into stable transcript cards
> - The benefit is that long-running agent work becomes easier to
supervise without changing the task/comment control-plane model

## Linked Issues or Issue Description

No public GitHub issue exists for this exact change.

Problem/motivation:

- During long-running sandboxed agent work, the issue UI can appear idle
even though the agent is actively syncing, running tools, or producing
incremental output.
- Operators need realtime feedback at the issue-thread layer, not only
after opening raw logs or waiting for the final heartbeat result.
- Related public context: #1808 previously added live-run status dots to
Projects; #4362 touches heartbeat wakeup behavior but is not a duplicate
of this runtime/UI feedback change.

## What Changed

- Added sandbox run-log streaming support and defaulted sandbox-capable
local adapters into the richer live-feedback path.
- Surfaced environment/sandbox sync progress through heartbeat runtime
status with bounded, redacted snippets.
- Added live issue-thread cache patching so visible active runs update
as progress events arrive.
- Folded repeated ACPX `tool_call` updates into one transcript card
instead of stacking duplicate cards.
- Updated adapter docs and added focused regression coverage for sandbox
log streaming, runtime status, ACPX parsing, live updates, transcript
rendering, and issue chat messages.

## Verification

- `pnpm install --frozen-lockfile`
- `pnpm exec vitest run ui/src/context/LiveUpdatesProvider.test.ts`
- `pnpm exec vitest run
server/src/services/heartbeat-run-runtime-status.test.ts
server/src/__tests__/heartbeat-runtime-state.test.ts
ui/src/context/LiveUpdatesProvider.test.ts`
- `pnpm exec vitest run
packages/adapter-utils/src/execution-target-sandbox.test.ts
packages/adapter-utils/src/sandbox-managed-runtime.test.ts
server/src/services/heartbeat-run-runtime-status.test.ts
server/src/__tests__/agent-live-run-routes.test.ts
server/src/__tests__/heartbeat-runtime-state.test.ts
packages/adapters/acpx-local/src/ui/parse-stdout.test.ts
ui/src/context/LiveUpdatesProvider.test.ts
ui/src/components/transcript/RunTranscriptView.test.tsx
ui/src/lib/issue-chat-messages.test.ts
ui/src/components/IssueChatThread.test.tsx`
- GitHub PR workflow on head `8397953e7b41ccd42e5d9457ee7e4dfb996e4ec5`:
`verify`, build, typecheck/release-registry, e2e, general shards,
serialized server shards, and canary dry run passed.
- Greptile Review on head `8397953e7b41ccd42e5d9457ee7e4dfb996e4ec5`:
Confidence Score 5/5, no unresolved review threads.

## Risks

- Live issue-thread cache patching could miss an edge case for a route
shape not covered by tests.
- Surfacing active-run snippets needs continued care around redaction;
this PR keeps snippets bounded and adds redaction-focused coverage.
- More frequent active-run UI refreshes could expose performance issues
on very large issue threads, though updates are scoped to visible
run/query caches.

## Model Used

OpenAI GPT-5 via Codex, operating as a tool-enabled coding agent with
shell, git, and repository-editing capabilities. Context window size is
not exposed in this runtime.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-07-02 22:21:56 -07:00

171 lines
6.1 KiB
TypeScript

// Shared, throttled progress reporting for execution-target sync/restore.
//
// Transports (sandbox / SSH) own the byte counting and call `report()` as bytes
// move; orchestrators own the per-phase label and direction. The reporter
// throttles emits so a long transfer doesn't flood the log: a line is emitted
// only when the percentage crosses a step boundary (default every 10%) or once
// at least `minIntervalMs` has elapsed since the last emit. The terminal
// completion line is always emitted via `complete()` (or when `report()` reaches
// the known total).
/** A sink for fully-formatted progress lines (newline included). */
export type RuntimeProgressSink = (line: string) => void | Promise<void>;
export type RuntimeProgressPhase =
| "Syncing"
| "Restoring"
| "Importing git history"
| "Exporting git history";
export type RuntimeProgressDirection = "to" | "from";
export type RuntimeProgressTarget = "sandbox" | "ssh";
export type RuntimeStatusPhase =
| "git_sync"
| "config_sync"
| "adapter_startup"
| "restore"
| "export"
| "finalize";
export interface RuntimeStatusUpdate {
phase: RuntimeStatusPhase;
message: string;
currentToolName?: string | null;
lastAssistantSnippet?: string | null;
lastEventAt?: Date | string | null;
}
export type RuntimeStatusSink = (update: RuntimeStatusUpdate) => void | Promise<void>;
export interface RuntimeProgressReporterOptions {
sink: RuntimeProgressSink;
phase: RuntimeProgressPhase;
/** Optional per-phase label, e.g. "workspace" or an asset key. */
label?: string;
direction: RuntimeProgressDirection;
target: RuntimeProgressTarget;
/** Emit when the percentage crosses this step. Default 10. */
stepPercent?: number;
/** Emit when at least this many ms have elapsed since the last emit. Default 2000. */
minIntervalMs?: number;
/** Injectable clock for deterministic tests. Default `Date.now`. */
now?: () => number;
}
export interface RuntimeProgressReporter {
/**
* Report progress. Throttled: only emits on a step crossing or after
* `minIntervalMs`. When `totalBytes` is known and `doneBytes` reaches it, the
* terminal 100% line is emitted and the reporter is marked complete.
*/
report(doneBytes: number, totalBytes: number | null): Promise<void>;
/**
* Emit the terminal completion line if it hasn't been emitted yet. Idempotent.
*/
complete(doneBytes?: number, totalBytes?: number | null): Promise<void>;
/**
* Emit a terminal failure line if no terminal line has been emitted yet, so a
* failed transfer leaves an explicit marker instead of a dangling percentage.
* Idempotent and mutually exclusive with `complete()`.
*/
fail(doneBytes?: number, totalBytes?: number | null): Promise<void>;
}
const BYTES_PER_MB = 1024 * 1024;
function formatMb(bytes: number): string {
return (Math.max(0, bytes) / BYTES_PER_MB).toFixed(1);
}
function clampPercent(value: number): number {
if (!Number.isFinite(value)) return 0;
return Math.min(100, Math.max(0, Math.round(value)));
}
export function createRuntimeProgressReporter(
options: RuntimeProgressReporterOptions,
): RuntimeProgressReporter {
const stepPercent = options.stepPercent && options.stepPercent > 0 ? options.stepPercent : 10;
const minIntervalMs =
options.minIntervalMs && options.minIntervalMs > 0 ? options.minIntervalMs : 2000;
const now = options.now ?? Date.now;
const prefix = `[paperclip] ${options.phase}${options.label ? ` ${options.label}` : ""} ${options.direction} ${options.target}`;
let lastEmitAt: number | null = null;
let lastStep = -1;
let lastDoneBytes = 0;
let lastTotalBytes: number | null = null;
let completed = false;
function buildLine(doneBytes: number, totalBytes: number | null): string {
if (totalBytes != null && totalBytes > 0) {
const pct = clampPercent((doneBytes / totalBytes) * 100);
return `${prefix}: ${pct}% (${formatMb(doneBytes)}/${formatMb(totalBytes)} MB)\n`;
}
return `${prefix}: ${formatMb(doneBytes)} MB\n`;
}
function buildFailLine(doneBytes: number, totalBytes: number | null): string {
if (totalBytes != null && totalBytes > 0) {
const pct = clampPercent((doneBytes / totalBytes) * 100);
return `${prefix}: failed at ${pct}% (${formatMb(doneBytes)}/${formatMb(totalBytes)} MB)\n`;
}
return `${prefix}: failed after ${formatMb(doneBytes)} MB\n`;
}
async function emit(doneBytes: number, totalBytes: number | null): Promise<void> {
lastEmitAt = now();
if (totalBytes != null && totalBytes > 0) {
lastStep = Math.floor(((doneBytes / totalBytes) * 100) / stepPercent);
}
await options.sink(buildLine(doneBytes, totalBytes));
}
return {
async report(doneBytes, totalBytes) {
lastDoneBytes = doneBytes;
lastTotalBytes = totalBytes;
if (completed) return;
const elapsedOk = lastEmitAt == null || now() - lastEmitAt >= minIntervalMs;
if (totalBytes != null && totalBytes > 0) {
const terminal = doneBytes >= totalBytes;
const step = Math.floor(((doneBytes / totalBytes) * 100) / stepPercent);
const stepOk = step > lastStep;
if (terminal || stepOk || elapsedOk) {
await emit(doneBytes, totalBytes);
}
if (terminal) completed = true;
return;
}
// Unknown total: no step boundaries, throttle purely on elapsed time.
if (elapsedOk) {
await emit(doneBytes, totalBytes);
}
},
async complete(doneBytes, totalBytes) {
if (completed) return;
completed = true;
const total = totalBytes !== undefined ? totalBytes : lastTotalBytes;
const done =
doneBytes !== undefined
? doneBytes
: total != null && total > 0
? total
: lastDoneBytes;
await options.sink(buildLine(done, total));
},
async fail(doneBytes, totalBytes) {
if (completed) return;
completed = true;
const total = totalBytes !== undefined ? totalBytes : lastTotalBytes;
const done = doneBytes !== undefined ? doneBytes : lastDoneBytes;
await options.sink(buildFailLine(done, total));
},
};
}