Files
PaperClipAI/server/src/services/issue-monitors.ts
T
DottaandPaperclip 8cfedd7df8 fix(runner): restore task monitors and durable timed waits (#15446)
## Thinking Path

> - Paperclip manages AI agents and their task execution.
> - Agents need a durable way to return to work after a delayed check.
> - The issue monitor scheduler already provides a one-shot wake for an
assignee.
> - Native runners reject generic execution-policy writes and had no
bound monitor tool.
> - Scheduling alone is insufficient because native completion also
needs to accept a timed wait.
> - This pull request adds an authorized monitor tool and connects it to
completion and the existing scheduler.
> - An agent can now schedule its next check, end the run, and resume on
the same task.

## Linked Issues or Issue Description

**What happened?**

A native runner could not set its own task monitor. `call_api` correctly
rejected execution-policy writes, while `schedule_wake` had no
production binding. `paperclip_finish` also rejected monitor waits.

**Expected behavior**

A standard native run can set a one-shot monitor on its current task or
another accessible task assigned to the same agent. After a confirmed
schedule on the current task, it can yield. The scheduler later delivers
`issue_monitor_due`.

**Steps to reproduce**

1. Start a standard native task.
2. Ask the agent to check the task again later and end its current run.
3. Inspect available tools and try the generic issue execution-policy
update.
4. Observe the missing native tool and the lifecycle-write denial.

Related PRs: #14680 concerns monitor notes in the shared wake prompt.
#11919 changes attempt-limit scope. This PR adds native scheduling and
completion authority and retains the existing cumulative attempt bounds.
It does not depend on either PR.

## What Changed

- Add provider-neutral `set_task_monitor` with a default current-task
target, future timestamp, required notes, existing bounds, and explicit
clearing.
- Check company, task visibility, ownership, runtime permissions, work
mode, and active-run authority. Preserve review-only restrictions.
Reject the reserved server-owned quota-recovery name before saving or
accepting a native wait.
- Commit the monitor, audit event, and retry receipt together. Retry
receipts survive a successor run without re-arming cleared or consumed
timers.
- Permit `paperclip_finish` to yield to a persisted monitor. Recheck
ownership and the schedule when committing final disposition. Release
execution without an immediate continuation.
- Preserve due monitors during native execution. Fence wake admission
and consumption against replacement, clearing, reassignment, and
completion. Preserve unrelated review policy.
- Expose scheduled and consumed monitor instructions in task context.
Update provider schemas, Rust validation, generated contracts, and
execution documentation.
- Add an opt-in live Codex smoke script with isolated data and explicit
run/session/runner/process evidence.

## Verification

- Repository `pnpm -r typecheck` and `pnpm build` passed after rebase.
Server typecheck passed again after review fixes. All CI test shards
pass on `3def77b1b`, including runner TypeScript/Rust, server,
serialized server, workspace, and browser tests. All CI gates are green,
including the canary dry run. Greptile is 5/5 on the same commit with
zero unresolved threads.
- The local monolithic `pnpm test:run`, started before the rebase, was
interrupted after current-head CI test coverage passed. It is not
counted as a standalone full-suite pass; the focused local regression
suites passed.
- Targeted server tests cover scheduling, replacement, clearing, policy
preservation, cumulative bounds, cross-run retries, permissions,
provider-neutral discovery, review restrictions, completion authority,
and scheduler/finalizer races.
- Runner contract/catalog/semantic tests and Rust terminal-tool tests
cover the new operation and monitor completion.
- Live Codex test passed twice (latest live run on `e0bcd63e6`) in a
temporary database and workspace, with a 300,000 ms warm window. First
run `67bd7709-c089-4d4a-9d2b-0d6b618a34b0` yielded at
`2026-10-07T13:15:03.274Z`. Second run
`97c323f9-595a-4cc5-a007-db5a2fbb937c` started at `13:15:30.952Z`,
received `issue_monitor_due`, and completed the same task. Exactly one
monitor wake was recorded.
- Both live runs used native session
`7b1dd753-1c9b-4e7a-b22f-a125dbc3748c`, runner
`e90d9a1b-3502-4ee3-b15e-edc024c555d4`, provider session
`01a11680-6d01-70c0-9a55-db7246ed66c3`, and PID `64218` with the same
process start time. This proves warm reuse for that local Codex test,
not only successful scheduling.
- Reproduce the paid live test with `node --import
./server/node_modules/tsx/dist/loader.mjs
server/scripts/smoke-native-task-monitor.ts --run`, with the installed
Codex binary on `PATH` and a valid local login.

## Risks

- The scheduler now defers monitor dispatch while the task has an active
native run. A stuck run still depends on the existing recovery
lifecycle.
- Idempotency uses the existing run ledger; no table or migration is
added.
- Other providers share the tested tool and completion contracts. Only
Codex received a live model test.
- Existing `call_api` lifecycle restrictions remain enforced. Monitor
waits do not bypass task blockers, reviews, or approvals.

## Model Used

OpenAI Codex, GPT-6 family, with tool use, code execution, and
TypeScript/Rust editing. The session does not expose the exact deployed
model ID or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-07 08:41:24 -05:00

167 lines
6.5 KiB
TypeScript

import { issueExecutionMonitorPolicySchema, PROVIDER_QUOTA_MONITOR_SERVICE_NAME } from "@paperclipai/shared";
import type { issues } from "@paperclipai/db";
import { z } from "zod";
import { forbidden, unprocessable } from "../errors.js";
import type { accessService } from "./access.js";
import { authorizationDeniedDetails, type AuthorizationActor } from "./authorization.js";
import { applyIssueMonitorPolicyTransition, normalizeIssueExecutionPolicy, parseIssueExecutionState, redactIssueMonitorExternalRef, setIssueExecutionPolicyMonitorScheduledBy } from "./issue-execution-policy.js";
type NormalizedExecutionPolicy = ReturnType<typeof normalizeIssueExecutionPolicy>;
export function monitorPoliciesEqual(
left: NormalizedExecutionPolicy | null,
right: NormalizedExecutionPolicy | null,
) {
return (
JSON.stringify(left?.monitor ?? null) ===
JSON.stringify(right?.monitor ?? null)
);
}
export function applyActorMonitorScheduledBy(
policy: NormalizedExecutionPolicy | null,
actorType: "agent" | "user",
) {
return setIssueExecutionPolicyMonitorScheduledBy(
policy,
actorType === "user" ? "board" : "assignee",
);
}
export async function assertCanManageIssueMonitor(
accessSvc: Pick<ReturnType<typeof accessService>, "decide">,
req: { actor: AuthorizationActor },
companyId: string,
assigneeAgentId: string | null,
monitorChanged: boolean,
) {
if (!monitorChanged) return;
if (req.actor.type === "board") return;
const runtimeDecision = await accessSvc.decide({
actor: req.actor,
action: "runtime:manage",
resource: { type: "company", companyId },
});
if (!runtimeDecision.allowed) {
throw forbidden(
runtimeDecision.explanation,
authorizationDeniedDetails(runtimeDecision),
);
}
if (
req.actor.type === "agent" &&
req.actor.agentId &&
req.actor.agentId === assigneeAgentId
)
return;
throw forbidden(
"Only the assignee agent or a board user can manage issue monitors",
);
}
export function summarizeIssueMonitor(
issue: {
monitorNextCheckAt?: Date | null;
monitorLastTriggeredAt?: Date | null;
monitorAttemptCount?: number | null;
monitorNotes?: string | null;
monitorScheduledBy?: string | null;
executionState?: unknown;
},
policy: NormalizedExecutionPolicy | null,
) {
const state = parseIssueExecutionState(issue.executionState);
return {
nextCheckAt:
issue.monitorNextCheckAt?.toISOString() ??
policy?.monitor?.nextCheckAt ??
null,
lastTriggeredAt:
issue.monitorLastTriggeredAt?.toISOString() ??
state?.monitor?.lastTriggeredAt ??
null,
attemptCount:
issue.monitorAttemptCount ?? state?.monitor?.attemptCount ?? 0,
notes:
policy?.monitor?.notes ??
issue.monitorNotes ??
state?.monitor?.notes ??
null,
scheduledBy:
issue.monitorScheduledBy ??
policy?.monitor?.scheduledBy ??
state?.monitor?.scheduledBy ??
null,
kind: policy?.monitor?.kind ?? state?.monitor?.kind ?? null,
serviceName:
policy?.monitor?.serviceName ?? state?.monitor?.serviceName ?? null,
externalRef: redactIssueMonitorExternalRef(
policy?.monitor?.externalRef ?? state?.monitor?.externalRef ?? null,
),
timeoutAt: policy?.monitor?.timeoutAt ?? state?.monitor?.timeoutAt ?? null,
maxAttempts:
policy?.monitor?.maxAttempts ?? state?.monitor?.maxAttempts ?? null,
recoveryPolicy:
policy?.monitor?.recoveryPolicy ?? state?.monitor?.recoveryPolicy ?? null,
status: state?.monitor?.status ?? (policy?.monitor ? "scheduled" : null),
clearReason: state?.monitor?.clearReason ?? null,
};
}
export const setTaskMonitorSchema = z.object({
taskId: z.string().uuid().optional(),
idempotencyKey: z.string().trim().min(1).max(240),
monitor: issueExecutionMonitorPolicySchema.omit({ scheduledBy: true }).extend({
notes: z.string().trim().min(1).max(500),
}).strict().refine(monitor => monitor.serviceName !== PROVIDER_QUOTA_MONITOR_SERVICE_NAME, {
message: "This serviceName is reserved for server-owned quota recovery; use a different service name for an ordinary task check",
path: ["serviceName"],
}).nullable(),
}).strict();
/** Caller holds the issue lock; merge only monitor state, never review policy. */
export function prepareIssueMonitorUpdate(
issue: typeof issues.$inferSelect,
monitor: z.infer<typeof setTaskMonitorSchema>["monitor"],
actor: AuthorizationActor,
now = new Date(),
) {
if (monitor && Date.parse(monitor.nextCheckAt) <= now.getTime()) {
throw unprocessable("Monitor nextCheckAt must be in the future");
}
if (monitor?.timeoutAt && Date.parse(monitor.timeoutAt) <= Date.parse(monitor.nextCheckAt)) {
throw unprocessable("Monitor timeoutAt must be later than nextCheckAt");
}
const previousPolicy = normalizeIssueExecutionPolicy(issue.executionPolicy);
const policy = applyActorMonitorScheduledBy(normalizeIssueExecutionPolicy({
...previousPolicy, monitor,
}), actor.type === "board" ? "user" : "agent");
const transition = applyIssueMonitorPolicyTransition({
issue, policy, previousPolicy, requestedAssigneePatch: {},
actor: { agentId: actor.agentId ?? null, userId: actor.userId ?? null },
monitorExplicitlyUpdated: true,
});
// Normalization supplies monitor defaults, but must not rewrite unrelated
// review/authorization settings (including forward-compatible policy fields).
const storedPolicy = { ...(issue.executionPolicy ?? {}) };
if (policy?.monitor) storedPolicy.monitor = policy.monitor;
else delete storedPolicy.monitor;
return { ...transition.patch, executionPolicy: Object.keys(storedPolicy).length ? storedPolicy : null } as Partial<typeof issues.$inferInsert>;
}
/** Timestamp presence alone is not authority to park a native task. */
export function eligibleIssueMonitorWait(
issue: typeof issues.$inferSelect,
agentId: string,
now = new Date(),
): string | null {
if (issue.assigneeAgentId !== agentId || issue.assigneeUserId ||
!["in_progress", "in_review"].includes(issue.status) || !issue.monitorNextCheckAt) return null;
const monitor = normalizeIssueExecutionPolicy(issue.executionPolicy)?.monitor;
if (!monitor || monitor.serviceName === PROVIDER_QUOTA_MONITOR_SERVICE_NAME ||
Date.parse(monitor.nextCheckAt) !== issue.monitorNextCheckAt.getTime() ||
(monitor.timeoutAt && Date.parse(monitor.timeoutAt) <= now.getTime()) ||
(monitor.maxAttempts != null && issue.monitorAttemptCount >= monitor.maxAttempts)) return null;
return issue.monitorNextCheckAt.toISOString();
}