Files
PaperClipAI/packages/shared/src/validators/search.ts
T
DottaandPaperclip ae77908618 feat(search): add bulk extract endpoint (#9507)
## Thinking Path

> - Paperclip is the open source control plane people use to coordinate
AI-agent companies
> - Agents and operators need company-scoped search to discover relevant
issue history safely
> - The interactive search endpoint intentionally returns compact
excerpts and low pagination caps for UI use
> - Automation that inventories repeated references, such as
pull-request URLs, needs exhaustive distinct matches without loading
full issue objects into an LLM context
> - Client-provided regular expressions would create an unsafe and
expensive query surface, so extraction must remain literal with
server-owned expansion modes
> - This pull request adds a bounded agent-oriented extraction endpoint
with explicit truncation
> - The benefit is deterministic, compact bulk discovery across issues,
comments, and documents while preserving company authorization and rate
limits

## Linked Issues or Issue Description

### Subsystem affected

`server/` REST API and `packages/shared/` contracts.

### Problem or motivation

The existing interactive company search caps issue pagination and
snippets, so automation cannot reliably enumerate every distinct literal
or pull-request URL across issue descriptions, comments, and linked
documents without fetching large full issue payloads.

### Proposed solution

Add `GET /api/companies/:companyId/search/extract` with escaped literal
matching, optional server-owned URL token expansion,
issue/comment/document scopes, status/date filters, higher issue-level
pagination caps, compact source references, and explicit
pagination/match truncation flags.

### Alternatives considered

Reusing `GET /issues?q=` would return unnecessarily large issue objects;
increasing interactive-search snippet limits would make the UI API
heavier; accepting arbitrary client regex would expose avoidable
database cost and ReDoS risk.

### Roadmap alignment

`ROADMAP.md` does not currently list a conflicting company-search or
bulk-extraction initiative. GitHub searches found no directly
duplicative open issue or pull request.

## What Changed

- Added shared query validation and response contracts for literal and
URL extraction.
- Added a company-scoped extraction service that pages issues, gathers
matching issue/comment/document sources, expands URL tokens,
deduplicates values, and reports truncation explicitly.
- Added the authenticated route using the existing company-search
authorization decision and rate limiter.
- Added targeted Vitest coverage for URL extraction, multi-source
dedupe, date/status filters, match caps, cross-company denial, and rate
limiting.
- Documented the extraction surface in the implementation specification.

## Verification

- `pnpm exec vitest run
server/src/__tests__/company-search-extract-service.test.ts
server/src/__tests__/company-search-extract-routes.test.ts
server/src/__tests__/company-search-rate-limit-routes.test.ts
server/src/__tests__/company-search-service.test.ts` — 30 tests passed.
- `pnpm --filter @paperclipai/shared typecheck` — passed.
- `pnpm --filter @paperclipai/server typecheck` — passed.
- `git diff --check` — passed.

## Risks

- Bulk substring search can scan large text columns. The endpoint
mitigates this with a minimum literal length, bounded issue pagination,
a 20-distinct-match cap per issue, explicit truncation, existing
company-search rate limiting, and no client-provided regex.
- URL expansion uses a fixed server-owned pattern plus an escaped
literal. A security review is requested as part of PR review to confirm
the pattern and abuse controls.
- No database migration or existing API response shape changes are
included.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- OpenAI Codex CLI coding agent; exact runtime model ID and
context-window size were not exposed to the session. Tool-enabled code
execution and repository editing were used with medium reasoning effort.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-07-15 19:05:06 -05:00

260 lines
9.7 KiB
TypeScript

import { z } from "zod";
import { ISSUE_PRIORITIES, ISSUE_STATUSES } from "../constants.js";
import { isUuidLike } from "../agent-url-key.js";
import {
COMPANY_SEARCH_EXTRACT_KINDS,
COMPANY_SEARCH_EXTRACT_SCOPES,
COMPANY_SEARCH_SCOPES,
COMPANY_SEARCH_SORTS,
} from "../types/search.js";
export const COMPANY_SEARCH_MAX_QUERY_LENGTH = 200;
export const COMPANY_SEARCH_MAX_TOKENS = 8;
export const COMPANY_SEARCH_DEFAULT_LIMIT = 20;
export const COMPANY_SEARCH_MAX_LIMIT = 50;
export const COMPANY_SEARCH_MAX_OFFSET = 200;
export const COMPANY_SEARCH_EXTRACT_DEFAULT_LIMIT = 100;
export const COMPANY_SEARCH_EXTRACT_MAX_LIMIT = 200;
export const COMPANY_SEARCH_EXTRACT_MAX_OFFSET = 5_000;
export const COMPANY_SEARCH_EXTRACT_MAX_MATCHES_PER_ISSUE = 20;
const UPDATED_WITHIN_RE = /^[1-9]\d{0,2}(h|d|w|m)$/;
function firstQueryValue(value: unknown): unknown {
return Array.isArray(value) ? value[0] : value;
}
function queryValues(value: unknown): unknown[] {
if (value === undefined || value === null) return [];
return Array.isArray(value) ? value : [value];
}
function parseOptionalString(value: unknown, ctx: z.RefinementCtx, field: string): string | undefined {
const raw = firstQueryValue(value);
if (raw === undefined || raw === null) return undefined;
if (typeof raw !== "string" && typeof raw !== "number") {
ctx.addIssue({ code: z.ZodIssueCode.custom, message: `${field} must be a string` });
return undefined;
}
const normalized = String(raw).trim();
return normalized.length > 0 ? normalized : undefined;
}
function parseIntegerQuery(
value: unknown,
ctx: z.RefinementCtx,
field: string,
fallback: number,
min: number,
max: number,
): number {
const raw = firstQueryValue(value);
if (raw === undefined || raw === null || raw === "") return fallback;
const text = typeof raw === "number" ? String(raw) : typeof raw === "string" ? raw.trim() : "";
if (!/^-?\d+$/.test(text)) {
ctx.addIssue({ code: z.ZodIssueCode.custom, message: `${field} must be an integer` });
return fallback;
}
const numeric = Number.parseInt(text, 10);
if (!Number.isInteger(numeric) || numeric < min || numeric > max) {
const range = min === 0 ? `between 0 and ${max}` : `between ${min} and ${max}`;
ctx.addIssue({ code: z.ZodIssueCode.custom, message: `${field} must be ${range}` });
return fallback;
}
return numeric;
}
function parseEnumList<T extends string>(
value: unknown,
ctx: z.RefinementCtx,
field: string,
allowed: readonly T[],
): T[] {
const allowedSet = new Set<string>(allowed);
const values: T[] = [];
for (const rawEntry of queryValues(value)) {
if (typeof rawEntry !== "string") {
ctx.addIssue({ code: z.ZodIssueCode.custom, message: `${field} must be a comma-separated string` });
continue;
}
for (const rawItem of rawEntry.split(",")) {
const item = rawItem.trim();
if (!item) continue;
if (!allowedSet.has(item)) {
ctx.addIssue({ code: z.ZodIssueCode.custom, message: `${field} contains an unsupported value` });
continue;
}
if (!values.includes(item as T)) values.push(item as T);
}
}
return values;
}
function parseOptionalUuid(value: unknown, ctx: z.RefinementCtx, field: string): string | undefined {
const normalized = parseOptionalString(value, ctx, field);
if (normalized === undefined) return undefined;
if (!isUuidLike(normalized)) {
ctx.addIssue({ code: z.ZodIssueCode.custom, message: `${field} must be a UUID` });
return undefined;
}
return normalized;
}
function parseAssigneeAgentId(value: unknown, ctx: z.RefinementCtx): string | null | undefined {
const normalized = parseOptionalString(value, ctx, "assigneeAgentId");
if (normalized === undefined) return undefined;
if (normalized.toLowerCase() === "null") return null;
if (!isUuidLike(normalized)) {
ctx.addIssue({ code: z.ZodIssueCode.custom, message: "assigneeAgentId must be a UUID or 'null'" });
return undefined;
}
return normalized;
}
function parseUpdatedAfter(value: unknown, ctx: z.RefinementCtx): string | undefined {
const normalized = parseOptionalString(value, ctx, "updatedAfter");
if (normalized === undefined) return undefined;
const date = new Date(normalized);
if (Number.isNaN(date.getTime())) {
ctx.addIssue({ code: z.ZodIssueCode.custom, message: "updatedAfter must be a valid date" });
return undefined;
}
return date.toISOString();
}
function parseUpdatedWithin(value: unknown, ctx: z.RefinementCtx): string | undefined {
const normalized = parseOptionalString(value, ctx, "updatedWithin");
if (normalized === undefined) return undefined;
if (!UPDATED_WITHIN_RE.test(normalized)) {
ctx.addIssue({ code: z.ZodIssueCode.custom, message: "updatedWithin must be a duration like 24h, 7d, 4w, or 3m" });
return undefined;
}
return normalized;
}
export const companySearchQuerySchema = z.object({
q: z.unknown()
.optional()
.transform((value, ctx) => (parseOptionalString(value, ctx, "q") ?? "").slice(0, COMPANY_SEARCH_MAX_QUERY_LENGTH)),
scope: z.unknown()
.optional()
.transform((value, ctx) => {
const normalized = parseOptionalString(value, ctx, "scope") ?? "all";
if (!(COMPANY_SEARCH_SCOPES as readonly string[]).includes(normalized)) {
ctx.addIssue({ code: z.ZodIssueCode.custom, message: "scope must be a supported search scope" });
return "all";
}
return normalized as (typeof COMPANY_SEARCH_SCOPES)[number];
}),
limit: z.unknown()
.optional()
.transform((value, ctx) => parseIntegerQuery(value, ctx, "limit", COMPANY_SEARCH_DEFAULT_LIMIT, 1, COMPANY_SEARCH_MAX_LIMIT)),
offset: z.unknown()
.optional()
.transform((value, ctx) => parseIntegerQuery(value, ctx, "offset", 0, 0, COMPANY_SEARCH_MAX_OFFSET)),
status: z.unknown()
.optional()
.transform((value, ctx) => parseEnumList(value, ctx, "status", ISSUE_STATUSES)),
priority: z.unknown()
.optional()
.transform((value, ctx) => parseEnumList(value, ctx, "priority", ISSUE_PRIORITIES)),
assigneeAgentId: z.unknown()
.optional()
.transform((value, ctx) => parseAssigneeAgentId(value, ctx)),
assigneeUserId: z.unknown()
.optional()
.transform((value, ctx) => parseOptionalString(value, ctx, "assigneeUserId")),
projectId: z.unknown()
.optional()
.transform((value, ctx) => parseOptionalUuid(value, ctx, "projectId")),
labelId: z.unknown()
.optional()
.transform((value, ctx) => parseOptionalUuid(value, ctx, "labelId")),
updatedWithin: z.unknown()
.optional()
.transform((value, ctx) => parseUpdatedWithin(value, ctx)),
updatedAfter: z.unknown()
.optional()
.transform((value, ctx) => parseUpdatedAfter(value, ctx)),
sort: z.unknown()
.optional()
.transform((value, ctx) => {
const normalized = parseOptionalString(value, ctx, "sort") ?? "relevance";
if (!(COMPANY_SEARCH_SORTS as readonly string[]).includes(normalized)) {
ctx.addIssue({ code: z.ZodIssueCode.custom, message: "sort must be relevance, updated, created, or priority" });
return "relevance";
}
return normalized as (typeof COMPANY_SEARCH_SORTS)[number];
}),
});
export type CompanySearchQuery = z.infer<typeof companySearchQuerySchema>;
export const companySearchExtractQuerySchema = z.object({
contains: z.unknown().transform((value, ctx) => {
const normalized = parseOptionalString(value, ctx, "contains");
if (!normalized || normalized.length < 2) {
ctx.addIssue({ code: z.ZodIssueCode.custom, message: "contains must be at least 2 characters" });
return "";
}
if (normalized.length > COMPANY_SEARCH_MAX_QUERY_LENGTH) {
ctx.addIssue({
code: z.ZodIssueCode.custom,
message: `contains must be at most ${COMPANY_SEARCH_MAX_QUERY_LENGTH} characters`,
});
}
return normalized.slice(0, COMPANY_SEARCH_MAX_QUERY_LENGTH);
}),
kind: z.unknown()
.optional()
.transform((value, ctx) => {
const normalized = parseOptionalString(value, ctx, "kind") ?? "literal";
if (!(COMPANY_SEARCH_EXTRACT_KINDS as readonly string[]).includes(normalized)) {
ctx.addIssue({ code: z.ZodIssueCode.custom, message: "kind must be literal or url" });
return "literal";
}
return normalized as (typeof COMPANY_SEARCH_EXTRACT_KINDS)[number];
}),
scope: z.unknown()
.optional()
.transform((value, ctx) => {
const normalized = parseOptionalString(value, ctx, "scope") ?? "all";
if (!(COMPANY_SEARCH_EXTRACT_SCOPES as readonly string[]).includes(normalized)) {
ctx.addIssue({ code: z.ZodIssueCode.custom, message: "scope must be all, issues, comments, or documents" });
return "all";
}
return normalized as (typeof COMPANY_SEARCH_EXTRACT_SCOPES)[number];
}),
limit: z.unknown()
.optional()
.transform((value, ctx) => parseIntegerQuery(
value,
ctx,
"limit",
COMPANY_SEARCH_EXTRACT_DEFAULT_LIMIT,
1,
COMPANY_SEARCH_EXTRACT_MAX_LIMIT,
)),
offset: z.unknown()
.optional()
.transform((value, ctx) => parseIntegerQuery(value, ctx, "offset", 0, 0, COMPANY_SEARCH_EXTRACT_MAX_OFFSET)),
status: z.unknown()
.optional()
.transform((value, ctx) => parseEnumList(value, ctx, "status", ISSUE_STATUSES)),
updatedWithin: z.unknown()
.optional()
.transform((value, ctx) => parseUpdatedWithin(value, ctx)),
updatedAfter: z.unknown()
.optional()
.transform((value, ctx) => parseUpdatedAfter(value, ctx)),
}).superRefine((value, ctx) => {
if (value.updatedWithin && value.updatedAfter) {
ctx.addIssue({
code: z.ZodIssueCode.custom,
message: "updatedWithin and updatedAfter cannot be used together",
});
}
});
export type CompanySearchExtractQuery = z.infer<typeof companySearchExtractQuerySchema>;