mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-11 05:31:46 +02:00
## Thinking Path > - Paperclip is the open source control plane people use to coordinate AI-agent companies > - Agents and operators need company-scoped search to discover relevant issue history safely > - The interactive search endpoint intentionally returns compact excerpts and low pagination caps for UI use > - Automation that inventories repeated references, such as pull-request URLs, needs exhaustive distinct matches without loading full issue objects into an LLM context > - Client-provided regular expressions would create an unsafe and expensive query surface, so extraction must remain literal with server-owned expansion modes > - This pull request adds a bounded agent-oriented extraction endpoint with explicit truncation > - The benefit is deterministic, compact bulk discovery across issues, comments, and documents while preserving company authorization and rate limits ## Linked Issues or Issue Description ### Subsystem affected `server/` REST API and `packages/shared/` contracts. ### Problem or motivation The existing interactive company search caps issue pagination and snippets, so automation cannot reliably enumerate every distinct literal or pull-request URL across issue descriptions, comments, and linked documents without fetching large full issue payloads. ### Proposed solution Add `GET /api/companies/:companyId/search/extract` with escaped literal matching, optional server-owned URL token expansion, issue/comment/document scopes, status/date filters, higher issue-level pagination caps, compact source references, and explicit pagination/match truncation flags. ### Alternatives considered Reusing `GET /issues?q=` would return unnecessarily large issue objects; increasing interactive-search snippet limits would make the UI API heavier; accepting arbitrary client regex would expose avoidable database cost and ReDoS risk. ### Roadmap alignment `ROADMAP.md` does not currently list a conflicting company-search or bulk-extraction initiative. GitHub searches found no directly duplicative open issue or pull request. ## What Changed - Added shared query validation and response contracts for literal and URL extraction. - Added a company-scoped extraction service that pages issues, gathers matching issue/comment/document sources, expands URL tokens, deduplicates values, and reports truncation explicitly. - Added the authenticated route using the existing company-search authorization decision and rate limiter. - Added targeted Vitest coverage for URL extraction, multi-source dedupe, date/status filters, match caps, cross-company denial, and rate limiting. - Documented the extraction surface in the implementation specification. ## Verification - `pnpm exec vitest run server/src/__tests__/company-search-extract-service.test.ts server/src/__tests__/company-search-extract-routes.test.ts server/src/__tests__/company-search-rate-limit-routes.test.ts server/src/__tests__/company-search-service.test.ts` — 30 tests passed. - `pnpm --filter @paperclipai/shared typecheck` — passed. - `pnpm --filter @paperclipai/server typecheck` — passed. - `git diff --check` — passed. ## Risks - Bulk substring search can scan large text columns. The endpoint mitigates this with a minimum literal length, bounded issue pagination, a 20-distinct-match cap per issue, explicit truncation, existing company-search rate limiting, and no client-provided regex. - URL expansion uses a fixed server-owned pattern plus an escaped literal. A security review is requested as part of PR review to confirm the pattern and abuse controls. - No database migration or existing API response shape changes are included. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex CLI coding agent; exact runtime model ID and context-window size were not exposed to the session. Tool-enabled code execution and repository editing were used with medium reasoning effort. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
260 lines
9.7 KiB
TypeScript
260 lines
9.7 KiB
TypeScript
import { z } from "zod";
|
|
import { ISSUE_PRIORITIES, ISSUE_STATUSES } from "../constants.js";
|
|
import { isUuidLike } from "../agent-url-key.js";
|
|
import {
|
|
COMPANY_SEARCH_EXTRACT_KINDS,
|
|
COMPANY_SEARCH_EXTRACT_SCOPES,
|
|
COMPANY_SEARCH_SCOPES,
|
|
COMPANY_SEARCH_SORTS,
|
|
} from "../types/search.js";
|
|
|
|
export const COMPANY_SEARCH_MAX_QUERY_LENGTH = 200;
|
|
export const COMPANY_SEARCH_MAX_TOKENS = 8;
|
|
export const COMPANY_SEARCH_DEFAULT_LIMIT = 20;
|
|
export const COMPANY_SEARCH_MAX_LIMIT = 50;
|
|
export const COMPANY_SEARCH_MAX_OFFSET = 200;
|
|
export const COMPANY_SEARCH_EXTRACT_DEFAULT_LIMIT = 100;
|
|
export const COMPANY_SEARCH_EXTRACT_MAX_LIMIT = 200;
|
|
export const COMPANY_SEARCH_EXTRACT_MAX_OFFSET = 5_000;
|
|
export const COMPANY_SEARCH_EXTRACT_MAX_MATCHES_PER_ISSUE = 20;
|
|
|
|
const UPDATED_WITHIN_RE = /^[1-9]\d{0,2}(h|d|w|m)$/;
|
|
|
|
function firstQueryValue(value: unknown): unknown {
|
|
return Array.isArray(value) ? value[0] : value;
|
|
}
|
|
|
|
function queryValues(value: unknown): unknown[] {
|
|
if (value === undefined || value === null) return [];
|
|
return Array.isArray(value) ? value : [value];
|
|
}
|
|
|
|
function parseOptionalString(value: unknown, ctx: z.RefinementCtx, field: string): string | undefined {
|
|
const raw = firstQueryValue(value);
|
|
if (raw === undefined || raw === null) return undefined;
|
|
if (typeof raw !== "string" && typeof raw !== "number") {
|
|
ctx.addIssue({ code: z.ZodIssueCode.custom, message: `${field} must be a string` });
|
|
return undefined;
|
|
}
|
|
const normalized = String(raw).trim();
|
|
return normalized.length > 0 ? normalized : undefined;
|
|
}
|
|
|
|
function parseIntegerQuery(
|
|
value: unknown,
|
|
ctx: z.RefinementCtx,
|
|
field: string,
|
|
fallback: number,
|
|
min: number,
|
|
max: number,
|
|
): number {
|
|
const raw = firstQueryValue(value);
|
|
if (raw === undefined || raw === null || raw === "") return fallback;
|
|
const text = typeof raw === "number" ? String(raw) : typeof raw === "string" ? raw.trim() : "";
|
|
if (!/^-?\d+$/.test(text)) {
|
|
ctx.addIssue({ code: z.ZodIssueCode.custom, message: `${field} must be an integer` });
|
|
return fallback;
|
|
}
|
|
const numeric = Number.parseInt(text, 10);
|
|
if (!Number.isInteger(numeric) || numeric < min || numeric > max) {
|
|
const range = min === 0 ? `between 0 and ${max}` : `between ${min} and ${max}`;
|
|
ctx.addIssue({ code: z.ZodIssueCode.custom, message: `${field} must be ${range}` });
|
|
return fallback;
|
|
}
|
|
return numeric;
|
|
}
|
|
|
|
function parseEnumList<T extends string>(
|
|
value: unknown,
|
|
ctx: z.RefinementCtx,
|
|
field: string,
|
|
allowed: readonly T[],
|
|
): T[] {
|
|
const allowedSet = new Set<string>(allowed);
|
|
const values: T[] = [];
|
|
for (const rawEntry of queryValues(value)) {
|
|
if (typeof rawEntry !== "string") {
|
|
ctx.addIssue({ code: z.ZodIssueCode.custom, message: `${field} must be a comma-separated string` });
|
|
continue;
|
|
}
|
|
for (const rawItem of rawEntry.split(",")) {
|
|
const item = rawItem.trim();
|
|
if (!item) continue;
|
|
if (!allowedSet.has(item)) {
|
|
ctx.addIssue({ code: z.ZodIssueCode.custom, message: `${field} contains an unsupported value` });
|
|
continue;
|
|
}
|
|
if (!values.includes(item as T)) values.push(item as T);
|
|
}
|
|
}
|
|
return values;
|
|
}
|
|
|
|
function parseOptionalUuid(value: unknown, ctx: z.RefinementCtx, field: string): string | undefined {
|
|
const normalized = parseOptionalString(value, ctx, field);
|
|
if (normalized === undefined) return undefined;
|
|
if (!isUuidLike(normalized)) {
|
|
ctx.addIssue({ code: z.ZodIssueCode.custom, message: `${field} must be a UUID` });
|
|
return undefined;
|
|
}
|
|
return normalized;
|
|
}
|
|
|
|
function parseAssigneeAgentId(value: unknown, ctx: z.RefinementCtx): string | null | undefined {
|
|
const normalized = parseOptionalString(value, ctx, "assigneeAgentId");
|
|
if (normalized === undefined) return undefined;
|
|
if (normalized.toLowerCase() === "null") return null;
|
|
if (!isUuidLike(normalized)) {
|
|
ctx.addIssue({ code: z.ZodIssueCode.custom, message: "assigneeAgentId must be a UUID or 'null'" });
|
|
return undefined;
|
|
}
|
|
return normalized;
|
|
}
|
|
|
|
function parseUpdatedAfter(value: unknown, ctx: z.RefinementCtx): string | undefined {
|
|
const normalized = parseOptionalString(value, ctx, "updatedAfter");
|
|
if (normalized === undefined) return undefined;
|
|
const date = new Date(normalized);
|
|
if (Number.isNaN(date.getTime())) {
|
|
ctx.addIssue({ code: z.ZodIssueCode.custom, message: "updatedAfter must be a valid date" });
|
|
return undefined;
|
|
}
|
|
return date.toISOString();
|
|
}
|
|
|
|
function parseUpdatedWithin(value: unknown, ctx: z.RefinementCtx): string | undefined {
|
|
const normalized = parseOptionalString(value, ctx, "updatedWithin");
|
|
if (normalized === undefined) return undefined;
|
|
if (!UPDATED_WITHIN_RE.test(normalized)) {
|
|
ctx.addIssue({ code: z.ZodIssueCode.custom, message: "updatedWithin must be a duration like 24h, 7d, 4w, or 3m" });
|
|
return undefined;
|
|
}
|
|
return normalized;
|
|
}
|
|
|
|
export const companySearchQuerySchema = z.object({
|
|
q: z.unknown()
|
|
.optional()
|
|
.transform((value, ctx) => (parseOptionalString(value, ctx, "q") ?? "").slice(0, COMPANY_SEARCH_MAX_QUERY_LENGTH)),
|
|
scope: z.unknown()
|
|
.optional()
|
|
.transform((value, ctx) => {
|
|
const normalized = parseOptionalString(value, ctx, "scope") ?? "all";
|
|
if (!(COMPANY_SEARCH_SCOPES as readonly string[]).includes(normalized)) {
|
|
ctx.addIssue({ code: z.ZodIssueCode.custom, message: "scope must be a supported search scope" });
|
|
return "all";
|
|
}
|
|
return normalized as (typeof COMPANY_SEARCH_SCOPES)[number];
|
|
}),
|
|
limit: z.unknown()
|
|
.optional()
|
|
.transform((value, ctx) => parseIntegerQuery(value, ctx, "limit", COMPANY_SEARCH_DEFAULT_LIMIT, 1, COMPANY_SEARCH_MAX_LIMIT)),
|
|
offset: z.unknown()
|
|
.optional()
|
|
.transform((value, ctx) => parseIntegerQuery(value, ctx, "offset", 0, 0, COMPANY_SEARCH_MAX_OFFSET)),
|
|
status: z.unknown()
|
|
.optional()
|
|
.transform((value, ctx) => parseEnumList(value, ctx, "status", ISSUE_STATUSES)),
|
|
priority: z.unknown()
|
|
.optional()
|
|
.transform((value, ctx) => parseEnumList(value, ctx, "priority", ISSUE_PRIORITIES)),
|
|
assigneeAgentId: z.unknown()
|
|
.optional()
|
|
.transform((value, ctx) => parseAssigneeAgentId(value, ctx)),
|
|
assigneeUserId: z.unknown()
|
|
.optional()
|
|
.transform((value, ctx) => parseOptionalString(value, ctx, "assigneeUserId")),
|
|
projectId: z.unknown()
|
|
.optional()
|
|
.transform((value, ctx) => parseOptionalUuid(value, ctx, "projectId")),
|
|
labelId: z.unknown()
|
|
.optional()
|
|
.transform((value, ctx) => parseOptionalUuid(value, ctx, "labelId")),
|
|
updatedWithin: z.unknown()
|
|
.optional()
|
|
.transform((value, ctx) => parseUpdatedWithin(value, ctx)),
|
|
updatedAfter: z.unknown()
|
|
.optional()
|
|
.transform((value, ctx) => parseUpdatedAfter(value, ctx)),
|
|
sort: z.unknown()
|
|
.optional()
|
|
.transform((value, ctx) => {
|
|
const normalized = parseOptionalString(value, ctx, "sort") ?? "relevance";
|
|
if (!(COMPANY_SEARCH_SORTS as readonly string[]).includes(normalized)) {
|
|
ctx.addIssue({ code: z.ZodIssueCode.custom, message: "sort must be relevance, updated, created, or priority" });
|
|
return "relevance";
|
|
}
|
|
return normalized as (typeof COMPANY_SEARCH_SORTS)[number];
|
|
}),
|
|
});
|
|
|
|
export type CompanySearchQuery = z.infer<typeof companySearchQuerySchema>;
|
|
|
|
export const companySearchExtractQuerySchema = z.object({
|
|
contains: z.unknown().transform((value, ctx) => {
|
|
const normalized = parseOptionalString(value, ctx, "contains");
|
|
if (!normalized || normalized.length < 2) {
|
|
ctx.addIssue({ code: z.ZodIssueCode.custom, message: "contains must be at least 2 characters" });
|
|
return "";
|
|
}
|
|
if (normalized.length > COMPANY_SEARCH_MAX_QUERY_LENGTH) {
|
|
ctx.addIssue({
|
|
code: z.ZodIssueCode.custom,
|
|
message: `contains must be at most ${COMPANY_SEARCH_MAX_QUERY_LENGTH} characters`,
|
|
});
|
|
}
|
|
return normalized.slice(0, COMPANY_SEARCH_MAX_QUERY_LENGTH);
|
|
}),
|
|
kind: z.unknown()
|
|
.optional()
|
|
.transform((value, ctx) => {
|
|
const normalized = parseOptionalString(value, ctx, "kind") ?? "literal";
|
|
if (!(COMPANY_SEARCH_EXTRACT_KINDS as readonly string[]).includes(normalized)) {
|
|
ctx.addIssue({ code: z.ZodIssueCode.custom, message: "kind must be literal or url" });
|
|
return "literal";
|
|
}
|
|
return normalized as (typeof COMPANY_SEARCH_EXTRACT_KINDS)[number];
|
|
}),
|
|
scope: z.unknown()
|
|
.optional()
|
|
.transform((value, ctx) => {
|
|
const normalized = parseOptionalString(value, ctx, "scope") ?? "all";
|
|
if (!(COMPANY_SEARCH_EXTRACT_SCOPES as readonly string[]).includes(normalized)) {
|
|
ctx.addIssue({ code: z.ZodIssueCode.custom, message: "scope must be all, issues, comments, or documents" });
|
|
return "all";
|
|
}
|
|
return normalized as (typeof COMPANY_SEARCH_EXTRACT_SCOPES)[number];
|
|
}),
|
|
limit: z.unknown()
|
|
.optional()
|
|
.transform((value, ctx) => parseIntegerQuery(
|
|
value,
|
|
ctx,
|
|
"limit",
|
|
COMPANY_SEARCH_EXTRACT_DEFAULT_LIMIT,
|
|
1,
|
|
COMPANY_SEARCH_EXTRACT_MAX_LIMIT,
|
|
)),
|
|
offset: z.unknown()
|
|
.optional()
|
|
.transform((value, ctx) => parseIntegerQuery(value, ctx, "offset", 0, 0, COMPANY_SEARCH_EXTRACT_MAX_OFFSET)),
|
|
status: z.unknown()
|
|
.optional()
|
|
.transform((value, ctx) => parseEnumList(value, ctx, "status", ISSUE_STATUSES)),
|
|
updatedWithin: z.unknown()
|
|
.optional()
|
|
.transform((value, ctx) => parseUpdatedWithin(value, ctx)),
|
|
updatedAfter: z.unknown()
|
|
.optional()
|
|
.transform((value, ctx) => parseUpdatedAfter(value, ctx)),
|
|
}).superRefine((value, ctx) => {
|
|
if (value.updatedWithin && value.updatedAfter) {
|
|
ctx.addIssue({
|
|
code: z.ZodIssueCode.custom,
|
|
message: "updatedWithin and updatedAfter cannot be used together",
|
|
});
|
|
}
|
|
});
|
|
|
|
export type CompanySearchExtractQuery = z.infer<typeof companySearchExtractQuerySchema>;
|