A monorepo presents a specific problem for taint analysis. The vulnerability you care about usually starts in one package and ends in another, and the two packages are deliberately decoupled — that is the entire point of the architecture. A handler in `apps/api` receives request data. A sink in `packages/db` executes it. Neither package imports the other.
Build the cone, not the file
The fix is to compute the reverse dependency cone of every changed file, then parse only that. In a well-factored monorepo this is usually a small percentage of the repository, but it captures every path that could reach the change.
Crucially, the cone must be computed from the *importer* graph, not the import graph. Scanning what a change imports tells you what it depends on. Scanning what imports it tells you what can reach it — which is the direction that matters for regression review.
- Invert the import graph once and cache it per commit SHA.
- For each changed file, walk importers transitively to a configurable depth.
- Parse only the resulting cone; skip packages with no path to the change.
- Recompute the cache only when the build manifest changes, not per push.
Taint state has to be explicit
Once you have the cone, you need to decide what to track. We use a small lattice: untainted, tainted, and sanitised-with-evidence. That third state is the one that matters in practice, because it is what stops the analysis from drowning in false positives.
A sanitiser call is only accepted as a sanitiser if it appears on every path to the sink. If any path reaches the sink untainted, the finding stands. This strictness is what keeps the analysis from drowning in false positives.
type Taint = "untainted" | "tainted" | "sanitised";
function join(a: Taint, b: Taint): Taint {
// The weakest guarantee wins at a join point.
if (a === "untainted" || b === "untainted") return "untainted";
if (a === "sanitised" && b === "sanitised") return "sanitised";
return "tainted";
}Where the time actually goes
Parsing dominates, and it is the part worth optimising. Incremental parsing — reusing the AST for files that did not change — speeds up cold reviews on large monorepos. The second win was deferring model inference until after the deterministic pass, so the model only ever sees candidates that already survived pattern matching.
The result is an architecture built for sub-minute reviews, on repositories large enough that a full re-scan would take minutes.