Fuzz target selection has always been a knowledge problem. You pick harness sites by asking where input transforms in ways the author might have under-specified: parsers, deserializers, arithmetic on external values, state machines driven by wire data. The implicit assumption is that a human author made deliberate decisions at those boundaries, and the fuzzer's job is to stress-test those decisions.
That assumption breaks when the code was generated by a language model.
What Actually Changes With AI-Generated Code
An LLM generates code by predicting structurally plausible completions given a prompt and context. It does not reason about threat models. It does not carry forward any author's intuition about which paths are dangerous. What it produces is code that looks reviewed: syntactically conventional, idiomatically correct, compiles cleanly. The risk concentrates somewhere different than it does in human-written code.
In human-written code, dangerous boundaries tend to cluster near the parts the author was least confident about. The sections revised multiple times. The comments that apologize. The late additions made under deadline pressure. These are discoverable heuristics, and code review history and coverage both correlate with them.
In AI-generated code, there is no such signal. A function that handles a subtle integer promotion edge case looks exactly like one that handles it correctly. Both are fluent. Both pass a style linter. The model did not hesitate, and there is no hesitation to observe. Conventional fuzz target selection heuristics rely on locating human uncertainty. When the code has no human uncertainty baked into it, those heuristics are systematically mis-aimed.
Where Coverage Blind Spots Migrate
A concrete illustrative example makes this tangible. Suppose a spec says:
parse_record(buf: &[u8]) -> Result<Record, ParseError>Precondition:bufcontains a length-prefixed UTF-8 string followed by a four-byte little-endian integer. The string length field is a two-byte unsigned integer. ReturnsErron any malformed input.
A human author writing this would probably worry about the length field exceeding the buffer, the string failing UTF-8 validation, and the trailing integer being misaligned. They might add explicit guards for some of these and quietly rely on a downstream library for others. A coverage-guided fuzzer aimed at this function would eventually find the gaps.
An LLM completing the same function from the same spec might produce:
pub fn parse_record(buf: &[u8]) -> Result<Record, ParseError> {
if buf.len() < 2 {
return Err(ParseError::TooShort);
}
let str_len = u16::from_le_bytes([buf[0], buf[1]]) as usize;
let str_bytes = buf.get(2..2 + str_len).ok_or(ParseError::TooShort)?;
let s = std::str::from_utf8(str_bytes).map_err(|_| ParseError::InvalidUtf8)?;
let int_offset = 2 + str_len;
if buf.len() < int_offset + 4 {
return Err(ParseError::TooShort);
}
let value = i32::from_le_bytes(buf[int_offset..int_offset + 4].try_into().unwrap());
Ok(Record { name: s.to_owned(), value })
}
This looks correct on first read. The spec says the string field is UTF-8, and there is a guard for that. What the spec does not say, and what the model was never asked to check, is whether str_len values near usize::MAX cause an integer overflow in 2 + str_len before the slice bounds check fires. On a 32-bit target, or after a future refactor that changes str_len's type, that unwrap() on try_into() is also a latent panic site.
A human author would likely have had a nagging feeling about the length arithmetic. The model had no such feeling. A coverage heuristic looking for "functions modified recently under pressure" or "functions with low review depth" would not flag this function at all.
Deriving Fuzz Targets From the Spec and Call Graph Together
The right answer is to stop treating author intuition as a proxy for risk location. The primary signal should be the gap between the spec's stated preconditions and the code's actual guard clauses.
This requires two artifacts that are often treated as independent:
The specification. Formal preconditions if you have them, but also structured comments, OpenAPI schemas, protobuf definitions, type signatures with invariants, and any document that states what the function should accept and reject. Even a brief doc comment is a spec.
The call graph with data-flow edges. Specifically, which functions are reachable from a trust boundary (a network socket, a file read, an IPC endpoint, a CLI argument parser) and how far each function sits from that boundary.
The heuristic that follows from combining these two sources:
For each function reachable within k hops of a trust boundary, enumerate the preconditions stated in its spec. For each precondition, check whether a corresponding guard clause exists in the function body. Rank functions by the number of preconditions that lack corresponding guards, weighted by distance from the trust boundary (closer is higher risk). Treat functions above a threshold as mandatory harness sites.
This is not fully automatable today without tooling investment, but the manual version is tractable as a triage exercise:
- Extract all public or exported functions reachable from your trust boundaries using a static call graph tool, or a grep-based approximation for a first pass.
- For each function, list its documented preconditions. If there are no documented preconditions for a function that accepts external input, that itself is a finding: the spec gap is the risk.
- Scan the function body for guard clauses: early returns on length checks, range checks, type validation, explicit panics with explanatory messages. Count how many spec preconditions have a visible corresponding guard.
- Rank by
(preconditions_without_guards) / (hops_from_trust_boundary + 1). Higher scores mean the function is close to untrusted input and its behavior is under-specified relative to its guards.
Functions at the top of this ranked list are your mandatory harness sites. Functions with zero documented preconditions that accept a raw byte slice or an untyped string near a trust boundary go directly to the top regardless of score. The missing spec is itself evidence of a blind spot.
Why This Heuristic Works for AI-Generated Code Specifically
The heuristic is not new in principle. Verification engineers have always wanted to align harness sites with under-specified behavior. What is new is the justification for making it the primary mechanism rather than a secondary check.
For human-written code, author intuition is a weak but real signal. You can read the commit history, look at which tests were written defensively, see where comments hedged. Noisy proxies, but not zero.
For AI-generated code, those proxies return noise. The commit history shows a single large commit from a generation step. There is no author to interview. The tests may themselves be AI-generated and share the same blind spots as the implementation.
The spec-to-guard-gap heuristic is grounded in artifacts that are independent of who or what wrote the code. The spec was written by a human, or extracted from a schema a human designed. The guard clauses are observable in the AST. Neither depends on author knowledge.
Practical Constraints and Honest Limitations
This approach assumes you have a spec. For many teams adopting AI code generation, the prompt is the spec, and it may be informal. If your preconditions live only in a chat history, you will need to reconstruct them before this heuristic applies. That reconstruction step is itself valuable: if you cannot state the preconditions for a function that processes external input, you do not yet know what to fuzz for.
The heuristic also does not eliminate the need for coverage feedback. Ranking tells you where to aim harnesses. Coverage-guided fuzzing tells you whether those harnesses are reaching the risky paths once deployed. Treat the two as sequential: use the spec-gap ranking to decide which functions to harness, then use coverage to verify the harnesses are exercising the interior branches you care about.
Finally, this is not a substitute for reviewing AI-generated code. It is a backstop for the review gaps that are statistically certain to exist at scale. If your team is generating thousands of lines per week with an LLM, complete review coverage is not achievable. The heuristic gives the verification layer a principled way to concentrate effort where the risk is highest, independent of whether review happened.
The Triage Algorithm, Summarized
To apply this concretely when onboarding a new AI-generated module:
- Map trust boundaries. Identify every point where the module accepts data from outside the process.
- Build a reachability set. Starting from each trust boundary, enumerate all functions reachable within a configurable hop limit (start with 5).
- Extract stated preconditions. For each function in the reachability set, list preconditions from docstrings, type contracts, or adjacent schema definitions.
- Audit guard clauses. For each precondition, determine whether a corresponding guard is present in the function body. A guard is any code that causes an early return or panic before the unsafe operation implied by the precondition.
- Score and rank.
score = unguarded_preconditions / (hop_distance + 1). Sort descending. - Flag zero-precondition functions near trust boundaries. These are spec absences. In AI-generated code, spec absences are not benign omissions; they are locations where the model filled in behavior that was never specified.
- Write harnesses for all functions above your threshold. The threshold is a policy decision, but starting with all functions scoring above zero and within two hops is a defensible default.
The goal is a harness selection process that does not depend on knowing what the developer was thinking, because for a significant and growing fraction of your codebase, no developer was thinking it at all.
If you are building or evaluating verification infrastructure for pipelines that include AI-generated code, OpenThunder Dev covers the tooling layer in depth.