LLMs Know a Little About Where Bugs Are, but That Doesn't Make Fuzzing Faster [llm-guided-fuzzing]
LLMs Know a Little About Where Bugs Are, but That Doesn't Make Fuzzing Faster [llm-guided-fuzzing]
Technical report, October 2026. Mostly negative results from using a language model’s hidden states to decide where a coverage-guided fuzzer should spend its time.
Reading transcripts of LLM agents hunting for bugs, I was surprised how often the agent says where the bug is before it has written a single test. It reads a function, points out that a length is never checked against the buffer it indexes, and only then writes the test that proves it. If a model can see that in the code, a fuzzer should be able to use it. Coverage-guided fuzzers such as AFL and libFuzzer treat all new coverage the same, so even a cheap score saying “this function looks buggy” could tell them where to spend their mutations.
So I built llmfuzz, a fuzzer that reads a language model’s hidden states for every function of a C/C++ program, turns them into a “looks buggy” score, and spends more of its budget where the score is high. That did not make fuzzing faster, and most of the later experiments were attempts to get a stronger signal out of the model. I tried bigger models, call-graph and AST context, and letting the model reason or write a doc comment before it answers. Finally I tested it on bugs fixed after the model was released, which it cannot have memorized.
The model does know a little about where bugs are, and on the recent bugs this is not memorization. But the effect is small. A function scores somewhat higher while it contains the bug than after the fix, yet different functions differ from each other by much more than that, and the plain rule “longer functions are more suspicious” ranked the bugs better. Even a perfect predictor would have made the fuzzer at most about twice as fast on the bugs I tried, because the fuzzer reached the buggy functions early and spent most of its time looking for the input that triggers the bug.
Summary
- A probe on Qwen3-4B’s hidden states separates vulnerable from non-vulnerable functions of similar length with an AUC of 0.57. On 13 known bugs in Google’s fuzzer-test-suite it puts the buggy function at a median percentile of 27, where ranking by function size alone gives 15.
- Weighting the fuzzer by the LLM’s scores did not make it faster on Linux. On two Linux machines the time to the first crash was 1.00× and 1.58× that of unweighted fuzzing (geometric mean of the per-target medians), while weights from function size alone gave 0.35×. The one large speedup, 8× on macOS, came from a single function that the LLM happened to rank 2nd.
- A perfect ranking helps at most about 2×. Putting the true buggy function first made woff2 crash 1.9× sooner and lcms no sooner, and random scores on the other functions erased even that.
- Context did not help. On four models from 4B to 27B parameters, none of ten kinds of context (callers, callees, call stacks, AST-selected definitions, argument slices and more) did better than the same amount of random code from the same file.
- Reasoning first helps a little. When Qwen3.6-27B first writes a doc comment for the function, it ranks the bugs higher (median percentile 35 → 20), and a comment written for another function makes it worse (50). No form of reasoning beat ranking by size, and the 27B’s reasoning openly recognizes famous bugs.
- On 148 single-function fixes committed after the model’s release, Qwen3.6-27B gives the buggy version a higher score than the fixed version 76% of the time, which neither code size nor the number of if-checks explains. Its own Yes/No answers are right for both versions in only 11% of the pairs.
How llmfuzz works
Scoring code with the J-lens
llmfuzz cuts a program into chunks, one per function, and splits functions longer than 150 lines into 150-line chunks. It runs each chunk through the model once, with a review question at the end:
// File: cmsintrp.c
... code of the chunk ...
/* Code review. Q: Is there a bug in the function above? A:Instead of letting the model answer, llmfuzz reads the Jacobian lens (J-lens) at the last position. The J-lens comes from Anthropic’s Verbalizable Representations Form a Global Workspace in Language Models (arXiv:2607.15495). It maps a layer’s residual stream into the final layer’s coordinates with the average Jacobian between the two, \(J_\ell = \mathbb{E}[\partial h_{\text{final}} / \partial h_\ell]\), and decodes the result with the model’s own unembedding: \(\mathrm{lens}_\ell(h_\ell) = \mathrm{lm\_head}(\mathrm{norm}(J_\ell h_\ell))\). Roughly, it shows what each layer is getting ready to say. I used the lenses that Neuronpedia has fit for Qwen3 and Qwen3.6.
I take two scores from the lens. The first needs no training: \(\log P(\text{yes}) - \log P(\text{no})\), averaged over the last quarter of the layers. I call it the zero-shot score. The second exists only for Qwen3-4B. It is a logistic-regression probe over the lens probabilities of 179 concept tokens (“overflow”, “bounds”, “unchecked”, …) at 12 layers, trained on 3,000 functions from PrimeVul, a dataset of C/C++ functions taken from vulnerability-fixing commits. Each vulnerable training function is paired with a non-vulnerable function from the same project whose length is within 15%, so the probe cannot get away with learning that long functions are vulnerable.
How the scores change fuzzing
The fuzzer is a libFuzzer-compatible SanitizerCoverage runtime of about 1,100 lines of C. llmfuzz standardizes the scores within each program and maps them to per-edge weights \(w = \mathrm{clip}(e^{z}, 0.5, 8)\). The weights are used in two places. A seed’s energy, which decides how often it is picked for mutation, is a weighted sum over the coverage features it was first to discover, \(\sum w^{\gamma} / \sqrt{1 + \text{hits}/256}\), so seeds that reach suspicious code are mutated more. Comparison-operand feedback (libFuzzer’s “value profile”) is turned on only in the top 10% of the code. One pick in ten ignores the weights, so a wrong ranking cannot shut the rest of the program out. Without a weight file neither mechanism is active, so the baseline is the same engine. That baseline holds its own against libFuzzer on the same harnesses, seeds and budget: it is faster on lcms and pcre2, slower on libcxxabi and about even on the rest.
How strong is the signal?
On PrimeVul’s test split the probe does better than chance, but not by much, and no better than counting lines.
| Test set | Probe | Zero-shot | Length only |
|---|---|---|---|
| Length-matched pairs | 0.569 | 0.548 | 0.5 |
| Natural distribution | 0.615 | 0.685 | 0.790 |
On the length-matched pairs, where length alone is at chance, the probe still scores above 0.5, so the hidden states carry some information beyond length. On the natural distribution that information is small next to length itself. Longer functions are more likely to contain vulnerabilities, and the zero-shot score does better there mainly because it correlates with length.
The fuzzing benchmark is seven targets from Google’s fuzzer-test-suite (FTS): libcxxabi, pcre2, lcms, woff2, OpenSSL 1.0.1f with Heartbleed, re2 and libxml2. They contain 13 known bug functions, and each target has 79 to 4,355 chunks.
The LLM puts 4 of the 13 bug functions in the top 10% of their target. Sometimes it is right where size is wrong: it ranks libcxxabi’s parse_floating_number 2nd of 79 chunks, while size puts it in the middle. It can also be far off, and over all 13 bugs ranking by size alone does better (median percentile 14.6 against 27.0).
Fuzzing with LLM weights
Every configuration runs the same engine and binary; only the weight file changes. A trial is a single process with a 900-second budget. I measure the time to the first crash, counting timeouts and out-of-memory errors as crashes and trials that find nothing as 900 s. Each target gets 6 to 10 trials per configuration.
| Weights | Linux | Linux, 2nd machine | macOS |
|---|---|---|---|
| LLM | 1.00× (2/5) | 1.58× (0/5) | 0.30× (3/3) |
| Size only | 0.35× (5/5) | – | 0.86× (2/3) |
| LLM + size | 1.14× (1/5) | – | 0.27× (3/3) |
| LLM shuffled | 1.36× (3/5) | – | 0.55× (1/3) |
| libFuzzer | 0.86× (3/5) | – | – |
Almost all of the macOS gain comes from libcxxabi, where the crash is in parse_floating_number, the function the LLM ranked 2nd. (The bug is a stack overflow that exists only on Apple arm64, where long double has 8 bytes but the demangler parses 16.) There, LLM weights found the bug in 0.6 s instead of 4.7 s (p = 0.001, Mann-Whitney U) and also beat randomly shuffled weights. On lcms the LLM weights helped on macOS (351 s against 828 s) and hurt on Linux (878 s against 486 s).
On Linux, weights from function size alone were the fastest, significantly so on libcxxabi (82 s against 837 s) and pcre2. Mixing in the LLM score undid most of that: with a quarter of LLM score added, libcxxabi was back at 796 s.
Wrong weights can be expensive. One shuffled ranking pushed lcms’s mutations toward slow inputs and cut the execution rate to a tenth (329 against 3,168 executions per second). It found the bug in 0 of 10 trials.
Two harder targets, libxml2 and re2, ran for 1,800 s. On libxml2 every weighted configuration (LLM, shuffled and size) did better than no weights, so there the weighting itself seems to help, and the LLM’s judgment has little to do with it. Final coverage of the code the LLM ranked highest was the same in every configuration; the weights only changed the order in which the fuzzer got there.
What a perfect ranking would buy
A ranking can fail because it is wrong, or because even a right one does not help. To tell the two apart, I replaced the LLM with an oracle that knows where the bug is and puts the crashing function at a chosen rank. In the simplest version, only the bug function is raised and all other functions get the same weight. In the noisy versions, the other functions get random scores, drawn again for every trial.
With the bug function first and the rest equal, woff2′s median went from 275 s to 143 s, 1.9× faster but not significant (p = 0.17). lcms got no faster. Its bug function sits in the inner loop of an interpolation routine, and ranking it first turned on comparison feedback there, which slowed execution by 26%.
Random scores on the other functions erase the gain. With the bug function still first but the rest random, woff2 was back at 257 s. With the bug function only somewhere in the top 25%, it was significantly slower than with no weights at all (736 s, p = 0.03).
Unweighted fuzzing reaches both bug functions early. The time goes into finding the particular combination of values that overflows the buffer, and putting more effort into inputs that reach the function does not produce it any sooner. Directed-fuzzing studies that use real bug locations found the same (LibAFLGo and AIJon, under related work). Even a perfect bug predictor, then, would have bought little on these bugs.
Adding context
Many vulnerabilities cannot be judged from one function. Whether an index is safe depends on what the callers pass in. So I put context in front of the function and repeated the ranking on four models: Qwen3-4B, and Qwen3-8B, Qwen3-14B and Qwen3.6-27B quantized to 4 bits. The call graph comes from the clang AST, starting at the fuzz harness.
The raw kinds of context, up to 3K tokens, are the first lines of every callee, the code around the call site in each caller, one call path from the harness down to the function, and the source text that precedes the function in its file. The selected kinds, up to 1.5K tokens per part, are the definitions of the macros, types, structures and globals the function uses, the statements in each caller that compute or check the call’s arguments (a backward slice by variable name), the callees’ signatures and comments, and a fixed-length window of surrounding code. As a control, the function gets random code from the same file, as long as the definitions.
| Context | Qwen3-4B | Qwen3-8B | Qwen3-14B | Qwen3.6-27B |
|---|---|---|---|---|
| Callees | 5–5 | 4–6 | 5–5 | 4–5 |
| Callers | 3–10 | 7–5 | – | – |
| Call stack | 4–5 | 9–2 | – | – |
| Call stack and callees | 7–6 | 8–5 | 8–4 | 7–5 |
| Preceding file text | 4–6 | 7–6 | 9–3 | 9–3 |
| Definitions | 5–6 | 5–5 | 7–3 | 5–4 |
| Caller argument slices | 6–4 | 6–4 | 6–3 | 5–4 |
| Definitions and slices | 7–4 | 8–5 | 10–2 | 8–4 |
| Callee signatures | 2–5 | 3–3 | 4–2 | 3–4 |
| Fixed window | 6–7 | 8–5 | 9–3 | 5–7 |
| Random code (control) | 6–4 | 7–4 | 9–0 | 6–3 |
Only two cells are significant. Both are for the 14B, and one of them is the random-code control. Given the function alone, the 14B ranked these bugs worse than chance, so almost any added code helps it, relevant or not. The best-looking number, a median percentile of 9.3 for the 27B with the preceding file text, became 30.2 with a fixed-length window instead. How much text comes before a function depends on where it sits in the file, and the 9.3 was probably an artifact of that.
There is an easy mistake to make here, and my first analysis made it. Context shifts the score of every chunk that receives it: same-length random code lowers the yes-no log-odds by 1.5 to 2.8. A bug function for which no context could be extracted keeps its unshifted score, and next to the shifted comparison chunks it looks more suspicious than it is. That first analysis reported a median percentile of 14.6 for the 4B with callees. Comparing only bug functions and comparison chunks that all received context gives 28.7.
Letting the model reason first
Everything so far reads a single forward pass, but the agents that impressed me reason before they answer. Maybe the knowledge only shows up once the model works through the code. I tried three ways of asking, in the chat format, on the same chunks.
- Answer directly. The question is “Does this function contain a bug (for example an out-of-bounds read or write, a use-after-free, an integer overflow or a logic error)? Answer with one word: Yes or No.” The lens is read where the answer starts.
- Think first. The model reasons, and then the lens is read at the same place. With a 512-token limit almost every chain of thought was cut off (91% for the 14B, all of them for the 27B), so I regenerated complete reasoning with vLLM, up to 8,192 tokens, and read the lens after it.
- Doc comment first. In a separate prompt, the model writes a short comment documenting the function: what it does, the preconditions on its arguments, and the invariants on buffer sizes and index ranges. It is told not to judge whether the code is correct. The comment goes above the code, and the model then answers directly. As a control, every function gets the comment the model wrote for another function of the same target.
| How the model is asked | Qwen3-14B | Qwen3.6-27B |
|---|---|---|
| Answer directly | 17.9% | 35.2% |
| Think first, ≤ 512 tokens | 27.8% (7–6) | 24.1% (5–8) |
| Think first, complete | 15.4% (7–6) | 18.3% (4–8)* |
| Doc comment first | 17.9% (7–5) | 20.1% (10–1) |
| Another function’s comment | – | 50.0% |
Writing a doc comment first helps the 27B. Its own comment moves the median from 35.2% to 20.1% (10 bugs up, 1 down, p = 0.01). Another function’s comment is worse than none (50.0%; own against other: 11–0, p = 0.001), so an extra comment by itself does not help. The comments often state the violated condition outright. For Heartbleed the model wrote payload (from message) must be <= s->s3->rrec.length - 3, which is exactly the missing check. The 14B gains nothing from its own comments.
Complete reasoning gives the 14B some information beyond size. Answering directly in this format, the 14B’s score mostly reflects function length: with size regressed out, the bugs sit at the 66th percentile, worse than chance. After complete reasoning they sit at the 33rd (12 up, 1 down, p = 0.003). Its own Yes/No answers also become more selective. It says Yes to 80% of the bug chunks and about half of the others, down from 100% and 86%. In its reasoning it names the actual cause for Heartbleed and for libxml2′s xmlDictComputeFastQKey (an index len - (plen + 2) that can go negative), and gives wrong reasons for most of the other bugs.
Reasoning that was cut off did not help either model, and complete reasoning did not help the 27B. None of these setups beat ranking by size: against size alone, the comment-first 27B is at 8–5, the 14B with complete reasoning at 5–8, and the 27B with complete reasoning at 5–8, none of them significant.
The 27B’s reasoning shows a worse problem. On these bugs it often recognizes the code instead of reasoning about it: “This is exactly Heartbleed.” “I found a reference to a bug in LittleCMS 2.0.” “Usually, these prompts come from a dataset of known bugs.” All 13 FTS bugs date from 2014 to 2017, so any of the results above may be partly memory.
Buggy and fixed versions of the same function
To separate knowing from remembering, I switched to pairs. Each pair is one function in two versions, just before and just after the commit that fixed a bug. The model is asked about each version separately, with the same prompt. A model that knows where the bug is should give the buggy version the higher score, and ideally say Yes to it and No to the fixed one.
The 148 new pairs are single-function memory-safety fixes from 30 active C/C++ projects (OpenSSL, curl, QEMU, FFmpeg, SQLite, PHP, radare2, libheif, Wireshark and others), committed between 2026-05-02 and 2026-10-08. Qwen3.6-27B was released on 2026-04-21, so it cannot have seen these fixes. I kept commits whose message describes a memory-safety fix and that change one function by at most 20 lines, then dropped 8 more after reading their messages. The 100 older pairs come from PrimeVul’s paired test split, at most 6 per project, and may well be in the training data.
The simplest rule wins, because fixes add code. “The version with less code is the buggy one” is right 85% of the time on both sets. That only works in a paired test: when you rank the functions of a code base, there is no fixed version to compare against.
The 27B knows something beyond that rule, and on the new bugs it cannot be memory. Answering directly, it scores the buggy version higher in 76% of the new pairs. It still does in 75% of the 24 pairs where the fix adds no code, and in 73% of the 82 pairs where the fix adds no if-check, where those rules are at or below chance. It does better on the new bugs than on PrimeVul (60%), plausibly because PrimeVul’s labels are noisy and the new fixes are mostly local missing checks. The 14B shows a signal only after complete reasoning (66%).
Its Yes/No answers barely tell the two versions apart. Answering directly, the 27B gets both versions right in 11% of the pairs and says No to both in 51%. With complete reasoning it gets both right in 22%, but it also says Yes to both in 66%, and the buggy version scores higher in only 62% of the pairs. Reasoning mostly made the model more willing to call code buggy.
When it gets a pair fully right, its reasoning often names the actual bug. In the 33 new pairs where the reasoning 27B said Yes to the buggy and No to the fixed version, its explanations mostly match the fix: a double free of bs in OpenSSL’s OCSP check, the underflow of get_width() - x0 in libheif, the overflow of sz * 4 in radare2′s r_str_escape_raw, an integer overflow in expat’s xcsdup. Across all pairs, its reasoning quotes a line from the fix site for 36% of the buggy versions and 14% of the fixed ones. Counting only fixes that insert lines, so that the quoted lines exist in both versions, the gap narrows to 40% against 29%.
So what I saw in the agent transcripts holds up. A strong model can spot a bug it has never seen from the code alone, and explain it. The transcripts just don’t show how often the same model is equally sure about code that is fine.
Why this does not make a better fuzzer
Putting the results together, I don’t expect bigger models or better prompts to fix this.
- The score shifts up a little when the same function contains a bug. A fuzzer has to compare thousands of different functions, and functions differ in size, complexity and style by much more than that shift. This is why every ranking in this report loses to function size, and why the model’s Yes/No answers mean little on their own.
- Finding the function is not the bottleneck. On these bugs the oracle experiment puts the value of a perfect ranking at about 2× at most. The fuzzer reaches the buggy functions early; what remains is constructing the values that trigger the bug.
- Wrong weights can cost more than right weights gain. A bad ranking can trap the fuzzer on slow inputs or turn on expensive feedback in a hot loop, and a ranking that was merely good (the bug in the top 25%) was already slower than no weights.
Other uses of the model might still work.
- Let the model state a concrete bug, this line under this condition, and have a test or a fuzzer confirm or refute it. That is where the model is at its best. A wrong guess costs one test, and a right one points straight at the trigger.
- Use the before and after that every commit provides. Whether the paired signal can flag commits that introduce a bug is untested here, and the oracle result would still limit how much fuzzing time it saves.
- Have the model write input generators and trigger conditions. In the literature, that is where the clear wins for LLMs in fuzzing come from.
Related work
Probing hidden states for vulnerabilities
Probes for C/C++ vulnerability detection are weak and often pick up surface features. Probing the Prefill (2026) trains an MLP probe on the last-layer states of models up to 27B and reaches an F1 of 17.6 to 19.5 on PrimeVul, below the 24.5 of a fine-tuned small model; it reports no AUC and does not control for length. Activation Probes Surface Code-Security Signals (ICML 2026 workshop) reaches 61 to 67% pairwise accuracy on Python and reports that the C/C++-dominated categories are near chance. In the ablations of SAGE (ISSTA 2026), a linear probe gets about 74% balanced accuracy on PrimeVul without length control and about 50% on the time-split PreciseBugs. Code Correctness Is Linearly Decodable reports an AUC of 0.88. When I recomputed it from the authors’ released data, problem difficulty alone gave 0.82, and the probe kept about 0.65 once difficulty and length were controlled for.
Function-level vulnerability detection is near its ceiling
In PrimeVul’s (ICSE 2025) paired setting, which requires both the vulnerable and the fixed version to be classified correctly, GPT-4 scores 12.9% where random guessing scores 22.7%. Weissberg et al. (ICSE 2026) find that a classifier on 23 code metrics reaches an F1 of 20.3 against 20.7 for the best fine-tuned model, and that LLM judgments causally follow those metrics. In SAGE’s tables, frontier models reach an MCC of 0.08 to 0.14 on PrimeVul and about 0 on new data.
Context
Risse et al. (ISSTA 2025) find that 43 to 52% of real vulnerabilities depend on called functions and 40 to 47% on the arguments callers pass. Pasting raw caller and callee code makes things worse: all six settings of CPRVul degrade, and in Lira et al. (EASE 2026) GPT-4.1-mini on C drops from 75.6% to 51.0%. VulEval improves only with oracle context, the functions that the fix also touched. Selected or summarized context helps a little, as in LLMxCPG (USENIX Security 2025) and PacVD. I am not aware of earlier work combining context with hidden-state probes; here, selected context was no better than random code.
Guiding fuzzers by bug likelihood
Guiding a fuzzer by predicted bug likelihood is older than LLMs: V-Fuzz, TortoiseFuzz (NDSS 2020), SAVIOR (S&P 2020), ParmeSan (USENIX Security 2020) and AFLChurn (CCS 2021). LLM4Fuzz multiplies an LLM’s per-function score into seed energy, almost the same mechanism as llmfuzz, and its own ablation favors the complexity component over the vulnerability component. Several systems in the AIxCC final used LLMs to rank suspicious functions for their fuzzers (see the AIxCC SoK). Attention Distance (ICSE 2026) guides directed fuzzing with LineVul’s attention and is 3.43× faster than AFLGo, but has no random or code-metric control. SoK: Where to Fuzz? (AsiaCCS 2024) compares target-selection methods on 1,621 OSS-Fuzz crashes: code metrics rank best, and LineVul is close to random.
Localization is not the bottleneck
With real bug locations on Magma, no directed strategy in LibAFLGo (EuroS&P 2025) beat undirected LibAFL. AIJon (2026) places LLM-written IJON annotations at the real bug locations and triggers 34 bugs, against 38 for AFL++. Perera et al. (TSE 2023, TOSEM 2024) ran “how accurate must a defect predictor be” experiments with oracles and controlled noise for search-based testing in Java; fuzzing does not yet have a complete version of that study.
Where LLMs help fuzzing
The clear wins come from generating inputs. G2Fuzz (USENIX Security 2025) has an LLM write input generators that AFL++ then mutates, for under $0.2 of model calls per 24 hours. When the target is known, Locus (ICSE 2026) has an LLM synthesize trigger predicates as feedback and is 41.6× faster on Magma. HyLLfuzz uses an LLM in place of a constraint solver when the fuzzer stalls. All three call the model rarely. Guiding every mutation with a neural network, on the other hand, did not beat AFL++ under a fair evaluation (Revisiting Neural Program Smoothing, ESEC/FSE 2023).
If you try this
- Include a baseline that only looks at size or complexity. It beat the LLM here, and it often does in the literature.
- When you add context, add a control with random code of the same length, and compare only chunks that all received context. Context shifts the score, and comparing chunks with and without it creates effects that are not there.
- Measure what a perfect predictor is worth before building one. Here it was about 2×.
- Test on bugs fixed after the model’s release, and compare buggy and fixed versions of the same function instead of ranking a handful of famous bugs.
- With 13 bugs, medians swing wildly. Definitions as context moved the 27B’s median from 27% to 6%, with 5 bugs up and 4 down. Use paired tests.
Limitations
- The ranking and fuzzing results rest on 13 bug functions in 7 FTS targets. Fuzzing trials ran 900 s (1,800 s for libxml2 and re2) with 6 to 10 trials per configuration, on my own engine; I did not reproduce them on AFL++.
- The engine’s seed energy ignored execution time. Adding AFL’s speed factor afterwards raised throughput 2 to 11 times on SQLite and libxml2, so a stronger engine would find these bugs sooner, and the relative order of the weightings might change.
- The oracle experiment covers 2 targets with 8 trials each.
- Qwen3-8B, 14B and 27B were used zero-shot (without a trained probe) and quantized to 4 bits. The complete reasoning was generated with a different 4-bit checkpoint (AWQ) than the one used for scoring, each function’s reasoning was sampled once, and 20% of the 27B’s reasoning on the FTS subsample (30% in the paired test) hit the 8,192-token limit.
- The lenses were fit on wikitext.
- The new-bug pairs were selected from commit messages. A few of the “fixes” may only be hardening, and PrimeVul’s labels are known to be noisy.
About this report
The fuzzer, the analysis scripts and the raw results of every trial are in the llmfuzz repository (bench/ and bench/results/). The lenses are Neuronpedia’s jacobian-lens collection; the models are Qwen3 and Qwen3.6. I wrote most of the code and ran most of the experiments with Claude Code.