By Louis Abraham and Nicolas Devatine 5 min read

We are releasing LLM Fuzz CI, the first CI tool that lets LLM agents attack the tests you wrote, and only reports inputs that fail them.

LLM Fuzz CI is a GitHub Action that lets a coding agent read your code and write adversarial inputs for the tests you mark. Your normal test run replays them, so every report is a real failure of your own assertions.

Code agents have become really good at programming, including at finding bugs (do not get me started about agents creating bugs). Sadly, agents are also extremely good at logorrhea producing torrents of text, flooding programmers with an endless stream of information of seemingly utmost importance. Security scanners are even worse: they drown CISOs in CVE reports rated “medium”, which programmers quickly dismiss as minor, and they never test the business logic anyway.

We tasked ourselves with solving this problem by creating a tool that follows these principles:

  1. the programmer should define (even with the help of agents) the behavior that should be verified using tests. (*)
  2. any bug the tool finds is a real violation of your tests: the agent only writes inputs, and your normal test run replays them. No false positive.
  3. the tool should be launched on code changes and/or on a regular basis (as packages get updates, models get smarter and non-obvious CVEs get published).

(*) As more and more code will be written by LLM, possibly most of it, the role of programmers will actually stay the same: converting business requirements into specs, and producing systems that abide those specs. Both steps need technical skill: writing unambiguous specs (a few rounds of prompting and some technical claude-linguo), defining a correct way to verify them, and spotting any gap between the original intention and the produced system. Even with the best coding agents, non-technical humans still struggle with both steps: without the technical language, they cannot discuss their own specs well.

The solution was pretty natural: have the programmer mark tests as targets and then run an agent to try to fuzz them.

How it works

You mark a test, set a spend limit with budget_usd, and optionally restrict the agent to some input keys with params. We supports Python with pytest and JavaScript with vitest.

import pytest

@pytest.mark.llm_fuzz(budget_usd=0.5)
def test_foo(llm_fuzz_case):
    result = foo(**llm_fuzz_case.input)
    assert "<script>" not in result
import { expect } from "vitest";
import { fuzzTest } from "llm-fuzz-ci";

fuzzTest("foo escapes its input", { budgetUsd: 0.5 }, (input) => {
  expect(foo(input.value)).not.toContain("<script>");
});

The agent will be tasked with replacing llm_fuzz_case.input / input.value with a value that will make the test fail.

You can then configure a CI action that will launch the tool in two phases, on two separate machines. First, a coding agent (by default Codex with an OpenAI or OpenRouter key, but we also support Claude Code) reads the code and the marked tests, tries to break them, and when succeeding, exports the inputs.

Second, a fresh runner checks out your code again, downloads those inputs, and replays them through your normal test run. The agent is free to do whatever it wants on the first machine: run your code, install things, edit files to understand them. None of it reaches the second one. The only thing that crosses is a file of inputs.

Finally, every run writes a summary, an artifact with the inputs and the agent’s trace, and optionally creates a GitHub issue and can fail the build.

You can find more information in the README.md.

How LLM Fuzzing is going to change testing

Fuzz testing, or fuzzing, originally means throwing a lot of unexpected inputs at code and watching what breaks. Classic fuzzers start from a seed input and mutate bytes at random, sometimes guided by coverage, then look for crashes, timeouts or surprising results. That works well for low-level code. It struggles with higher-level code: the fuzzer does not know what your function is supposed to do, which parameters interact, or how to tell a logic bug from a legitimate corner case. Checks that depend on meaning slip through because the random generator never hits the interesting combinations. An LLM agent can read the code and the test, and aim straight at those combinations.

Mixing LLM and fuzzing is not new. OSS-Fuzz-Gen from Google and PromptFuzz have LLMs write fuzz targets for C and C++ libraries. Garak and promptfoo fuzz LLMs themselves, not your code. XBOW (hosted pentesting with verified exploits), OpenAI’s Aardvark and Google’s Big Sleep hunt for vulnerabilities on their own, without your tests as the judge. But LLM Fuzz CI is the first tool that:

  • runs in CI, so you can schedule regular checks or gate releases
  • uses a separate step to generate inputs, which (mostly) prevents reward hacking
  • and plugs into pytest and vitest, so programmers define what to test in a familiar way

Unlike classic security scanners, LLM Fuzz CI has the ability to find advanced vulnerabilities by actually reading and understanding the code. Unlike “PR-review” tools, no noise is created: this is a tool for programmers to enforce guarantees of their code.

Generating arbitrary inputs enables use cases not covered by standard fuzzing tools, like attacking LLM agents through their prompts. We first used this tool to test an accounting agent on a testing branch, and discovered that when prompted with a valid account id different from the user’s, it would launch a tool that erroneously allowed access to data of the other account. Smart testing agents are thus able to find and demonstrate complex bugs (even using prompt engineering) from end-to-end tests and not just unit tests. This makes LLM Fuzz CI a nice complement to another testing concept we released last year: cached stubs turn end-to-end pipelines into fast, reproducible tests by caching non-deterministic steps and external calls.

A note on open source releases

Code has become a commodity, and any recent coding agent should be able to replicate our repo.

Rather than a tragedy, this is fine: nobody needs to care about our particular implementation. It is easy to replicate, or even copy (please do), into the frameworks that will spread it. This new age puts even more emphasis on concepts. The core of innovation has always been novel ideas. Technical implementation was both a gate, because it took technical skills to demonstrate those ideas, and a fun way to spend time (at least for me).

More than a tool, this release is about a novel, simple and useful paradigm of how to use code agents well to protect the quality of software and the time of maintainers.