NixBench

tasks
29
areas
9
evaluators
29
current trials
72

Can AI coding agents write Nix that actually passes?

Objective repository-repair tasks scored by hidden shell evaluators—not by whether the output merely looks plausible.

Current release: 29-task corpus · one hidden evaluator per task · uncertainty shown when repeat trials exist

Compare the signal first. Inspect the scatter second.

Ordered effort paths lead the view. Select a model for uncertainty and effort labels, or reveal every trial to inspect the underlying variation.

Corpus
Evidence
Y-axis
5 models · 16 configurations · 72 trials16 means shown14/16 replicated

Configuration means, with uncertainty

Mean tasks passed against mean agent seconds per task. Paths connect ordered effort configurations; select a model to reveal its 95% Student's t intervals and effort labels. The task axis focuses on 6–29 to make the observed differences legible. The time axis uses logarithmic spacing because observed runtimes span more than one order of magnitude. Individual trials are hidden in this summary view.

Focused: 629 tasksLog time axis↑ more tasks← less time
95% CI / observed
GPT-5.6 Sol via Codex CLI29-task corpus · 5 recorded trials · 4m timeoutmediumn=524.0 / 29
22.5–25.5observed 222538.5s18m 36s / corpus0
GPT-5.6 Sol via Codex CLI29-task corpus · 5 recorded trials · 4m timeoutxhighn=524.0 / 29
23.1–24.9observed 232560.7s29m 20s / corpus0
GPT-5.6 Sol via Codex CLI29-task corpus · 5 recorded trials · 4m timeoutlown=523.6 / 29
22.2–25.0observed 222527.4s13m 14s / corpus0
GPT-5.6 Sol via Codex CLI29-task corpus · 5 recorded trials · 4m timeouthighn=523.6 / 29
22.9–24.3observed 232447.4s22m 53s / corpus0
GPT-5.6 Luna via Codex CLI29-task corpus · 5 recorded trials · 4m timeoutmediumn=523.2 / 29
22.2–24.2observed 222430.5s14m 45s / corpus0
GPT-5.6 Sol via Codex CLI29-task corpus · 5 recorded trials · 4m timeoutmaxn=523.2 / 29
20.1–26.3observed 192581.4s39m 19s / corpus3
GPT-5.6 Luna via Codex CLI29-task corpus · 5 recorded trials · 4m timeoutlown=523.0 / 29
22.1–23.9observed 222426.6s12m 51s / corpus1
GPT-5.6 Terra via Codex CLI29-task corpus · 5 recorded trials · 4m timeouthighn=522.8 / 29
20.8–24.8observed 202438.4s18m 34s / corpus0
GPT-5.6 Terra via Codex CLI29-task corpus · 5 recorded trials · 4m timeoutxhighn=522.8 / 29
21.8–23.8observed 222448.8s23m 35s / corpus0
GPT-5.6 Luna via Codex CLI29-task corpus · 5 recorded trials · 4m timeouthighn=522.6 / 29
21.2–24.0observed 212444.4s21m 29s / corpus0
GPT-5.6 Luna via Codex CLI29-task corpus · 5 recorded trials · 4m timeoutxhighn=522.4 / 29
21.7–23.1observed 222353.9s26m 2s / corpus0
GPT-5.6 Luna via Codex CLI29-task corpus · 5 recorded trials · 4m timeoutmaxn=522.4 / 29
20.7–24.1observed 212475.0s36m 16s / corpus4
GPT-5.6 Terra via Codex CLI29-task corpus · 5 recorded trials · 4m timeoutlown=522.0 / 29
21.1–22.9observed 212325.4s12m 16s / corpus0
GPT-5.6 Terra via Codex CLI29-task corpus · 5 recorded trials · 4m timeoutmediumn=521.8 / 29
20.2–23.4observed 202328.4s13m 44s / corpus0
Gemma 4 26B-A4B QAT Q4_0 via OpenCode + llama.cpp29-task corpus · 1 recorded trial · 1h timeoutdefaultsingle run10.0 / 29
unavailableobserved 10101869.1s15h 3m / corpus3
Ternary Bonsai 27B Q2_0 via OpenCode + llama.cpp29-task corpus · 1 recorded trial · 1h timeoutdefaultsingle run7.0 / 29
unavailableobserved 772478.7s19h 58m / corpus4
  1. GPT-5.6 Sol via Codex CLI29-task corpus · 4m timeoutmedium
    Evidence
    n=5
    Mean tasks
    24.0/29
    95% CI
    22.5–25.5
    Seconds / task
    38.5s
    Observed 2225 tasks · 0 timeouts
  2. GPT-5.6 Sol via Codex CLI29-task corpus · 4m timeoutxhigh
    Evidence
    n=5
    Mean tasks
    24.0/29
    95% CI
    23.1–24.9
    Seconds / task
    60.7s
    Observed 2325 tasks · 0 timeouts
  3. GPT-5.6 Sol via Codex CLI29-task corpus · 4m timeoutlow
    Evidence
    n=5
    Mean tasks
    23.6/29
    95% CI
    22.2–25.0
    Seconds / task
    27.4s
    Observed 2225 tasks · 0 timeouts
  4. GPT-5.6 Sol via Codex CLI29-task corpus · 4m timeouthigh
    Evidence
    n=5
    Mean tasks
    23.6/29
    95% CI
    22.9–24.3
    Seconds / task
    47.4s
    Observed 2324 tasks · 0 timeouts
  5. GPT-5.6 Luna via Codex CLI29-task corpus · 4m timeoutmedium
    Evidence
    n=5
    Mean tasks
    23.2/29
    95% CI
    22.2–24.2
    Seconds / task
    30.5s
    Observed 2224 tasks · 0 timeouts
  6. GPT-5.6 Sol via Codex CLI29-task corpus · 4m timeoutmax
    Evidence
    n=5
    Mean tasks
    23.2/29
    95% CI
    20.1–26.3
    Seconds / task
    81.4s
    Observed 1925 tasks · 3 timeouts
  7. GPT-5.6 Luna via Codex CLI29-task corpus · 4m timeoutlow
    Evidence
    n=5
    Mean tasks
    23.0/29
    95% CI
    22.1–23.9
    Seconds / task
    26.6s
    Observed 2224 tasks · 1 timeouts
  8. GPT-5.6 Terra via Codex CLI29-task corpus · 4m timeouthigh
    Evidence
    n=5
    Mean tasks
    22.8/29
    95% CI
    20.8–24.8
    Seconds / task
    38.4s
    Observed 2024 tasks · 0 timeouts
  9. GPT-5.6 Terra via Codex CLI29-task corpus · 4m timeoutxhigh
    Evidence
    n=5
    Mean tasks
    22.8/29
    95% CI
    21.8–23.8
    Seconds / task
    48.8s
    Observed 2224 tasks · 0 timeouts
  10. GPT-5.6 Luna via Codex CLI29-task corpus · 4m timeouthigh
    Evidence
    n=5
    Mean tasks
    22.6/29
    95% CI
    21.2–24.0
    Seconds / task
    44.4s
    Observed 2124 tasks · 0 timeouts
  11. GPT-5.6 Luna via Codex CLI29-task corpus · 4m timeoutxhigh
    Evidence
    n=5
    Mean tasks
    22.4/29
    95% CI
    21.7–23.1
    Seconds / task
    53.9s
    Observed 2223 tasks · 0 timeouts
  12. GPT-5.6 Luna via Codex CLI29-task corpus · 4m timeoutmax
    Evidence
    n=5
    Mean tasks
    22.4/29
    95% CI
    20.7–24.1
    Seconds / task
    75.0s
    Observed 2124 tasks · 4 timeouts
  13. GPT-5.6 Terra via Codex CLI29-task corpus · 4m timeoutlow
    Evidence
    n=5
    Mean tasks
    22.0/29
    95% CI
    21.1–22.9
    Seconds / task
    25.4s
    Observed 2123 tasks · 0 timeouts
  14. GPT-5.6 Terra via Codex CLI29-task corpus · 4m timeoutmedium
    Evidence
    n=5
    Mean tasks
    21.8/29
    95% CI
    20.2–23.4
    Seconds / task
    28.4s
    Observed 2023 tasks · 0 timeouts
  15. Gemma 4 26B-A4B QAT Q4_0 via OpenCode + llama.cpp29-task corpus · 1h timeoutdefault
    Evidence
    single run
    Mean tasks
    10.0/29
    95% CI
    unavailable
    Seconds / task
    1869.1s
    Observed 1010 tasks · 3 timeouts
  16. Ternary Bonsai 27B Q2_0 via OpenCode + llama.cpp29-task corpus · 1h timeoutdefault
    Evidence
    single run
    Mean tasks
    7.0/29
    95% CI
    unavailable
    Seconds / task
    2478.7s
    Observed 77 tasks · 4 timeouts

Corpora are intentionally separated and time is normalized per task. The focused y-axis is explicitly labelled; Full scale restores the zero baseline. Lines show configuration order from lower to higher effort; they do not imply continuous scaling or monotonic treatment. See the reproducibility method and local OpenCode run provenance. Raw run IDs are shown in trial tooltips.

Current trial environments: 2 agent versions · 2 hosts · 2 repository revisions · network unknown · timeout budgets 240s, 3600s.

Plausible Nix often fails at evaluation time.

The benchmark gives agents a copied starter tree, a prompt, and no access to the hidden evaluator. It rewards final worktree behavior, not a fluent explanation of what the code should do.

Original repair tasks

Tasks are written for this corpus rather than lifted from merged patches, which keeps the answer out of the visible prompt.

Nix-specific failure surfaces

The corpus covers flakes, modules, overlays, derivations, fetchers, Home Manager, shell escaping, and package contracts.

Hand-written checks

Each task has a shell evaluator that checks behavior with small fake package sets and libraries instead of relying on LLM judging.

Diff-backed runs

Every run records logs, timings, pass state, score JSON, and the final diff so failures can be inspected after the benchmark ends.

29 small repositories, one hidden evaluator each.

All 29 tasks

Respect NixOS, Home Manager, and nix-darwin boundaries

Keep module outputs separated instead of leaking options across systems.

Patch Python CUDA package inputs

Repair Python/CUDA packaging without falling back to generic Linux path guesses.

Compose module paths from arguments

Build paths with Nix values while avoiding string interpolation traps.

Debug network symptoms without false leads

Explain the observed NixOS service behavior without chasing a plausible but wrong network diagnosis.

Manage home files declaratively

Use Home Manager file and XDG options rather than imperative setup.

Pin a GitHub source fetcher

Preserve the fixed-output fetcher contract with a commit pin and SRI hash.

7

easy tasks for syntax, lookup, stale options, and small contracts

16

medium repairs across flakes, containers, issue reports, overlays, packaging, and shell integration

6

hard tasks for modules, overlays, portals, Rust purity, and Python/CUDA package inputs

The agent edits a worktree. The evaluator scores the result.

A simple, inspectable protocol separates generation from evaluation.

Read the protocol
  1. copy

    Starter files and the prompt enter a clean temporary workdir.

  2. edit

    The agent reads NIXBENCH_PROMPT.md and modifies only local files.

  3. check

    A hidden shell evaluator scores the final tree after the agent exits.

  4. record

    Logs, timing, score JSON, and the final diff are written under results/.

Add evidence, not another claim.

Use the open harness, preserve the run artifacts, and add a comparable row to the benchmark.

Open the run guide