How best-of-K sampling works for LLM code generation
Best-of-K asks a model for several programs and keeps the best one. Finding a correct draft is the easy half. Recognising it is where most of the work goes, ours included.
Ask a language model for the same function twice and you will often get two different programs. Sometimes one is right and the other is wrong. Best-of-K sampling, also called best-of-N, takes advantage of that: draw several drafts, check them, keep the best.
why asking twice helps
A language model writes code one token at a time, and at each step it chooses among likely continuations with some randomness. Raise the temperature and the choices spread out. Lower it and the model repeats its favourite path. Either way, the output is a sample from a distribution over programs.
That distribution often contains correct programs that the model does not produce every time. A model that solves a task in three attempts out of ten knows how to solve it and does so inconsistently. Best-of-K puts that knowledge to use without retraining anything.
Code suits the method unusually well, because a program can be run. Prose rarely offers a check as direct as executing a function and comparing its output. The founding paper for this line of work, the Codex report from 2021, built its whole evaluation on that property: hand-written unit tests that a draft either passes or fails.
measuring it with pass@k
The Codex report also defined the measure most people still use. pass@k is the share of problems where at least one of k drafts passes the tests. (The report calls each draft a sample.)
Measuring it directly, by drawing exactly k drafts, gives noisy numbers. So the report's estimator draws more, n per problem, counts the c that pass, and asks a cleaner question. If you picked k of those n at random, how likely is it that none would be correct?
pass@k = 1 − C(n − c, k) / C(n, k)
Here C(a, b) counts the ways to choose b items from a. The fraction is the chance that a random set of k drafts holds no correct one, and one minus it is the chance that it holds at least one.
Take ten drafts of one function, three of them correct.
| k | pass@k | worked out |
|---|---|---|
| 1 | 0.30 | 1 − 7/10 |
| 2 | 0.53 | 1 − 21/45 |
| 3 | 0.71 | 1 − 35/120 |
| 5 | 0.92 | 1 − 21/252 |
| 8 | 1.00 | only 7 drafts are wrong |
Three drafts take the chance of holding a correct one from 30% to about 71%. That arithmetic assumes the drafts differ from one another. If all ten were copies of a single program, every row would read 0.30.
coverage and selection
pass@k answers one question: is there a correct program somewhere in the pile? Call that coverage. A system that has to hand back a single answer faces a second question: can it tell which program is the correct one? Call that selection.
pass@k assumes perfect selection. It counts a problem as solved if any draft passes the hidden tests, which only an oracle could know in advance. That is why AlphaCode, DeepMind's 2022 competitive-programming system, reported a stricter measure called n@k: the share of problems solved using n submissions chosen from k samples.
The Codex report shows how far apart the two can be. For its strongest model on its own benchmark:
A correct program was in the pile for more than three problems in four. Picking by the model's own confidence delivered fewer than half. Most research in this area tries to close that gap, and ATLAS has to close it too.
a short history of choosers
Read in order, the work on best-of-K is a series of better ways to choose.
The simplest chooser ranks drafts by the model's average token probability. It costs nothing and it beats a random pick, but the Codex numbers above show how far it falls short of an oracle.
AlphaCode took a different route in 2022. It sampled at enormous scale, up to millions of programs per problem, then ran them on the small examples printed in each problem statement. That filter removed about 99% of samples, and for roughly one problem in ten nothing survived at all. Among the survivors, AlphaCode generated fresh test inputs, grouped the programs that produced identical outputs, and submitted one program from each of the largest groups. The reasoning is that programs which behave alike are probably the same program, and a large group suggests a common, correct idea.
Shi and colleagues (2022) made that behavioural idea general: run every draft on a few inputs and prefer the one whose results agree most with the others. CodeT, the same year, asked the model to write the tests too, then scored each draft by the tests it passed and by how many other drafts agreed with it. This needs no human tests. Its weak point is that, to write the expected output of a test, the model must already solve the problem.
Self-consistency (2022) brought the same intuition to reasoning: sample many chains of thought and take the most common final answer. A trained model can also score drafts without running them at all, and OpenAI's work on process supervision (2023) showed how much a well-trained verifier can lift results on mathematics.
Large Language Monkeys (2024) found that coverage keeps climbing over four orders of magnitude of samples, often in a straight line on a log scale, while voting and reward models level off after a few hundred. PlanSearch, the same year, observed that models repeatedly produce "highly similar, yet incorrect generations", then raised diversity by searching over natural-language plans before writing any code.
where it breaks
Best-of-K helps only when drafts fail in different ways. If the model holds one mistaken belief about a problem, ten drafts express the same mistake ten times, and voting can even let three copies of a wrong program outvote one correct program.
Training can make this worse, because work that makes a model better at its first answer can narrow the range of answers it gives. Kirk and colleagues (2023) found that reinforcement learning from human feedback "significantly reduces output diversity". Yue and colleagues (2025) found that models trained with verifiable rewards beat their base models at small k, while "the base models achieve a higher pass@k score when k is large".
A chooser that trusts a thin test suite will happily pick a wrong program that passes it. EvalPlus (2023) extended the tests of a popular code benchmark eightyfold and watched pass@k fall by up to 19.3 to 28.9%.
Stroebl, Kapoor and Narayanan (2024) worked through what happens when a check sometimes passes wrong code. Suppose each draft is correct with probability p, and the check passes a wrong draft with probability f. Resampling until something passes converges on an accuracy of p / (p + (1 − p) · f), and no amount of extra compute moves it. With p = 0.3 and f = 0.1, the ceiling is about 0.81. They also found that the best number of attempts is often fewer than ten.
The chooser can become the target too, as Gao, Schulman and Hilton (2022) found when they studied best-of-n selection against a learned reward model. As n grew, the true quality of the chosen answer first rose and then fell, while the proxy score kept climbing. A scorer is only a model of quality, so if you rank enough drafts with it, you start selecting for its mistakes.
how ATLAS does it, in release 3.1.6
ATLAS uses best-of-K inside its V3 pipeline, where it calls the drafts candidates, and only where the extra work is likely to pay.
Not every write qualifies. Code and HTML files of ten lines or more go to the pipeline. So does an edit that leaves a code file at 80 lines or more, or, in Python, with complex branching, and the pipeline then rewrites the whole file. Lines that are only inserted never go. Configuration, data, styles, documentation, shell scripts and files under ten lines are written directly. If the agent rewrites a file right after a failing run of it, that rewrite also skips the pipeline. Each pipeline call has a time limit, three minutes by default.
The pipeline starts with a probe. It writes one fresh candidate and checks it, and if that candidate passes, the pipeline stops there, with K = 1. The model's own version of the file goes into the prompt as a reference to improve on.
If the probe fails, a gate decides how many candidates to spend: three, five or eight, never fewer than three, and the failed probe counts as one of them. The gate reads the lens's score of the probe, and a probe that looks far from good code earns more candidates. In 3.1.6 that score barely changes from one input to the next, so the gate spends exactly three every time (#281). The gate also lowers K when the remaining time cannot afford more.
PlanSearch fills the new slots. It first asks the model for distinct sets of constraints, each pointing toward a different family of algorithm, then writes one plan per set, then one program per plan. A second method, DivSampling, prepends a varied role, instruction or style to the prompt, and fills any slot PlanSearch leaves empty.
Then every candidate is checked in a scratch sandbox. Most non-Python files get a syntax or compile check. In 3.1.6, C, C++, Swift, JSX, Vue and Svelte files, like any other language without its own check, are checked as if they were Python, so their candidates fail. A Python file is, by default, loaded and then tested with up to five tests the model wrote, and it is rejected if it passes fewer than half. A Python file is only compiled and linted when words such as "game", "menu" or "flask" appear in it or in the project files. If ATLAS detects a build command for the project, it also runs, on a temporary copy.
A first-round candidate that passed can still be rejected. In Python code, a structural check rejects a call to a name with no definition, import or builtin. An experimental cross-file check is off by default. The lens is meant to veto code it scores as a collapse toward a stub, but in 3.1.6's default setup that veto cannot fire.
Among the survivors, the candidate with the lowest lens energy wins. How well that score picks the best of several passing candidates has not yet been measured.
If nothing survives, the pipeline takes the failure the lens scores best and analyses it from up to four angles, making at most three repairs. Each gets the full check, and the first that passes is written. If the time left looks enough for one more round, a refinement loop tries again.
The last step is wrong in release 3.1.6 and earlier. When no candidate passes even after repair, ATLAS writes the best-scoring candidate anyway and tells the agent the edit was verified. If every candidate is vetoed, the model's own unchecked version receives the same "verified" message. It is a known issue, and the fix is planned for 3.2.0. Until then, review files the pipeline rewrote, and run your project's own tests before you rely on a "done".
The flaw is best-of-K's own selection step, applied where it has no evidence to work with. Among candidates that passed their checks, ranking picks a winner. Among candidates that all failed, it picks the least bad failure, and the system should report that nothing passed.
what this changes for us
Extra candidates cost time, so on easy tasks the pipeline stops at the probe. When it drafts more, they only help if they explore different ideas, which is why it searches over plans first and varies its prompts. More candidates raise coverage. Selection still depends on the chooser, and every chooser has blind spots: thin tests, correlated agreement, a scorer that can be gamed. That is why the record a system keeps should say which candidate passed which check, and say when nothing passed.
pass@k describes a pile of drafts with an oracle standing beside it. A tool delivers one answer, so the fair measure is the one AlphaCode used: how often the answer you hand back is right.
Our own score got this wrong. A reader raised a related point on our tracker (#3): our best-of-three score should not be called pass@1. We relabelled it. Later we found that it counted a task as solved when any of three drafts passed the examples printed in the problem, which is the oracle count that pass@k makes, and we withdrew it.
sources
- Chen et al. (2021), Evaluating Large Language Models Trained on Code, the Codex report: pass@k and its estimator.
- Li et al. (2022), Competition-Level Code Generation with AlphaCode, Science 378.
- Shi et al. (2022), Natural Language to Code Translation with Execution.
- Chen et al. (2022), CodeT: Code Generation with Generated Tests.
- Wang et al. (2022), Self-Consistency Improves Chain of Thought Reasoning in Language Models.
- Lightman et al. (2023), Let's Verify Step by Step.
- Brown et al. (2024), Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.
- Wang et al. (2024), Planning In Natural Language Improves LLM Search For Code Generation (PlanSearch).
- Wang et al. (2025), On the Effect of Sampling Diversity in Scaling LLM Inference (DivSampling).
- Kirk et al. (2023), Understanding the Effects of RLHF on LLM Generalisation and Diversity.
- Yue et al. (2025), Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
- Liu et al. (2023), Is Your Code Generated by ChatGPT Really Correct? (EvalPlus).
- Stroebl, Kapoor and Narayanan, The Limits of Inference Scaling Through Resampling.
- Gao, Schulman and Hilton (2022), Scaling Laws for Reward Model Overoptimization.
- ATLAS 3.1.6: the V3 pipeline and the per-file classification; the known issue; the label discussion in #3.
run it yourself
ATLAS is open source under AGPL-3.0, and it runs on your own hardware.