5 october 2026 · explainer · inferstep lab

Test-time scaling for LLMs: history, methods and limits

Test-time scaling means doing more work on each request with the same model. Where the idea came from, the four ways to do it, what limits each, and how ATLAS spends its own budget, including one place where our naming got ahead of our code.

A harbour at night in blue and gold, with ships and a few faint lights.
James McNeill Whistler, Nocturne: Blue and Gold — Southampton Water (1872), Art Institute of Chicago. Public domain.

Compute goes into a language model at two moments. Training comes before the model answers anyone, and it fixes the weights. Inference happens every time the model answers, with the weights already fixed, and the only question then is how much work to do on this one request. Test-time scaling is the study of that second moment, and it is where ATLAS works: it runs on your own hardware with a model you choose, so the one budget it can always shape is the work done per request.

two kinds of compute

Train-time compute covers pre-training, fine-tuning and reinforcement learning, and all three end with a fixed set of weights. Test-time compute is spent per request, on more tokens, more candidates, more checking or more refinement.

The two can stand in for each other. In 2021, Andy Jones showed with board-game agents that "the test-time and train-time compute available to an agent can be traded off while maintaining performance." When OpenAI introduced its o1 models in 2024, it described their accuracy as improving both with "more reinforcement learning (train-time compute)" and with "more time spent thinking (test-time compute)." Snell and colleagues (2024) asked the practical question: given a budget, how should one trade inference-time compute against a larger model?

For a local tool that trade is lopsided, because the hardware you own fixes the model.

Two dotted figures: on the left, many paths fan out from one point; on the right, a single path winds inward toward a target.
Two ways to spend: many paths at once, chosen between afterwards, or one path revised step by step.

four families

We sort the methods into four families, and real systems combine them.

Parallel sampling draws several independent answers, then chooses one by majority vote, with a learned verifier, or by running tests. Our note on best-of-K covers this family in depth.

Sequential refinement takes one answer, gets feedback on it, and writes a better one. The feedback can come from the model's own critique, from a compiler, or from failing tests.

Longer reasoning lets the model write out more intermediate steps before it commits to an answer, and some methods control how long that reasoning runs. Search with a verifier grows partial solutions step by step, keeps the promising branches and prunes the rest, guided by a score for each step.

independent attempts dependent steps model judges outside check voting self-consistency longer reasoning chain of thought, o1, R1, s1 self-critique sample and check trained verifiers, AlphaCode, best-of-K repair and search Self-Debugging, Tree of Thoughts, step-scored search
A conceptual map of the families: whether the attempts depend on each other, and whether something outside the model judges them.

a short history

The idea is older than large language models. Spending extra effort at decode time began with beam search, which keeps several partial hypotheses alive at every step and chooses among them at the end.

year work what it added
1976 Lowerre, HARPY speech recogniser the search later called beam search
2014 Sutskever, Vinyals and Le beam search for neural sequence models
2019 Holtzman et al., nucleus sampling decoding choices change quality, with the same model
2021 Chen et al., Codex repeated sampling as a strategy for code
2021 Cobbe et al., verifiers for math sample many, rank with a trained verifier
2022 Wei et al., chain of thought intermediate reasoning steps improve answers
2022 Li et al., AlphaCode large-scale sampling, filtering on behaviour
2022 Wang et al., self-consistency sample reasoning paths, take the majority answer
2023 Yao et al., Tree of Thoughts search over reasoning steps, with look-ahead
2023 Lightman et al., process supervision score each step, as well as the final answer
2024 Snell et al., compute-optimal scaling allocate compute by how hard each prompt is
2024 OpenAI, o1 reasoning learned through reinforcement learning
2025 DeepSeek-R1 longer thinking that emerges during RL training
2025 Muennighoff et al., s1 budget forcing: control thinking length directly

The table mixes two kinds of work: generating more (samples, steps, search) and judging better (verifiers, step scores, tests). Several entries pair them. Cobbe's verifiers and AlphaCode both sample many and then judge, because extra generation is only useful if something can pick from it.

The last three entries are a different kind of result. With o1 and DeepSeek-R1, longer reasoning became something a model learns during training. R1's report describes a model that "learns to allocate more thinking time to a problem by reevaluating its initial approach." Then s1 showed how to control that length from the outside with a small trick. To stop the model thinking, append the end-of-thinking marker. To make it think longer, suppress the marker and append the word "Wait".

what limits it

More samples raise coverage, but something still has to pick. Brown and colleagues (2024) kept sampling across four orders of magnitude, and the share of problems solved by some sample kept climbing. Without an automatic check to pick the right one, voting and reward models levelled off after a few hundred samples.

An imperfect check puts a ceiling on the whole approach. If a check sometimes passes wrong answers, resampling cannot push accuracy past a limit set by that error rate, however much compute you spend (Stroebl, Kapoor and Narayanan, 2024).

Snell and colleagues found that the best way to spend depends on how hard the prompt is: easy questions gain most from sequential refinement, hard ones need a balance of refinement and parallel sampling, and on questions beyond what the base model can do at all, a larger model trained for longer is likely the better investment. Their compute-optimal strategy, which adapts per prompt, was more than four times as efficient as plain best-of-N on their math benchmark.

Forcing longer thinking helps and then flattens, s1 found, and suppressing the end-of-thinking marker too often sends the model into repetitive loops.

Refinement works only when the feedback carries information. Huang and colleagues (2023) found that models "struggle to self-correct their responses without external feedback", and sometimes get worse after trying. Self-Debugging (2023), which feeds the model real execution results, matched or beat baselines that generated more than ten times as many candidates.

how ATLAS spends its budget

In release 3.1.6, ATLAS combines two of the families: parallel sampling with checks, and repair driven by real error output.

One request, from the prompt to the written file: three plans, the model, a checked tool call, the probe, more candidates and their check.

For the loop that chooses tools, ATLAS turns the model's thinking channel off, because those calls must produce strictly structured output. If a model reasons anyway and has not started its answer, the stream is cut after about six thousand tokens of reasoning and the model is asked again. An answer that starts repeating the same phrase word for word is cut too. Budget control here means limiting reasoning, in the same spirit as s1's cap.

For each substantive turn, before work begins, ATLAS drafts three short plans, one after another, at different temperatures. It scores them with simple rules, such as rewarding a plan that includes a verification step, and puts the winner in front of the model.

The V3 pipeline runs when the model writes a code or HTML file of ten lines or more, and on edits that leave a file large or complex. Short files, configuration and documentation skip it. So does a rewrite made right after a failing run of the same file, where a quick fix beats a multi-minute search.

The pipeline first drafts one candidate, the probe, and checks it: with generated tests for a Python task with clear inputs and outputs, and otherwise with a syntax or compile check. If it passes, the work stops there. An easy task costs one candidate. Stopping at the first stage whose check passes is the most adaptive thing the pipeline does.

If the probe fails, a gate decides how many more to draft. It reads the lens's score of the probe and allocates three, five or eight candidates in all, counting the first. This is the same idea as Snell's per-prompt allocation. On a 3.1.6 install the lens scores every probe almost the same, so the gate allocates three each time (#281). The architecture document on main records that whether lens-driven allocation beats a fixed rule "is unmeasured."

ATLAS names one of its components Budget Forcing, after the s1 paper. The name is ahead of the code. In the live 3.1.6 product, the pipeline's calls keep the model's thinking channel off, and nothing appends "Wait" or cuts thinking short. The tiers set how many candidates are drafted, and the length cap on some of them, and leave the depth of reasoning per candidate untouched. Our released documentation still describes the tiers as thinking budgets, and it is wrong on this point.

When nothing passes, the pipeline takes the failure the lens scores best and drafts up to three repairs, each from a different angle, checking each against the sandbox, and a further refinement loop runs only if the remaining time can afford one full round. Both are driven by the actual error output, the kind of feedback Self-Debugging used.

Each pipeline call has a three-minute limit. If it runs out, ATLAS writes the model's own version of the file, after syntax and structural checks. For a new file, it reports that it did so.

one pipeline call, within 3 minutes probe: 1 candidate plus tests written for the task gate: K = 3, 5 or 8 never fewer than 3 PlanSearch: 1 + 2(K − 1) calls constraints, then a plan and code per candidate, thinking off checks and lens scores sandbox runs, one embedding each repair: up to 3 drafts each from a different angle refinement: 2 rounds at most only if time allows a full round
Where one pipeline call spends its model calls in release 3.1.6. An easy task stops after the probe; a hard one uses about twenty model calls at three candidates, and about thirty at eight.

where that leaves us

The probe is cheap, and it saves the most on easy work. What we allocate beyond it should follow evidence of difficulty and beat a fixed allocation of the same total compute. Ours has not been measured against one, and in 3.1.6 it is a fixed allocation of three in practice.

Budget Forcing should perform the technique it is named after, or carry a name that matches what it does. Today it does neither.

Test-time compute substitutes for a stronger model only where the model already has some success. Past that point, extra attempts mostly produce extra failures, and a larger model is likely the better investment, as Snell and colleagues found.

sources

run it yourself

ATLAS is open source under AGPL-3.0, and it runs on your own hardware.