Test-time scaling for LLMs: history, methods and limits
Test-time scaling means doing more work on each request with the same model. Where the idea came from, the four ways to do it, what limits each, and how ATLAS spends its own budget, including one place where our naming got ahead of our code.
Compute goes into a language model at two moments. Training comes before the model answers anyone, and it fixes the weights. Inference happens every time the model answers, with the weights already fixed, and the only question then is how much work to do on this one request. Test-time scaling is the study of that second moment, and it is where ATLAS works: it runs on your own hardware with a model you choose, so the one budget it can always shape is the work done per request.
two kinds of compute
Train-time compute covers pre-training, fine-tuning and reinforcement learning, and all three end with a fixed set of weights. Test-time compute is spent per request, on more tokens, more candidates, more checking or more refinement.
The two can stand in for each other. In 2021, Andy Jones showed with board-game agents that "the test-time and train-time compute available to an agent can be traded off while maintaining performance." When OpenAI introduced its o1 models in 2024, it described their accuracy as improving both with "more reinforcement learning (train-time compute)" and with "more time spent thinking (test-time compute)." Snell and colleagues (2024) asked the practical question: given a budget, how should one trade inference-time compute against a larger model?
For a local tool that trade is lopsided, because the hardware you own fixes the model.
four families
We sort the methods into four families, and real systems combine them.
Parallel sampling draws several independent answers, then chooses one by majority vote, with a learned verifier, or by running tests. Our note on best-of-K covers this family in depth.
Sequential refinement takes one answer, gets feedback on it, and writes a better one. The feedback can come from the model's own critique, from a compiler, or from failing tests.
Longer reasoning lets the model write out more intermediate steps before it commits to an answer, and some methods control how long that reasoning runs. Search with a verifier grows partial solutions step by step, keeps the promising branches and prunes the rest, guided by a score for each step.
a short history
The idea is older than large language models. Spending extra effort at decode time began with beam search, which keeps several partial hypotheses alive at every step and chooses among them at the end.
| year | work | what it added |
|---|---|---|
| 1976 | Lowerre, HARPY speech recogniser | the search later called beam search |
| 2014 | Sutskever, Vinyals and Le | beam search for neural sequence models |
| 2019 | Holtzman et al., nucleus sampling | decoding choices change quality, with the same model |
| 2021 | Chen et al., Codex | repeated sampling as a strategy for code |
| 2021 | Cobbe et al., verifiers for math | sample many, rank with a trained verifier |
| 2022 | Wei et al., chain of thought | intermediate reasoning steps improve answers |
| 2022 | Li et al., AlphaCode | large-scale sampling, filtering on behaviour |
| 2022 | Wang et al., self-consistency | sample reasoning paths, take the majority answer |
| 2023 | Yao et al., Tree of Thoughts | search over reasoning steps, with look-ahead |
| 2023 | Lightman et al., process supervision | score each step, as well as the final answer |
| 2024 | Snell et al., compute-optimal scaling | allocate compute by how hard each prompt is |
| 2024 | OpenAI, o1 | reasoning learned through reinforcement learning |
| 2025 | DeepSeek-R1 | longer thinking that emerges during RL training |
| 2025 | Muennighoff et al., s1 | budget forcing: control thinking length directly |
The table mixes two kinds of work: generating more (samples, steps, search) and judging better (verifiers, step scores, tests). Several entries pair them. Cobbe's verifiers and AlphaCode both sample many and then judge, because extra generation is only useful if something can pick from it.
The last three entries are a different kind of result. With o1 and DeepSeek-R1, longer reasoning became something a model learns during training. R1's report describes a model that "learns to allocate more thinking time to a problem by reevaluating its initial approach." Then s1 showed how to control that length from the outside with a small trick. To stop the model thinking, append the end-of-thinking marker. To make it think longer, suppress the marker and append the word "Wait".
what limits it
More samples raise coverage, but something still has to pick. Brown and colleagues (2024) kept sampling across four orders of magnitude, and the share of problems solved by some sample kept climbing. Without an automatic check to pick the right one, voting and reward models levelled off after a few hundred samples.
An imperfect check puts a ceiling on the whole approach. If a check sometimes passes wrong answers, resampling cannot push accuracy past a limit set by that error rate, however much compute you spend (Stroebl, Kapoor and Narayanan, 2024).
Snell and colleagues found that the best way to spend depends on how hard the prompt is: easy questions gain most from sequential refinement, hard ones need a balance of refinement and parallel sampling, and on questions beyond what the base model can do at all, a larger model trained for longer is likely the better investment. Their compute-optimal strategy, which adapts per prompt, was more than four times as efficient as plain best-of-N on their math benchmark.
Forcing longer thinking helps and then flattens, s1 found, and suppressing the end-of-thinking marker too often sends the model into repetitive loops.
Refinement works only when the feedback carries information. Huang and colleagues (2023) found that models "struggle to self-correct their responses without external feedback", and sometimes get worse after trying. Self-Debugging (2023), which feeds the model real execution results, matched or beat baselines that generated more than ten times as many candidates.
how ATLAS spends its budget
In release 3.1.6, ATLAS combines two of the families: parallel sampling with checks, and repair driven by real error output.
For the loop that chooses tools, ATLAS turns the model's thinking channel off, because those calls must produce strictly structured output. If a model reasons anyway and has not started its answer, the stream is cut after about six thousand tokens of reasoning and the model is asked again. An answer that starts repeating the same phrase word for word is cut too. Budget control here means limiting reasoning, in the same spirit as s1's cap.
For each substantive turn, before work begins, ATLAS drafts three short plans, one after another, at different temperatures. It scores them with simple rules, such as rewarding a plan that includes a verification step, and puts the winner in front of the model.
The V3 pipeline runs when the model writes a code or HTML file of ten lines or more, and on edits that leave a file large or complex. Short files, configuration and documentation skip it. So does a rewrite made right after a failing run of the same file, where a quick fix beats a multi-minute search.
The pipeline first drafts one candidate, the probe, and checks it: with generated tests for a Python task with clear inputs and outputs, and otherwise with a syntax or compile check. If it passes, the work stops there. An easy task costs one candidate. Stopping at the first stage whose check passes is the most adaptive thing the pipeline does.
If the probe fails, a gate decides how many more to draft. It reads the lens's score of the probe and allocates three, five or eight candidates in all, counting the first. This is the same idea as Snell's per-prompt allocation. On a 3.1.6 install the lens scores every probe almost the same, so the gate allocates three each time (#281). The architecture document on main records that whether lens-driven allocation beats a fixed rule "is unmeasured."
ATLAS names one of its components Budget Forcing, after the s1 paper. The name is ahead of the code. In the live 3.1.6 product, the pipeline's calls keep the model's thinking channel off, and nothing appends "Wait" or cuts thinking short. The tiers set how many candidates are drafted, and the length cap on some of them, and leave the depth of reasoning per candidate untouched. Our released documentation still describes the tiers as thinking budgets, and it is wrong on this point.
When nothing passes, the pipeline takes the failure the lens scores best and drafts up to three repairs, each from a different angle, checking each against the sandbox, and a further refinement loop runs only if the remaining time can afford one full round. Both are driven by the actual error output, the kind of feedback Self-Debugging used.
Each pipeline call has a three-minute limit. If it runs out, ATLAS writes the model's own version of the file, after syntax and structural checks. For a new file, it reports that it did so.
where that leaves us
The probe is cheap, and it saves the most on easy work. What we allocate beyond it should follow evidence of difficulty and beat a fixed allocation of the same total compute. Ours has not been measured against one, and in 3.1.6 it is a fixed allocation of three in practice.
Budget Forcing should perform the technique it is named after, or carry a name that matches what it does. Today it does neither.
Test-time compute substitutes for a stronger model only where the model already has some success. Past that point, extra attempts mostly produce extra failures, and a larger model is likely the better investment, as Snell and colleagues found.
sources
- Jones (2021), Scaling Scaling Laws with Board Games.
- Sutskever, Vinyals and Le (2014), Sequence to Sequence Learning with Neural Networks.
- Holtzman et al. (2019), The Curious Case of Neural Text Degeneration.
- Chen et al. (2021), Evaluating Large Language Models Trained on Code.
- Cobbe et al. (2021), Training Verifiers to Solve Math Word Problems.
- Wei et al. (2022), Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.
- Li et al. (2022), Competition-Level Code Generation with AlphaCode.
- Wang et al. (2022), Self-Consistency Improves Chain of Thought Reasoning in Language Models.
- Yao et al. (2023), Tree of Thoughts.
- Lightman et al. (2023), Let's Verify Step by Step.
- Snell et al. (2024), Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.
- OpenAI (2024), OpenAI o1 System Card.
- DeepSeek-AI (2025), DeepSeek-R1.
- Muennighoff et al. (2025), s1: Simple test-time scaling.
- Brown et al. (2024), Large Language Monkeys.
- Stroebl, Kapoor and Narayanan, The Limits of Inference Scaling Through Resampling.
- Huang et al. (2023), Large Language Models Cannot Self-Correct Reasoning Yet.
- Chen, Lin, Schärli and Zhou (2023), Teaching Large Language Models to Self-Debug.
- Madaan et al. (2023), Self-Refine.
- ATLAS 3.1.6: the architecture document, Plan Mode, and the released note on allocation.
run it yourself
ATLAS is open source under AGPL-3.0, and it runs on your own hardware.