The geometric lens: can hidden states spot broken code?
The geometric lens scores code by where it lands in the model's own representation, before anything runs. We do not yet know whether the idea works, and in the current release the lens gives no usable signal. Here is the idea, the machinery, what we have seen so far, and how we plan to find out.
When a model writes a function, it also forms an internal picture of that function: a long list of numbers, its hidden states, that places the code somewhere in the model's own representation. The geometric lens asks whether that position already tells us something useful. Does correct code land in different places from broken code, and can a small learned scorer read the difference before anything runs? We do not know yet. This is a hypothesis we are still testing.
the idea
Language models know more about their own outputs than their text reveals, and several results from the last few years show it.
Kadavath and colleagues (2022) showed that models can estimate, with some skill, whether their own answers are correct. Azaria and Mitchell (2023) trained a small classifier on a model's hidden states and found it predicted whether a statement was true better than the model's own probability for that sentence. Burns and colleagues (2022) found truth-related structure in hidden states without using labels at all. Then Marks and Tegmark (2023) showed that, for simple true and false statements, truth appears as a direction in the model's representation space. That is the sense of "geometry" in the lens's name: position and direction in the space where the model represents what it reads.
Ribeiro and colleagues (2025) then did the same for code. They contrasted the hidden states of correct and incorrect programs written for the same task, found a direction that separates them, and used it to choose between candidates without running a single test. Their signal beat both the model's own likelihood and its stated confidence.
The lens adds an older framing from machine learning to that line of thought. An energy-based model, in LeCun's classic account (2006), learns a function that gives low energy to good configurations and high energy to bad ones. Applied to code, the picture is a landscape in which correct programs sit downhill and broken ones uphill.
Our hypothesis is this. If passing and failing code land in measurably different places in the model's own representation, a small scorer trained on those places can rank candidates cheaply, with no second model and no execution.
how it works in ATLAS today
In release 3.1.6, when the pipeline has a candidate, it sends the code, and only the code, to the same model server that wrote it. The server returns one vector: the average of the model's hidden states across the program's tokens. For the Gemma model in our registry, that vector holds 3,840 numbers. A program longer than the server's batch, 1,024 tokens by default, cannot be embedded at all. Because the lens uses the model's own representation, it needs no second model.
Two small scorers read that vector, both on the CPU, though each score still costs one pass through the model server. C(x) is a small neural network that turns the vector into a single number, its energy, where lower is better. It is trained on code the model itself wrote, labelled pass or fail. The atlas lens build command uses a pairwise ranking loss that pushes every failing sample's energy above every passing one by a margin. The Gemma C(x) behind the committed calibration was trained by regression. G(x) compresses the vector to 128 numbers and passes them to a gradient-boosted tree classifier, which returns a probability that the code passes.
A lens bundle is trained against one model's representation, so its scores belong to that model alone. On any other model, the docs say, they are "wrong — not just suboptimal". Each bundle names its model, and if the server is serving a different one, the lens reports itself disabled and scores nothing.
The lens was designed to do four things. The first is the budget: when the first candidate fails, the lens's reading of it sets three, five or eight candidates in all. The second is ranking: among candidates that passed their checks, the lowest energy wins. Our note on best-of-K describes that step.
The third job is a veto on stubs. A candidate can run cleanly and still do nothing useful, and the code records the case that motivated the veto: a ten-line dashboard file holding a single heading, which passed the sandbox while the lens's score sagged far below normal.
The fourth is to watch the agent while it works. Every write is scored, and an edit scores only its new text. One severe score adds a "Lens severe-quality alert" to the agent's context, and two low scores in a row add "Lens regression detected".
Raw energies mean little until they are calibrated, because they depend on the model and the training run. A calibration file maps them onto a scale from 0 to 1 with a logistic curve set from that model's average passing and failing energies: its midpoint is their average, and its steepness is four divided by the gap between them. A second file sets G(x)'s thresholds from the low percentiles of passing scores. "Severe" is the 5th percentile, so about one good sample in twenty falls below it.
That is the design, and a 3.1.6 install with the default settings gets much less of it: the budget comes out as three every time, and neither the veto nor the alerts can fire.
a worked example
Suppose the task is to write median(xs), and the problem's examples use only lists of odd length. Two candidates come back.
The first sorts the list and, when the length is even, averages the two middle values. The second returns sorted(xs)[len(xs) // 2], which is right for odd lengths and wrong for even ones. Both pass every example, because no example has an even length.
The tests cannot separate them, so the choice falls to the lens. If the first candidate sits lower in energy, it wins and the code is right. If the energies are reversed, the second wins, and nothing downstream will notice until someone passes in an even-length list.
Ranking matters only among candidates that already passed their checks. The illustration is ours, and the numbers a real lens would assign are unknown until measured.
where we are
A 3.1.6 install loads calibration files that ship with the repository, and the lens reports itself calibrated, with its interventions active. That report does not hold up. The calibration was fit to the C(x) on our development server, which differs from the published one that installs download, and nothing checks that the two belong together. And under the model server's default pooling setting, per-token scores cannot be computed at all. The service falls back to a neutral 0.5 and keeps reporting itself healthy, so the veto and the agent alerts can never fire. The development line had the same fault until a recent fix. A degraded component should say so in what it returns as well as in its log.
We scored five very different inputs on an install with the same layout as the release: correct code, buggy code, a class, plain prose and repeated stubs. G(x) came back identical to the last digit, 0.6142, and the calibrated C(x) was about 0 for all five. With scores that flat, every first candidate gets the same budget of three. So the lens in releases 3.1.3 to 3.1.6 reports itself healthy and gives no usable signal, and that is now a listed known issue (#281).
On the development line the pooling setting is fixed, and so is a third cause, a vector-scale fault. The mismatched calibration is still open: the issue asks for one matched, validated set of lens files per model, checked when it loads. Whether an install built from the development line behaves correctly has not been measured yet.
Ranking among passing candidates uses the raw energy. Whether that carries any signal in 3.1.6 is doubtful, and a candidate too long to embed scores 0.0 and wins the ranking.
On main, the architecture document, corrected after the 3.1.6 tag, says: "Whether lens-driven allocation beats fixed or randomly assigned tiers is unmeasured." A small comparison on tasks the system was tuned on found no difference its sample could resolve between the development lens, calibrated and with different C(x) weights, and the published lens without calibration files, which is not the layout an install has. Tuned tasks can show defects. This question needs held-out tasks.
The training labels are weaker than they look. The training data came from benchmark runs whose grader ran only the examples printed in each problem. A "pass" in that data means "passed the printed examples", which is a lower bar than correct.
Length got in the way twice, and we found both problems on our development build. They are in release 3.1.6, and the fixes are on the development line, unreleased.
- The veto read the lowest per-token score in a candidate. In a long program, the lowest point of a noisy trace sinks whatever the content, so the minimum could not separate good code from bad, and the veto never fired. The development line's pipeline veto now reads the average score against a hand-set cutoff of 0.52, when the bundle supplies one. The agent alerts still read the minimum.
- On clean code, C(x) rises with the logarithm of the text's length. The development line now scores C(x) against a baseline for that length when setting the budget, while ranking still uses the raw energy.
We should have expected this. Azaria and Mitchell noted the same weakness in a model's sentence probabilities, which also depend on length. Any score computed over text should first be checked against the text's length, and we are adding that check to how we evaluate the lens.
how we plan to measure it
The issue that scopes the next lens says it "must be measured on held-out tasks of the kinds of work it judges (#238), not on LiveCodeBench." Held-out means tasks the lens has never been tuned on.
Our core experiment compares the lens with plain alternatives: the same tasks and the same budget, with the lens deciding in one arm and a fixed rule deciding in the other. That answers the architecture document's question, whether lens-driven allocation beats fixed or random allocation.
The labels have to change too, because a lens trained on "passed the printed examples" learns that bar. Retraining on graded outcomes is part of the same work.
That issue also lists the requirements for the next lens: one consistent signal per write; thresholds per model, with a provenance record for every hand-set value; drift checks that actually run; no "likely correct" on code that fails; and learning from new outcomes without special cases.
Beyond the lens, the roadmap pairs the sandbox, which is deterministic, with the lens, which is probabilistic, into one graded status for each result. A sandbox failure stays a veto. The lens adds judgement where execution cannot see, such as a stub that runs without doing the task.
On the development line, for 3.2.0, the lens is required and has no silent fallback: if it cannot score, ATLAS stops and says why. An uncalibrated lens still counts as able to score. The decision is recorded on the development line (ADR 0011).
We still think the hypothesis is sound. The released lens cannot test it, and the evidence from our development build is thin. The next measurement will show whether the signal justifies the machinery, and we will report it with its method, here and on the tracker, whichever way it comes out. The rule we have to meet is in our own design record: "uncalibrated or missing artifacts must not steer selection with wrong numbers."
sources
- Kadavath et al. (2022), Language Models (Mostly) Know What They Know.
- Azaria and Mitchell (2023), The Internal State of an LLM Knows When It's Lying.
- Burns et al. (2022), Discovering Latent Knowledge in Language Models Without Supervision.
- Marks and Tegmark (2023), The Geometry of Truth.
- Ribeiro et al. (2025), On LLMs' Internal Representation of Code Correctness.
- LeCun et al. (2006), A Tutorial on Energy-Based Learning.
- Ni et al. (2023), LEVER: Learning to Verify Language-to-Code Generation with Execution.
- Lightman et al. (2023), Let's Verify Step by Step.
- ATLAS 3.1.6: the Geometric Lens; the released architecture text on allocation; ADR 0005, lens optionality.
- Issues: #281 (calibration files), #237 (the next lens), #188 (the roadmap), ADR 0011 (the lens becomes required).
- Development line: the per-token veto and the length baseline.
run it yourself
ATLAS is open source under AGPL-3.0, and it runs on your own hardware.