5 october 2026 · explainer · inferstep lab

The geometric lens: can hidden states spot broken code?

The geometric lens scores code by where it lands in the model's own representation, before anything runs. We do not yet know whether the idea works, and in the current release the lens gives no usable signal. Here is the idea, the machinery, what we have seen so far, and how we plan to find out.

A great wave curling over three boats, with Mount Fuji small in the distance.
Katsushika Hokusai, Under the Wave off Kanagawa (1830–33), woodblock print, Art Institute of Chicago. Public domain.

When a model writes a function, it also forms an internal picture of that function: a long list of numbers, its hidden states, that places the code somewhere in the model's own representation. The geometric lens asks whether that position already tells us something useful. Does correct code land in different places from broken code, and can a small learned scorer read the difference before anything runs? We do not know yet. This is a hypothesis we are still testing.

the idea

Language models know more about their own outputs than their text reveals, and several results from the last few years show it.

Kadavath and colleagues (2022) showed that models can estimate, with some skill, whether their own answers are correct. Azaria and Mitchell (2023) trained a small classifier on a model's hidden states and found it predicted whether a statement was true better than the model's own probability for that sentence. Burns and colleagues (2022) found truth-related structure in hidden states without using labels at all. Then Marks and Tegmark (2023) showed that, for simple true and false statements, truth appears as a direction in the model's representation space. That is the sense of "geometry" in the lens's name: position and direction in the space where the model represents what it reads.

Ribeiro and colleagues (2025) then did the same for code. They contrasted the hidden states of correct and incorrect programs written for the same task, found a direction that separates them, and used it to choose between candidates without running a single test. Their signal beat both the model's own likelihood and its stated confidence.

The lens adds an older framing from machine learning to that line of thought. An energy-based model, in LeCun's classic account (2006), learns a function that gives low energy to good configurations and high energy to bad ones. Applied to code, the picture is a landscape in which correct programs sit downhill and broken ones uphill.

A dotted landscape of contour lines with two deep basins; a few bright points sit at different heights, some low in the basins and some on the slopes.
The energy landscape as a metaphor: candidates low in a basin look like code that passed; candidates on the slopes look like code that failed. The shipped model is a learned score on a single vector, and this picture is only a way to think about it.

Our hypothesis is this. If passing and failing code land in measurably different places in the model's own representation, a small scorer trained on those places can rank candidates cheaply, with no second model and no execution.

how it works in ATLAS today

In release 3.1.6, when the pipeline has a candidate, it sends the code, and only the code, to the same model server that wrote it. The server returns one vector: the average of the model's hidden states across the program's tokens. For the Gemma model in our registry, that vector holds 3,840 numbers. A program longer than the server's batch, 1,024 tokens by default, cannot be embedded at all. Because the lens uses the model's own representation, it needs no second model.

Two small scorers read that vector, both on the CPU, though each score still costs one pass through the model server. C(x) is a small neural network that turns the vector into a single number, its energy, where lower is better. It is trained on code the model itself wrote, labelled pass or fail. The atlas lens build command uses a pairwise ranking loss that pushes every failing sample's energy above every passing one by a margin. The Gemma C(x) behind the committed calibration was trained by regression. G(x) compresses the vector to 128 numbers and passes them to a gradient-boosted tree classifier, which returns a probability that the code passes.

A lens bundle is trained against one model's representation, so its scores belong to that model alone. On any other model, the docs say, they are "wrong — not just suboptimal". Each bundle names its model, and if the server is serving a different one, the lens reports itself disabled and scores nothing.

candidate code (the text only) sent to the same model server one vector: the average token state 3,840 numbers for the Gemma model C(x) small network → energy G(x) 128 numbers, trees → P(pass) calibration files, per model energy → a score from 0 to 1 P(pass) → severe and low thresholds used for how many drafts, which passer wins, a veto on stubs, alerts in the agent loop
Inside the lens in release 3.1.6. The ranking reads the averaged vector; the veto and the alerts are meant to read each token's score and take the lowest. With the default settings, the veto and the alerts cannot fire on a 3.1.6 install, and the budget comes out as three every time.

The lens was designed to do four things. The first is the budget: when the first candidate fails, the lens's reading of it sets three, five or eight candidates in all. The second is ranking: among candidates that passed their checks, the lowest energy wins. Our note on best-of-K describes that step.

The third job is a veto on stubs. A candidate can run cleanly and still do nothing useful, and the code records the case that motivated the veto: a ten-line dashboard file holding a single heading, which passed the sandbox while the lens's score sagged far below normal.

The fourth is to watch the agent while it works. Every write is scored, and an edit scores only its new text. One severe score adds a "Lens severe-quality alert" to the agent's context, and two low scores in a row add "Lens regression detected".

Raw energies mean little until they are calibrated, because they depend on the model and the training run. A calibration file maps them onto a scale from 0 to 1 with a logistic curve set from that model's average passing and failing energies: its midpoint is their average, and its steepness is four divided by the gap between them. A second file sets G(x)'s thresholds from the low percentiles of passing scores. "Severe" is the 5th percentile, so about one good sample in twenty falls below it.

1.0 0.5 0 8.5 9.5 10.5 11.5 12.5 raw energy (lower is better) pass mean 9.25 → 0.12 fail mean → 0.88 uncalibrated: always 0.5
The calibration curve committed for the Gemma model in release 3.1.6. The passing mean always lands near 0.12 and the failing mean near 0.88; the curve's steepness comes from the gap between them. These values were derived for a C(x) trained on our development server, which differs from the one installs download.

That is the design, and a 3.1.6 install with the default settings gets much less of it: the budget comes out as three every time, and neither the veto nor the alerts can fire.

a worked example

Suppose the task is to write median(xs), and the problem's examples use only lists of odd length. Two candidates come back.

The first sorts the list and, when the length is even, averages the two middle values. The second returns sorted(xs)[len(xs) // 2], which is right for odd lengths and wrong for even ones. Both pass every example, because no example has an even length.

The tests cannot separate them, so the choice falls to the lens. If the first candidate sits lower in energy, it wins and the code is right. If the energies are reversed, the second wins, and nothing downstream will notice until someone passes in an even-length list.

The design: each candidate that passed takes its place on the landscape, and the lowest one is kept. In release 3.1.6 the scores barely change from one input to the next, so no candidate stands out.

Ranking matters only among candidates that already passed their checks. The illustration is ours, and the numbers a real lens would assign are unknown until measured.

where we are

A 3.1.6 install loads calibration files that ship with the repository, and the lens reports itself calibrated, with its interventions active. That report does not hold up. The calibration was fit to the C(x) on our development server, which differs from the published one that installs download, and nothing checks that the two belong together. And under the model server's default pooling setting, per-token scores cannot be computed at all. The service falls back to a neutral 0.5 and keeps reporting itself healthy, so the veto and the agent alerts can never fire. The development line had the same fault until a recent fix. A degraded component should say so in what it returns as well as in its log.

We scored five very different inputs on an install with the same layout as the release: correct code, buggy code, a class, plain prose and repeated stubs. G(x) came back identical to the last digit, 0.6142, and the calibrated C(x) was about 0 for all five. With scores that flat, every first candidate gets the same budget of three. So the lens in releases 3.1.3 to 3.1.6 reports itself healthy and gives no usable signal, and that is now a listed known issue (#281).

On the development line the pooling setting is fixed, and so is a third cause, a vector-scale fault. The mismatched calibration is still open: the issue asks for one matched, validated set of lens files per model, checked when it loads. Whether an install built from the development line behaves correctly has not been measured yet.

Ranking among passing candidates uses the raw energy. Whether that carries any signal in 3.1.6 is doubtful, and a candidate too long to embed scores 0.0 and wins the ranking.

On main, the architecture document, corrected after the 3.1.6 tag, says: "Whether lens-driven allocation beats fixed or randomly assigned tiers is unmeasured." A small comparison on tasks the system was tuned on found no difference its sample could resolve between the development lens, calibrated and with different C(x) weights, and the published lens without calibration files, which is not the layout an install has. Tuned tasks can show defects. This question needs held-out tasks.

The training labels are weaker than they look. The training data came from benchmark runs whose grader ran only the examples printed in each problem. A "pass" in that data means "passed the printed examples", which is a lower bar than correct.

Length got in the way twice, and we found both problems on our development build. They are in release 3.1.6, and the fixes are on the development line, unreleased.

We should have expected this. Azaria and Mitchell noted the same weakness in a model's sentence probabilities, which also depend on length. Any score computed over text should first be checked against the text's length, and we are adding that check to how we evaluate the lens.

how we plan to measure it

The issue that scopes the next lens says it "must be measured on held-out tasks of the kinds of work it judges (#238), not on LiveCodeBench." Held-out means tasks the lens has never been tuned on.

Our core experiment compares the lens with plain alternatives: the same tasks and the same budget, with the lens deciding in one arm and a fixed rule deciding in the other. That answers the architecture document's question, whether lens-driven allocation beats fixed or random allocation.

The labels have to change too, because a lens trained on "passed the printed examples" learns that bar. Retraining on graded outcomes is part of the same work.

That issue also lists the requirements for the next lens: one consistent signal per write; thresholds per model, with a provenance record for every hand-set value; drift checks that actually run; no "likely correct" on code that fails; and learning from new outcomes without special cases.

Beyond the lens, the roadmap pairs the sandbox, which is deterministic, with the lens, which is probabilistic, into one graded status for each result. A sandbox failure stays a veto. The lens adds judgement where execution cannot see, such as a stub that runs without doing the task.

On the development line, for 3.2.0, the lens is required and has no silent fallback: if it cannot score, ATLAS stops and says why. An uncalibrated lens still counts as able to score. The decision is recorded on the development line (ADR 0011).

We still think the hypothesis is sound. The released lens cannot test it, and the evidence from our development build is thin. The next measurement will show whether the signal justifies the machinery, and we will report it with its method, here and on the tracker, whichever way it comes out. The rule we have to meet is in our own design record: "uncalibrated or missing artifacts must not steer selection with wrong numbers."

sources

run it yourself

ATLAS is open source under AGPL-3.0, and it runs on your own hardware.