A.T.L.A.S.
Adaptive Test-time Learning and Autonomous Specialization
A coding agent that runs on your own GPU, with the model you choose. For new code files and larger edits, it drafts one or more versions and tests them in a sandbox before it writes the file.
install
Linux (Ubuntu, Debian, RHEL, Fedora, Rocky or Alma), or macOS on Apple Silicon, with Python 3.9 or newer. The installer adds Docker if it is missing; Podman also works.
The ready-made NVIDIA image is built for RTX 50-series cards. Older NVIDIA cards need a one-time local build, which uses about 17 GB more disk while it runs.
curl -fsSL https://raw.githubusercontent.com/inferstep/ATLAS/main/scripts/atlas-bootstrap.sh | bashupgrading from 3.1.3 or older: re-run the install command, or git pull, then atlas upgrade
the idea
A model small enough to run on one GPU will sometimes write a reply the software cannot parse, call a function that does not exist, or say the work is done while a test still fails.
ATLAS keeps the model you have and never changes its weights. The extra work goes into the system around it: an agent loop that checks what the model asks to do, and a pipeline called V3 that tests a new code file before it is written. All of it runs on your machine, with no hosted model and no key from a model provider.
the parts
| part | what it does |
|---|---|
| terminal interface | Where you type. It asks you before a new file is written or a command runs. |
| agent loop | Stands between you and the model. Nothing the model asks for runs until the loop has checked it. |
| plan step | Gives the model a to-do list before real work starts. |
| V3 pipeline | Drafts candidates for one code file, checks each one, and chooses among those that pass. |
| sandbox | A container where commands and candidates run, with a read-only system and limits on memory, CPU and processes. |
| geometric lens | Scores code from the model's own hidden states. In release 3.1.6 it gives no usable signal. |
| model server | llama.cpp, serving the model you chose. |
one request, step by step
Suppose you ask ATLAS to write a small Python program in a new file, and to make sure it works. The picture at the top of this page follows the whole path. Here it is one step at a time.
1. A plan comes first
Before a turn of real work, three short plans are drafted and scored by simple rules. The winner goes to the model as its to-do list. A short message or a plain question skips this step.
2. Each tool call is checked
By default a grammar holds every reply to one of three shapes: a tool call, a message, or "done". Before a tool call runs, the loop checks it. You are asked before a new file is written, a file is deleted, or a command runs, and a short list of destructive commands is always refused. Commands run in the sandbox, for thirty seconds by default.
3. Each write is routed
Configuration, data, documentation and any file under ten lines are written directly. A code or HTML file of ten lines or more goes to the V3 pipeline. Small edits are written directly too. An edit reaches the pipeline only when it leaves the file at 80 lines or more, or, in Python, with complex branching.
4. The pipeline drafts candidates and checks each one
The pipeline writes its own draft of the file, called a candidate. The first one is the probe, written with the model's own version in the prompt as a reference to improve on. If the probe passes its check, the pipeline stops there. If it fails, more candidates are drafted from different plans, at least three in all, because more candidates only help when they fail in different ways. Each one is checked in the sandbox. Most files get a syntax or compile check. A Python file is loaded and tested with up to five tests that the model writes for it, unless it looks like an interactive program.
5. One candidate is chosen
Among the candidates that passed, the geometric lens picks one. It reads the model's own hidden states for each candidate and places it on a landscape that is low where code looks like code that passed. The picture shows the design. In release 3.1.6 the scores barely change between inputs (#281), so no candidate stands out.
6. If nothing passes, the pipeline tries a repair
The pipeline takes the failed candidate that scored best, studies the failure, and drafts up to three repairs. The first repair that passes is written. If the pipeline runs out of its three minutes, the model's own version is written after syntax and structural checks.
In release 3.1.6, if nothing passes even after repair, the best-scoring candidate is written anyway, and the model is told the edit was verified. This is a known issue. The fix is planned for 3.2.0.
7. A "done" is questioned
Before a file is written, write gates check its syntax and, in Python, look for calls to names that nothing defines. When the model says "done", the finish gate looks at the whole turn. It sends the turn back when nothing has passed earlier and either a check the model ran is failing or a fix you asked for has nothing that verifies it. Each of its checks does this up to three times, and then lets the turn finish.
what does not work yet
The 3.1 line is experimental. Review what ATLAS writes, and run your own tests before you rely on a "done". We have not yet measured whether each part improves the result, and we will publish the numbers when we have them.
- A candidate that failed its checks can be written and reported as verified (3.1.0 to 3.1.6). The repair step has the details.
- C, C++, Swift, JSX, Vue and Svelte files, and any other language without its own check, are checked as if they were Python, so their candidates fail.
- The lens gives no usable signal (3.1.3 to 3.1.6), so the count is three every time.
- The pipeline does not see your request and never checks the model's own version, so a candidate can pass its checks and still drop something you asked for.
- Your project's own tests are not run. The model's tests call only the first function in a file.
- A turn that uses up its tries at the finish gate ends as normal, with no warning.
- The sandbox can reach the network unless you set
ATLAS_SANDBOX_NET_INTERNAL=true. - Go, C and C++ programs build in the sandbox but cannot run there (#163), and front-end code is never run in a browser (#127).
- Each new request carries over only the text of earlier messages.
The first four are fixed on the development line and planned for 3.2.0. For the lens, the fix makes its scores respond to the input. Whether those scores improve the result is still to be measured. The release page keeps the current list.
read more
The architecture document for 3.1.6 is the full reference. It differs from the code in a few places. Where the two disagree, this page follows the code.
go further
Read the source, pick a starter issue, or ask a question. Security problems go to a private report.