first project · open source · agpl-3.0 · release 3.1.6 "maia"

A.T.L.A.S.

Adaptive Test-time Learning and Autonomous Specialization

A coding agent that runs on your own GPU, with the model you choose. For new code files and larger edits, it drafts one or more versions and tests them in a sandbox before it writes the file.

One request, from the prompt to the written file.
16 GB+of GPU memory
~20 GBof disk (22 GB with AMD ROCm)
1user, on your own machine

install

Linux (Ubuntu, Debian, RHEL, Fedora, Rocky or Alma), or macOS on Apple Silicon, with Python 3.9 or newer. The installer adds Docker if it is missing; Podman also works.

The ready-made NVIDIA image is built for RTX 50-series cards. Older NVIDIA cards need a one-time local build, which uses about 17 GB more disk while it runs.

curl -fsSL https://raw.githubusercontent.com/inferstep/ATLAS/main/scripts/atlas-bootstrap.sh | bash

upgrading from 3.1.3 or older: re-run the install command, or git pull, then atlas upgrade

the idea

A model small enough to run on one GPU will sometimes write a reply the software cannot parse, call a function that does not exist, or say the work is done while a test still fails.

ATLAS keeps the model you have and never changes its weights. The extra work goes into the system around it: an agent loop that checks what the model asks to do, and a pipeline called V3 that tests a new code file before it is written. All of it runs on your machine, with no hosted model and no key from a model provider.

the parts

partwhat it does
terminal interfaceWhere you type. It asks you before a new file is written or a command runs.
agent loopStands between you and the model. Nothing the model asks for runs until the loop has checked it.
plan stepGives the model a to-do list before real work starts.
V3 pipelineDrafts candidates for one code file, checks each one, and chooses among those that pass.
sandboxA container where commands and candidates run, with a read-only system and limits on memory, CPU and processes.
geometric lensScores code from the model's own hidden states. In release 3.1.6 it gives no usable signal.
model serverllama.cpp, serving the model you chose.

one request, step by step

Suppose you ask ATLAS to write a small Python program in a new file, and to make sure it works. The picture at the top of this page follows the whole path. Here it is one step at a time.

1. A plan comes first

Before a turn of real work, three short plans are drafted and scored by simple rules. The winner goes to the model as its to-do list. A short message or a plain question skips this step.

2. Each tool call is checked

By default a grammar holds every reply to one of three shapes: a tool call, a message, or "done". Before a tool call runs, the loop checks it. You are asked before a new file is written, a file is deleted, or a command runs, and a short list of destructive commands is always refused. Commands run in the sandbox, for thirty seconds by default.

3. Each write is routed

Configuration, data, documentation and any file under ten lines are written directly. A code or HTML file of ten lines or more goes to the V3 pipeline. Small edits are written directly too. An edit reaches the pipeline only when it leaves the file at 80 lines or more, or, in Python, with complex branching.

4. The pipeline drafts candidates and checks each one

The pipeline writes its own draft of the file, called a candidate. The first one is the probe, written with the model's own version in the prompt as a reference to improve on. If the probe passes its check, the pipeline stops there. If it fails, more candidates are drafted from different plans, at least three in all, because more candidates only help when they fail in different ways. Each one is checked in the sandbox. Most files get a syntax or compile check. A Python file is loaded and tested with up to five tests that the model writes for it, unless it looks like an interactive program.

5. One candidate is chosen

Among the candidates that passed, the geometric lens picks one. It reads the model's own hidden states for each candidate and places it on a landscape that is low where code looks like code that passed. The picture shows the design. In release 3.1.6 the scores barely change between inputs (#281), so no candidate stands out.

6. If nothing passes, the pipeline tries a repair

The pipeline takes the failed candidate that scored best, studies the failure, and drafts up to three repairs. The first repair that passes is written. If the pipeline runs out of its three minutes, the model's own version is written after syntax and structural checks.

In release 3.1.6, if nothing passes even after repair, the best-scoring candidate is written anyway, and the model is told the edit was verified. This is a known issue. The fix is planned for 3.2.0.

7. A "done" is questioned

Before a file is written, write gates check its syntax and, in Python, look for calls to names that nothing defines. When the model says "done", the finish gate looks at the whole turn. It sends the turn back when nothing has passed earlier and either a check the model ran is failing or a fix you asked for has nothing that verifies it. Each of its checks does this up to three times, and then lets the turn finish.

what does not work yet

The 3.1 line is experimental. Review what ATLAS writes, and run your own tests before you rely on a "done". We have not yet measured whether each part improves the result, and we will publish the numbers when we have them.

The first four are fixed on the development line and planned for 3.2.0. For the lens, the fix makes its scores respond to the input. Whether those scores improve the result is still to be measured. The release page keeps the current list.

read more

The architecture document for 3.1.6 is the full reference. It differs from the code in a few places. Where the two disagree, this page follows the code.

go further

Read the source, pick a starter issue, or ask a question. Security problems go to a private report.