Teton Code

An AI coding agent that routes work across a range of models — so you spend frontier-model money only where frontier-model intelligence matters.

Install

brew install atelier-fashion/tap/teton

Then just open the CLI:

teton

The first run notices there is no daemon and offers to register teton-code with launchd for you (brew services start teton) so it starts now and survives reboots — press return to accept, or decline and it starts an unmanaged daemon for the session. The manual command keeps working, before or instead of the offer.

If Homebrew answers that the tap is not trusted when the service is registered, run brew trust atelier-fashion/tap once — the install command above is self-authorizing because it names the tap in full, and the short name teton is not.

Latest release v0.1.36. Two binaries — teton (CLI) and teton-code (daemon). No Rust toolchain, no cmake, no build step. Leaving is one command too: teton uninstall removes the service, the model and its data, the binaries, and the tap — after showing you the plan and asking once.

Two promises, both made visible

Cost control

Each phase of work goes to the tier you choose: architecture to a frontier model, implementation to a cheap one executing from well-specified task artifacts, mechanical I/O to the local daemon. A live cost meter shows per-phase attribution and measured savings against an all-frontier baseline — so the routing pays for itself in numbers you can read, not in claims.

Privacy boundaries

Mark paths as local-only and their content never leaves your machine. The rule is enforced at the daemon's single egress point — one place where anything can go out — and verified by egress-capture tests, not vibes.

Base camp and summits

Base camp: a slim local model, always with you

Teton Code probes your machine, benchmarks the best-fitting local model, and runs it as a persistent daemon. It carries the always-on cheap tier: routing, summarization, commit messages, secret redaction, and offline fallback.

Summits: bring your own models

Register Anthropic or any OpenAI-compatible endpoint — DeepSeek, Kimi, Ollama, vLLM — with your own keys. A mountain range is a range: of peaks, of sizes, of routes. So is your model lineup.

What runs where

Local inference support, per platform.
Platform Local inference
macOS, Apple Silicon (arm64) Metal GPU acceleration
macOS, Intel (x86_64) CPU only
Linux x86_64 (glibc) CPU only
Windows Not supported

Remote models run the same everywhere; only the local base-camp model depends on your hardware. On Linux, brew services is not a v1 claim — run teton-code yourself, or write a systemd user unit.

The model arrives on first run, with your consent

The install is small because it ships no weights. The first time you run teton, it proposes a local model matched to your hardware and names that model's exact download size and RAM floor before fetching anything — today's catalog spans roughly 1 GB for the smallest model to about 18 GB for the largest, and which one you are offered depends on your machine. Nothing downloads until you say yes.