AI that never leaves your device
EditVolt ships with on-device models for chat, the agent, inline completions, and codebase search. No API key, no account, no round-trip to the cloud — and it keeps working with no internet at all.
Frontier-style assistance, computed locally
On first run, EditVolt offers a local chat model sized for your hardware. Everything downstream — reasoning, completion, and retrieval — happens on your CPU, GPU, or Apple Neural Engine.
Sized for your hardware, chosen in one click
EditVolt inspects your RAM and disk and recommends a chat model that will run smoothly. Choose model opens the full catalogue with the footprint of each option, so you always know what fits before you download.
- No sign-up and no key — chat and the agent work out of the box
- Swap tiers any time from EditVolt Settings → AI Models
- Downloads are consented, resumable, and show progress in the status bar
✓ Detected: 32 GB RAM · Apple M-series
● Gemma 4 12B — recommended
○ Gemma 4 E4B — lighter footprint
○ Nemotron Nano 12B v2 — alt lane
Downloading Gemma 4 12B · 7.0 GB · 62%
Downloaded once, checked, and kept on disk
Model weights are fetched over HTTPS from Hugging Face and verified with a SHA-256 hash before use.
They live under ~/.editvolt/models — delete the folder and
they're gone. Point the download source at an internal mirror if your org requires it.
- SHA-256 verification on every downloaded model
- Repoint
product.models.baseUrlto an internal mirror - Restrict which licenses you accept with
product.models.allowedLicenses
├─ gemma-4-12b/ 7.0 GB · sha256 ✓
├─ code/ completion engine
└─ embed/ MiniLM-L6-v2
✓ HTTPS only · hash-verified
✓ 0 bytes sent to the network
Chat, agent, completions, and search — all local
The same local-first architecture powers the whole IDE. Inline completions run on a bundled engine, and your codebase is embedded on-device so semantic search never uploads a line of code.
- Inline completions on a bundled local engine — offline, zero tokens sent
- Codebase indexed with a bundled embedding model (MiniLM-L6-v2)
- Agent tools, @Workspace answers, and search all query the local index
✓ Chat & agent — Gemma 4 12B
✓ Completions — bundled code engine
✓ Embeddings — MiniLM-L6-v2
✓ Works with the network cable pulled
On-device chat models
Every model below runs with no key and no account. The default is Gemma 4 12B; lighter tiers trade some quality for speed and a smaller footprint, and Qwen3.8 27B goes the other way if your machine has the memory for it.
| Model | Size | Best for |
|---|---|---|
| Gemma 4 E2B | 3.3 GB | Snappy on most machines; lightest footprint |
| Gemma 4 E4B | 5.2 GB | Middle tier — balance of speed and quality |
| Gemma 4 12B | 7.0 GB | Best quality — the default chat tier |
| NVIDIA Nemotron Nano 12B v2 | 7.5 GB | Alternative 12B lane |
| Qwen3.8 27B | 19.0 GB | Strongest open coding model — needs 32 GB+ RAM |
product.models.chatTier — auto, e2b, e4b, 12b, or off. Inline-completion size is controlled separately by product.models.localTier.Privacy and reliability, by architecture
Running locally isn't just a privacy setting — it changes what the product can promise.
Your code stays put
Prompts, files, and embeddings are processed on your machine. Nothing is sent to EditVolt — we run no servers and never see your code.
Works offline
Once a model is downloaded, chat, the agent, and completions keep working on a plane, in a SCIF, or with the network cable pulled.
No per-token bill
Local inference has no metered cost. Use it as much as you like — there are no usage quotas and nothing to meter.
Strict privacy mode
One setting forces every AI request — chat, agent, and embeddings — onto on-device models and refuses any escalation to the cloud.
Air-gap friendly
Repoint downloads to an internal mirror and restrict accepted licenses — a clean fit for regulated and disconnected environments.
Cloud when you want it
Prefer a frontier model for a heavy task? Connect your own Claude, Gemini, or OpenAI-compatible key — local stays the default.
On-device, answered
Do I need an account or API key to use on-device models?
No. Chat and the agent work with no key and no account — signing in is entirely optional and unrelated to AI features. On first run EditVolt offers a local chat model sized for your machine, and inline completions run on a bundled local engine with no key either.
Where do the models come from, and are they safe to download?
Weights are fetched over HTTPS from Hugging Face by default and verified with a SHA-256 hash before use.
You can repoint product.models.baseUrl to an internal mirror, and downloads always require your
explicit consent.
How much disk and memory do I need?
It depends on the tier — from 3.3 GB for Gemma 4 E2B, up to 7.5 GB for Nemotron Nano 12B v2, or 19.0 GB for Qwen3.8 27B on a machine with 32 GB+ of RAM. EditVolt recommends a model that fits your hardware, and you can drop to a lighter tier at any time.
Is my code used to train these models?
No. The models run locally and are not trained on your code. Indexing happens on your machine, and with strict privacy mode enabled nothing leaves your device at all.
Can I still use a cloud model when I want to?
Yes. Connect your own Claude, Gemini, Hugging Face, or any OpenAI-compatible key under EditVolt Settings → AI Models. Requests go straight to the provider with no middleman — but on-device stays the default.
Run AI on your own terms ⚡
Download EditVolt and pair with a model that never leaves your machine.
⚡ Download EditVolt