🔒 Runs entirely on your machine

AI that never leaves your device

EditVolt ships with on-device models for chat, the agent, inline completions, and codebase search. No API key, no account, no round-trip to the cloud — and it keeps working with no internet at all.

How it works

Frontier-style assistance, computed locally

On first run, EditVolt offers a local chat model sized for your hardware. Everything downstream — reasoning, completion, and retrieval — happens on your CPU, GPU, or Apple Neural Engine.

① Pick a model

Sized for your hardware, chosen in one click

EditVolt inspects your RAM and disk and recommends a chat model that will run smoothly. Choose model opens the full catalogue with the footprint of each option, so you always know what fits before you download.

  • No sign-up and no key — chat and the agent work out of the box
  • Swap tiers any time from EditVolt Settings → AI Models
  • Downloads are consented, resumable, and show progress in the status bar
⚡ On-device models

Detected: 32 GB RAM · Apple M-series
Gemma 4 12B — recommended
○ Gemma 4 E4B — lighter footprint
○ Nemotron Nano 12B v2 — alt lane

Downloading Gemma 4 12B · 7.0 GB · 62%
② Verified & stored locally

Downloaded once, checked, and kept on disk

Model weights are fetched over HTTPS from Hugging Face and verified with a SHA-256 hash before use. They live under ~/.editvolt/models — delete the folder and they're gone. Point the download source at an internal mirror if your org requires it.

  • SHA-256 verification on every downloaded model
  • Repoint product.models.baseUrl to an internal mirror
  • Restrict which licenses you accept with product.models.allowedLicenses
~/.editvolt/models/
├─ gemma-4-12b/ 7.0 GB · sha256 ✓
├─ code/ completion engine
└─ embed/ MiniLM-L6-v2

HTTPS only · hash-verified
0 bytes sent to the network
③ Runs on every surface

Chat, agent, completions, and search — all local

The same local-first architecture powers the whole IDE. Inline completions run on a bundled engine, and your codebase is embedded on-device so semantic search never uploads a line of code.

  • Inline completions on a bundled local engine — offline, zero tokens sent
  • Codebase indexed with a bundled embedding model (MiniLM-L6-v2)
  • Agent tools, @Workspace answers, and search all query the local index
⚡ Surfaces · local

Chat & agent — Gemma 4 12B
Completions — bundled code engine
Embeddings — MiniLM-L6-v2
Works with the network cable pulled
The catalogue

On-device chat models

Every model below runs with no key and no account. The default is Gemma 4 12B; lighter tiers trade some quality for speed and a smaller footprint, and Qwen3.8 27B goes the other way if your machine has the memory for it.

Model Size Best for
Gemma 4 E2B 3.3 GB Snappy on most machines; lightest footprint
Gemma 4 E4B 5.2 GB Middle tier — balance of speed and quality
Gemma 4 12B 7.0 GB Best quality — the default chat tier
NVIDIA Nemotron Nano 12B v2 7.5 GB Alternative 12B lane
Qwen3.8 27B 19.0 GB Strongest open coding model — needs 32 GB+ RAM
⚡ Set your tier with product.models.chatTierauto, e2b, e4b, 12b, or off. Inline-completion size is controlled separately by product.models.localTier.
Why on-device

Privacy and reliability, by architecture

Running locally isn't just a privacy setting — it changes what the product can promise.

🔒

Your code stays put

Prompts, files, and embeddings are processed on your machine. Nothing is sent to EditVolt — we run no servers and never see your code.

✈️

Works offline

Once a model is downloaded, chat, the agent, and completions keep working on a plane, in a SCIF, or with the network cable pulled.

💸

No per-token bill

Local inference has no metered cost. Use it as much as you like — there are no usage quotas and nothing to meter.

🛡️

Strict privacy mode

One setting forces every AI request — chat, agent, and embeddings — onto on-device models and refuses any escalation to the cloud.

🏢

Air-gap friendly

Repoint downloads to an internal mirror and restrict accepted licenses — a clean fit for regulated and disconnected environments.

🔌

Cloud when you want it

Prefer a frontier model for a heavy task? Connect your own Claude, Gemini, or OpenAI-compatible key — local stays the default.

FAQ

On-device, answered

Do I need an account or API key to use on-device models?

No. Chat and the agent work with no key and no account — signing in is entirely optional and unrelated to AI features. On first run EditVolt offers a local chat model sized for your machine, and inline completions run on a bundled local engine with no key either.

Where do the models come from, and are they safe to download?

Weights are fetched over HTTPS from Hugging Face by default and verified with a SHA-256 hash before use. You can repoint product.models.baseUrl to an internal mirror, and downloads always require your explicit consent.

How much disk and memory do I need?

It depends on the tier — from 3.3 GB for Gemma 4 E2B, up to 7.5 GB for Nemotron Nano 12B v2, or 19.0 GB for Qwen3.8 27B on a machine with 32 GB+ of RAM. EditVolt recommends a model that fits your hardware, and you can drop to a lighter tier at any time.

Is my code used to train these models?

No. The models run locally and are not trained on your code. Indexing happens on your machine, and with strict privacy mode enabled nothing leaves your device at all.

Can I still use a cloud model when I want to?

Yes. Connect your own Claude, Gemini, Hugging Face, or any OpenAI-compatible key under EditVolt Settings → AI Models. Requests go straight to the provider with no middleman — but on-device stays the default.

Run AI on your own terms ⚡

Download EditVolt and pair with a model that never leaves your machine.

⚡ Download EditVolt