Offline voice assistant, NO dedicated GPU, less 15Gb RAM
Most voice assistants are walkie-talkies. You press, speak, wait. This one keeps the microphone open while it answers, so you can interrupt mid-sentence, change your mind, or correct yourself and get a reaction within one audio buffer. Everything runs locally. No cloud, no accounts, no round-trips, skills and tool-calling supported
What's interesting here
Barge-in has two intensities. A short murmur pauses playback so you can finish; a firm interruption drops the rest of the reply. The history keeps what was already said plus a marker that you cut in, and a quiet "mhmm" doesn't interrupt anything.
Turn-ending is a layered decision, not a timer. A per-frame speech detector, a pause threshold, and a classifier over the last few seconds of your utterance each vote on whether you're done talking. When they disagree, a four-token judge prompt answers LISTEN or REPLY and breaks the tie.
A tool gate sits between the model and your machine. Every round of proposed tool calls is checked against your actual last sentence; anything you didn't ask for is denied with a spoken explanation. This killed a whole class of phantom tool calls that appeared after interruptions.
Two brains share one conversation. One streams from a locally served model over a standard chat API. The other spawns a coding-agent CLI as a subprocess and inherits its tools, skills, and per-session memory. A checkbox swaps them at runtime.
History is interrupt-safe. Incomplete tool exchanges are dropped when a turn is cut, so skills aren't re-read every turn. Before each request the history is trimmed to fit the context window, oldest turns first, never the one being answered.
The model mirrors your language: Italian gets Italian, English gets English, and it follows you when you switch mid-conversation, turn by turn.
How it's wired
One hub process owns the 32 ms frame loop, the state machine, all queues, conversation history, and playback pacing. Around it sit three small services, each a configuration of one vendored native server with its own HTTP endpoint and health check.
Transcription runs full-duplex over a single socket: the client pushes audio and reads streamed partials on the same connection, hand-rolling the raw protocol because request/response can't do this. Synthesis streams audio in 20 ms slices that the hub paces against a clock, which is what lets pause, resume, and flush from a barge-in land within one buffer period. Turn-scoring feeds the hub's votes. The reasoning core is external; the launcher checks its health endpoint and refuses to start otherwise.
mic (16 kHz) ─► VAD ─► STT :8091 ─► HUB turn-end ─► LLM :8080 / agent CLI
│
sentence splitter ─► TTS :8092 ─► paced player
▲
barge-in hooksOne socket to the UI carries text events and binary audio in both directions; the newest client wins. Tools include datetime, timers, a notepad, and shell commands behind an approval policy (off / ask / auto), a whitelist, and rejection of chaining or substitution metacharacters. Skills are instruction folders enabled from the UI and never auto-discovered: only the ones you tick reach the prompt, read once per conversation, and their scripts skip the approval prompt. Every setting is an env var from one file; the UI mutates the same object at runtime, and only skill selection persists across restarts.