How it works
What happens between your word and the reply.
A wake word on the device, speech to text while you talk, a fast path or a local model, skills, memory and a voice of its own, all on hardware you own.

The voice pipeline: one spoken turn
Pick a real turn and press play: every bar is a stage of that turn, at the moment it really happened.
This turn: first sound 2.2 s after the last word.
- sound
- thinking
- fast path
- skill call
- end of speech detected
- silence + a small turn model agree
- first sound
- 2.2 s after the last word
After the turn
- The microphone re-opens: just answer, no wake word.
- Talk over the reply and it stops and listens.
- Later, when idle, it reflects in the background: it updates its mood and remembers new facts.
Timings come from the pipeline trace of each captured turn, on a capable desktop machine. In the project's own benchmarks on a hub-class machine, a conversation starts speaking about 1.5 s after you stop, and a known command in 600 to 800 ms.
- You speak
- The wake word and voice activity detection run on the device; your speech is captured until a pause is detected. captured on the device until you pause
- It understands
- With a live microphone the text builds while you are still talking. These turns were sent as recordings, so here it runs right after the last word. the transcript is ready right after you stop
- It decides
- A known command skips the model; anything else goes to it. The routing section below shows the decision with your own phrases. a known command runs at once; otherwise the model takes it
- It thinks
- The prompt is assembled from the persona, the current mood, facts about you, your notes and the last turns. Words then stream out as they are generated. builds the prompt, then writes the reply
- It acts
- The model sees a short list of the skills that fit the request. If it calls one, the node checks that skill's policy, runs the call (several at once if needed) and replies from the skill's own answer or from a second pass of the model that reads the result. only when the model decides to call a skill
- It talks
- As soon as a sentence is complete it is turned into speech, while the model keeps writing the next one. sentence by sentence
- You hear it
- Playback starts with the first sentence; the rest queue behind it without gaps. playback starts with the first sentence
- end of speech detected
- silence + a small turn model agree
Fast path or model?
Every request first looks for a fast path: a known command is handled at once, with no model involved. Everything else goes to the language model, which sees the skills that fit and decides whether to use one or simply answer. Try a phrase and watch which way it goes.
Is it a known command?
Yes: Home Assistant does it: turn off the kitchen light. On a real node the device responds right away
The command templates are checked against a test corpus: a sentence that is not a command never matched one.
Memory: what Domia remembers
Five layers, all kept on the node. All of them can reach the prompt: recent turns always, facts and knowledge when they are relevant, episodes and the user model once per session. Reflection writes facts, episodes and the user model when it is idle.
Everything stays on the node. Nothing is persisted on a peer; context travels inside a delegated request.
- Recent turns
- A rolling window of the latest exchanges from the current session. It is always in the prompt, so a follow-up resolves without repeating yourself.
- Facts
- Extracted after the turn by a reflection pass in the background, never by the reply itself. Recalled by relevance to what you just said.
- Knowledge base
- Notes you author in the console or through the API. They are searched by relevance and added to the prompt when they match the request.
- Episodes
- A short summary written by reflection once a session ends. The latest ones are loaded into the prompt once per session, not searched on every turn.
- User model
- A short picture of you: interests, tendencies and preferences. Reflection updates it between sessions.
Deployments: one device, a hub, or a mesh
Every Domia node runs the same software; a console template decides whether it is a standalone companion, a full hub for a whole property, or one of several hubs sharing the work. Every arrow below stays inside your network.
One hub listens through a speaker in every room.
- Domia node
- The same software as every other node. Its role comes from the template you apply in the console, not from a different build.
- Identity
- An identity has its own voice, persona and memory.
- Delegated identity
- Milo's persona and context travel inside the request; nothing is stored on hub B.
- Wake word
- The wake word is spotted before anything else runs.
- Speech to text
- Streaming, on this node, one of eight engines.
- Routing
- A known command takes the fast path and runs without the model. Anything else goes to the model, which gets the skills that fit and decides whether to call one or answer.
- Language model
- A local model served on your own hardware; nothing leaves your network.
- Text to speech
- One of six engines, one voice per identity.
- Memory
- Per-identity memory stored on this node. Nothing is stored on a peer.
- Microphone
- Any built-in or USB microphone; audio stays on the machine.
- Speaker
- The reply is synthesised here and played on the same device.
- Home Assistant Voice PE
- Off-the-shelf hardware running its factory firmware; nothing to change.
- Browser room
- A room that browsers and apps can join.
- Your own board
- Built on Domia's reference protocol, for satellites you make yourself.
- Wyoming satellite
- Works with any satellite that speaks Wyoming.
- Hub-class peer
- A peer with headroom; neighbours can borrow its language model.
- A smaller machine
- Speech to text and text to speech run on board; its language model runs on hub B.
- Any device, same behaviour
- Every kind of satellite is handled the same way once connected; swap hardware freely.
- Apps can talk to it too
- Apps that speak the OpenAI Realtime API connect straight to the node and go through the same stages.
Satellites: a microphone in every room
Four kinds of satellite, one behaviour: the device listens and plays, and everything else runs on your Domia node.
Choose:
What this device does
Home Assistant Voice PE
the node connects to the device
On the device
- wake word
- microphone
- speaker
- LEDs
Between device and node
- microphone audio
- reply audio in the format the device advertises
- announcements
Good to know
- Follow-up mode works on the factory firmware and is on by default.
- The device plays only the format it advertises, so the node converts every reply for it.
- The node notices when a device stops answering.
- Any of these devices, same behaviour
- Every kind of satellite is handled the same way: the device listens and plays, the node does the rest.
Skills: what it can do
Built-in tools, ready-made skills for Home Assistant and Music Assistant, any MCP server you add, and routines that chain them.
Switch a group off to see what changes:
The bolt marks what the fast path answers without the model. “Ask first” marks the few actions that confirm before they run; you can change that for each skill.
What is on from the start. Built-in tools work from the first turn. Home Assistant, Music Assistant and other MCP servers start working once you connect them in the console. Each one can be switched off per identity.
- Not shown to the model
The engines behind a turn
Every stage is a swappable engine chosen in the configuration. These are the ones Domia ships with or talks to.
Works with
- Home Assistant
- Music Assistant
- MCP servers
- ESPHome
- Wyoming
- LiveKit
Speech engines run inside the node or in a local server beside it, and the language model runs in a local server.
How to run it on your hardwareWhat stays on your network
See it in the code
Everything on this page is open source. Read it, run it, change it.