Skip to content

How it works

What happens between your word and the reply.

A wake word on the device, speech to text while you talk, a fast path or a local model, skills, memory and a voice of its own, all on hardware you own.

The Domia character behind a small glowing node, under an arc of four steps: listen, understand, think, speak

The voice pipeline: one spoken turn

Pick a real turn and press play: every bar is a stage of that turn, at the moment it really happened.

This turn: first sound 2.2 s after the last word.

Pick a turn
  • sound
  • thinking
  • fast path
  • skill call

end of speech detected
silence + a small turn model agree
first sound
2.2 s after the last word
  1. You speak
    microphone · captured on the device until you pause
  2. It understands
    speech to text · the transcript is ready right after you stop
  3. It decides
    routing · a known command runs at once; otherwise the model takes it
  4. It thinks
    language model · builds the prompt, then writes the reply
  5. It acts
    skipped in this turn
  6. It talks
    text to speech · sentence by sentence
  7. You hear it
    the reply · playback starts with the first sentence

After the turn

  • The microphone re-opens: just answer, no wake word.
  • Talk over the reply and it stops and listens.
  • Later, when idle, it reflects in the background: it updates its mood and remembers new facts.

Timings come from the pipeline trace of each captured turn, on a capable desktop machine. In the project's own benchmarks on a hub-class machine, a conversation starts speaking about 1.5 s after you stop, and a known command in 600 to 800 ms.

You speak
The wake word and voice activity detection run on the device; your speech is captured until a pause is detected. captured on the device until you pause
It understands
With a live microphone the text builds while you are still talking. These turns were sent as recordings, so here it runs right after the last word. the transcript is ready right after you stop
It decides
A known command skips the model; anything else goes to it. The routing section below shows the decision with your own phrases. a known command runs at once; otherwise the model takes it
It thinks
The prompt is assembled from the persona, the current mood, facts about you, your notes and the last turns. Words then stream out as they are generated. builds the prompt, then writes the reply
It acts
The model sees a short list of the skills that fit the request. If it calls one, the node checks that skill's policy, runs the call (several at once if needed) and replies from the skill's own answer or from a second pass of the model that reads the result. only when the model decides to call a skill
It talks
As soon as a sentence is complete it is turned into speech, while the model keeps writing the next one. sentence by sentence
You hear it
Playback starts with the first sentence; the rest queue behind it without gaps. playback starts with the first sentence
end of speech detected
silence + a small turn model agree
See the timings of every turn in the console

Fast path or model?

Every request first looks for a fast path: a known command is handled at once, with no model involved. Everything else goes to the language model, which sees the skills that fit and decides whether to use one or simply answer. Try a phrase and watch which way it goes.

Try a phrase

turn off the kitchen light

Is it a known command?

Yes: Home Assistant does it: turn off the kitchen light. On a real node the device responds right away

Yes

Handled without the model

Matched against the command templates your skills bring.

Home Assistant does it: turn off the kitchen light.

On a real node the device responds right away

No

The model decides

It sees the skills that fit the request and either calls one or answers in its own words.

The command templates are checked against a test corpus: a sentence that is not a command never matched one.

Memory: what Domia remembers

Five layers, all kept on the node. All of them can reach the prompt: recent turns always, facts and knowledge when they are relevant, episodes and the user model once per session. Reflection writes facts, episodes and the user model when it is idle.

  1. in the prompt

    The last exchanges, in the prompt.

    "And the kitchen too." Domia knows which light you mean.

    written bya turn

  2. in the prompt

    Things you told it, in the prompt when relevant.

    "The dog is called Nala and sleeps in the kitchen."

    written byreflection

  3. in the prompt

    What you wrote for it (wifi, house rules), in the prompt when relevant.

    "Guest wifi is casa-guest. Bins go out on Tuesday night."

    written byyou

  4. in the prompt, once per session

    Summaries of past sessions, in the prompt once per session.

    "Last week you planned a dinner for six and asked for a shopping list."

    written byreflection

  5. in the prompt, once per session

    Who you are to it: interests, habits and how familiar you are.

    "Prefers short answers and no small talk."

    written byreflection

Everything stays on the node. Nothing is persisted on a peer; context travels inside a delegated request.

Recent turns
A rolling window of the latest exchanges from the current session. It is always in the prompt, so a follow-up resolves without repeating yourself.
Facts
Extracted after the turn by a reflection pass in the background, never by the reply itself. Recalled by relevance to what you just said.
Knowledge base
Notes you author in the console or through the API. They are searched by relevance and added to the prompt when they match the request.
Episodes
A short summary written by reflection once a session ends. The latest ones are loaded into the prompt once per session, not searched on every turn.
User model
A short picture of you: interests, tendencies and preferences. Reflection updates it between sessions.
See the Memories screen in the console

Deployments: one device, a hub, or a mesh

Every Domia node runs the same software; a console template decides whether it is a standalone companion, a full hub for a whole property, or one of several hubs sharing the work. Every arrow below stays inside your network.

Choose:

One hub listens through a speaker in every room.

  • Domia nodehub · hub-class machine · Nova, Sage, Ember · memory · stored on this node
  • Home Assistant Voice PEfactory firmware · binds to Nova
  • Browser roombrowser or app · binds to Sage
  • Your own boardbuilt by you · Domia's reference protocol · binds to Ember
  • Wyoming satellitecommunity devices · binds to Nova
Domia node
The same software as every other node. Its role comes from the template you apply in the console, not from a different build.
Identity
An identity has its own voice, persona and memory.
Delegated identity
Milo's persona and context travel inside the request; nothing is stored on hub B.
Wake word
The wake word is spotted before anything else runs.
Speech to text
Streaming, on this node, one of eight engines.
Routing
A known command takes the fast path and runs without the model. Anything else goes to the model, which gets the skills that fit and decides whether to call one or answer.
Language model
A local model served on your own hardware; nothing leaves your network.
Text to speech
One of six engines, one voice per identity.
Memory
Per-identity memory stored on this node. Nothing is stored on a peer.
Microphone
Any built-in or USB microphone; audio stays on the machine.
Speaker
The reply is synthesised here and played on the same device.
Home Assistant Voice PE
Off-the-shelf hardware running its factory firmware; nothing to change.
Browser room
A room that browsers and apps can join.
Your own board
Built on Domia's reference protocol, for satellites you make yourself.
Wyoming satellite
Works with any satellite that speaks Wyoming.
Hub-class peer
A peer with headroom; neighbours can borrow its language model.
A smaller machine
Speech to text and text to speech run on board; its language model runs on hub B.
Any device, same behaviour
Every kind of satellite is handled the same way once connected; swap hardware freely.
Apps can talk to it too
Apps that speak the OpenAI Realtime API connect straight to the node and go through the same stages.

Satellites: a microphone in every room

Four kinds of satellite, one behaviour: the device listens and plays, and everything else runs on your Domia node.

Your Domia node

Speech to text, routing, the model, the voice and memory all run here. Every kind of satellite is handled the same way: the device listens and plays, the node does the rest.

Choose:

What this device does

Home Assistant Voice PE

the node connects to the device

On the device

  • wake word
  • microphone
  • speaker
  • LEDs

Between device and node

  • microphone audio
  • reply audio in the format the device advertises
  • announcements

Good to know

  • Follow-up mode works on the factory firmware and is on by default.
  • The device plays only the format it advertises, so the node converts every reply for it.
  • The node notices when a device stops answering.

Apps can talk to it too

Apps built for the OpenAI Realtime API connect to the node and go through the same stages.

Any of these devices, same behaviour
Every kind of satellite is handled the same way: the device listens and plays, the node does the rest.
Set up a satellite

Skills: what it can do

Built-in tools, ready-made skills for Home Assistant and Music Assistant, any MCP server you add, and routines that chain them.

Switch a group off to see what changes:

The bolt marks what the fast path answers without the model. “Ask first” marks the few actions that confirm before they run; you can change that for each skill.

Built-in

On from the start
13 tools every node ships with. 6 of them exist only for the fast path and are not shown to the model.
  • timefast path
  • datefast path
  • timerfast path
  • reminderfast path
  • alarmfast path
  • remember
  • forgetask first

Home Assistant

Once connected
An MCP-backed skill for lights, switches, covers, climate and scenes. A few of its tools are blocked because a built-in tool already covers them.
  • lightsfast path
  • switchesfast path
  • coversfast pathask first
  • climatefast path
  • fansfast path
  • vacuumfast path
  • mowerfast path
  • lock and unlockask first
  • alarm panelsask first
  • sirensask first
  • live state

Music Assistant

Once connected
An MCP-backed skill. Playback and volume take the fast path; searches and queueing go through the model.
  • play
  • pausefast path
  • skipfast path
  • volumefast path
  • what's playingfast path
  • search

Your MCP server

Once connected
Any tool, no code on the node.

Point a node at any MCP server and its tools become skills.

A server can ship phrases and hints for its tools; the node keeps the final say on what may run.

Routines

On from the start
One tool that runs several steps in order.

"Good night"

  1. Turn off the living room lights
  2. Pause the music
  3. Set an alarm for seven

Up to 8 ordered steps, no routines inside routines. The strictest step policy applies to the whole routine, and it answers with one reply.

What is on from the start. Built-in tools work from the first turn. Home Assistant, Music Assistant and other MCP servers start working once you connect them in the console. Each one can be switched off per identity.

Not shown to the model
Only the fast path can reach this.

The engines behind a turn

Every stage is a swappable engine chosen in the configuration. These are the ones Domia ships with or talks to.

  • Wake word

    A small keyword spotter that works on the device that listens.

    • sherpa-onnx keyword spotting
  • Speech to text

    Transcribes while you speak, inside the node or through a local speech server.

    • Whisper
    • Moonshine
    • Zipformer
    • Parakeet
    • Nemotron
    • NeMo-Speech
    • OpenAI-compatible server
  • Turn detection

    Silence plus a small model decide when you have finished, not a fixed timeout.

    • Silero VAD
    • Smart Turn
  • Language model

    A language model served on your own hardware. Reflection can use a second, smaller model so it stays out of the way of the voice.

    • Ollama
    • llama.cpp
    • OpenAI-compatible server
  • Text to speech

    One voice per identity, set in the configuration.

    • Kokoro
    • Pocket
    • VITS
    • Kitten
    • Matcha
    • Supertonic
  • Memory and search

    Search over what Domia remembers runs on the node. Facts, episodes and the knowledge base live in a local database on the node.

    • SQLite
    • local embeddings

Works with

  • Home Assistant
  • Music Assistant
  • MCP servers
  • ESPHome
  • Wyoming
  • LiveKit

Speech engines run inside the node or in a local server beside it, and the language model runs in a local server.

How to run it on your hardware

What stays on your network

  • Stays on your network
    Audio, transcripts and memory never touch a third party. There are no accounts and no cloud fallback.
  • Mesh secured by default
    Nodes need a shared secret before they talk to each other. Encrypted transport, secret rotation and network access lists are opt-in.
  • Each Domia keeps its own memory
    Every identity has its own facts, notes and mood. What one room hears is not shared with another unless you decide so, and memory can be switched off per identity.

See it in the code

Everything on this page is open source. Read it, run it, change it.