Building voice AI for someone on a mouthpiece ventilator

Share
Building voice AI for someone on a mouthpiece ventilator
Back in the month of January 2026, Our voice assistant worked in the office. It did not work in our first client's flat.
As Our client uses a mouthpiece ventilator — a gooseneck arm holding a mouthpiece near his face, from which he takes assisted breaths. He speaks between them. The system had passed every test we ran on ourselves. In his living room it caught roughly one command in ten, and confidently transcribed his wake phrase as something he had never said.

What the ventilator does to speech

Phrases are bounded by breath, not syntax — the phrase ends when the air does, often mid-clause. Amplitude decays across each phrase, so a sentence is a descending ramp rather than a flat signal. The pause between phrases is respiratory, not semantic: longer than a conversational pause, and meaning the opposite of finished. And the machine cycles audibly, in the same amplitude band as the quiet end of his speech.

Each of those breaks a different default.

Measure before you buy hardware

After three days of swapping microphones we did what we should have done first. Sixty recordings, his room, his volume, ventilator running, RMS logged per utterance.

The spread was 107 to 2850. Same speaker, same session.

That reframed everything. We had been treating this as a quiet-signal problem — better mic, get closer, add gain. It was never a level problem. It was a variance problem, and almost every fix for the first makes the second worse.

The lavalier proved it. Clipped to his collar, close to the mouth, exactly what instinct wants: 6712 on phrase onsets, 145 on tails. Clipping at one end, below the noise floor at the other, sometimes in the same sentence. Proximity amplifies loud parts more than it rescues quiet ones. A desk mic at fixed distance won — not a better microphone, just one that removed a source of variance instead of adding one.

Two defaults with no valid setting

Our capture pipeline gated speech on a fixed amplitude threshold. Set it low enough to catch his phrase tails and it triggered continuously on the ventilator. Set it high enough to reject the machine and it cut the last third of every sentence. There was no correct value, because his speech and the noise floor overlap. Any system deciding speech-versus-silence on amplitude alone fails this speaker regardless of tuning.

Endpointing had the same shape. Voice activity detection closes an utterance after a fixed silence, typically 500–800 ms. His breath pauses are longer. One sentence became three fragments, each transcribed separately, each turned into a fluent unrelated sentence by a model doing its best with a scrap. We extended the timeout to his respiratory rhythm and accepted the responsiveness cost.

The model was not mishearing him. It was substituting.

Whisper transcribed his wake phrase as "Hey Bob." Reliably. Zero correct out of ten.

Whisper's decoder is a language model. Given weak acoustic evidence it does not return uncertainty — it returns the most probable token sequence under its prior. "Hey Neo" is not probable. "Hey Bob" is. So it produced "Hey Bob," fluently, with no low-confidence flag attached.

The fix was five words:

python

initial_prompt = "Neo. Hey Neo. Neo is a voice assistant."

0/10 to 10/10 on retest, 29 of 30 across a wider sample. Days of hardware work, and the defect was that the model had never been told the word existed.

There is a worse version of this. On very low-amplitude, out-of-vocabulary audio, bare Whisper produced racial slurs — repeatably, not as a fluke. Same mechanism: thin evidence, harder fallback to the prior, and the prior contains everything that was in the training data. For a device that speaks aloud in someone's living room this is unacceptable, and the people most exposed are exactly those whose speech is quietest and least like the training distribution. Vocabulary prompt, confidence gate, output filter — all three, not one.

Where it landed

Desk mic at fixed distance, rolling noise-floor estimation, breath-tuned endpointing, vocabulary prompt, confidence gating: 97% transcription accuracy at his natural speaking volume.

What generalises

  • Measure before buying hardware. Log RMS per utterance.
  • Distinguish level from variance. They demand opposite fixes.
  • Treat endpoint timeouts as an accessibility setting, not a performance knob.
  • Never act on an ungated transcription. Low-amplitude audio is where fluent wrong answers come from.

None of these defaults are neutral. A 700 ms timeout encodes an assumption about how long a body takes between phrases. An RMS gate encodes an assumption about how much air a person has. A decoder prior encodes an assumption about which words are worth expecting. None were written to exclude anyone. All of them do.

They persist because the people they exclude are not in the benchmark sets. A model can score well on every standard evaluation and still be unusable by someone with a ventilator arm in front of his face — and nothing in the evaluation will tell you, because he was never in it.


Neo is a wellbeing, safety-alerting and care-coordination companion. It is not a diagnostic or therapeutic device and makes no clinical claims. This post describes speech recognition engineering only.

Read more