Media pipeline
How call audio becomes text, and how a reply becomes sound in the other person's ear, with the latency and quality rules in between.
Audio moves in two directions through PhoneGate. Receive (RX) carries the caller's voice to the server. Transmit (TX) carries the reply back into the call.
Receive: voice to text
- Capture. After the call starts, a
StreamRecorderon the phone opens the voice downlink and writes 16 kHz mono PCM to its output. - Frames. The daemon sends 20 ms frames to the server. The send queue is capped at 80 ms, so a slow link drops old audio instead of building up delay.
- Voice activity. A WebRTC voice activity detector requires a continuous speech onset and at least 100 ms of voice. Clicks and rustling are rejected by energy, noise floor and zero-crossing rate.
- Recognition. When Groq keys are configured, one Whisper request goes over a warm keep-alive connection. A local
faster-whisper tinymodel is the fallback. - Trust check. The recogniser's output is judged by the real
avg_logprob,no_speech_prob,compression_ratio, speech rate and a list of known silence hallucinations. The panel shows the computed confidence, and the reasons for a rejection are written to the log.
Transmit: text to voice
- Synthesis. A local Piper voice creates 22.05 kHz mono PCM with no network wait. Edge TTS stays available as a backup.
- Pre-delivery. After a pause in dialling, and for short phrases, the PCM is compressed ahead of time, sent to the phone and loaded into an
AudioTrack. - Playback. A persistent
UplinkPlayerServiceswitches the Samsung audio HAL on only while transmitting and releases it immediately, so it never blocks receiving.
caller ─▶ downlink ─▶ 20 ms frames ─▶ VAD ─▶ Whisper ─▶ transcript
reply ─▶ Piper ─▶ PCM ─▶ AudioTrack ─▶ uplink HAL (only while speaking) ─▶ caller
Design notes
- Bounded queues over completeness. A phone call is a live conversation. Late audio is worse than lost audio, so every buffer has a hard limit.
- Local first. Synthesis and the fallback recogniser run on the server, so a call still works when an external provider is down.
- Honest confidence. Recognition results carry a score computed from real model signals, not a guess, and weak results are discarded rather than shown as fact.
Hardware dependent
The transmit path relies on how Samsung's audio HAL routes audio to the cellular uplink. It was developed against one handset and is documented as an experiment, not as a portable technique.