Case study · Rimbalo
Speech recognition on the phone, in airplane mode
Rimbalo is professional dictation for Android and iOS, written in Kotlin Multiplatform. Every step from microphone to text runs on the device, and the app never opens a network connection.
Two hard constraints. The Android build has no network permission, and a build task fails if one appears. And there is no generative model: the app writes nothing the user didn't say.
The pipeline
- Voice activity detection. Silero VAD on LiteRT cuts the microphone stream into utterances at 700 ms of silence, or at 19 s.
- Features. A Kotlin port of NeMo's log-mel preprocessor turns each utterance into the model's input.
- Encoder. NVIDIA's FastConformer-TDT (Parakeet), converted to LiteRT. On Android through the
CompiledModelAPI; on iOS through a Swift bridge to LiteRT. - Predictor and joint networks. Two small graphs called once per token, at under 2.5 ms a call.
- Decoding. A Kotlin TDT decoder, greedy or beam search, biased toward the user's own word list so client names and citations come out spelled their way. It reproduces NeMo's tokens, frames and probabilities, and a parity test holds it there in CI.
The models are not in the app binary. They arrive as store asset packs (Play Asset Delivery on Android, Background Assets on iOS): about 140 MB for English, and a larger shared model for the 24 other languages.
Choosing the model
The multilingual Parakeet v3 has 627 M parameters. NVIDIA's English-only model has 115 M. Both were converted and scored on the same 186 LibriSpeech clips, at every quantisation, before anything was built on top of them.
| Model | Encoder | Download | Word error rate |
|---|---|---|---|
| v3, 25 languages, full precision | — | — | 1.74% |
| v3, mixed precision | 424 MB | ≈460–500 MB | 1.66% |
| English 110m, full precision | — | — | 1.94% |
| English 110m, fp16 | 227 MB | ≈249 MB | 1.94% |
| English 110m, int8 (chosen) | 112 MB | ≈134 MB | 2.07% |
| English 110m, 4-bit | 71 MB | ≈93 MB | 2.11% |
English-only costs about 0.2 points of word error rate, and the extra errors are mostly proper names. Those are exactly what a user's word list fixes, so the smaller model plus biasing beats the bigger model's download. The smaller model is about twice as sensitive to quantisation, which is why it ships at int8 and not 4-bit: the 4-bit version changes ordinary words ("formerly" to "formally", "ranch" to "branch").
What broke on the phones, and the fixes
The encoder converted cleanly and then failed on the devices in three different ways. Each fix was proven numerically identical to NeMo before conversion, and the export script refuses to convert otherwise.
Same wrong answer on two different GPUs
At fp16, the Pixel 7's OpenCL backend and the iPhone's Metal backend both returned garbage, with a cosine similarity to PyTorch of −0.0343 on both, identical to 16 digits. Two unrelated GPU stacks producing the same wrong number rules out a driver bug and points at the graph. The cause: on real speech the residual stream reaches about 10⁴, and LayerNorm squares it, past fp16's limit of 65,504. The fix runs LayerNorm on x/256 with eps/256². LayerNorm is scale-invariant and 256 is a power of two, so the fp32 output is bit-identical, while the largest square shrinks 65,536 times. Cosine similarity went from −0.03 to 0.9997.
Integer and boolean ops the GPU delegates don't have
Every padding mask was derived inside the graph from an integer length, which Metal and OpenCL can't run and the iOS
CPU path rejected outright. The encoder was rewritten to take its masks as float inputs: convolution masking becomes
x · keep, attention becomes s · keep + (1 − keep) · (−10000). The graph has no integer or
boolean logic left, and its output matches NeMo with a maximum absolute difference of 0.0.
Weights the iOS runtime couldn't read
The converter stored 638 of the encoder's constants as external buffers after the FlatBuffer, which the iOS runtime of the day could not load. A post-processing step moves them inline, 16-byte aligned, and verifies the outputs are unchanged. It now runs on every export.
A dictation session on three phones
The int8 English model ships on the CPU through XNNPACK, with a weight cache and a per-device thread limit. Every quantised variant decodes 20 s of speech in under 0.4 s on either phone's CPU, so the GPU isn't needed. To test it, 54 real recordings, 39 minutes of speech, were played through the app's own VAD and recognizer at speaking pace as one continuous 44-minute session.
| Pixel 1 (2016) | Pixel 7 | iPhone SE 3 | |
|---|---|---|---|
| Wait for the last text after release: median / worst | 1.77 / 4.11 s | 0.85 / 1.82 s | 0.29 / 0.57 s |
| Real-time factor (decode time ÷ audio time) | 0.22–0.24 | 0.10–0.14 | 0.036–0.039 |
| CPU cores used, on average | 0.56 | 0.24 | 0.07 |
| Peak memory | 528 MB | 605 MB | 325 MB |
| Model load, cache warm | 545 ms | 751 ms | 138 ms |
Every phone kept up with speech for the whole session, and memory stayed flat. The Pixel 7 reached light thermal throttling at minute 8. A second run, unplugged, showed the cause was fast charging, not dictation.
The model that came out
Rimbalo's File Note pack turns a lawyer's dictation into a structured file note. The first design had an on-device LLM write its summaries, so three candidates ran on LiteRT-LM with schema-constrained JSON decoding: Qwen3-1.7B, Qwen3-4B-Instruct and Phi-4-mini. They were scored on synthetic dictations, including sensitive ones.
- Qwen3-1.7B never refused. It sanitised instead: of 34 narrative sections about sensitive matters, 21 softened or left out the facts the note exists to record, and 5 contradicted the dictation.
- Qwen3-4B, the model chosen, still put a false statement in 22% of routine summaries (95% CI 16–31%), and left out a decision or the main issue in 26% more. Prompt variants aimed at those errors didn't fix them.
Nothing on the phone could flag those errors before a lawyer signed the note, so the generative model was removed. Every section is now built by rules from what the lawyer said, measured on held-out corpora. The model download and the 8 GB memory floor for packs went away, and the app can't put words in the user's mouth.
Diamond Edge