Diamond Edge

Case study · Rimbalo

Speech recognition on the phone, in airplane mode

Rimbalo is professional dictation for Android and iOS, written in Kotlin Multiplatform. Every step from microphone to text runs on the device, and the app never opens a network connection.

Two hard constraints. The Android build has no network permission, and a build task fails if one appears. And there is no generative model: the app writes nothing the user didn't say.

The pipeline

  1. Voice activity detection. Silero VAD on LiteRT cuts the microphone stream into utterances at 700 ms of silence, or at 19 s.
  2. Features. A Kotlin port of NeMo's log-mel preprocessor turns each utterance into the model's input.
  3. Encoder. NVIDIA's FastConformer-TDT (Parakeet), converted to LiteRT. On Android through the CompiledModel API; on iOS through a Swift bridge to LiteRT.
  4. Predictor and joint networks. Two small graphs called once per token, at under 2.5 ms a call.
  5. Decoding. A Kotlin TDT decoder, greedy or beam search, biased toward the user's own word list so client names and citations come out spelled their way. It reproduces NeMo's tokens, frames and probabilities, and a parity test holds it there in CI.

The models are not in the app binary. They arrive as store asset packs (Play Asset Delivery on Android, Background Assets on iOS): about 140 MB for English, and a larger shared model for the 24 other languages.

Choosing the model

The multilingual Parakeet v3 has 627 M parameters. NVIDIA's English-only model has 115 M. Both were converted and scored on the same 186 LibriSpeech clips, at every quantisation, before anything was built on top of them.

ModelEncoderDownloadWord error rate
v3, 25 languages, full precision——1.74%
v3, mixed precision424 MB≈460–500 MB1.66%
English 110m, full precision——1.94%
English 110m, fp16227 MB≈249 MB1.94%
English 110m, int8 (chosen)112 MB≈134 MB2.07%
English 110m, 4-bit71 MB≈93 MB2.11%

English-only costs about 0.2 points of word error rate, and the extra errors are mostly proper names. Those are exactly what a user's word list fixes, so the smaller model plus biasing beats the bigger model's download. The smaller model is about twice as sensitive to quantisation, which is why it ships at int8 and not 4-bit: the 4-bit version changes ordinary words ("formerly" to "formally", "ranch" to "branch").

What broke on the phones, and the fixes

The encoder converted cleanly and then failed on the devices in three different ways. Each fix was proven numerically identical to NeMo before conversion, and the export script refuses to convert otherwise.

Same wrong answer on two different GPUs

At fp16, the Pixel 7's OpenCL backend and the iPhone's Metal backend both returned garbage, with a cosine similarity to PyTorch of −0.0343 on both, identical to 16 digits. Two unrelated GPU stacks producing the same wrong number rules out a driver bug and points at the graph. The cause: on real speech the residual stream reaches about 10⁴, and LayerNorm squares it, past fp16's limit of 65,504. The fix runs LayerNorm on x/256 with eps/256². LayerNorm is scale-invariant and 256 is a power of two, so the fp32 output is bit-identical, while the largest square shrinks 65,536 times. Cosine similarity went from −0.03 to 0.9997.

Integer and boolean ops the GPU delegates don't have

Every padding mask was derived inside the graph from an integer length, which Metal and OpenCL can't run and the iOS CPU path rejected outright. The encoder was rewritten to take its masks as float inputs: convolution masking becomes x · keep, attention becomes s · keep + (1 − keep) · (−10000). The graph has no integer or boolean logic left, and its output matches NeMo with a maximum absolute difference of 0.0.

Weights the iOS runtime couldn't read

The converter stored 638 of the encoder's constants as external buffers after the FlatBuffer, which the iOS runtime of the day could not load. A post-processing step moves them inline, 16-byte aligned, and verifies the outputs are unchanged. It now runs on every export.

Along the way: Tensor G2 has no NPU path in LiteRT 2.2 (asking for one silently falls back to CPU), the iOS runtime is not published as a package and has to be assembled from the repository's prebuilts, and LiteRT's two Android libraries share a namespace that AGP 9 rejects.

A dictation session on three phones

The int8 English model ships on the CPU through XNNPACK, with a weight cache and a per-device thread limit. Every quantised variant decodes 20 s of speech in under 0.4 s on either phone's CPU, so the GPU isn't needed. To test it, 54 real recordings, 39 minutes of speech, were played through the app's own VAD and recognizer at speaking pace as one continuous 44-minute session.

Pixel 1 (2016)Pixel 7iPhone SE 3
Wait for the last text after release: median / worst1.77 / 4.11 s0.85 / 1.82 s0.29 / 0.57 s
Real-time factor (decode time ÷ audio time)0.22–0.240.10–0.140.036–0.039
CPU cores used, on average0.560.240.07
Peak memory528 MB605 MB325 MB
Model load, cache warm545 ms751 ms138 ms

Every phone kept up with speech for the whole session, and memory stayed flat. The Pixel 7 reached light thermal throttling at minute 8. A second run, unplugged, showed the cause was fast charging, not dictation.

The model that came out

Rimbalo's File Note pack turns a lawyer's dictation into a structured file note. The first design had an on-device LLM write its summaries, so three candidates ran on LiteRT-LM with schema-constrained JSON decoding: Qwen3-1.7B, Qwen3-4B-Instruct and Phi-4-mini. They were scored on synthetic dictations, including sensitive ones.

Nothing on the phone could flag those errors before a lawyer signed the note, so the generative model was removed. Every section is now built by rules from what the lawyer said, measured on held-out corpora. The model download and the 8 GB memory floor for packs went away, and the app can't put words in the user's mouth.

Rimbalo's product page · Back to Diamond Edge