Open source · Evaluation
Where does the small model stop being good enough?
Small on-device models are worse than frontier models, but few teams can say where the line falls for the feature they're building. Quality and on-device cost are each measured well, and separately. These two projects measure them together.
AI Harness
A Kotlin Multiplatform app for Android, iOS and desktop with three LLM backends behind one streaming interface: Claude, Gemini, and local GGUF models running through llama.cpp via the llamatik bridge. Models download from Hugging Face onto the phone.
On top of that sits a headless evaluation runner. It runs a graded suite case by case against one model and writes a run record for each case: the score, the output, time to first token, wall time, and the device's state: battery at start and end, thermal status and peak memory, read on Android.
- One model layer, two consumers. The chat screen and the runner share the same backends, so what is measured is what users would run.
- Constrained decoding on the phone. Local models are held to a JSON schema as they generate, so they return well-formed structured output.
- Driven from a laptop. Suites go onto the phone with
adb push, run through an instrumentation test, and the run records come back withadb pull. A low pass rate is a finding, not a test failure: the test fails only when the rig itself couldn't run. - Deterministic scoring. No LLM judge sits between the measurement and the answer.
The task
A measurement needs a real task. The suite being written is natural-language control of a BLE-connected appliance: "Set it to 225 and tell me when it's there." It is narrow enough that a 3B model has a real chance, its outputs are structured so they can be graded exactly, and on-device is justified rather than hypothetical: a grill in a backyard with no signal, and a latency budget a cloud round trip can't meet.
The first comparison planned: Claude and Gemini as the frontier baseline against Gemma 3n, Qwen 3B and Llama 3.2 3B on the phone, each at Q4 and Q8.
llm-eval-ci
Prompt changes are code changes, and code changes need regression tests. llm-eval-ci makes "did that prompt edit make things worse?" a number, and makes a regression fail the build.
name: capitals
model: claude-opus-5
system: Answer with the city name only.
scorers:
- name: exact_match
cases:
- id: france
prompt: What is the capital of France?
expected: Paris
tags: [europe]
- Graded YAML suites, run concurrently by an async runner with a bounded semaphore.
- Deterministic scorers (exact match, contains, not contains, regex), extended with one decorated function.
- Pass rate, cost and latency per configuration, broken down by tag; refusals and errors reported as their own statuses.
- GitHub Actions runs the tests on Python 3.11 to 3.13 and an offline smoke run. A manual live evaluation is
gated by
--min-pass-rate. - AI Harness writes the same run-record format from the phone, so on-device and cloud runs compare directly.
The same discipline decided a product: in Rimbalo, measurement showed an on-device LLM putting false statements in 22% of summaries, and the generative model was removed.
Diamond Edge