Diamond Edge

Open source · Evaluation

Where does the small model stop being good enough?

Small on-device models are worse than frontier models, but few teams can say where the line falls for the feature they're building. Quality and on-device cost are each measured well, and separately. These two projects measure them together.

AI Harness

A Kotlin Multiplatform app for Android, iOS and desktop with three LLM backends behind one streaming interface: Claude, Gemini, and local GGUF models running through llama.cpp via the llamatik bridge. Models download from Hugging Face onto the phone.

On top of that sits a headless evaluation runner. It runs a graded suite case by case against one model and writes a run record for each case: the score, the output, time to first token, wall time, and the device's state: battery at start and end, thermal status and peak memory, read on Android.

The task

A measurement needs a real task. The suite being written is natural-language control of a BLE-connected appliance: "Set it to 225 and tell me when it's there." It is narrow enough that a 3B model has a real chance, its outputs are structured so they can be graded exactly, and on-device is justified rather than hypothetical: a grill in a backyard with no signal, and a latency budget a cloud round trip can't meet.

The first comparison planned: Claude and Gemini as the frontier baseline against Gemma 3n, Qwen 3B and Llama 3.2 3B on the phone, each at Q4 and Q8.

Status: the backends, the runner, telemetry and the run records are built and tested. The graded suite is being written, so there are no published numbers yet.

AI Harness on GitHub

llm-eval-ci

Prompt changes are code changes, and code changes need regression tests. llm-eval-ci makes "did that prompt edit make things worse?" a number, and makes a regression fail the build.

name: capitals
model: claude-opus-5
system: Answer with the city name only.
scorers:
  - name: exact_match
cases:
  - id: france
    prompt: What is the capital of France?
    expected: Paris
    tags: [europe]

llm-eval-ci on GitHub

The same discipline decided a product: in Rimbalo, measurement showed an on-device LLM putting false statements in 22% of summaries, and the generative model was removed.

Back to Diamond Edge