Engineering

Running On-Device Models on Android

An on-device model is useful when it still works inside limits of battery, latency, and privacy.

Running On-Device Models on Android

The first demo shows that a phone can answer offline. The design question is when it should refuse.

Set the budget first

The budget includes model size, time to first token, heat, and memory. Without a budget, on-device inference only moves a server problem into a pocket.

Retrieval still matters

A smaller model needs accurate context. Local notes, recent conversation, and a clear instruction usually help more than a few extra billion parameters.

suspend fun answerOnDevice(question: String): String {
    val context = retriever.search(question, limit = 4)
    return model.generate(
        system = "Answer only from the supplied notes.",
        user = context.joinToString("n") + "nn" + question
    )
}

Make failure visible

A timeout, a fall back to the cloud, or a plain refusal should be part of the product, not a line in a log.

On this page

Discussion

Leave a comment

Comments are published after moderation. Name and email are required.