The first demo shows that a phone can answer offline. The design question is when it should refuse.
Set the budget first
The budget includes model size, time to first token, heat, and memory. Without a budget, on-device inference only moves a server problem into a pocket.
Retrieval still matters
A smaller model needs accurate context. Local notes, recent conversation, and a clear instruction usually help more than a few extra billion parameters.
suspend fun answerOnDevice(question: String): String {
val context = retriever.search(question, limit = 4)
return model.generate(
system = "Answer only from the supplied notes.",
user = context.joinToString("n") + "nn" + question
)
}
Make failure visible
A timeout, a fall back to the cloud, or a plain refusal should be part of the product, not a line in a log.
Discussion