Can an app use AI without sending your data to the cloud?
Yes. A language model small enough to run on a phone can summarise, classify, extract structured data and rewrite text without a single network call, and both Apple and Google now ship such a model inside the operating system where any app can call it for free. The hard constraint is not capability — it is device coverage. On hardware too old for the built-in model, it simply is not there, so anything built this way needs a decided fallback rather than a hopeful one.
Can an app use AI without sending data to a server?
Yes. An app can run a language model locally, on the device’s own processor, so the text or image being analysed never leaves the hardware. This is called on-device inference, and it removes the network round trip completely: there is no API endpoint, no request body carrying user data, no provider-side log, and no third-party breach that can expose something a user typed into your app. For a feature touching health, financial or private-message data, that is a materially different privacy position from a cloud call rather than a difference in wording.
The distinction worth holding onto is that this is an architectural property, not a policy one. A cloud provider promising not to train on your data is making a promise; a model running on the device is not in a position to break one, because the data never reaches anybody who could.
What does on-device AI mean in practice?
There are two routes, and they have very different costs. The first is to call a model the operating system already provides, which costs nothing in app size and nothing per use, but only works on devices whose OS ships one. The second is to bundle your own model weights inside the app and run them with a local inference runtime, which works on far more devices but makes the app substantially larger.
Most products that ship this well use both: the OS model where it exists, your own small bundled model or a cloud call where it does not. Deciding which fallback applies to which device class is the actual engineering decision, and it is better made before the feature is designed than after.
Which on-device model APIs exist on iOS and Android today?
Both platforms now expose a built-in model to third-party apps, and a handful of cross-platform runtimes exist for shipping your own.
iOS — the Foundation Models framework
Apple’s Foundation Models framework, introduced in iOS 26, gives apps a Swift API onto the same roughly 3-billion-parameter on-device model that powers Apple Intelligence. There is no API key, no per-token cost and no network dependency, so a feature built on it keeps working on a plane. Apple also publishes an adapter toolkit for training LoRA adapters that specialise the built-in model for one app’s task, which is how you get domain-specific behaviour without shipping a model of your own.
Android — Gemini Nano through ML Kit
On Android the on-device model is Gemini Nano, run by AICore, an operating system service rather than something bundled in your app. The supported route to it is ML Kit’s GenAI APIs: task-shaped APIs for summarisation, proofreading, rewriting and image description, plus a Prompt API for custom prompts. Device support is explicitly a floor rather than universal — the published baseline at the time of writing is Pixel 8 and later and Galaxy S24 and later, widening over time.
Either platform — shipping your own weights
If the built-in model is not available or not suitable, local inference runtimes such as llama.cpp, ONNX Runtime, LiteRT and MLC will run a model you choose on either platform, and Core ML will do it on Apple hardware. This buys you full control of the model and removes the device floor, at the price of app size and of maintaining an inference stack yourself.
What can a small on-device model actually do well?
A 3-billion-parameter model is a genuinely useful tool for bounded, text-shaped work, and a poor substitute for a frontier model on anything open-ended. The tasks it handles reliably:
- Summarising something the user already has — a note, a message thread, an article.
- Classifying and tagging: routing a captured note to the right list, labelling an expense, detecting intent.
- Extracting structure from loose text, which is the quiet workhorse — turning “dentist tuesday half four” into a dated, titled task.
- Rewriting and proofreading, where tone and grammar matter and facts are supplied by the user.
- Short, grounded question answering over content already on the device.
What it does badly is equally specific, and worth stating plainly: multi-step reasoning, long documents that exceed a small context window, anything depending on broad world knowledge, and anything where a confidently wrong answer is expensive. A small local model does not know much; it is good at operating on text you give it. Features designed around that constraint work well, and features that ignore it produce the demo that impresses once and annoys thereafter.
What does on-device AI cost?
There is no per-token bill, which is the headline and genuinely changes the economics of a feature used constantly by every user. The costs move elsewhere instead, and the one that surprises people is app size.
The arithmetic is easy to do and worth doing before committing. A 3-billion-parameter model quantised to 4 bits is about 1.5 GB of weights — three billion values at half a byte each. That is far too large to bundle into a typical mobile app, which is the single best reason to prefer the operating system’s model where it exists: calling it costs zero bytes of download, because the model is already on the device as part of the OS.
The remaining costs are engineering ones. Two code paths instead of one, a fallback that has to be tested on hardware that lacks the built-in model, and output evaluation that is harder than with a cloud model because you cannot change the model under a deployed app — on the OS route, the model updates when the OS does, on someone else’s schedule.
When is on-device AI the wrong choice?
Often, and it is worth being direct about it. Send the work to a cloud model when the task needs real reasoning or breadth of knowledge, when inputs are long, when you need every user on identical model behaviour regardless of their hardware, when you must change model behaviour quickly after launch without shipping an app update, or when the data is not sensitive and the privacy argument is therefore not buying anything.
The expensive mistake is choosing on-device because it sounds better rather than because the data demands it. It constrains what the feature can do, doubles the paths you maintain, and ties you to a model you cannot upgrade on your own timetable. Where the data genuinely cannot leave the device, every one of those costs is worth paying. Where it can, you have taken on all of them for a line in a press release.
How does Percuno approach on-device AI?
On-device AI is the specialism Percuno is built around, and the claim rests on employment history rather than on Percuno’s own client work, which does not exist yet. Kip Sarapas — Percuno’s founder and only engineer — works on on-device and local LLM integration in mobile products in his current role as a mobile engineering lead, which is where the experience behind this article comes from. Percuno is a new company with no clients and no case studies, and this article is not one.
It is also why Superset, Percuno’s own AI workout and food-prep planning app, is being built by someone who already knows how to keep that data off a server. An app that knows what you eat and how you train is a clear case where the privacy argument is real rather than decorative. Superset is in development and not yet released.
For client work, this is the piece most teams cannot staff: a product manager can ask for “AI, but this data cannot leave the device”, and comparatively few people can deliver that sentence. If that is the problem you have, the services page describes how projects work, and hello@percuno.com reaches the engineer directly.
Frequently asked questions
Can AI features work without an internet connection?
- Yes, if the model runs on the device. An on-device model needs no network connection, so the feature keeps working offline. Features that call a cloud model do not.
Is on-device AI more private than a cloud API?
- Yes, and the difference is architectural rather than contractual. With on-device inference the data never leaves the hardware, so there is no provider log, no transmission and no third party who could be breached. A cloud provider can only promise not to misuse data it has already received.
How large is an on-device language model?
- A 3-billion-parameter model quantised to 4 bits is roughly 1.5 GB of weights, which is usually too large to bundle in a mobile app. Calling the model built into iOS or Android instead costs nothing in app size, because it is already part of the operating system.
Which devices support on-device AI?
- Recent ones only. On Android, the published baseline for Gemini Nano at the time of writing is Pixel 8 and later and Galaxy S24 and later. Apple exposes its on-device model through the Foundation Models framework from iOS 26. Older hardware needs a fallback, either a bundled model or a cloud call.
Is an on-device model as capable as a cloud model?
- No. A model small enough to run on a phone is good at bounded text work such as summarising, classifying, extracting structure and rewriting. It is not a substitute for a frontier cloud model on multi-step reasoning, long documents or broad world knowledge.
Can Percuno build an AI feature that keeps data on the device?
- Yes. On-device and local LLM integration is the specialism Percuno is built around, based on its founder’s work on local model integration in mobile products. Percuno is a new company with no clients yet, so there are no case studies to show.
Other articles
- Should you hire one senior developer or an agency?A decision framework rather than a pitch: when one developer is better, when an agency is better, and how to manage the bus-factor risk if you pick one person.
- What do you do when a software project stalls?The rebuild-or-continue decision, what a codebase assessment actually checks, and the list of accounts and keys to collect from the previous developer before the relationship ends.
- Which task manager works on both iPhone and Windows?A straight answer on which apps actually cover mismatched devices, why Things does not, and what to check before you trust an app with your lists.