Insights / AI
Where should an AI model run: cloud, edge, or device?
Choose an inference location from latency, privacy, connectivity, update, and cost constraints—not from a hardware trend.
Kiran Bandarupalli · 2 Oct 2026 · 2 min read

Inference location is an architectural choice with consequences for product behavior and operations. The smallest model that meets a measured need is often a better starting point than choosing a location because a device has a new accelerator or a cloud provider offers a convenient endpoint.
Cloud inference
A hosted model is a good fit when connectivity is dependable, centralized updates matter, and the product can tolerate network latency. The team owns request security, provider availability, data handling, rate limits, and per-request cost. Design a timeout and a useful non-AI fallback; otherwise a provider outage becomes an application outage.
Edge inference
A gateway or local server can serve several devices while keeping some data near its source. This can reduce round trips and support a site that loses its upstream link. It introduces fleet management: the team must patch the host, distribute compatible models, monitor resource pressure, and recover machines that are physically difficult to reach.
On-device inference
Running on the device can reduce network dependence and limit what leaves the hardware. The constraints are memory, energy, thermals, model size, and the diversity of device versions. Updates need signing, rollback, and compatibility checks. A model that works on a development board may not fit the final power budget.
Compare the whole lifecycle
- Measure end-to-end latency at realistic input sizes and network conditions.
- Decide which data may leave the device and how long it is retained.
- Estimate compute, connectivity, distribution, and support costs together.
- Test updates, partial failure, stale model versions, and offline operation.
Prototype the most constrained option first, but measure a baseline for the others. The right answer may be hybrid: local filtering and urgent decisions, with asynchronous cloud analysis for tasks that tolerate delay. Document the trade-off and revisit it when workload evidence changes.