Uploaded February 2026 | Updated September 2026, 2 days ago
In this session, Artur Morys-Magiera explores what it actually takes to run LLM inference directly on mobile devices in React Native applications.
The talk covers:
Why teams move inference on-device: reliability, privacy, and latency
Model size constraints and OS-provisioned alternatives
GPU, NPU, and CPU acceleration tradeoffs across iOS and Android
Debugging performance issues in abstraction layers like OpenCL
Comparing TF Lite, ONNX, ExecuTorch, MLC, llama.cpp, and Apple-based solutions
Practical optimizations including quantization and compilation-time improvements
If you’re evaluating on-device AI in a React Native app, this is a grounded overview of what works today, and what still requires careful tradeoffs.
Follow Callstack on X 🐦 https://x.com/callstackio
In this session, Artur Morys-Magiera explores what it actually takes to run LLM inference directly on mobile devices in React Native applications.
The talk covers:
Why teams move inference on-device: reliability, privacy, and latency
Model size constraints and OS-provisioned alternatives
GPU, NPU, and CPU acceleration tradeoffs across iOS and Android
Debugging performance issues in abstraction layers like OpenCL
Comparing TF Lite, ONNX, ExecuTorch, MLC, llama.cpp, and Apple-based solutions
Practical optimizations including quantization and compilation-time improvements
If you’re evaluating on-device AI in a React Native app, this is a grounded overview of what works today, and what still requires careful tradeoffs.
Follow Callstack on X 🐦 https://x.com/callstackio










