Building Sidd AI, a fully offline personal AI assistant for Android running Gemma 4 INT2 locally with Whisper STT and TTS.
Fully offline AI assistant on Android — Gemma 4 INT2, Whisper STT, local TTS

Demo Videos
Sidd AI — voice conversation demo
Sidd AI — productivity automation demo
How It Works
Sidd AI runs entirely on-device with no internet required. Speech is captured and transcribed locally using Google Whisper (STT). The transcribed text is sent to a Gemma 4 INT2 quantized model that runs inference directly on the phone's GPU/NPU via GGML. The model's response is then converted back to speech using an on-device TTS engine, creating a fully offline voice conversation loop. All processing stays on the device — no data ever leaves the phone.
Technology Used
Gemma 4 INT2 quantized model for on-device LLM inference, Google Whisper for offline speech-to-text (STT), on-device TTS engine for text-to-speech, Kotlin + Jetpack Compose for the Android UI, GGML/llama.cpp for quantized model execution, Android NNAPI for hardware-accelerated inference, and Room database for local conversation history.
Why It's Better
Unlike cloud-based AI assistants (Google Assistant, Siri, ChatGPT), Sidd AI requires zero internet connectivity and processes everything locally. This means complete privacy — no voice data or conversations are sent to any server. The INT2 quantized Gemma 4 model delivers surprisingly capable reasoning while fitting within mobile memory constraints. Running Whisper STT and TTS on-device eliminates the latency of network round-trips, making conversations feel instant and natural.