Offline AI Chat Private puts a local AI assistant on your Android phone. Download a model once, then chat, write, summarize, translate, study, brainstorm, and work with text without a constant internet connection. Prompts, answers, and attachments are processed on your device; chats are not sent to a cloud AI service for inference.
Choose a model for your phone and task. Free starter models are Tiny (SmolLM2) for lower-memory devices and basic text chat, Quick (Gemma 4 E2B) for fast everyday replies with compatible image and audio input, Smart (Gemma 4 E4B) for harder prompts and longer sessions, and Compact (Phi-4 Mini) as a text-first option for CPU-focused use. Download sizes are approximate; keep extra free storage and choose a lighter model if your phone is slow or low on memory.
Premium unlocks Gemma 4 12B, Qwen2-VL-2B, DeepSeek R1 1.5B, Qwen2.5 1.5B, Qwen2.5 Coder 3B, and Qwen3 0.6B, and chat export as a Markdown file. All saved chats are free.
Attach TXT or Markdown documents for free, or use Premium to ask questions about text-based PDF documents. PDF text is extracted locally; scanned or password-protected PDFs are not supported.
On first run, the app can check your device and recommend a suitable model. You can skip the recommendation and choose any compatible model.
Tiny (SmolLM2) - 143 MB; Quick (Gemma 4 E2B) - 2.6 GB; Smart (Gemma 4 E4B) - 3.7 GB; Compact (Phi-4 Mini) - 3.8 GB; Premium: Gemma 4 12B - 6.6 GB; Qwen2-VL-2B - 1.8 GB; DeepSeek R1 1.5B - 1.8 GB; Qwen2.5 1.5B - 1.6 GB; Qwen2.5 Coder 3B - 3.4 GB; Qwen3 0.6B - 614 MB.
After a model is downloaded, AI inference works offline. Internet is used for app installation and updates, model downloads, Google Play purchases and restores, and sending limited usage analytics and crash diagnostics when connected. Auto uses CPU by default for stability. Compatible phones can use GPU acceleration when selected or when experimental GPU probing is enabled; image analysis requires a compatible vision model and a working GPU runtime.