🚀 Ollama Local AI is an independent, privacy-first Ollama Android app for running, connecting, and routing Large Language Models (LLMs) from your phone.
⚠️ Important notice: Ollama Local AI is not affiliated with, sponsored by, or endorsed by Ollama. Ollama is a trademark of its respective owner.
Turn your Android device into a secure local-network OpenAI-compatible LLM proxy for IDEs, coding assistants, scripts, and custom AI tools.
📱 Connect Cursor, VS Code, Antigravity, Windsurf, or any OpenAI-compatible client to your phone’s LAN IP. Manage local models, llama.cpp, cloud providers, and custom endpoints from one Android app instead of configuring multiple endpoints on your development machine.
🔒 Local-first privacy
Run local AI models on-device with no internet required. Local model data, provider settings, and API keys are stored on your Android device. When you choose a cloud provider, requests go directly to the provider you configure—without a third-party relay from this app.
⚡ Key features
🧠 Local LLM inference: Run local models on your phone with llama.cpp engine support.
💬 Chat mode: Chat with local models or configured cloud providers such as OpenAI and Claude, then switch providers when needed.
🎨 Canvas mode: Build and test websites inside chat while the AI constructs them live.
🌐 WebX Preview: Preview websites created by Ollama or local AI from any browser on your network.
🖥️ Cross-device WebUI: Open the chat and management console from your phone, tablet, or PC.
🖥️ Zero-config local hosting: Use the embedded NanoHTTPD server directly on Android.
🔌 OpenAI-compatible API: Use standard endpoints such as /v1/chat/completions, /v1/models, and /health with compatible IDEs, scripts, and tools.
🔄 Provider switching: Save and switch between Ollama, llama.cpp, NVIDIA, OpenAI, Claude, Hugging Face, and custom API endpoints.
🚦 Smart proxy routing: Use provider timeouts, rate limits, session-stable routing, local pools, failover, and round-robin request distribution.
📡 LAN master/worker pairing: Link Android phones and devices into a private local AI network.
🌐 Multi-device clusters: Connect devices into a local LLM inference cluster.
🤖 Multi-agent orchestration: Coordinate agents, providers, tools, and tasks.
🎞️ Live agent stages: Follow agent progress and execution stages in real time.
📊 Traffic Observatory: Monitor request flow, routing, performance statistics, and proxy diagnostics.
🧪 Tool Lab: Test AI tools, provider connections, API requests, and workflows.
🖥️ Web Console: Manage routing, chat, providers, tools, and diagnostics in your browser.
🎨 Material 3 interface: Use a clean, responsive Android UI designed for local AI development.
🔋 Efficient routing: Lazy token-bucket rate limiting helps control traffic and avoid unnecessary background work.
🔐 Encrypted storage: Protect provider credentials with AES-256 encrypted EncryptedSharedPreferences.
🧵 Responsive performance: Coroutines and offloaded I/O keep the Compose interface fluid during routing and network work.
🛠️ Built for developers
Use Ollama Local AI with Cursor, VS Code, Antigravity, Windsurf, OpenAI-compatible clients, automation scripts, and custom AI workflows. Chat, code, route requests, preview web apps, connect devices, and observe AI traffic from one Android application.
🛣️ More in beta
Additional local AI integrations, provider support, developer tools, and performance improvements are in development.
🔥 Build a flexible local LLM workflow with Ollama Local AI—your Android hub for local models, cloud routing, IDE connectivity, and private network AI development.
Ollama Local AI is an independent third-party application and is not affiliated with or endorsed by Ollama.