Run Local AI Models
Offline On Your Phone.
TokForge runs GGUF and MNN language models directly on your handset. No cloud, no account, no subscription, and it keeps working in airplane mode. Chat, roleplay with characters, ask questions about your own documents, send it a picture, generate images, and hear replies out loud, all on-device. See what real phones actually do on the public speed leaderboard.
Scan this with your Android phone to open TokForge on Google Play. On iPhone, the beta is invite-free through TestFlight. Prefer to read first? Check whether your phone can run a given model, or start with running AI offline on Android.
See It Running Offline
Airplane mode on. No Wi-Fi. No data. Still generating.
AI Personas & Characters NEW
Build reusable personas, bind one per conversation, and import character cards or whole chats from SillyTavern, Layla, and PocketPal. Each character feels different because each one is.
Auto-Optimized for Your Device
Three inference backends and five GPU paths. TokForge detects your hardware and picks the fastest config automatically. No tuning required.
Small Models Feel Instant HOT
Sub-1B models answer at conversational speed on modern handsets. See what current phones actually reach on the live leaderboard, and TQ4 TurboQuant keeps their context memory lean.
Chat With Your Documents
Attach PDFs, DOCX, or EPUB files. TokForge summarizes, indexes, and searches them so your AI can answer grounded in your documents, all on-device.
Hear Your AI Speak
Offline text-to-speech with 11 natural voices and adjustable speed. Powered by Kokoro TTS. No internet, no latency, no data sent anywhere.
Generate Images Offline NEW
Stable Diffusion runs on your phone. No cloud, no API fees. SD1.5-Turbo, LCM, SD3.x, and an optional NPU Fast tier. Lock a face or subject across images with reference-image identity (IP-Adapter), plus LoRA and multi-subject support.
Send Your AI a Picture NEW
Attach a photo and ask about it. Qwen3-VL and Gemma-4 Omni run natively; SmolVLM ships as a sidecar for models without built-in vision. Text, vision, audio, and video, all on-device.
On-Device AI Agents NEW
Build no-code agents that call tools and run multi-step tasks, all on your phone. Ships with built-in agents and a simple builder, with triggers and bounded loops.
Roleplay & Group Chat NEW
A dedicated roleplay mode with lorebooks, multi-character group chat, and rotation modes. Bring your whole cast into one scene.
See TokForge in Action
The redesigned TokForge, running locally on iPhone & Android: characters, image creation, agents, and more, all on-device.











Real Phone Speeds, Measured
Every speed figure we publish comes from a benchmark run on a real handset and submitted from inside the app. We do not publish estimates.
Phone inference speed moves with the app version, the llama.cpp and MNN revision, the quantization, the thermal state of the handset and the SoC itself, so a number baked into a web page goes stale within weeks. That is why the leaderboard is the source of record here rather than a table on this page. Each row carries the device, the SoC, the model, the backend, the app version and the date it was submitted, so you can find a phone close to yours instead of reading an average that describes nobody.
What actually sets the speed
RAM decides which models will load at all. The SoC and its memory bandwidth decide how fast tokens come out. Model size and quantization trade quality against both. A small model on a mid-range phone will comfortably out-run a large model on a flagship, which is why the leaderboard is grouped by size class rather than presented as one ranking.
Three engines, chosen per device
TokForge ships llama.cpp for GGUF, MNN, and an optional connection to your own OpenAI-compatible server. It routes per device and per model rather than assuming one is best. GPU is not automatically faster than CPU on a phone: on several profiles in our own fleet the CPU path wins, and the app is built to pick the faster one instead of defending a preference.
Vulkan is a lab lane, not a default
Vulkan is built, tested and visible in the app, but it stays lab-only unless a route proves correctness and beats CPU or OpenCL on a specific device and model. MNN Vulkan on Mali-G925 has unresolved numeric corruption, so release builds route it to CPU automatically, and MNN Vulkan on Adreno is unstable. The only Vulkan path that ships on by default is GGUF Vulkan with cooperative matrix on D9400 and D9300 class Mali.
Sustained speed is lower than the first turn
A cold phone is fast and a warm phone throttles. Long sessions settle to a steady state below the opening burst, and that steady state is set by the power and thermal envelope of the handset rather than by the model. Benchmarks that quote only a cold first run overstate what a real conversation feels like, so leaderboard rows record the thermal state they were taken in.
Draft-verified decoding
On supported Android devices TokForge can pair a small draft model with your main model and verify its guesses in a single pass. Accepted guesses come out faster, rejected ones are discarded, and the text is the same either way. The gain depends heavily on the device, the model and the kind of text being written, and the app turns it off where it does not help. See how draft-verified decoding works.
Reproduce it yourself
ForgeLab is built into the app. It runs the same warmup and repeat structure we use internally, on your handset, and can submit the result to the public board. If our numbers and your numbers disagree, your phone is the one that counts. Read the benchmark methodology.
Everything you need for local AI chat.
Built for privacy-conscious users, roleplay enthusiasts, and developers who want full control.
TurboQuant: Leaner Context Memory HOT
TQ4 stores the attention (KV) cache in a compact quantized form, freeing RAM for longer chats on small models. Useful for long conversations and quick back-and-forth on RAM-tight devices.
Generate Images Offline NEW
Stable Diffusion runs on your phone. No cloud, no API fees. SD1.5-Turbo on Adreno, SD1.5-LCM on MNN, SD3.x, and an optional NPU Fast tier on rooted devices. Lock a face or subject across images with reference-image identity (IP-Adapter), plus GPU-LoRA, multi-subject masks, and batch generation. Create while in airplane mode.
Send Your AI a Picture NEW
Attach a photo and ask about it. Qwen3-VL and Gemma-4 Omni run natively; SmolVLM ships as a 500 MB sidecar for models without built-in vision. Text, vision, audio, and video, all on-device.
Personas, Roleplay & Group Chat NEW
Create reusable personas and bind one per conversation, then jump into a dedicated roleplay mode with lorebooks, multi-character group chat, and rotation. Import characters and whole chat histories from SillyTavern, Layla, and PocketPal.
On-Device AI Agents NEW
Build no-code agents that call tools and run multi-step jobs entirely on-device: bounded loops, triggers, and a batch image-gen agent. Ships with built-in agents and a simple builder.
Model Hub & Leaderboard NEW
Browse and download models from Hugging Face inside the app, then see how your phone stacks up on a public phone-speed leaderboard. Hugging Face revision pinning supported.
Your Phone, Optimized Automatically
Three inference engines (MNN, GGUF, Remote API) and multiple GPU paths. TokForge profiles your hardware on first launch and picks the fastest config: Snapdragon, Dimensity, Exynos, or Tensor. You can also connect to a remote server for bigger models.
Chat With Your Documents
Attach PDFs, Word docs, EPUBs, or plain text. TokForge indexes and summarizes them, then your AI answers questions grounded in the actual content, all processed on-device, nothing uploaded anywhere.
Hear Your AI Talk Back
11 natural voices with adjustable speed via Kokoro TTS, plus the new ZipVoice decoder. Streams per-sentence as the model generates. First audio in seconds, not minutes. Voice input too.
Engineered for speed and control.
-
Personas + character cards: Build a persona library, bind one per chat, and import cards or whole histories from SillyTavern, Layla & PocketPal
-
Three inference engines: MNN, GGUF, and Remote API. GPU-accelerated and CPU-optimized paths auto-selected per chipset
-
Hardware profiler: Detects your chipset, GPU, and RAM to recommend the best config
-
150+ API endpoints: Full remote control from any device on your network: run benchmarks, manage models, change settings, generate images (batch + reference), run agents, upload documents
-
Benchmark database: Save results, compare across devices, export and share your configs
Chat + inference pipeline
Local-first by design.
Transparent by default.
- No cloud required
- No analytics / no telemetry
- Remote API is opt-in and user-configured
- Web search is opt-in and off by default. Chats stay local unless you enable it
- Models download directly from Hugging Face
- License + model card shown per download (including LoRA and image diffusion bundles)
- Generated images stay on your device. No upload, no watermark service
- iPhone & iPad (iOS 16+) and Android 8.0+ (ARM64)
Get TokForge for iPhone & Android
TokForge is free on Google Play (Android) open testing and TestFlight (iPhone & iPad) beta. Install directly or join the community to help shape the future of private mobile AI.
Join the Beta
No telemetry. No background reporting. Your data stays on your device unless you explicitly opt in.
Work Out What Your Phone Can Do
Before you download anything, these answer the two questions people actually arrive with: will it run, and how fast.
Can my phone run it?
A RAM calculator built for phones rather than desktop graphics cards. It applies the same gate the app applies, and it shows the measurements the estimate rests on instead of a confident single number.
Model pages, with real numbers
Qwen3.5 0.8B, 2B, 4B and 9B, and Gemma 4 E2B. Download size, the RAM floor the app enforces, and a live table of what real handsets have recorded for each one.
MNN vs GGUF on Android
What the two formats are, which models exist in which, why the app picks one for your device, and why this page does not quote you a speed multiplier.