Run Local AI Models
Offline On Your Phone.

TokForge runs GGUF and MNN language models directly on your handset. No cloud, no account, no subscription, and it keeps working in airplane mode. Chat, roleplay with characters, ask questions about your own documents, send it a picture, generate images, and hear replies out loud, all on-device. See what real phones actually do on the public speed leaderboard.

Works in airplane mode Zero telemetry Free, no subscription Android & iPhone
TokForge running on an Android phone: character library with Rex and Luna, plus on-device image generation and agents
The app's home screen on Android. Everything on it runs on the handset.

See It Running Offline

Airplane mode on. No Wi-Fi. No data. Still generating.

Tap to play demo
✈ Recorded on-device in airplane mode
No Wi-Fi, no mobile data, no server round-trip

AI Personas & Characters NEW

Build reusable personas, bind one per conversation, and import character cards or whole chats from SillyTavern, Layla, and PocketPal. Each character feels different because each one is.

Auto-Optimized for Your Device

Three inference backends and five GPU paths. TokForge detects your hardware and picks the fastest config automatically. No tuning required.

Small Models Feel Instant HOT

Sub-1B models answer at conversational speed on modern handsets. See what current phones actually reach on the live leaderboard, and TQ4 TurboQuant keeps their context memory lean.

Chat With Your Documents

Attach PDFs, DOCX, or EPUB files. TokForge summarizes, indexes, and searches them so your AI can answer grounded in your documents, all on-device.

Hear Your AI Speak

Offline text-to-speech with 11 natural voices and adjustable speed. Powered by Kokoro TTS. No internet, no latency, no data sent anywhere.

Generate Images Offline NEW

Stable Diffusion runs on your phone. No cloud, no API fees. SD1.5-Turbo, LCM, SD3.x, and an optional NPU Fast tier. Lock a face or subject across images with reference-image identity (IP-Adapter), plus LoRA and multi-subject support.

Send Your AI a Picture NEW

Attach a photo and ask about it. Qwen3-VL and Gemma-4 Omni run natively; SmolVLM ships as a sidecar for models without built-in vision. Text, vision, audio, and video, all on-device.

On-Device AI Agents NEW

Build no-code agents that call tools and run multi-step tasks, all on your phone. Ships with built-in agents and a simple builder, with triggers and bounded loops.

Roleplay & Group Chat NEW

A dedicated roleplay mode with lorebooks, multi-character group chat, and rotation modes. Bring your whole cast into one scene.

Real Phone Speeds, Measured

Every speed figure we publish comes from a benchmark run on a real handset and submitted from inside the app. We do not publish estimates.

Live Phone-Speed Leaderboard prefill, decode, time-to-first-token, peak memory and app version, per device and per model, updated continuously.
View live leaderboard →

Phone inference speed moves with the app version, the llama.cpp and MNN revision, the quantization, the thermal state of the handset and the SoC itself, so a number baked into a web page goes stale within weeks. That is why the leaderboard is the source of record here rather than a table on this page. Each row carries the device, the SoC, the model, the backend, the app version and the date it was submitted, so you can find a phone close to yours instead of reading an average that describes nobody.

What actually sets the speed

RAM decides which models will load at all. The SoC and its memory bandwidth decide how fast tokens come out. Model size and quantization trade quality against both. A small model on a mid-range phone will comfortably out-run a large model on a flagship, which is why the leaderboard is grouped by size class rather than presented as one ranking.

Three engines, chosen per device

TokForge ships llama.cpp for GGUF, MNN, and an optional connection to your own OpenAI-compatible server. It routes per device and per model rather than assuming one is best. GPU is not automatically faster than CPU on a phone: on several profiles in our own fleet the CPU path wins, and the app is built to pick the faster one instead of defending a preference.

Vulkan is a lab lane, not a default

Vulkan is built, tested and visible in the app, but it stays lab-only unless a route proves correctness and beats CPU or OpenCL on a specific device and model. MNN Vulkan on Mali-G925 has unresolved numeric corruption, so release builds route it to CPU automatically, and MNN Vulkan on Adreno is unstable. The only Vulkan path that ships on by default is GGUF Vulkan with cooperative matrix on D9400 and D9300 class Mali.

Sustained speed is lower than the first turn

A cold phone is fast and a warm phone throttles. Long sessions settle to a steady state below the opening burst, and that steady state is set by the power and thermal envelope of the handset rather than by the model. Benchmarks that quote only a cold first run overstate what a real conversation feels like, so leaderboard rows record the thermal state they were taken in.

Draft-verified decoding

On supported Android devices TokForge can pair a small draft model with your main model and verify its guesses in a single pass. Accepted guesses come out faster, rejected ones are discarded, and the text is the same either way. The gain depends heavily on the device, the model and the kind of text being written, and the app turns it off where it does not help. See how draft-verified decoding works.

Reproduce it yourself

ForgeLab is built into the app. It runs the same warmup and repeat structure we use internally, on your handset, and can submit the result to the public board. If our numbers and your numbers disagree, your phone is the one that counts. Read the benchmark methodology.

Speed data source: the live leaderboard, submitted from the app on real devices. Methodology →

Everything you need for local AI chat.

Built for privacy-conscious users, roleplay enthusiasts, and developers who want full control.

TurboQuant: Leaner Context Memory HOT

TQ4 stores the attention (KV) cache in a compact quantized form, freeing RAM for longer chats on small models. Useful for long conversations and quick back-and-forth on RAM-tight devices.

Generate Images Offline NEW

Stable Diffusion runs on your phone. No cloud, no API fees. SD1.5-Turbo on Adreno, SD1.5-LCM on MNN, SD3.x, and an optional NPU Fast tier on rooted devices. Lock a face or subject across images with reference-image identity (IP-Adapter), plus GPU-LoRA, multi-subject masks, and batch generation. Create while in airplane mode.

Send Your AI a Picture NEW

Attach a photo and ask about it. Qwen3-VL and Gemma-4 Omni run natively; SmolVLM ships as a 500 MB sidecar for models without built-in vision. Text, vision, audio, and video, all on-device.

Personas, Roleplay & Group Chat NEW

Create reusable personas and bind one per conversation, then jump into a dedicated roleplay mode with lorebooks, multi-character group chat, and rotation. Import characters and whole chat histories from SillyTavern, Layla, and PocketPal.

On-Device AI Agents NEW

Build no-code agents that call tools and run multi-step jobs entirely on-device: bounded loops, triggers, and a batch image-gen agent. Ships with built-in agents and a simple builder.

Model Hub & Leaderboard NEW

Browse and download models from Hugging Face inside the app, then see how your phone stacks up on a public phone-speed leaderboard. Hugging Face revision pinning supported.

Your Phone, Optimized Automatically

Three inference engines (MNN, GGUF, Remote API) and multiple GPU paths. TokForge profiles your hardware on first launch and picks the fastest config: Snapdragon, Dimensity, Exynos, or Tensor. You can also connect to a remote server for bigger models.

Chat With Your Documents

Attach PDFs, Word docs, EPUBs, or plain text. TokForge indexes and summarizes them, then your AI answers questions grounded in the actual content, all processed on-device, nothing uploaded anywhere.

Hear Your AI Talk Back

11 natural voices with adjustable speed via Kokoro TTS, plus the new ZipVoice decoder. Streams per-sentence as the model generates. First audio in seconds, not minutes. Voice input too.

Read the docs

Engineered for speed and control.

  • Personas + character cards: Build a persona library, bind one per chat, and import cards or whole histories from SillyTavern, Layla & PocketPal
  • Three inference engines: MNN, GGUF, and Remote API. GPU-accelerated and CPU-optimized paths auto-selected per chipset
  • Hardware profiler: Detects your chipset, GPU, and RAM to recommend the best config
  • 150+ API endpoints: Full remote control from any device on your network: run benchmarks, manage models, change settings, generate images (batch + reference), run agents, upload documents
  • Benchmark database: Save results, compare across devices, export and share your configs
1. Pick a persona, character, or agent, or start a blank chat (text, image, or vision)
2. TokForge detects your hardware & picks the fastest config
3. GPU-accelerated inference → real-time token streaming (or image generation)
4. Rich rendering with reasoning blocks, citations & markdown
5. Memory learns from the conversation in the background

Chat + inference pipeline

Local-first by design.
Transparent by default.

  • No cloud required
  • No analytics / no telemetry
  • Remote API is opt-in and user-configured
  • Web search is opt-in and off by default. Chats stay local unless you enable it
  • Models download directly from Hugging Face
  • License + model card shown per download (including LoRA and image diffusion bundles)
  • Generated images stay on your device. No upload, no watermark service
  • iPhone & iPad (iOS 16+) and Android 8.0+ (ARM64)

Get TokForge for iPhone & Android

TokForge is free on Google Play (Android) open testing and TestFlight (iPhone & iPad) beta. Install directly or join the community to help shape the future of private mobile AI.

v3.6.0 beta: AI personas with roleplay & group chat, on-device no-code agents, an in-app model hub with public leaderboard, image generation with reference-image identity (SD1.5 / LCM / SD3.x / optional NPU), on-device vision (Qwen3-VL + Gemma-4 Omni), TurboQuant KV-cache compression, document search with citations, streaming TTS, persistent memory, optional opt-in web search, 150+ API endpoints, and more, now on iPhone & iPad too. Free on Google Play (Android) and TestFlight (iPhone & iPad).

Join the Beta

No spam. We'll only email you about beta access.

No telemetry. No background reporting. Your data stays on your device unless you explicitly opt in.