Documentation
TokForge is free. Android is live on Google Play; iPhone and iPad are in public beta on TestFlight. No account, and it works with the network off once a model is downloaded.
Reading on a computer? Check which models fit your phone, or see real measured speeds on the phone-speed leaderboard.
TokForge runs open large language models directly on your phone. Chat, roleplay, and group chat; on-device image generation; text-to-speech and voice input; document memory and retrieval; and an optional local HTTP control API for automation. Inference runs on the device. After you download a model, the core chat experience works with no account and no server round-trip.
The app ships on two platforms: Android (live on Google Play) and iPhone and iPad (public beta on TestFlight). Most features exist on both; the differences are called out below.
Get the app
Android. Install TokForge from Google Play like any other app. Updates arrive through Play.
iPhone and iPad. Install Apple TestFlight, then open the TokForge invite link and tap Install. iOS is in public beta on TestFlight; there is no App Store listing yet.
On first launch the app profiles your device and recommends a model to match its memory and chipset. Download the model, open a chat, and start. See the Quickstart for a screen-by-screen walkthrough.
What you need
- Android: Android 8.0 or newer (API 26), 64-bit ARM. 4 GB of RAM minimum, 8 GB recommended. Room for at least one model (a few hundred MB to several GB).
- iPhone and iPad: iOS or iPadOS 16 or newer. The app shows the models that fit your device's memory.
- The Android app download is a little over 100 MB. Models are downloaded separately from inside the app and can be deleted anytime.
What you can do
Chat, roleplay, and group chat Android and iOS
- Chat with streaming tokens and a live tok/s counter, markdown and code rendering, and collapsible reasoning (
<think>) blocks for models that support it. - Roleplay with characters, and group chat with two or more characters taking turns. Each member can have its own read-aloud voice, mute state, and talkativeness.
- Regenerate replies, browse alternate responses, and edit messages in place.
Characters, personas, and lorebooks Android and iOS
- Character hub with built-in personalities plus import of your own cards (Character Card V2 PNG or JSON, and common flat card formats).
- chub.ai import: search and pull character cards from inside the app.
- Personas describe who you are to the model, applied per conversation.
- Lorebooks / world info inject the right background details only when their keywords come up.
Memory, RAG, and knowledge Android and iOS
- Persistent memory: the app extracts facts in the background so a character remembers you across conversations.
- Document RAG: attach PDFs, DOCX, EPUB, or text and the model answers grounded in them, using hybrid keyword plus semantic retrieval.
- Knowledge graph: extracted facts and their relationships, browsable and editable.
- Opt-in web search: off by default. When you turn it on, the model can pull in fresh results for a turn.
Images, voice, and vision Android and iOS
- On-device image generation from a chat prompt. Android uses Stable Diffusion routes (with optional NPU acceleration on supported Snapdragon phones); iPhone uses Apple CoreML on the Neural Engine. Reference images, LoRAs, styles, and batches are supported where the device allows.
- Text-to-speech to hear replies aloud, with high-quality offline voices, plus voice cloning from a short sample.
- Voice input: speak instead of typing, transcribed on-device.
- Vision: send a picture and ask about it.
Engines and models Android and iOS
- Android engines: MNN, which runs on CPU, with OpenCL and Vulkan GPU paths depending on the phone; llama.cpp (GGUF) on CPU, with OpenCL and Vulkan GPU paths; and a remote-API option to point at your own endpoint. The app picks a sensible default automatically.
- iPhone engines: llama.cpp on Metal (the main engine), plus MLX models on iOS 17 or newer. You can also point at your own remote endpoint.
- Model downloader with Hugging Face search built in: find, download, and switch models from inside the app, with progress, resume, and disk-space checks.
AutoForge, leaderboard, backup Android and iOS
- AutoForge benchmarks a model on your device (prefill and decode tok/s, time-to-first-token, memory) and can auto-tune settings for your hardware.
- Leaderboard: opt in to submit a privacy-clean benchmark row and see how devices compare at leaderboard.tokforge.ai. Submissions carry no chat content.
- Backup and restore your characters, conversations, and settings to a file.
Which model fits your device
TokForge recommends a model based on your device on first launch, so you do not have to choose blind. The tables below are the plain-English version of that guidance. Bigger models are smarter but need more memory and run slower. The app never silently overloads a device: if a model would run out of memory it refuses and offers a Load anyway option, and a model that would run slowly on your phone is labelled Large & slow rather than hidden.
Android, by RAM
| Device RAM | Comfortable pick | Notes |
|---|---|---|
| 4 GB | Qwen3.5 0.8B (about 0.6 GB) | Small and fast. Good for quick chat on entry devices. |
| 6 GB | Qwen3.5 2B (about 1.4 GB) | Quick everyday chat. Qwen3 1.7B and Llama 3.2 3B are also offered. |
| 8 GB | Qwen3.5 4B (about 2.8 GB) or Gemma 4 E2B (about 3.6 GB) | The everyday sweet spot. |
| 12 GB | Qwen3.5 9B (about 6.8 GB) or Llama 3.1 8B (about 5.3 GB) | Noticeably stronger answers with headroom for longer chats. |
| 16 GB and up | Qwen3.5 9B (about 6.8 GB) | The flagship pick with more headroom. Phones with 24 GB can step up to Qwen3.5 27B (about 17.6 GB). |
iPhone and iPad, by RAM
| Device RAM | Comfortable pick | Notes |
|---|---|---|
| 4 GB | A 1B-class model such as Llama 3.2 1B | Tightest tier. Small models only. |
| 6 GB | A 3B such as Llama 3.2 3B or Qwen2.5 3B | Small models are quick. Image generation uses Apple CoreML. |
| 8 GB | Qwen3.5 4B, or a compact build of Qwen3.5 9B | The sweet spot: 4B models run well, and a select larger model fits. |
| 12 GB | Qwen3.5 9B, or larger models such as a 14B or a 30B mixture-of-experts | Flagship tier. The largest models page from storage, so they run slower. |
Rule of thumb: pick the model the app suggests first, then try the next size up if it stays responsive. A 4B or smaller model is the reliable everyday choice on almost any modern phone.
Dig deeper
Quickstart
Install TokForge, download your first model, and start chatting, with screenshots for each step.
API Reference
The on-device control API for both platforms: enable it, authenticate, and drive inference, models, image generation, memory, agents, and benchmarking over local HTTP.
AutoForge & Methodology
How TokForge measures on-device speed: optimization tiers, AutoForge config sweeps, and reproducible, cross-device benchmarking.
Speculative Decoding
An optional speed-up that uses a small draft model. Off by default, and turned on automatically only where it measured faster.
Memory System
Persistent per-character memory with background extraction, hybrid retrieval, knowledge graphs, and document import.
Guides 15
Step-by-step walkthroughs for offline AI setup, character cards, document RAG, AutoForge tuning, API automation, and more.
Privacy Policy
Inference runs on-device. No accounts and no analytics in the app. Read the full policy.
Terms of Service
Usage terms for the TokForge platform and beta program.