TokForge TokForge
Home / Docs

Documentation

Android 1.3.6.1 on Google Play. iPhone and iPad: public beta on TestFlight. On-device chat, roleplay, group chat, image generation, voice, memory, and a local control API. This hub covers what the app does, how to get it, which model fits your phone, and where to go deeper.

Run this on your own phone

TokForge is free. Android is live on Google Play; iPhone and iPad are in public beta on TestFlight. No account, and it works with the network off once a model is downloaded.

Get it on Google Play iPhone beta on TestFlight Discord community

Reading on a computer? Check which models fit your phone, or see real measured speeds on the phone-speed leaderboard.

TokForge runs open large language models directly on your phone. Chat, roleplay, and group chat; on-device image generation; text-to-speech and voice input; document memory and retrieval; and an optional local HTTP control API for automation. Inference runs on the device. After you download a model, the core chat experience works with no account and no server round-trip.

The app ships on two platforms: Android (live on Google Play) and iPhone and iPad (public beta on TestFlight). Most features exist on both; the differences are called out below.

Get the app

Android. Install TokForge from Google Play like any other app. Updates arrive through Play.

iPhone and iPad. Install Apple TestFlight, then open the TokForge invite link and tap Install. iOS is in public beta on TestFlight; there is no App Store listing yet.

On first launch the app profiles your device and recommends a model to match its memory and chipset. Download the model, open a chat, and start. See the Quickstart for a screen-by-screen walkthrough.

What you need

What you can do

Chat, roleplay, and group chat Android and iOS

Characters, personas, and lorebooks Android and iOS

Memory, RAG, and knowledge Android and iOS

Images, voice, and vision Android and iOS

Engines and models Android and iOS

AutoForge, leaderboard, backup Android and iOS

Which model fits your device

TokForge recommends a model based on your device on first launch, so you do not have to choose blind. The tables below are the plain-English version of that guidance. Bigger models are smarter but need more memory and run slower. The app never silently overloads a device: if a model would run out of memory it refuses and offers a Load anyway option, and a model that would run slowly on your phone is labelled Large & slow rather than hidden.

Android, by RAM

Device RAMComfortable pickNotes
4 GBQwen3.5 0.8B (about 0.6 GB)Small and fast. Good for quick chat on entry devices.
6 GBQwen3.5 2B (about 1.4 GB)Quick everyday chat. Qwen3 1.7B and Llama 3.2 3B are also offered.
8 GBQwen3.5 4B (about 2.8 GB) or Gemma 4 E2B (about 3.6 GB)The everyday sweet spot.
12 GBQwen3.5 9B (about 6.8 GB) or Llama 3.1 8B (about 5.3 GB)Noticeably stronger answers with headroom for longer chats.
16 GB and upQwen3.5 9B (about 6.8 GB)The flagship pick with more headroom. Phones with 24 GB can step up to Qwen3.5 27B (about 17.6 GB).

iPhone and iPad, by RAM

Device RAMComfortable pickNotes
4 GBA 1B-class model such as Llama 3.2 1BTightest tier. Small models only.
6 GBA 3B such as Llama 3.2 3B or Qwen2.5 3BSmall models are quick. Image generation uses Apple CoreML.
8 GBQwen3.5 4B, or a compact build of Qwen3.5 9BThe sweet spot: 4B models run well, and a select larger model fits.
12 GBQwen3.5 9B, or larger models such as a 14B or a 30B mixture-of-expertsFlagship tier. The largest models page from storage, so they run slower.
Rule of thumb: pick the model the app suggests first, then try the next size up if it stays responsive. A 4B or smaller model is the reliable everyday choice on almost any modern phone.

Dig deeper