TokForge TokForge
Home / Docs / How-to 1.3.8

TokForge 1.3.8: How-To

Android · TokForge 1.3.8 (beta)

Beta: this applies to TokForge 1.3.8, available on the Google Play beta track and as the Discord APK. The Play production version is 1.3.6.1.

This guide explains, feature by feature, what you can do and what you will see in 1.3.8. Most of the speed work is automatic: where a section says "Nothing to do", the app handles it and the section tells you what you will notice. The terminal has its own guide (TokForge Terminal: User Guide).

Where Settings paths are given, the group names are the ones on the Settings screen: "Speed & engine", "Connections", "Privacy & about" and so on. Where a path starts with "Settings (Advanced mode)", first set "Settings Mode" at the top of Settings to "Advanced".

First run and the model it picks · The simple model list and the Advanced switch · Measured on your phone · Pause background checks · CPU prompt, GPU reply and the one-copy split · Speed features · GPU picks you make yourself · Agents tab · Backup and restore · Characters and personas · Documents · Remote API · Upgrading from 1.3.6.1 or older · Other changes you can see

First run and the model it picks

New installs start with one question instead of a long model list.

  1. Open TokForge and tap "Get Started".
  2. "What do you want to do?" Pick one: "Chat and questions", "Writing and stories", "Code and math" or "Make pictures".
  3. TokForge shows "Recommended for your phone" with one model, its download size, a speed estimate and a network note.
  4. Tap "Get recommended model (... GB)" and confirm the download.
  5. When the model is ready, start your first chat.

What it picks

What the card tells you

Other buttons on the card

Speed boost tick (Snapdragon 8 Gen 3 and 8 Elite Gen 5 phones only): the card may show a ticked line "Includes speed boost for your phone (+... GB)". Untick it if you do not want the extra download. The model downloads first, so you can start chatting while the boost finishes. See "Speed features" below.

Good to know

The simple model list and the Advanced switch

The model list now sorts models by how well they suit your phone:

The "Fits" badge now asks the same memory question as loading, so a model that is too big says "Too big" up front instead of "Fits" followed by a refused load.

The Advanced switch

Also new in the catalog

Searching Hugging Face (Browse) uses its own badges ("Fits your phone", "Tight fit, may be slow") that only look at memory. Check the file size against your free storage before you download from there.

Measured on your phone

TokForge now checks each speed trick on your own phone before using it: the answers must match the normal route, and the trick must really be faster on your phone. This screen shows what it found and lets you overrule it.

Where: Settings (Advanced mode) > "Speed & engine" > "Performance" > "Measured on your phone".

The master switch: "Use speed features this phone has proven". Turn it off and every automatic speed feature is off. Turn it back on to let the rows below decide.

Each row is one feature, with a status line and three buttons, "Off", "Auto" and "On". Rows can include "GPU route (GGUF)", "GPU route (MNN)", "Qwen3.5: GPU reads the prompt", "One model copy for CPU and GPU", "Prompt lookup", "Multi-token prediction (MNN)", "Multi-token prediction (GGUF)", "Draft model speculation", "NPU reads long prompts", "Gemma 4: GPU reads the prompt", "MoE models: CPU and GPU share the work" and "Image generation on the GPU".

What the status lines mean

What the buttons mean

"Changes apply the next time a model loads."

Also in "Performance": "Measure CPU and GPU for each model" (on by default): "Runs a short speed test the first time a model loads and uses the faster route. Results stay on this phone."

Results are tied to the app version, the system update, the GPU driver and the model file. When any of those change, your phone checks again by itself.

Reading the route lines

You can see what your phone decided in several places.

While a check runs, a banner at the top of the app says:

In My Models, on the loaded model's card

Notices after a check

In the chat, tap "Tune", then look under "Hardware":

In AutoForge (Settings (Advanced mode) > "Speed & engine" > "Performance" > the "AutoForge" card, then the "Tune" tab), MNN models have a "Backend status" row. Its labels mean:

GGUF models have no "Backend status" row; AutoForge shows a plain line such as "Running on your processor." or "Running on your phone's graphics chip."

Known limit: the AutoForge row can still read "Certified CPU route ... the certified default for this device" while a GPU check is pending, or when the real reason is the model family. The model card line in My Models is the better guide.

After updating, each phone measures its routes once while idle and cool (this can take 15 to 25 minutes per model in the background). A phone that keeps running hot may not finish; it keeps its current setup.

Pause background checks

Where: Settings (Advanced mode) > "Speed & engine" > "Performance" > "Pause background checks".

"When on, TokForge does not run its speed test or GPU check by itself while the phone is idle. Checks you start still run."

What it holds: the automatic GPU check and the automatic speed test. While it is on, the model card says "GPU check paused: background checks are off in Settings. Tap Check now to run it."

What it does not hold: a check you start yourself, with "Check now" or from AutoForge.

You do not need this switch. Without it, automatic checks already run only while the app sits idle in the foreground, the phone is cool, the battery is not low (or the phone is charging) and power saving is off. They wait while a download runs and stop as soon as the running step ends when you type, switch models or leave the app; a chat you send during a check can wait a short while. If a check cannot put your model back afterwards, you see "The speed check could not load ... again. Load it again to keep chatting."

Known limit: after you turn "Pause background checks" off again, a speed check that was waiting may not start on its own. A check you start yourself still works: "Check now" on the model's card in My Models runs the GPU check, and AutoForge runs a speed test.

CPU prompt, GPU reply and the one-copy split

Nothing to do. Your phone decides.

What it is: a reply has two parts, reading your prompt and writing the reply. TokForge can give each part to a different chip:

TokForge uses a mix only where your phone's own check passes and the measurement wins. Otherwise the model stays on the CPU, because it measured faster, the phone ran hot, memory was short, or the phone has no usable GPU route. The recommended Qwen3.5 model can use the "GPU prompt, CPU reply" mix on some phones.

How to see it

For MNN models you can also ask for the mix yourself: "Tune" > "Hardware" > "GPU reads my prompt" ("GPU reads your prompt, CPU writes the reply. Long prompts start sooner. Takes effect the next time the model loads."). It still has to pass your phone's check.

Memory: while a GGUF model loads, TokForge now lets go of file pages it has already copied, so loading takes less memory. Speed checks also skip any route that would not fit your phone's free memory or its per-app limits.

Speed features

Nothing to do. These turn on only where a complete test on your phone shows a gain.

What they are

None of them applies to the recommended Qwen3.5 model.

Where they show: each has a row with "Off", "Auto" and "On" in Settings (Advanced mode) > "Speed & engine" > "Performance" > "Measured on your phone". Prompt lookup also has its own row in "Performance" ("Applies the next time a model loads. Auto turns it on only on phones and models where it was measured faster.").

On some MediaTek phones, the draft setting for Qwen3 8B and 14B is no longer on by default, because it measured slower there. You can still turn it on with "On" in the "Draft model speculation" row.

Other opt-ins: the Gemma 4 mix ("Gemma 4: GPU reads the prompt") and Vulkan for picture generation ("Image generation on the GPU") stay off unless you set them to "On" in "Measured on your phone". (Gemma 4 E4B stays on the CPU on every GPU.)

NPU speed boost (Snapdragon 8 Gen 3 and 8 Elite Gen 5)

An optional download lets the phone's AI chip (the NPU) read your prompt, so first replies to long prompts come sooner. It works with Qwen3 0.6B, 1.7B and 4B Abliterated. The CPU still writes the reply. It stays off unless you turn it on.

Turning it on at first run: leave "Includes speed boost for your phone (+... GB)" ticked on the first-run card. The card says "Your phone checks it before using it. It turns on the next time you open TokForge."

Turning it on later:

  1. Load one of the models above.
  2. In the chat, tap "Tune", then look under "Hardware".
  3. Turn on "Faster long prompts (NPU)". The line under it shows the download size. Confirm the download.
  4. Wait for "Ready. From the next time this model loads, prompts of 512 tokens or more are read by the NPU."
  5. Close and reopen TokForge when it says "Restart TokForge to start using the NPU for this model."

The switch only shows on these two chip families, for a model that has a boost download.

What you may see

The self-check: the first time the boost loads, it checks itself against the CPU. If the check fails you see "In a quick check the NPU read a prompt differently from the CPU, so this model reads prompts on the CPU on this phone. TokForge checks again after an app or system update." Your chats keep working on the CPU.

To free the space, use "Remove NPU files (... GB)" in the same section.

GPU picks you make yourself

You can still choose the engine yourself in the chat's "Tune" > "Hardware" > "Backend". "Auto (measured)" lets your phone decide; that is the default.

What happens when you pick a GPU route:

  1. A GPU setup you pick yourself gets the GPU check before it starts. The banner says "Checking the GPU for this model (first time only, one to two minutes)...".
  2. If your phone passes, your pick runs.
  3. If your phone turns it down, TokForge tells you why in a notice, runs the model on the CPU, and the speed panel shows the reason under your pick.

Notices you may see

Where a GPU pick is always turned down in 1.3.8:

If the GPU stalls: TokForge moves that model to the CPU and, once nothing of yours is running, restarts itself in the background to free the GPU. A GGUF model that stalls on the GPU before its first word can still take several minutes to give up.

To go back, pick "Auto (measured)" or CPU in the same "Backend" setting.

Agents tab

Agents now show what they can do, show each step while they work, and keep their results.

Where: on the Home screen, the "Agents" card ("Hand off a small job and let it run itself"). On a phone-width screen it can be off to the side: swipe the Explore row or tap "See all".

The list

The three presets

Running one:

  1. Tap a preset.
  2. Type your request in the box ("Type what you want, or tap an example"), or tap one of the example requests to fill it. Examples include "When was the Eiffel Tower finished?", "Write a short email asking my landlord to fix the heating." and "Summarize this in five bullet points."
  3. For "Summarize", tap "Choose a file" or paste the text. If the box is empty it asks: "Choose a file, or paste the text you want summarized."
  4. Tap "Run".
  5. Watch the steps, for example "Step 1 of 2: deciding what to do", "Found 3 results on DuckDuckGo" and "Writing the answer". The timer line says how long it has run and when it stops on its own.
  6. Tap "Stop" to end it early.

Leaving a running agent asks "Stop this run?" ("Leaving this screen stops the agent.") with "Stop and leave" or "Keep running".

Using the result

A long file is marked as shortened, for you ("This file is long. The agent reads the first part only.") and for the model.

History: "Recent runs" lists your latest runs. Open one to see its steps and result, or tap "Run again". A run cut short by the app closing shows "Stopped: the app closed during the run".

If you edited a preset, "Reset to default" brings it back.

Errors are in plain words, for example "No model is loaded. Load a chat model in Models, then try again." or "The model is busy with another task. Try again when it finishes." Tap "Details" for more.

What the model line means: the run screen tells you how well the loaded model suits agents.

Known limit: on phones where the model list marks Qwen3.5 4B "Too big", the Agents screen can still offer "Get Qwen3.5 4B", and there is nothing to get. Keep the model you have; agents still work with it.

Also good to know

Backup and restore

Where: Settings > "Privacy & about" > "Data & Backup".

What a backup holds, in the app's words: "Backup exports your conversations, characters, settings, and model metadata." The models themselves are not in it.

Make a backup:

  1. Tap "Create Backup".
  2. Choose where to save the file.
  3. Wait for "Backing up..." to finish. The line under the buttons then reads "Backup created: ..." with the number of conversations and characters in the file.

Restore a backup:

  1. Tap "Restore Backup".
  2. Pick the backup file.
  3. Wait for "Restoring..." to finish. The line under the buttons reads "Restored: ... API keys must be re-entered. Restarting..." and TokForge closes.
  4. After a restore, TokForge restarts its background service but may not reopen its screen; open it again from the launcher.

"Restore can overwrite local state, so keep at least one known-good backup file."

If a restore is refused, the line reads "Restore failed: ..." with the reason.

What changed in 1.3.8

Importing chats ("Import chats" in the same section) is not the same as Restore. For now, importing a chat export adds every chat again as a copy instead of merging, and a character comes back without its card. Use Restore Backup to move everything to a new phone.

Characters and personas

Nothing new to set up. One known limit:

In a character chat with memory on, asking "What is my name?" can get a stock "I don't have your name recorded" reply instead of your persona's name. A memory shortcut answers questions that start like that without asking the model. Asking in other words, or turning memory off for that chat, may get the question to the model instead; we have not tested every wording.

Documents

  1. In a chat, attach a document.
  2. Watch the status under it. It shows two parts, "Exact: ..." and "Broad: ..." (for example "Exact: Ready" and "Broad: Summarizing").
  3. Wait until both say "Ready" ("Exact and broad doc questions are ready."). This can take a little while after you attach it.
  4. Ask your question.

Known limit: if you ask while TokForge is still preparing the document, the reply can say "The attached documents do not clearly specify that." even though the answer is in the file. Ask again once both parts say "Ready".

The chip under a reply says how TokForge searched the document, not that the answer is right. Read the excerpt shown with the reply to check.

Remote API

Use a model running on your own server or a service you choose, instead of the phone.

Set it up:

  1. Settings > "Connections" (tap "Manage connections") > "Remote API".
  2. "Base URL": type the server's address (an OpenAI-style address, often ending in /v1). TokForge reads the server's model list by itself; "Fetch models" asks again.
  3. "Model name": pick a model from the list.
  4. "API key (optional)": add a key if your server needs one. Keys are stored encrypted on your phone.
  5. "Parameter compatibility": leave it on "Auto" unless you know you need another setting.
  6. Tap "Test Connection". It checks that the address works, your key is accepted and models are found. A good result ends with "Connected. Selected model is available."

Use it in a chat:

  1. Open a chat and tap the engine chip at the top of the chat.
  2. Pick "Remote API". You see "Remote API backend selected." and the chip reads "Cloud".

The unencrypted warning: if the address starts with http:// (not https://) you see "Unencrypted connection. Data sent to this server is not encrypted." That is fine on your own home network. Use https:// for anything that crosses the internet.

What is sent: your messages and the chat you are in go to the server you set up. TokForge does not send them anywhere else. If the terminal is on and you approve a command in a chat that uses a remote model, the approval dialog adds "The command output will be sent to your remote model.", and that output goes to your server too.

You may see "Model included reasoning (hidden)" with some remote models even when thinking is off. It is only a notice.

Upgrading from 1.3.6.1 or older

Your chats, characters and models stay. These things change once:

Important: do not install an older TokForge (1.3.4.1 or earlier) over 1.3.8. It will not open, and the only way back is to uninstall, which erases your chats and settings. Make a backup before you change versions.

Other changes you can see

Chat

Screens

AutoForge needs the Developer API server for its tests. If it is off, the "Quick", "Thorough" and "Most thorough" cards say so. To turn it on:

  1. Settings > "Settings Mode" > "Advanced".
  2. Settings > "Connections" > "Developer tools (Advanced)".
  3. Turn on "Developer API server (Advanced)". Turn it off again when you are done; most people should leave it off.

Stability (nothing to do)

Downloads: if the network drops during a model download, the download stops instead of waiting. Tap "Retry" and it continues where it stopped.

Phone notes

The terminal (opt-in, first version) has its own guide: TokForge Terminal: User Guide.

Get 1.3.8

To get 1.3.8 from Google Play, open the TokForge listing and join the beta there. The Discord APK is the same version. Everyone else stays on 1.3.6.1, the Play production version.

Important: do not install an older TokForge (1.3.4.1 or earlier) over 1.3.8. It will not open, and the only way back is to uninstall, which erases your chats and settings.

The full list of changes is in the 1.3.8 changelog.