TokForge 1.3.8 changelog
One update instead of several: TokForge now measures what your own phone does well and uses it, and new users get a much simpler start. Version 1.3.7 was never released, so this page lists everything that changed since 1.3.6.1, area by area.
Get 1.3.8
To get 1.3.8 from Google Play, open the TokForge listing and join the beta there. The Discord APK is the same version. Everyone else stays on 1.3.6.1, the Play production version.
Important: do not install an older TokForge (1.3.4.1 or earlier) over 1.3.8. It will not open, and the only way back is to uninstall, which erases your chats and settings.
Step-by-step help: the 1.3.8 how-to and the terminal guide.
On this page
Getting started · Your phone picks the route · GPU, family by family · Memory · Faster on the CPU · Speed features, measured per phone · Chat and prompts · Agents tab · Terminal (opt-in, first version) · Settings and UI · Backup and restore · Fixes · Known limits
Getting started
- Easier first run. Tell TokForge what you want to do (chat, writing, code or pictures) and it picks one model that fits your phone, with its size, a speed estimate and a Wi-Fi note.
- What it picks. Chat and Code get a Qwen3.5 model sized to your phone's memory (0.8B, 2B or 4B). Writing gets Qwen3 4B Abliterated on phones with 8 GB or more.
- Speed boost on the first-run card. On Snapdragon 8 Gen 3 and 8 Elite Gen 5 phones, the card ticks the optional speed boost download where it fits, and you can untick it. The model downloads first, so you can start chatting while the boost finishes.
- Simpler model list. Models are sorted into Best for your phone, Fits and Too big, and the technical details sit behind one Advanced switch. If you already use TokForge, your model list stays as it was.
- One Advanced switch for the whole app hides engine words (MNN, GGUF, quant names, OpenCL, Vulkan, NPU). If you are upgrading, it starts switched on, so nothing moves for you.
- Plain descriptions. Catalog descriptions are rewritten in plain words in all 8 languages, and none of them promises the GPU any more, because your phone decides that.
- New Qwen3 1.7B build in the catalog. A copy of the old one you already installed stays.
Your phone picks the route
- GPU check while idle. For the usual automatic GPU setups, a new model starts right away and TokForge compares the GPU's answers with the CPU's once, while the app sits idle. A GPU setup you pick yourself, and a few newer ones, get that check before they start.
- A fairer check. It judges only the words the CPU is sure about, so a near tie no longer fails a GPU that answers correctly. If the check fails, the model stays on the CPU and its card says why in plain words.
- First check in its own process. The first GPU check of an MNN model runs in a separate helper, so a GPU crash or hang during it does not close the app. A slower GPU that is still working through its first check is not mistaken for a hang.
- CPU or GPU, by measurement. TokForge also times the CPU, the GPU and a mix of both (GPU reads the prompt and CPU writes the reply, or the other way round) for the model you loaded, and keeps whichever measured faster on your phone. Settings (Advanced mode) > Speed & engine > Performance > Measured on your phone shows what it found and lets you switch any of it off. Each route and speed feature there has its result in plain words, Off, Auto and On, and one switch covers all of them.
- Rechecks when something changes. Results are tied to the app version, the system update, the GPU driver and the model file, so a change to any of them gets a fresh check. Older fixed rules for particular phones are now only starting points.
- The recommended Qwen3.5 model can read long prompts on the GPU while the CPU writes the reply, where your phone's own check passes and the measurement wins. On our test fleet that was three phones of ten; the rest stay on the CPU because it measured faster, the phone ran hot, memory was short, or the phone has no usable GPU route.
- One copy for both. For GGUF models, the "CPU prompt, GPU reply" mix can share one copy of the model between CPU and GPU (more under Memory).
- Steadier speed checks. The automatic check now times whole replies, a few times each, before it picks, and a phone that runs hot gets a shorter check. Settings (Advanced mode) > Speed & engine > Performance > Pause background checks holds the automatic checks while it is on.
- Checks stay out of your way. Automatic checks run only while the app sits idle in the foreground, the phone is cool, the battery is not low (or the phone is charging) and power saving is off. They wait while a download runs and stop as soon as the running step ends when you type, switch models or leave the app; a chat you send during a check can wait a short while. If a check cannot put your model back afterwards, TokForge tells you to load it again.
- Clearer GPU status. The MNN route line in AutoForge > Tune now says whether its GPU route was checked, is still waiting for its check, or ran without the output check. While Pause background checks is on, it says the check is paused.
- AutoForge keeps your measured route. AutoForge no longer switches off the routes your phone measured. It still tunes threads, precision, cache and context size, and its result names the route in use.
GPU, family by family
Every GPU route below is used only where your phone's own check passes and the measurement wins. Otherwise the model stays on the CPU.
- Adreno 750 and 840 (Snapdragon 8 Gen 3 and 8 Elite Gen 5). MNN can measure the "GPU prompt, CPU reply" mix. MNN Vulkan replies on the Adreno 840 wait for the GPU more efficiently. MNN on OpenCL no longer gives wrong or empty replies when a chat continues on the Adreno 750.
- All Adreno 700 and 800 GPUs. GGUF on Vulkan now follows rules for each Adreno family that keep its answers correct (the older engine gave wrong output on some Adreno GPUs). GGUF on Vulkan is offered only where those rules apply, with crash protection while it loads.
- Mali and Immortalis (G715, G720, G925). MNN can measure the "GPU prompt, CPU reply" mix, with tuned Vulkan reply settings and a new Vulkan prompt kernel for Qwen3.5. GGUF on Vulkan uses a smaller work tile that suits these GPUs. On Tensor G4 phones such as the Pixel 9, an MNN OpenCL crash is fixed.
- Xclipse (Exynos). The same smaller GGUF tile and tuned Vulkan settings. MNN on OpenCL no longer gives wrong or empty replies when a chat continues on the Xclipse 940.
- Any other GPU. GGUF on Vulkan now reads your phone's real GPU name, so GPUs that were never on our list get measured too, with crash protection. A Vulkan pick you make yourself is honoured there after its check.
-
What stays on the CPU, and why.
- Immortalis-G720: MNN on OpenCL stays off for every model, because it stalled in our tests there.
- PowerVR (Pixel 10): the engine groundwork is in, but the GPU measured slower than the CPU, so these phones stay on the CPU, even when you pick Vulkan yourself.
- Adreno 650 (Snapdragon 865): its Vulkan version is too old for GGUF, and MNN on OpenCL hung during warm-up there. GGUF on Vulkan is not offered on Adreno GPUs older than the 700 series.
- Gemma 4 E4B stays on the CPU on every GPU, because it crashed a GPU driver in our tests.
- More models on the mix. Llama 3.2 3B on MNN is no longer refused the "GPU prompt, CPU reply" mix, so your phone can measure it too.
- On phones that limit how much GPU memory one app may use, a GPU setup that would go over that limit now runs on the CPU instead.
- If you pick "CPU prompt, GPU reply" and your phone cannot run it (for example, not enough memory, or the phone's GPU memory limit), TokForge now tells you why and runs the model on the CPU.
- The same now goes for "GPU prompt, CPU reply" picks and for GPU picks the phone turns down.
- If the GPU stalls, TokForge moves that model to the CPU and, once nothing of yours is running, restarts itself in the background to free the GPU.
- A GPU route that hangs during a speed check is marked as failed, and a chat you send takes the engine back after a short wait.
Memory
- One copy of the model for CPU and GPU. For GGUF models, the "CPU prompt, GPU reply" route can now share one copy of the model between the CPU and the GPU instead of loading two, where your phone's own check passes and the measurement wins. Two copies are kept only when they measure faster. The route line says when both share one copy.
- Weights no longer counted twice. While a GGUF model loads, TokForge now lets go of the file pages it has already copied for the CPU or GPU, so loading takes less memory.
- During a large model's first GPU check, TokForge now frees memory it no longer needs right away, so the system is less likely to close the app during that check.
- Speed checks respect memory limits. Each route the automatic check would try must first fit your phone's free memory and the system's per-app limits, or it is skipped rather than risking the app being closed.
- HyperOS. On Xiaomi phones with HyperOS, GPU mixes, speed checks and catalog badges respect the per-app memory limit, so a large model's GPU mix is refused rather than risking the app being closed.
- GPU memory sizing. MNN's GPU memory checks now use sizes measured for each GPU family.
Faster on the CPU
- Faster long prompts on the CPU. On phones whose processor supports it, such as the Pixel 9 and the Xiaomi 14T Pro, the first reply to a long prompt comes much sooner.
- Faster replies on the CPU. GGUF models now write their replies faster on the CPU on many phones, after a fix to how the processor's cores wait for each other.
- Four threads on Tensor G4. On Tensor G4 phones such as the Pixel 9, GGUF replies now use four CPU threads instead of two, which measured faster with the fix above.
Speed features, measured per phone
- Extra speed only where your phone shows a gain. Prompt lookup, multi-token prediction and the Llama draft pack each get a quick test on your phone, and TokForge turns them on only where a complete test shows a gain.
- What they are. Prompt lookup for MNN models, multi-token prediction for Qwen3.5 GGUF files you add from Browse (the catalog's Qwen3.5 download has no prediction head yet), and the Llama draft pack, which now works only with much larger Llama 3 models and is never downloaded on its own. None of them applies to the recommended Qwen3.5 model.
- They stay yours. Each has Off, Auto and On in Settings (Advanced mode) > Speed & engine > Performance > Measured on your phone.
- Draft defaults retired. On some MediaTek phones, the Qwen3 8B and 14B draft setting is no longer on by default, because it measured slower there. You can still turn it on.
- Speed boost on Snapdragon 8 Gen 3 and 8 Elite Gen 5. An optional download lets the phone's AI chip read your prompt, so first replies to long prompts come sooner with Qwen3 0.6B, 1.7B and 4B Abliterated. It turns on after the next app start.
- Opt-in and checked. The speed boost stays off unless you turn on Faster long prompts (NPU) or tick it on the first-run card. It checks itself on its first load and stays on the CPU if that check fails. Its download waits for Wi-Fi when you are on mobile data or Data Saver, and continues there.
- Other opt-ins. The Gemma 4 mix (GPU reads the prompt, CPU writes the reply) and Vulkan for picture generation stay off unless you turn them on in Settings (Advanced mode) > Speed & engine > Performance > Measured on your phone.
Chat and prompts
- Quicker follow-ups in Qwen chats (GGUF models, thinking off): TokForge now reuses the earlier part of the chat instead of reading all of it again on every turn.
- First message in a new chat. It is now kept exactly once (a quick first send could be doubled or dropped), and it waits a moment for the chat's saved settings instead of running with the defaults.
- Your whole message. MNN models now keep your whole current message when it fits. A message too long to fit keeps its start and end, and the chat tells you.
- Past messages are no longer shortened while they fit the chat's memory window. MNN chats still have a small window, so very long MNN chats can forget early facts.
- Long MNN chats: fixed a rounding problem in how the model keeps track of positions in the chat past about 2,000 tokens.
- If you stop a reply before it starts, that message no longer confuses the next answer.
- Honest route notices. When a route you picked has to run on the CPU, the speed panel also shows the reason under your pick.
- No more "No model loaded" dead end. If a chat's model is not loaded but another local model is, the chat now uses the loaded one and says so, instead of failing on every send.
- Your message comes first. TokForge no longer starts its background memory reflection while you are in a chat, so a new message is less likely to wait behind one.
- One approval per message. When a reply only repeats your message and TokForge retries it, a command you already approved is not asked for or run a second time.
Agents tab
- The Agents tab shows what each agent does and that it is working. Each agent has its own run screen: type a request, watch its steps, stop it, and use the result. There are three starter presets (Look it up, Write or rewrite, Summarize), and your recent runs can be run again.
- Example requests. Each preset has example requests you can tap to fill the box. Look it up searches the web and lists its sources, Write or rewrite drafts text, and Summarize takes a file or pasted text.
- Use the result. Select the text, open the sources, Copy, Share or Save as file.
- History. Your latest runs show under Recent runs. Open one to see its steps and result, or run it again.
- Plain errors. Errors are explained in plain words in your language, and leaving a running agent asks before it stops the run.
- Summarize asks for text when the box is empty, a file that had to be shortened is marked as shortened for you and for the model, and an edited preset can be reset.
- Steadier tool use. Agents that use tools now run with fixed, non-random settings and no longer make an extra forced search after they have answered. Agents without tools keep your own sampling settings.
Terminal (opt-in, first version)
- Terminal, first version, off by default. Settings Mode: Advanced, then Connections > Developer tools (Advanced) > Terminal (sandboxed): it turns on only after your phone passes its own check. In chat, a command runs only after you approve it, and the approval says that the output goes to the model.
- What it is. A command shell that works in its own scratch folder inside TokForge. Its tools (toybox and curl) come with the app: there is no package manager, and it does not download programs to run.
- Internet is off unless you turn on Allow internet in the terminal.
- Stop everything ends all commands at once and tells you nothing is left running. Ctrl+C stops the current command, and an extra key row has Ctrl, Esc, Tab and the arrows.
- Commands in chat. Ask the model to run something, or type a line starting with $, and an approval shows the exact command first. With a remote chat model, the approval also says the output will be sent to it. The result shows as a "Ran a command" card with the exit code and the output.
- "Command not run." If you choose Don't run, or do not answer within a minute, the chat shows a "Command not run" line before the model replies. A system pop-up no longer uses up that minute.
- Clear about what ran. The model is told the terminal's limits, and the result names the one command that ran, with an instruction not to claim anything else.
- What it cannot do. Install packages, ping, show text colors, or run a script file directly (use sh script.sh). Commands run from chat only, not from agents. Android 11 phones cannot run the terminal. On HyperOS, two system warning lines no longer clutter the output.
Settings and UI
- My Models lists the loaded model first, and the Settings screen adapts to wider screens. On tablets, its groups sit in a grid.
- Effective backend now sits right under the MNN and GGUF backend choices, and names a CPU and GPU mix in use instead of just the CPU. A working "GPU prompt, CPU reply" pick is labelled as such.
- Batch threads read Automatic when set to 0 and still show the count in use.
- Downloaded badge is short and sits below the model name in the catalog. The background download notice in My Models is shorter, and Details keeps the full text.
- AutoForge: when the Developer API server is off, the Quick, Thorough and Most thorough cards say it must be on before you tap.
- The setup screen's speed boost sentence names the real Settings switch, and several languages got translation fixes.
- Clearer wording: model names say Abliterated, and the app dropped fixed speed promises it could not back for your phone. The speed on the first-run card is an estimate.
Backup and restore
- Restore works on Android 15 and newer. Every backup was refused there with "Invalid backup: database failed integrity check". Restore now accepts good backups and still refuses a damaged one before your current data is touched. Android 14 and older were not affected. After a restore, TokForge restarts its background service but may not reopen its screen; open it again from the launcher.
- Backups include what you just wrote. A backup could miss chats, messages and characters saved in the last minutes before it, even though its message counted them. Now they are included and the counts match the file. If the database stays busy, the backup stops with a clear message instead of writing an incomplete file. Backups made with older versions may lack their newest chats, so make a fresh one after updating.
Fixes
Fixes for GPU routes, memory, chat and backups are in their sections above. The rest:
Crashes and stops
- The most common "TokForge isn't responding" stop is fixed. We confirmed the fix on Samsung and Xiaomi phones.
- When a GGUF model writes a reply on your phone's processor, the reply no longer gets priority over the app's own screen, so scrolling and typing get their fair share of the processor.
- If the GPU runs out of memory during an MNN reply, the reply stops with an error instead of closing the app.
GPU
- After TokForge closed unexpectedly, the first MNN reply on OpenCL no longer starts slowly.
Chat and models
- Mistral 7B v0.3 and Gemma 3 1B (MNN): a broken tokenizer file that made every prompt far longer than it should be is now repaired when the model loads.
UI
- The catalog's Fits badge now asks the same memory question as loading, so a model too big for your phone says Too big, with the memory it needs, instead of Fits followed by a refused load.
- App notices no longer hide behind three-button navigation.
Known limits
- The GPU check of a new model runs once while the app sits idle in the foreground, and a chat you send during it waits for it. When TokForge cannot make a CPU answer to compare with, the GPU route runs without that comparison. The route line says so.
- After updating, each phone measures its routes once while idle and cool (this can take 15 to 25 minutes per model in the background). A phone that keeps running hot may not finish; it keeps its current setup.
- A GGUF model that stalls on the GPU before its first word can still take several minutes to give up.
- If you used AutoForge on an MNN model before, its backend pick resets to Auto once, so the routes your phone measured can run. Its other settings stay, and you can pick a backend again in Settings or in the model's speed panel.
- The terminal is a first version. It stays off until you turn it on, it is available only where your phone passes its check, and Android 11 phones cannot run it.
- Xiaomi phones with HyperOS limit how much memory one app can use, so the largest GPU setups are skipped there, and on 12 GB HyperOS phones the 14B models are marked Too big.
- Pixel 10 (PowerVR) and Snapdragon 865 phones stay on the CPU for now. Large mixture-of-experts models may stay on the CPU too.
- On some MediaTek phones, such as the OnePlus Ace 5 Ultra, the GPU check runs inside the app instead of in a separate safe process. OpenCL works normally there.
- After you turn Pause background checks off again, a speed check that was waiting may not start on its own. A check you start yourself still works.
- In the terminal, here-documents (cat <<EOF) fail. Write files with echo or printf, and run scripts with sh.
- On a Samsung tablet with about 5 GB of memory, a 4B model can get the app closed by the system or make it very slow. Pick a smaller model there.
- Long MNN chats drop their oldest messages sooner than they need to, so a fact from early in a long chat can be forgotten, and the reply right after that takes longer.
- Importing a chat export adds every chat again as a copy instead of merging, and a character comes back without its card.
- If the network drops during a model download, the download stops instead of waiting. Tap Retry and it continues where it stopped.
- In a character chat with memory on, asking "What is my name?" can get a stock "I don't have your name recorded" reply instead of your persona's name.
- If you ask about a document you just attached while TokForge is still preparing it, the reply can say the document does not specify it. Ask again once it is ready.
Thanks for the reports. Keep them coming in our Discord.