TokForge 1.3.8: How-To
This guide explains, feature by feature, what you can do and what you will see in 1.3.8. Most of the speed work is automatic: where a section says "Nothing to do", the app handles it and the section tells you what you will notice. The terminal has its own guide (TokForge Terminal: User Guide).
Where Settings paths are given, the group names are the ones on the Settings screen: "Speed & engine", "Connections", "Privacy & about" and so on. Where a path starts with "Settings (Advanced mode)", first set "Settings Mode" at the top of Settings to "Advanced".
First run and the model it picks · The simple model list and the Advanced switch · Measured on your phone · Pause background checks · CPU prompt, GPU reply and the one-copy split · Speed features · GPU picks you make yourself · Agents tab · Backup and restore · Characters and personas · Documents · Remote API · Upgrading from 1.3.6.1 or older · Other changes you can see
First run and the model it picks
New installs start with one question instead of a long model list.
- Open TokForge and tap "Get Started".
- "What do you want to do?" Pick one: "Chat and questions", "Writing and stories", "Code and math" or "Make pictures".
- TokForge shows "Recommended for your phone" with one model, its download size, a speed estimate and a network note.
- Tap "Get recommended model (... GB)" and confirm the download.
- When the model is ready, start your first chat.
What it picks
- "Chat and questions" and "Code and math" get a Qwen3.5 model sized to your phone's memory (0.8B, 2B or 4B).
- "Writing and stories" gets Qwen3 4B Abliterated on phones with 8 GB of memory or more.
- If your phone is low on storage, the card says "A smaller pick, because your phone is low on space."
- If nothing fits for that choice, it says "Nothing fits this phone for that yet. Try another choice, or pick a model yourself."
What the card tells you
- Speed: "About ... words per second on your phone", or "on phones like yours (estimate)", or "Speed is measured on your phone after your first chat". It is an estimate, not a promise.
- Network: "You're on Wi-Fi", or "You're on mobile data. Wi-Fi is best for a ... GB download.", or "No internet right now. The download starts when you're back online."
Other buttons on the card
- "Show other options" lists a few more models for the same task.
- "Change what I want to do" goes back to the question.
- "Skip, I'll pick a model myself" goes straight to the model list.
Speed boost tick (Snapdragon 8 Gen 3 and 8 Elite Gen 5 phones only): the card may show a ticked line "Includes speed boost for your phone (+... GB)". Untick it if you do not want the extra download. The model downloads first, so you can start chatting while the boost finishes. See "Speed features" below.
Good to know
- Android may ask a few things during the first download: to keep the download going in the background, to allow notifications, and about battery optimization. Allowing them helps a long download finish.
- Right after your first model, TokForge also downloads a few helper models it uses for documents, attached pictures and voice (around 1 GB together). They show in My Models.
- Your first chat may show "... isn't tuned for this phone yet" with "Dismiss" and "Optimize". You can tap "Dismiss". TokForge measures the best route for your phone by itself while it is idle; AutoForge (the "Optimize" button) is extra, optional tuning.
The simple model list and the Advanced switch
The model list now sorts models by how well they suit your phone:
- "Best for your phone": the pick that suits your phone best.
- "Fits": runs on your phone.
- "Too big for your phone (needs ... GB)": your phone does not have enough memory. Its download button is greyed out.
- A few very large models show "Large & slow ยท opt-in": they can run, but slowly.
The "Fits" badge now asks the same memory question as loading, so a model that is too big says "Too big" up front instead of "Fits" followed by a refused load.
The Advanced switch
- In the model list, the "Advanced" switch reads "Show engine details, file formats and the checks measured on your phone".
- Settings > "Settings Mode" > "Advanced" is the same app-wide switch.
- With Advanced off, the list hides most engine words: MNN, GGUF, OpenCL, Vulkan and NPU. Some quant names still show in My Models and Browse. With it on, you see them all again, plus the technical details.
- If you are upgrading, it starts switched on, so your list looks as it did before.
- A few screens (for example the chat's settings sheet) can still show an engine name with Advanced off.
Also new in the catalog
- Descriptions are rewritten in plain words in all 8 languages. None of them promises the GPU any more, because your phone decides that (see the next section).
- Model names say "Abliterated" where a model has its refusals removed.
- A new Qwen3 1.7B build. If you already installed the old one, your copy stays.
Searching Hugging Face (Browse) uses its own badges ("Fits your phone", "Tight fit, may be slow") that only look at memory. Check the file size against your free storage before you download from there.
Measured on your phone
TokForge now checks each speed trick on your own phone before using it: the answers must match the normal route, and the trick must really be faster on your phone. This screen shows what it found and lets you overrule it.
Where: Settings (Advanced mode) > "Speed & engine" > "Performance" > "Measured on your phone".
The master switch: "Use speed features this phone has proven". Turn it off and every automatic speed feature is off. Turn it back on to let the rows below decide.
Each row is one feature, with a status line and three buttons, "Off", "Auto" and "On". Rows can include "GPU route (GGUF)", "GPU route (MNN)", "Qwen3.5: GPU reads the prompt", "One model copy for CPU and GPU", "Prompt lookup", "Multi-token prediction (MNN)", "Multi-token prediction (GGUF)", "Draft model speculation", "NPU reads long prompts", "Gemma 4: GPU reads the prompt", "MoE models: CPU and GPU share the work" and "Image generation on the GPU".
What the status lines mean
- "On: checked and faster on this phone": it passed and measured faster, so Auto uses it. (It may also say by how much.)
- "Off: not faster on this phone": it worked but did not win, so Auto leaves it off.
- "Off: failed the check on this phone": its answers did not match, so it stays off.
- "Not available for this phone or model": it does not apply here.
- "Checking now": a check is running.
- "Not checked yet": no result yet. Checks run when the phone is idle.
What the buttons mean
- "Auto" (the default): used only where this phone's check passed and the feature measured faster.
- "Off": never used.
- "On": you want it. It still has to pass the correctness check on your phone; a failed check keeps it off.
"Changes apply the next time a model loads."
Also in "Performance": "Measure CPU and GPU for each model" (on by default): "Runs a short speed test the first time a model loads and uses the faster route. Results stay on this phone."
Results are tied to the app version, the system update, the GPU driver and the model file. When any of those change, your phone checks again by itself.
Reading the route lines
You can see what your phone decided in several places.
While a check runs, a banner at the top of the app says:
- "Checking the GPU for this model (first time only, one to two minutes)..." A chat you send meanwhile waits for it.
- "Measuring CPU and GPU speed for this model..." with "Skip". "Skip" stops the speed check for that model. A chat you send while a check runs can wait a short while.
In My Models, on the loaded model's card
- "GPU check pending: it runs when the phone is idle, first time only." with "Check now". Tap "Check now" to run it right away.
- After the speed check, a short line says where the work runs and why, for example "... replies stay on CPU", "CPU is as fast as the GPU or faster here, so this model stays on CPU", "GPU failed a quick accuracy check here, so this model uses CPU" or "No GPU route to measure for this model here".
Notices after a check
- "GPU check passed: ... runs on the GPU on this phone."
- Or the reason it stays on the CPU, for example "In a quick check the GPU gave a different answer than the CPU, so this model runs on the CPU on this phone. TokForge checks again after an app or system update."
In the chat, tap "Tune", then look under "Hardware":
- "Backend" shows your choice. "Auto (measured)" means your phone's measurement decides.
- The "Effective" value under it says what really runs. A mix reads "GPU reads your prompt, CPU writes the reply." or "CPU reads your prompt, GPU writes the reply."
In AutoForge (Settings (Advanced mode) > "Speed & engine" > "Performance" > the "AutoForge" card, then the "Tune" tab), MNN models have a "Backend status" row. Its labels mean:
- Pending: "GPU route (check pending)" or "CPU route (GPU check pending)". The GPU check has not run yet; it runs once when the phone is idle. While "Pause background checks" is on, the sentence under it says the check is paused.
- Passed: "GPU verified on your device" or "Certified GPU route" with "This model passed TokForge's quick GPU check on this phone (the GPU gave the same answer as the CPU), so it runs on the GPU."
- Measured: "TokForge measured this model on this phone and the GPU route was faster, so it runs on the GPU." Or "CPU route (GPU tested on your device)": the GPU was tested and the CPU stayed faster.
- Family default: "CPU route (supported for this model)". Models of this family (such as Qwen3.5) keep a running state on the CPU, so the CPU is their normal route. You can still try the GPU from the backend setting.
- "GPU route (output check skipped)": it runs on the GPU without the comparison, because TokForge could not get a CPU answer to compare with. It tries again on a later load.
- "GPU turned off again (back on CPU)": the GPU had a problem, so the model went back to the CPU and is checked again later.
- "GPU experiment (your choice)": you picked the GPU yourself.
GGUF models have no "Backend status" row; AutoForge shows a plain line such as "Running on your processor." or "Running on your phone's graphics chip."
Known limit: the AutoForge row can still read "Certified CPU route ... the certified default for this device" while a GPU check is pending, or when the real reason is the model family. The model card line in My Models is the better guide.
After updating, each phone measures its routes once while idle and cool (this can take 15 to 25 minutes per model in the background). A phone that keeps running hot may not finish; it keeps its current setup.
Pause background checks
Where: Settings (Advanced mode) > "Speed & engine" > "Performance" > "Pause background checks".
"When on, TokForge does not run its speed test or GPU check by itself while the phone is idle. Checks you start still run."
What it holds: the automatic GPU check and the automatic speed test. While it is on, the model card says "GPU check paused: background checks are off in Settings. Tap Check now to run it."
What it does not hold: a check you start yourself, with "Check now" or from AutoForge.
You do not need this switch. Without it, automatic checks already run only while the app sits idle in the foreground, the phone is cool, the battery is not low (or the phone is charging) and power saving is off. They wait while a download runs and stop as soon as the running step ends when you type, switch models or leave the app; a chat you send during a check can wait a short while. If a check cannot put your model back afterwards, you see "The speed check could not load ... again. Load it again to keep chatting."
Known limit: after you turn "Pause background checks" off again, a speed check that was waiting may not start on its own. A check you start yourself still works: "Check now" on the model's card in My Models runs the GPU check, and AutoForge runs a speed test.
CPU prompt, GPU reply and the one-copy split
Nothing to do. Your phone decides.
What it is: a reply has two parts, reading your prompt and writing the reply. TokForge can give each part to a different chip:
- "GPU prompt, CPU reply": the GPU reads your prompt and the CPU writes the reply. It mostly helps long prompts.
- "CPU prompt, GPU reply": the CPU reads your prompt and the GPU writes the reply.
- For GGUF models, "CPU prompt, GPU reply" can share one copy of the model between the CPU and the GPU instead of loading two, which saves memory. Two copies are kept only where they measure faster.
TokForge uses a mix only where your phone's own check passes and the measurement wins. Otherwise the model stays on the CPU, because it measured faster, the phone ran hot, memory was short, or the phone has no usable GPU route. The recommended Qwen3.5 model can use the "GPU prompt, CPU reply" mix on some phones.
How to see it
- In the chat, "Tune" > "Hardware": the "Effective" value reads "GPU reads your prompt, CPU writes the reply." or "CPU reads your prompt, GPU writes the reply."
- In My Models, the model's line says when "both use one shared copy of the model".
- In "Measured on your phone": the rows "Qwen3.5: GPU reads the prompt" and "One model copy for CPU and GPU".
For MNN models you can also ask for the mix yourself: "Tune" > "Hardware" > "GPU reads my prompt" ("GPU reads your prompt, CPU writes the reply. Long prompts start sooner. Takes effect the next time the model loads."). It still has to pass your phone's check.
Memory: while a GGUF model loads, TokForge now lets go of file pages it has already copied, so loading takes less memory. Speed checks also skip any route that would not fit your phone's free memory or its per-app limits.
Speed features
Nothing to do. These turn on only where a complete test on your phone shows a gain.
What they are
- Prompt lookup (MNN models): speeds up replies that repeat text already in the chat, such as summaries, quotes, code edits and answers about a document. The model checks every guess.
- Multi-token prediction: works with Qwen3.5 GGUF files you add from Browse (the catalog's Qwen3.5 download has no prediction head yet).
- The Llama draft pack: works only with much larger Llama 3 models and is never downloaded on its own.
None of them applies to the recommended Qwen3.5 model.
Where they show: each has a row with "Off", "Auto" and "On" in Settings (Advanced mode) > "Speed & engine" > "Performance" > "Measured on your phone". Prompt lookup also has its own row in "Performance" ("Applies the next time a model loads. Auto turns it on only on phones and models where it was measured faster.").
On some MediaTek phones, the draft setting for Qwen3 8B and 14B is no longer on by default, because it measured slower there. You can still turn it on with "On" in the "Draft model speculation" row.
Other opt-ins: the Gemma 4 mix ("Gemma 4: GPU reads the prompt") and Vulkan for picture generation ("Image generation on the GPU") stay off unless you set them to "On" in "Measured on your phone". (Gemma 4 E4B stays on the CPU on every GPU.)
NPU speed boost (Snapdragon 8 Gen 3 and 8 Elite Gen 5)
An optional download lets the phone's AI chip (the NPU) read your prompt, so first replies to long prompts come sooner. It works with Qwen3 0.6B, 1.7B and 4B Abliterated. The CPU still writes the reply. It stays off unless you turn it on.
Turning it on at first run: leave "Includes speed boost for your phone (+... GB)" ticked on the first-run card. The card says "Your phone checks it before using it. It turns on the next time you open TokForge."
Turning it on later:
- Load one of the models above.
- In the chat, tap "Tune", then look under "Hardware".
- Turn on "Faster long prompts (NPU)". The line under it shows the download size. Confirm the download.
- Wait for "Ready. From the next time this model loads, prompts of 512 tokens or more are read by the NPU."
- Close and reopen TokForge when it says "Restart TokForge to start using the NPU for this model."
The switch only shows on these two chip families, for a model that has a boost download.
What you may see
- "Speed boost: downloading ...%" then "Speed boost: checking the files" then "Speed boost ready. It turns on the next time you open TokForge."
- On mobile data or Data Saver, the download waits: "Speed boost: waiting for Wi-Fi. It downloads by itself on Wi-Fi, not over mobile data." It continues once you are on Wi-Fi.
- If it stopped: "The speed boost did not finish. TokForge tries again the next time this model loads." In the switch's section, "Resume download" restarts it.
The self-check: the first time the boost loads, it checks itself against the CPU. If the check fails you see "In a quick check the NPU read a prompt differently from the CPU, so this model reads prompts on the CPU on this phone. TokForge checks again after an app or system update." Your chats keep working on the CPU.
To free the space, use "Remove NPU files (... GB)" in the same section.
GPU picks you make yourself
You can still choose the engine yourself in the chat's "Tune" > "Hardware" > "Backend". "Auto (measured)" lets your phone decide; that is the default.
What happens when you pick a GPU route:
- A GPU setup you pick yourself gets the GPU check before it starts. The banner says "Checking the GPU for this model (first time only, one to two minutes)...".
- If your phone passes, your pick runs.
- If your phone turns it down, TokForge tells you why in a notice, runs the model on the CPU, and the speed panel shows the reason under your pick.
Notices you may see
- "Even one shared copy of this model would not fit this phone's memory, so it runs on the CPU."
- "The CPU and GPU copies of this model would not both fit this phone's memory, so it runs on the CPU."
- "The GPU has not passed a quick check for this model on this phone yet, so it runs on the CPU."
- "The GPU cannot read this model's prompt on this phone, so it runs on the CPU."
- "The model loader could not find a usable Vulkan GPU, so this model runs on the CPU."
- "The GPU is slower than the CPU for this model on this phone, so it runs on the CPU. TokForge checks again after an app or system update."
Where a GPU pick is always turned down in 1.3.8:
- PowerVR GPUs (Pixel 10) stay on the CPU, even when you pick Vulkan.
- Gemma 4 E4B stays on the CPU on every GPU.
- GGUF on Vulkan is not offered on Adreno GPUs older than the 700 series (for example Snapdragon 865 phones).
- On Immortalis-G720, MNN on OpenCL stays off for every model.
- On phones that limit how much GPU memory one app may use, and on Xiaomi phones with HyperOS (per-app memory limit), a setup that would go over the limit runs on the CPU.
If the GPU stalls: TokForge moves that model to the CPU and, once nothing of yours is running, restarts itself in the background to free the GPU. A GGUF model that stalls on the GPU before its first word can still take several minutes to give up.
To go back, pick "Auto (measured)" or CPU in the same "Backend" setting.
Agents tab
Agents now show what they can do, show each step while they work, and keep their results.
Where: on the Home screen, the "Agents" card ("Hand off a small job and let it run itself"). On a phone-width screen it can be off to the side: swipe the Explore row or tap "See all".
The list
- "Agents do a task for you, step by step" explains the idea.
- "Start here" has the three presets.
- "Recent runs" shows your latest runs.
- "More ready-made agents" and "Your agents" follow.
The three presets
- "Look it up": "Searches the web and gives you a short answer with links." It searches DuckDuckGo and reads the titles and short previews of the top 3 results. It does not open full pages. "Uses the internet. Your question is sent to DuckDuckGo."
- "Write or rewrite": "Turns a few notes into an email, message or short post, or rewrites your text." "Works offline. Nothing leaves your phone." Check names, dates and facts before you send.
- "Summarize": "Reads a file or text you paste and gives you the key points." It reads PDF, Word, EPUB and text files, or pasted text, up to about the first 1,000 words. "Works offline. The file stays on your phone."
Running one:
- Tap a preset.
- Type your request in the box ("Type what you want, or tap an example"), or tap one of the example requests to fill it. Examples include "When was the Eiffel Tower finished?", "Write a short email asking my landlord to fix the heating." and "Summarize this in five bullet points."
- For "Summarize", tap "Choose a file" or paste the text. If the box is empty it asks: "Choose a file, or paste the text you want summarized."
- Tap "Run".
- Watch the steps, for example "Step 1 of 2: deciding what to do", "Found 3 results on DuckDuckGo" and "Writing the answer". The timer line says how long it has run and when it stops on its own.
- Tap "Stop" to end it early.
Leaving a running agent asks "Stop this run?" ("Leaving this screen stops the agent.") with "Stop and leave" or "Keep running".
Using the result
- Select any part of the text.
- Tap a link under "Sources" to open it.
- "Copy", "Share" or "Save as file".
A long file is marked as shortened, for you ("This file is long. The agent reads the first part only.") and for the model.
History: "Recent runs" lists your latest runs. Open one to see its steps and result, or tap "Run again". A run cut short by the app closing shows "Stopped: the app closed during the run".
If you edited a preset, "Reset to default" brings it back.
Errors are in plain words, for example "No model is loaded. Load a chat model in Models, then try again." or "The model is busy with another task. Try again when it finishes." Tap "Details" for more.
What the model line means: the run screen tells you how well the loaded model suits agents.
- "Qwen3.5 4B is loaded. It is the best tested model for agents."
- "... is loaded. Agents work with it, but Qwen3.5 4B gets more tasks right."
- "... is loaded. Larger models are often too slow to finish agent tasks on a phone. Qwen3.5 4B is a better fit."
- "This phone has less than 8 GB of memory. Agents may be slow or stop early."
- "Get Qwen3.5 4B" takes you to the models.
Known limit: on phones where the model list marks Qwen3.5 4B "Too big", the Agents screen can still offer "Get Qwen3.5 4B", and there is nothing to get. Keep the model you have; agents still work with it.
Also good to know
- Agents that use tools now run with fixed, non-random settings, which makes tool use steadier. Agents without tools keep your own sampling settings.
- Results can show formatting marks such as ** around words.
- Small models can get facts wrong in a summary. Check anything important.
- Agents cannot run terminal commands in 1.3.8.
Backup and restore
Where: Settings > "Privacy & about" > "Data & Backup".
What a backup holds, in the app's words: "Backup exports your conversations, characters, settings, and model metadata." The models themselves are not in it.
Make a backup:
- Tap "Create Backup".
- Choose where to save the file.
- Wait for "Backing up..." to finish. The line under the buttons then reads "Backup created: ..." with the number of conversations and characters in the file.
Restore a backup:
- Tap "Restore Backup".
- Pick the backup file.
- Wait for "Restoring..." to finish. The line under the buttons reads "Restored: ... API keys must be re-entered. Restarting..." and TokForge closes.
- After a restore, TokForge restarts its background service but may not reopen its screen; open it again from the launcher.
"Restore can overwrite local state, so keep at least one known-good backup file."
If a restore is refused, the line reads "Restore failed: ..." with the reason.
What changed in 1.3.8
- Restore works on Android 15 and newer. Before, every backup was refused there with "Invalid backup: database failed integrity check". Restore now accepts good backups and still refuses a damaged one before your current data is touched. Android 14 and older were not affected.
- Backups include what you just wrote. Chats, messages and characters saved in the last minutes before a backup are now included, and the counts match the file.
- If the database stays busy, the backup stops with "Backup failed: Database busy, so no backup was written. Try again in a moment." Wait a moment and tap "Create Backup" again.
- Backups made with older versions may lack their newest chats, so make a fresh one after updating.
Importing chats ("Import chats" in the same section) is not the same as Restore. For now, importing a chat export adds every chat again as a copy instead of merging, and a character comes back without its card. Use Restore Backup to move everything to a new phone.
Characters and personas
Nothing new to set up. One known limit:
In a character chat with memory on, asking "What is my name?" can get a stock "I don't have your name recorded" reply instead of your persona's name. A memory shortcut answers questions that start like that without asking the model. Asking in other words, or turning memory off for that chat, may get the question to the model instead; we have not tested every wording.
Documents
- In a chat, attach a document.
- Watch the status under it. It shows two parts, "Exact: ..." and "Broad: ..." (for example "Exact: Ready" and "Broad: Summarizing").
- Wait until both say "Ready" ("Exact and broad doc questions are ready."). This can take a little while after you attach it.
- Ask your question.
Known limit: if you ask while TokForge is still preparing the document, the reply can say "The attached documents do not clearly specify that." even though the answer is in the file. Ask again once both parts say "Ready".
The chip under a reply says how TokForge searched the document, not that the answer is right. Read the excerpt shown with the reply to check.
Remote API
Use a model running on your own server or a service you choose, instead of the phone.
Set it up:
- Settings > "Connections" (tap "Manage connections") > "Remote API".
- "Base URL": type the server's address (an OpenAI-style address, often ending in /v1). TokForge reads the server's model list by itself; "Fetch models" asks again.
- "Model name": pick a model from the list.
- "API key (optional)": add a key if your server needs one. Keys are stored encrypted on your phone.
- "Parameter compatibility": leave it on "Auto" unless you know you need another setting.
- Tap "Test Connection". It checks that the address works, your key is accepted and models are found. A good result ends with "Connected. Selected model is available."
Use it in a chat:
- Open a chat and tap the engine chip at the top of the chat.
- Pick "Remote API". You see "Remote API backend selected." and the chip reads "Cloud".
The unencrypted warning: if the address starts with http:// (not https://) you see "Unencrypted connection. Data sent to this server is not encrypted." That is fine on your own home network. Use https:// for anything that crosses the internet.
What is sent: your messages and the chat you are in go to the server you set up. TokForge does not send them anywhere else. If the terminal is on and you approve a command in a chat that uses a remote model, the approval dialog adds "The command output will be sent to your remote model.", and that output goes to your server too.
You may see "Model included reasoning (hidden)" with some remote models even when thinking is off. It is only a notice.
Upgrading from 1.3.6.1 or older
Your chats, characters and models stay. These things change once:
- No first-run screens: you go straight to Home.
- Advanced starts switched on, so your model list looks as it did. To try the simple list, turn off "Advanced" in the model list, or set Settings > "Settings Mode" to "Basic".
- AutoForge pick reset: if you used AutoForge on an MNN model before, its backend pick (CPU or GPU) resets to "Auto (measured)" once, so the routes your phone measures can run. Its other settings stay. The app does not show a message about it. To choose a backend again, use "Tune" > "Hardware" > "Backend" in the chat, or Settings. (Your saved AutoForge preset still holds the old pick, so applying it again can bring it back.)
- Notifications: on the first open after the update, Android may ask whether TokForge may send notifications. Either answer works.
- A one-time measurement: each phone measures its routes once while idle and cool (15 to 25 minutes per model in the background). Keep using the app normally.
- Make a fresh backup: backups from older versions may lack their newest chats.
Important: do not install an older TokForge (1.3.4.1 or earlier) over 1.3.8. It will not open, and the only way back is to uninstall, which erases your chats and settings. Make a backup before you change versions.
Other changes you can see
Chat
- Quicker follow-ups in Qwen chats (GGUF models, thinking off). Nothing to do.
- The first message in a new chat is kept exactly once and uses the chat's saved settings.
- MNN models keep your whole current message when it fits. A message too long to fit keeps its start and end, and the chat tells you.
- Past messages are no longer shortened while they fit the chat's memory window. Known limit: MNN chats still have a small window, so very long MNN chats can forget early facts. When the oldest messages are trimmed, the reply after the trim takes longer.
- If you stop a reply before it starts, that message no longer confuses the next answer.
- No more "No model loaded" dead end: if the chat's model is not loaded but another local model is, the chat uses the loaded one and says so.
- TokForge no longer starts its background memory work while you are in a chat, so your message comes first.
- Mistral 7B v0.3 and Gemma 3 1B (MNN): a broken tokenizer file is repaired when the model loads. Nothing to do.
Screens
- My Models lists the loaded model first. On tablets, Settings groups sit in a grid.
- "Effective" sits right under the MNN and GGUF backend choices and names a CPU and GPU mix when one is in use.
- Batch threads read "Automatic" when set to 0, and still show the count in use.
- The "Downloaded" badge is short and sits below the model name. Details keeps the full download notice.
- App notices no longer hide behind three-button navigation.
- The setup screen's speed boost sentence names the real Settings switch, and several languages got translation fixes.
AutoForge needs the Developer API server for its tests. If it is off, the "Quick", "Thorough" and "Most thorough" cards say so. To turn it on:
- Settings > "Settings Mode" > "Advanced".
- Settings > "Connections" > "Developer tools (Advanced)".
- Turn on "Developer API server (Advanced)". Turn it off again when you are done; most people should leave it off.
Stability (nothing to do)
- The most common "TokForge isn't responding" stop is fixed.
- While a GGUF model writes a reply on the processor, scrolling and typing now get their fair share of the processor.
- If the GPU runs out of memory during an MNN reply, the reply stops with an error instead of closing the app.
- After TokForge closed unexpectedly, the first MNN reply on OpenCL no longer starts slowly.
- Faster on the CPU: long prompts get their first reply sooner on phones whose processor supports it (such as the Pixel 9 and the Xiaomi 14T Pro), GGUF replies are faster on many phones, and Tensor G4 phones such as the Pixel 9 now use four CPU threads for GGUF replies.
Downloads: if the network drops during a model download, the download stops instead of waiting. Tap "Retry" and it continues where it stopped.
Phone notes
- On a Samsung tablet with about 5 GB of memory, a 4B model can get the app closed by the system or make it very slow. Pick a smaller model there.
- Xiaomi phones with HyperOS limit how much memory one app can use, so the largest GPU setups are skipped there, and on 12 GB HyperOS phones the 14B models are marked "Too big".
- Pixel 10 (PowerVR) and Snapdragon 865 phones stay on the CPU for now. Large mixture-of-experts models may stay on the CPU too.
- On some MediaTek phones, such as the OnePlus Ace 5 Ultra, the GPU check runs inside the app instead of in a separate helper process. OpenCL works normally there.
The terminal (opt-in, first version) has its own guide: TokForge Terminal: User Guide.
Get 1.3.8
To get 1.3.8 from Google Play, open the TokForge listing and join the beta there. The Discord APK is the same version. Everyone else stays on 1.3.6.1, the Play production version.
Important: do not install an older TokForge (1.3.4.1 or earlier) over 1.3.8. It will not open, and the only way back is to uninstall, which erases your chats and settings.
The full list of changes is in the 1.3.8 changelog.