Phone-Speed Leaderboard
Real on-device LLM benchmarks, submitted from the TokForge app on real phones, now including iPhone. Ranked by decode speed in tokens per second. No simulators and no cloud, just what these devices actually do.
How fast does a phone actually run a local LLM? This page answers that with measurements instead of estimates. Every row is a benchmark run submitted from the TokForge app on a real handset: how many tokens per second the phone writes (decode), how fast it reads the prompt (prefill), how long the wait is before the first word appears (time to first token), and the exact model, quantization and backend that produced the number.
If the question is which phone is fastest for local AI, whether an iPhone keeps up with an Android flagship on Qwen or Llama, what an SM8850, an SM8750, an MT6899 or an Apple A17 Pro does on a small model, or whether a 9B model stays usable on a phone at all, the table below is the evidence for the devices people have actually tested. Speeds vary with the model, the quantization, the backend and how warm the phone is, so every row carries all four rather than a single headline figure.
The board is live. The full dataset is also baked into this page as static HTML with a dated snapshot, so it can be read and quoted without running JavaScript, and it is available as JSON from the public API linked at the foot of the page.
TokForge is free. Android is live on Google Play; iPhone and iPad are in public beta on TestFlight. No account, and it works with the network off once a model is downloaded.
Looking for one model in particular? Qwen3.5 0.8B, 2B, 4B and 9B, and Gemma 4 E2B each have a page carrying only that model's rows. Not sure what fits? Can my phone run it.
The full benchmark dataset
Every run the public API returns, unfiltered, baked into this page as static HTML. This is the citable copy: it does not need JavaScript, it carries an explicit measurement date per row, and it is refreshed daily from the same source the live board reads. The table scrolls sideways: past the model column each row also carries parameter count, quantization, backend, prefill speed, decode speed, time to first token, RAM tier, OS build, the method the passes were combined with, the app version and the measurement date.
Snapshot: 2026-08-22 (UTC). 64 benchmark runs from 44 devices, 28 chips and 37 model builds. Oldest run measured 2026-06-10, newest 2026-08-22, submitted from TokForge 1.1.1. Data collected with the TokForge app on physical devices, no simulators and no cloud inference. The App version column carries the exact build behind every row. Refreshed daily from the live board.
| # | Device | SoC | Model | Params | Quant | Backend | Prefill t/s | Decode t/s | TTFT | RAM | OS | Method | App version | Measured (UTC) |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | REDMAGIC NX809J | SM8850 | Qwen3.5-0.8B-MNN | 0.8B | MNN | CPU | 274.5 | 78.3 | 149 ms | 23 GB | Android 16 | best3 | 1.0 | 2026-08-13 |
| 2 | Apple iPhone18,2 | Apple A19 Pro | qwen2.5-1.5b-instruct-q4_k_m.gguf | 1.5B | Q4_K_M | LLAMA.CPP | - | 64.4 | - | 11 GB | iOS 26.5 | warmup1_median2 | 16.163 (1830) | 2026-07-06 |
| 3 | motorola motorola edge 30 pro | SM8450 | Qwen3-0.6B-MNN | 0.6B | MNN | CPU | 195.2 | 53.7 | 224 ms | 11 GB | Android 14 | best3 | 1.0 | 2026-08-14 |
| 4 | REDMAGIC NX809J | SM8850 | Qwen3-1.7B-MNN | 1.7B | - | CPU | 45.1 | 50.6 | - | 23 GB | - | - | v3.6.0-beta43 | 2026-06-11 |
| 5 | POCO 2412DPC0AG | MT6899 | Qwen3-0.6B-MNN | 0.6B | MNN | CPU | 221.6 | 50.2 | 175 ms | 11 GB | Android 16 | best3 | 1.0 | 2026-08-12 |
| 6 | Apple iPhone16,2 | Apple A17 Pro | Qwen2.5-1.5B-Instruct-Q4_K_M | 1.5B | Q4_K_M | GGUF | 312.5 | 48.3 | - | 8 GB | - | - | 0.1.0 | 2026-06-20 |
| 7 | OnePlus KB2005 | SM8250 | Qwen3-0.6B-lk-alpha-20k-MNN | 0.6B | MNN | CPU | 136.2 | 43.9 | 345 ms | 11 GB | Android 16 | best3 | 1.1.1 | 2026-08-21 |
| 8 | nubia NX733J | SM8750 | Qwen3-0.6B-Q8_0 | 0.6B | Q8_0 | GGUF | 339.4 | 42.0 | 96 ms | 23 GB | Android 15 | best3 | v3.6.0-beta130 | 2026-07-16 |
| 9 | Apple iPhone16,2 | Apple A17 Pro | llama-3.2-1b-instruct-q4_k_m.gguf | 1B | Q4_K_M | LLAMA.CPP | 479.4 | 41.3 | - | 7 GB | iOS 18.7 | warmup1_median2 | 16.163 (1830) | 2026-07-06 |
| 10 | samsung SM-G988U | SM8250 | Qwen3.5-0.8B-MNN | 0.8B | - | CPU | - | 38.3 | 342 ms | 10 GB | Android 11 | median3 | v3.6.0-beta130 | 2026-07-14 |
| 11 | nubia NX733J | SM8750 | Qwen3-0.6B.Q4_K_M | 0.6B | 0.6B.Q4_K_M | GGUF | 262.4 | 36.4 | 160 ms | 23 GB | Android 15 | best3 | v3.6.0-beta130 | 2026-07-16 |
| 12 | docomo SO-53B | SM8350 | DeepSeek-R1-Distill-Qwen-0.5B-CoMa.IQ4_XS | 0.5B | CoMa.IQ4_XS | GGUF | 27.1 | 27.1 | - | 7 GB | - | - | v3.6.0-beta70 | 2026-06-16 |
| 13 | vivo V2270A | SM7325 | Qwen3.5-0.8B-MNN | 0.8B | MNN | CPU | 74.7 | 24.4 | 655 ms | 11 GB | Android 13 | best3 | v3.6.0-beta120 | 2026-07-08 |
| 14 | REDMAGIC NX809J | SM8850 | Qwen3-4B-abliterated-MNN | 4B | MNN | CPU | 55.0 | 23.5 | 773 ms | 23 GB | Android 16 | best3 | v3.6.0-beta146 | 2026-08-13 |
| 15 | HONOR BVL-AN00 | SM8650 | LFM2.5-1.2B-Instruct-Q8_0 | 1.2B | Q8_0 | GGUF | - | 20.3 | 203 ms | 15 GB | Android 16 | median3 | v3.6.0-beta139 | 2026-07-21 |
| 16 | Lenovo TB520FU | SM8650 | Qwen3-1.7B-Q4_K_M | 1.7B | Q4_K_M | GGUF | 19.1 | 19.1 | - | 15 GB | - | - | v3.6.0-beta39 | 2026-06-10 |
| 17 | POCO 2412DPC0AG | MT6899 | Gemma-4-E2B-it-MNN | 2B | - | CPU | - | 16.1 | 572 ms | 11 GB | Android 16 | median3 | 1.0 | 2026-08-12 |
| 18 | Lenovo TB390FU | SM8735P | Qwen3.5-4B-uncensored-MNN | 4B | MNN | CPU | 43.6 | 15.3 | 832 ms | 7 GB | Android 16 | best3 | v3.6.0-beta146 | 2026-08-11 |
| 19 | POCO 2412DPC0AG | MT6899 | Gemma-4-E2B-it-MNN | 2B | MNN | CPU | 49.4 | 14.8 | 725 ms | 11 GB | Android 16 | best3 | v3.6.0-beta144 | 2026-07-31 |
| 20 | Apple iPhone 17 Pro Max | Apple A19 Pro | Qwen3-30B-A3B (MoE, top-4 experts) | 30B MoE | INT4 W4A8 | MNN CPU (KLEIDIAI W4A8) | 10.6 | 14.4 | - | 12 GB | - | - | 16.108 | 2026-07-01 |
| 21 | samsung SM-S948N | SM8850 | Qwen3.5-9B-uncensored-MNN | 9B | - | CPU | - | 13.8 | 1327 ms | 15 GB | Android 16 | median3 | v3.6.0-beta140 | 2026-07-23 |
| 22 | REDMAGIC NX809J | SM8850 | Qwen3.5-9B-uncensored-MNN | 9B | MNN | CPU | 32.0 | 13.8 | 1233 ms | 23 GB | Android 16 | best3 | 1.0 | 2026-08-14 |
| 23 | samsung SM-G990W2 | SM8350 | Qwen3.5-2B-MNN | 2B | MNN | CPU | 13.7 | 13.7 | - | 5 GB | - | - | v3.6.0-beta105 | 2026-06-27 |
| 24 | REDMAGIC NX809J | SM8850 | Qwen3.6-35B-A3B-abliterated-MNN | 35B MoE | MNN | CPU | 13.6 | 13.6 | - | 23 GB | - | - | v3.6.0-beta75 | 2026-06-16 |
| 25 | HONOR BVL-AN00 | SM8650 | Qwen3.5-4B-uncensored-MNN | 4B | - | CPU | - | 13.5 | 671 ms | 15 GB | Android 16 | median3 | v3.6.0-beta139 | 2026-07-21 |
| 26 | HONOR BVL-N49 | SM8650 | Qwen3 | - | - | CPU | - | 11.5 | 2924 ms | 11 GB | Android 16 | median3 | v3.5.0-RC20.23.69 | 2026-07-09 |
| 27 | POCO 23049PCD8G | SM7475 | Qwen3.5-4B-uncensored-MNN | 4B | - | CPU | - | 11.5 | 986 ms | 7 GB | Android 16 | median3 | v3.6.0-beta144 | 2026-08-01 |
| 28 | REDMAGIC NX799J | SM8750 | DarkIdol-Llama-3.1-8B-Instruct-1.2-Uncensored-MNN | 8B | MNN | OPENCL | 23.2 | 11.3 | 2827 ms | 11 GB | Android 16 | best3 | v3.6.0-beta137 | 2026-07-21 |
| 29 | nubia NP05J | SM8750 | Qwen3.5-9B-uncensored-MNN | 9B | MNN | CPU | 11.3 | 11.3 | - | 11 GB | - | - | v3.6.0-beta105 | 2026-06-26 |
| 30 | Xiaomi 2509FPN0BC | SM8850 | Qwen3.5-4B-uncensored-MNN | 4B | MNN | CPU | 11.1 | 11.1 | - | 15 GB | - | - | v3.6.0-beta70 | 2026-06-21 |
| 31 | google Pixel 10 Pro XL | Tensor G5 | Llama-3.2-1B-Instruct-Uncensored.i1-Q4_K_M | 1B | Q4_K_M | GGUF | 108.0 | 11.0 | 333 ms | 15 GB | Android 17 | best3 | v3.6.0-beta145 | 2026-08-06 |
| 32 | realme RMX3370 | SM8250 | LFM2.5-1.2B-Thinking-Q4_K_M | 1.2B | Q4_K_M | GGUF | 54.8 | 10.8 | 702 ms | 11 GB | Android 13 | best3 | v3.6.0-beta143 | 2026-07-30 |
| 33 | POCO 2412DPC0AG | MT6899 | Qwen3.5-4B-MNN | 4B | - | CPU | - | 10.8 | 983 ms | 11 GB | Android 16 | median3 | v3.6.0-beta105 | 2026-08-12 |
| 34 | vivo V2463A | SM8750 | Qwen3.5-4B-uncensored-MNN | 4B | MNN | CPU | 39.7 | 10.6 | 1200 ms | 15 GB | Android 16 | best3 | 1.1 | 2026-08-16 |
| 35 | google Pixel 10 | Tensor G5 | Qwen3.5-4B-uncensored-MNN | 4B | - | CPU | 8.9 | 9.8 | - | 11 GB | - | - | v3.6.0-beta70 | 2026-06-16 |
| 36 | TECNO TECNO LJ9 | MT6897 | Qwen3.5-4B-uncensored-MNN | 4B | MNN | CPU | 9.2 | 9.2 | - | 11 GB | - | - | v3.5.0-RC20.23.81 | 2026-06-23 |
| 37 | google Pixel 10 Pro XL | Tensor G5 | Qwen3.5-4B-MNN | 4B | MNN | CPU | 14.6 | 9.0 | 137 ms | 15 GB | Android 17 | best3 | v3.6.0-beta145 | 2026-08-07 |
| 38 | OPPO CPH2305 | SM8450 | Gemma-4-E2B-it-MNN | 2B | MNN | CPU | 8.5 | 8.5 | - | 11 GB | - | - | v3.6.0-beta105 | 2026-07-02 |
| 39 | samsung SM-S928W | SM8650 | Qwen3-VL-4B-Instruct-Uncensored-abliterated.Q4_K_M | 4B | abliterated.Q4_K_M | GGUF | 32.7 | 8.4 | 983 ms | 11 GB | Android 16 | best3 | v3.6.0-beta146 | 2026-08-13 |
| 40 | BLU BOLD N2 | mt6833 | gemma-4-E2B-it-qat-UD-Q4_K_XL | 2B | Q4_K_XL | GGUF | 6.7 | 8.2 | - | 7 GB | - | - | v3.6.0-beta55 | 2026-06-15 |
| 41 | HONOR BVL-AN00 | SM8650 | Qwen3-4B-abliterated-MNN | 4B | - | CPU | - | 8.1 | 2819 ms | 15 GB | Android 16 | median3 | v3.6.0-beta139 | 2026-07-21 |
| 42 | POCO 2412DPC0AG | MT6899 | microsoft_Phi-4-mini-reasoning-Q4_K_M | 3.8B | Q4_K_M | GGUF | - | 8.0 | 993 ms | 11 GB | Android 16 | median3 | v3.6.0-beta146 | 2026-08-12 |
| 43 | Redmi 25053RT47C | SM8735 | Qwen3.5-4B-uncensored-MNN | 4B | - | CPU | 6.1 | 7.7 | - | 11 GB | - | - | v3.6.0-beta55 | 2026-06-14 |
| 44 | realme RMX5060 | MT6989 | Qwen3.5-9B-uncensored-MNN | 9B | MNN | CPU | 10.2 | 7.4 | 185 ms | 11 GB | Android 16 | best3 | 1.1.1 | 2026-08-22 |
| 45 | Redmi 2407FRK8EC | MT6989 | Qwen3-8B-abliterated-v2-MNN | 8B | MNN | CPU | 7.2 | 7.2 | - | 23 GB | - | - | v3.6.0-beta39 | 2026-06-10 |
| 46 | iQOO I2407 | SM7635 | Gemma-4-E2B-it-MNN | 2B | - | MNN | 5.9 | 6.7 | - | 11 GB | - | - | v3.6.0-beta105 | 2026-06-26 |
| 47 | HONOR BVL-N49 | SM8650 | Qwen3.5-9B-uncensored-MNN | 9B | MNN | CPU | - | 6.5 | 3180 ms | 11 GB | Android 16 | median3 | v3.5.0-RC20.23.69 | 2026-07-09 |
| 48 | POCO 2412DPC0AG | MT6899 | Phi-4-mini-reasoning-Q8_0 | 3.8B | Q8_0 | GGUF | - | 6.0 | 742 ms | 11 GB | Android 16 | median3 | v3.6.0-beta146 | 2026-08-12 |
| 49 | samsung SM-A546B | s5e8835 | Qwen3.5-4B-uncensored-MNN | 4B | - | CPU | - | 5.3 | 2800 ms | 7 GB | Android 16 | median3 | v3.6.0-beta145 | 2026-08-06 |
| 50 | google Pixel 9a | Tensor G4 | Qwen3.5-4B-MNN | 4B | MNN | CPU | 5.2 | 5.2 | - | 7 GB | - | - | v3.6.0-beta70 | 2026-06-25 |
| 51 | samsung SM-G988U | SM8250 | Qwen3-4B-Instruct-2507-Q4_K_M.gguf | 4B | - | GGUF | - | 4.8 | 1566 ms | 10 GB | Android 11 | median3 | v3.6.0-beta129 | 2026-07-13 |
| 52 | samsung SM-S918B | SM8550 | Qwen3.5-9B-uncensored-MNN | 9B | MNN | CPU | 15.3 | 4.8 | 2838 ms | 11 GB | Android 16 | best3 | v3.6.0-beta130 | 2026-07-16 |
| 53 | Xiaomi 2107113SG | SM8350 | starcoder2-3b.i1-Q6_K | 3B | Q6_K | GGUF | - | 4.7 | 4484 ms | 7 GB | Android 14 | median3 | v3.6.0-beta124 | 2026-07-11 |
| 54 | OnePlus PMB110 | MT6993 | Qwen3.5-9B-uncensored-MNN | 9B | MNN | CPU | 36.8 | 4.1 | 1487 ms | 15 GB | Android 16 | best3 | v3.6.0-beta146 | 2026-08-10 |
| 55 | POCO M2102J20SG | SM8150 | mlabonne_gemma-3-4b-it-abliterated-Q4_K_M | 4B | Q4_K_M | GGUF | 11.7 | 3.8 | 3006 ms | 7 GB | Android 13 | best3 | v3.6.0-beta143 | 2026-08-17 |
| 56 | vivo vivo 2004 | SM7150 | Qwen3.5-4B-uncensored-MNN | 4B | MNN | CPU | 8.1 | 3.8 | 5138 ms | 7 GB | Android 12 | best3 | v3.6.0-beta146 | 2026-08-12 |
| 57 | samsung SM-G988U | SM8250 | Qwen3-4B-Instruct-2507-Q4_K_M | 4B | Q4_K_M | GGUF | 15.8 | 3.8 | 2185 ms | 10 GB | Android 11 | best3 | v3.6.0-beta129 | 2026-07-13 |
| 58 | google Pixel 10 Pro XL | Tensor G5 | Qwen3-6B-Hivemind-Inst-Hrtic-Ablit-Uncensored-Q4_K_M-imat | 6B | imat | GGUF | 23.2 | 3.5 | 1475 ms | 15 GB | Android 17 | best3 | v3.6.0-beta145 | 2026-08-07 |
| 59 | samsung SM-A336E | s5e8825 | Qwen3.5-2B-MNN | 2B | MNN | CPU | 2.0 | 3.2 | 434 ms | 5 GB | Android 16 | best3 | v3.6.0-beta139 | 2026-07-22 |
| 60 | vivo V2314A | MT6896Z/CZA | Qwen3.5-4B-uncensored-MNN | 4B | MNN | CPU | 13.0 | 3.2 | 3817 ms | 11 GB | Android 15 | best3 | v3.6.0-beta140 | 2026-07-23 |
| 61 | Redmi 24069RA21C | SM8635 | Qwen3.5-4B-MNN | 4B | MNN | CPU | 2.7 | 2.7 | 327 ms | 11 GB | Android 16 | best3 | v3.6.0-beta146 | 2026-08-10 |
| 62 | Xiaomi 2509FPN0BC | SM8850 | Qwen3.5-9B-MNN | 9B | - | CPU | 2.1 | 2.3 | - | 15 GB | - | - | v3.6.0-beta55 | 2026-06-16 |
| 63 | nubia NX733J | SM8750 | Qwen3.5-27B-MNN | 27B | MNN | CPU | 2.0 | 1.7 | 1045 ms | 23 GB | Android 15 | best3 | v3.6.0-beta130 | 2026-07-16 |
| 64 | Apple iPhone15,3 | Apple A16 Bionic | x | - | - | LLAMA.CPP | - | 1.0 | - | 5 GB | - | - | 16.134 (1801) | 2026-07-06 |
How these numbers were produced
Each score comes from the benchmark built into the TokForge app, run on a physical phone that the owner chose to submit from. The prompt is fixed, the reply length is capped so every run generates the same amount of work (128 tokens on Android, 96 on iPhone), and the app runs several timed passes rather than one. The Method column records how those passes were combined: best3 is the best of three passes with the prefill figure taken from that same pass, median3 is the median of three, and warmup1_median2 is one warmup pass followed by the median of two. Rows submitted by app versions that predate the tag show a dash. Submissions are signed by the app before they are sent, which keeps invented numbers off the board, and they carry no personal data.
Each row is the best submitted score for one combination of phone chip, model, backend and quantization, so a device can appear more than once with different models. TTFT, peak memory, OS version and the method tag arrived with later app versions, so a dash means the run predates that field rather than a failed measurement. Nothing here is filtered for flattery: slow rows, old builds and odd model names all stay in.
To cite it: TokForge On-Device LLM Benchmark Dataset, snapshot date as shown above, https://tokforge.ai/leaderboard/. Machine-readable JSON for the same rows is at leaderboard.tokforge.ai/v1/board.
How to read this board
Decode t/s
How fast the phone writes its reply, in tokens per second. A token is roughly three quarters of a word. 15 or more feels like comfortable reading speed, 8 to 15 is usable, under 8 feels slow.
Prefill t/s
How fast the phone reads your prompt before it starts answering. Higher means less waiting before the first word appears. A dash means that run did not measure prefill.
TTFT
Time to first token: the pause between sending a message and seeing the first word of the answer, in milliseconds. Reported by newer app versions, so a dash just means an older submission.
Size classes
Models are grouped by total parameter count: Under 2B, 2-4B (2B up to 5B), 7-9B (5B up to 10B), and 14B+ (10B and up). MoE (mixture-of-experts) models get their own tab because only a fraction of their parameters is active per token, so they behave differently from dense models of the same total size. Models we cannot size yet appear under Other.
What counts as a run
Every score comes from the TokForge app's built-in benchmark on a real phone: a fixed prompt, a fixed reply length, and several timed passes. Each row shows the best submitted score for that phone chip, model, backend, and quantization. Scores are signed by the app before they are sent, which keeps drive-by fake numbers off the board and carries no personal data.
Submissions come straight from the app's Forge/Benchmark screen (opt-in) and are anti-spam signed. Raw API: /v1/board · standalone dashboard: leaderboard.tokforge.ai. Decode tokens/sec is sustained generation speed; higher is faster.