TokForge TokForge

Phone-Speed Leaderboard

Real on-device LLM benchmarks, submitted from the TokForge app on real phones, now including iPhone. Ranked by decode speed in tokens per second. No simulators and no cloud, just what these devices actually do.

How fast does a phone actually run a local LLM? This page answers that with measurements instead of estimates. Every row is a benchmark run submitted from the TokForge app on a real handset: how many tokens per second the phone writes (decode), how fast it reads the prompt (prefill), how long the wait is before the first word appears (time to first token), and the exact model, quantization and backend that produced the number.

If the question is which phone is fastest for local AI, whether an iPhone keeps up with an Android flagship on Qwen or Llama, what an SM8850, an SM8750, an MT6899 or an Apple A17 Pro does on a small model, or whether a 9B model stays usable on a phone at all, the table below is the evidence for the devices people have actually tested. Speeds vary with the model, the quantization, the backend and how warm the phone is, so every row carries all four rather than a single headline figure.

The board is live. The full dataset is also baked into this page as static HTML with a dated snapshot, so it can be read and quoted without running JavaScript, and it is available as JSON from the public API linked at the foot of the page.

Run this on your own phone

TokForge is free. Android is live on Google Play; iPhone and iPad are in public beta on TestFlight. No account, and it works with the network off once a model is downloaded.

Get it on Google Play iPhone beta on TestFlight

Looking for one model in particular? Qwen3.5 0.8B, 2B, 4B and 9B, and Gemma 4 E2B each have a page carrying only that model's rows. Not sure what fits? Can my phone run it.

Live 64 devices   Auto-refreshes every 60s
Loading the live board...Fetching the latest on-device benchmarks from the TokForge app.

The full benchmark dataset

Every run the public API returns, unfiltered, baked into this page as static HTML. This is the citable copy: it does not need JavaScript, it carries an explicit measurement date per row, and it is refreshed daily from the same source the live board reads. The table scrolls sideways: past the model column each row also carries parameter count, quantization, backend, prefill speed, decode speed, time to first token, RAM tier, OS build, the method the passes were combined with, the app version and the measurement date.

Snapshot: 2026-08-22 (UTC). 64 benchmark runs from 44 devices, 28 chips and 37 model builds. Oldest run measured 2026-06-10, newest 2026-08-22, submitted from TokForge 1.1.1. Data collected with the TokForge app on physical devices, no simulators and no cloud inference. The App version column carries the exact build behind every row. Refreshed daily from the live board.

#DeviceSoCModelParamsQuantBackendPrefill t/sDecode t/sTTFTRAMOSMethodApp versionMeasured (UTC)
1REDMAGIC NX809JSM8850Qwen3.5-0.8B-MNN0.8BMNNCPU274.578.3149 ms23 GBAndroid 16best31.02026-08-13
2Apple iPhone18,2Apple A19 Proqwen2.5-1.5b-instruct-q4_k_m.gguf1.5BQ4_K_MLLAMA.CPP-64.4-11 GBiOS 26.5warmup1_median216.163 (1830)2026-07-06
3motorola motorola edge 30 proSM8450Qwen3-0.6B-MNN0.6BMNNCPU195.253.7224 ms11 GBAndroid 14best31.02026-08-14
4REDMAGIC NX809JSM8850Qwen3-1.7B-MNN1.7B-CPU45.150.6-23 GB--v3.6.0-beta432026-06-11
5POCO 2412DPC0AGMT6899Qwen3-0.6B-MNN0.6BMNNCPU221.650.2175 ms11 GBAndroid 16best31.02026-08-12
6Apple iPhone16,2Apple A17 ProQwen2.5-1.5B-Instruct-Q4_K_M1.5BQ4_K_MGGUF312.548.3-8 GB--0.1.02026-06-20
7OnePlus KB2005SM8250Qwen3-0.6B-lk-alpha-20k-MNN0.6BMNNCPU136.243.9345 ms11 GBAndroid 16best31.1.12026-08-21
8nubia NX733JSM8750Qwen3-0.6B-Q8_00.6BQ8_0GGUF339.442.096 ms23 GBAndroid 15best3v3.6.0-beta1302026-07-16
9Apple iPhone16,2Apple A17 Prollama-3.2-1b-instruct-q4_k_m.gguf1BQ4_K_MLLAMA.CPP479.441.3-7 GBiOS 18.7warmup1_median216.163 (1830)2026-07-06
10samsung SM-G988USM8250Qwen3.5-0.8B-MNN0.8B-CPU-38.3342 ms10 GBAndroid 11median3v3.6.0-beta1302026-07-14
11nubia NX733JSM8750Qwen3-0.6B.Q4_K_M0.6B0.6B.Q4_K_MGGUF262.436.4160 ms23 GBAndroid 15best3v3.6.0-beta1302026-07-16
12docomo SO-53BSM8350DeepSeek-R1-Distill-Qwen-0.5B-CoMa.IQ4_XS0.5BCoMa.IQ4_XSGGUF27.127.1-7 GB--v3.6.0-beta702026-06-16
13vivo V2270ASM7325Qwen3.5-0.8B-MNN0.8BMNNCPU74.724.4655 ms11 GBAndroid 13best3v3.6.0-beta1202026-07-08
14REDMAGIC NX809JSM8850Qwen3-4B-abliterated-MNN4BMNNCPU55.023.5773 ms23 GBAndroid 16best3v3.6.0-beta1462026-08-13
15HONOR BVL-AN00SM8650LFM2.5-1.2B-Instruct-Q8_01.2BQ8_0GGUF-20.3203 ms15 GBAndroid 16median3v3.6.0-beta1392026-07-21
16Lenovo TB520FUSM8650Qwen3-1.7B-Q4_K_M1.7BQ4_K_MGGUF19.119.1-15 GB--v3.6.0-beta392026-06-10
17POCO 2412DPC0AGMT6899Gemma-4-E2B-it-MNN2B-CPU-16.1572 ms11 GBAndroid 16median31.02026-08-12
18Lenovo TB390FUSM8735PQwen3.5-4B-uncensored-MNN4BMNNCPU43.615.3832 ms7 GBAndroid 16best3v3.6.0-beta1462026-08-11
19POCO 2412DPC0AGMT6899Gemma-4-E2B-it-MNN2BMNNCPU49.414.8725 ms11 GBAndroid 16best3v3.6.0-beta1442026-07-31
20Apple iPhone 17 Pro MaxApple A19 ProQwen3-30B-A3B (MoE, top-4 experts)30B MoEINT4 W4A8MNN CPU (KLEIDIAI W4A8)10.614.4-12 GB--16.1082026-07-01
21samsung SM-S948NSM8850Qwen3.5-9B-uncensored-MNN9B-CPU-13.81327 ms15 GBAndroid 16median3v3.6.0-beta1402026-07-23
22REDMAGIC NX809JSM8850Qwen3.5-9B-uncensored-MNN9BMNNCPU32.013.81233 ms23 GBAndroid 16best31.02026-08-14
23samsung SM-G990W2SM8350Qwen3.5-2B-MNN2BMNNCPU13.713.7-5 GB--v3.6.0-beta1052026-06-27
24REDMAGIC NX809JSM8850Qwen3.6-35B-A3B-abliterated-MNN35B MoEMNNCPU13.613.6-23 GB--v3.6.0-beta752026-06-16
25HONOR BVL-AN00SM8650Qwen3.5-4B-uncensored-MNN4B-CPU-13.5671 ms15 GBAndroid 16median3v3.6.0-beta1392026-07-21
26HONOR BVL-N49SM8650Qwen3--CPU-11.52924 ms11 GBAndroid 16median3v3.5.0-RC20.23.692026-07-09
27POCO 23049PCD8GSM7475Qwen3.5-4B-uncensored-MNN4B-CPU-11.5986 ms7 GBAndroid 16median3v3.6.0-beta1442026-08-01
28REDMAGIC NX799JSM8750DarkIdol-Llama-3.1-8B-Instruct-1.2-Uncensored-MNN8BMNNOPENCL23.211.32827 ms11 GBAndroid 16best3v3.6.0-beta1372026-07-21
29nubia NP05JSM8750Qwen3.5-9B-uncensored-MNN9BMNNCPU11.311.3-11 GB--v3.6.0-beta1052026-06-26
30Xiaomi 2509FPN0BCSM8850Qwen3.5-4B-uncensored-MNN4BMNNCPU11.111.1-15 GB--v3.6.0-beta702026-06-21
31google Pixel 10 Pro XLTensor G5Llama-3.2-1B-Instruct-Uncensored.i1-Q4_K_M1BQ4_K_MGGUF108.011.0333 ms15 GBAndroid 17best3v3.6.0-beta1452026-08-06
32realme RMX3370SM8250LFM2.5-1.2B-Thinking-Q4_K_M1.2BQ4_K_MGGUF54.810.8702 ms11 GBAndroid 13best3v3.6.0-beta1432026-07-30
33POCO 2412DPC0AGMT6899Qwen3.5-4B-MNN4B-CPU-10.8983 ms11 GBAndroid 16median3v3.6.0-beta1052026-08-12
34vivo V2463ASM8750Qwen3.5-4B-uncensored-MNN4BMNNCPU39.710.61200 ms15 GBAndroid 16best31.12026-08-16
35google Pixel 10Tensor G5Qwen3.5-4B-uncensored-MNN4B-CPU8.99.8-11 GB--v3.6.0-beta702026-06-16
36TECNO TECNO LJ9MT6897Qwen3.5-4B-uncensored-MNN4BMNNCPU9.29.2-11 GB--v3.5.0-RC20.23.812026-06-23
37google Pixel 10 Pro XLTensor G5Qwen3.5-4B-MNN4BMNNCPU14.69.0137 ms15 GBAndroid 17best3v3.6.0-beta1452026-08-07
38OPPO CPH2305SM8450Gemma-4-E2B-it-MNN2BMNNCPU8.58.5-11 GB--v3.6.0-beta1052026-07-02
39samsung SM-S928WSM8650Qwen3-VL-4B-Instruct-Uncensored-abliterated.Q4_K_M4Babliterated.Q4_K_MGGUF32.78.4983 ms11 GBAndroid 16best3v3.6.0-beta1462026-08-13
40BLU BOLD N2mt6833gemma-4-E2B-it-qat-UD-Q4_K_XL2BQ4_K_XLGGUF6.78.2-7 GB--v3.6.0-beta552026-06-15
41HONOR BVL-AN00SM8650Qwen3-4B-abliterated-MNN4B-CPU-8.12819 ms15 GBAndroid 16median3v3.6.0-beta1392026-07-21
42POCO 2412DPC0AGMT6899microsoft_Phi-4-mini-reasoning-Q4_K_M3.8BQ4_K_MGGUF-8.0993 ms11 GBAndroid 16median3v3.6.0-beta1462026-08-12
43Redmi 25053RT47CSM8735Qwen3.5-4B-uncensored-MNN4B-CPU6.17.7-11 GB--v3.6.0-beta552026-06-14
44realme RMX5060MT6989Qwen3.5-9B-uncensored-MNN9BMNNCPU10.27.4185 ms11 GBAndroid 16best31.1.12026-08-22
45Redmi 2407FRK8ECMT6989Qwen3-8B-abliterated-v2-MNN8BMNNCPU7.27.2-23 GB--v3.6.0-beta392026-06-10
46iQOO I2407SM7635Gemma-4-E2B-it-MNN2B-MNN5.96.7-11 GB--v3.6.0-beta1052026-06-26
47HONOR BVL-N49SM8650Qwen3.5-9B-uncensored-MNN9BMNNCPU-6.53180 ms11 GBAndroid 16median3v3.5.0-RC20.23.692026-07-09
48POCO 2412DPC0AGMT6899Phi-4-mini-reasoning-Q8_03.8BQ8_0GGUF-6.0742 ms11 GBAndroid 16median3v3.6.0-beta1462026-08-12
49samsung SM-A546Bs5e8835Qwen3.5-4B-uncensored-MNN4B-CPU-5.32800 ms7 GBAndroid 16median3v3.6.0-beta1452026-08-06
50google Pixel 9aTensor G4Qwen3.5-4B-MNN4BMNNCPU5.25.2-7 GB--v3.6.0-beta702026-06-25
51samsung SM-G988USM8250Qwen3-4B-Instruct-2507-Q4_K_M.gguf4B-GGUF-4.81566 ms10 GBAndroid 11median3v3.6.0-beta1292026-07-13
52samsung SM-S918BSM8550Qwen3.5-9B-uncensored-MNN9BMNNCPU15.34.82838 ms11 GBAndroid 16best3v3.6.0-beta1302026-07-16
53Xiaomi 2107113SGSM8350starcoder2-3b.i1-Q6_K3BQ6_KGGUF-4.74484 ms7 GBAndroid 14median3v3.6.0-beta1242026-07-11
54OnePlus PMB110MT6993Qwen3.5-9B-uncensored-MNN9BMNNCPU36.84.11487 ms15 GBAndroid 16best3v3.6.0-beta1462026-08-10
55POCO M2102J20SGSM8150mlabonne_gemma-3-4b-it-abliterated-Q4_K_M4BQ4_K_MGGUF11.73.83006 ms7 GBAndroid 13best3v3.6.0-beta1432026-08-17
56vivo vivo 2004SM7150Qwen3.5-4B-uncensored-MNN4BMNNCPU8.13.85138 ms7 GBAndroid 12best3v3.6.0-beta1462026-08-12
57samsung SM-G988USM8250Qwen3-4B-Instruct-2507-Q4_K_M4BQ4_K_MGGUF15.83.82185 ms10 GBAndroid 11best3v3.6.0-beta1292026-07-13
58google Pixel 10 Pro XLTensor G5Qwen3-6B-Hivemind-Inst-Hrtic-Ablit-Uncensored-Q4_K_M-imat6BimatGGUF23.23.51475 ms15 GBAndroid 17best3v3.6.0-beta1452026-08-07
59samsung SM-A336Es5e8825Qwen3.5-2B-MNN2BMNNCPU2.03.2434 ms5 GBAndroid 16best3v3.6.0-beta1392026-07-22
60vivo V2314AMT6896Z/CZAQwen3.5-4B-uncensored-MNN4BMNNCPU13.03.23817 ms11 GBAndroid 15best3v3.6.0-beta1402026-07-23
61Redmi 24069RA21CSM8635Qwen3.5-4B-MNN4BMNNCPU2.72.7327 ms11 GBAndroid 16best3v3.6.0-beta1462026-08-10
62Xiaomi 2509FPN0BCSM8850Qwen3.5-9B-MNN9B-CPU2.12.3-15 GB--v3.6.0-beta552026-06-16
63nubia NX733JSM8750Qwen3.5-27B-MNN27BMNNCPU2.01.71045 ms23 GBAndroid 15best3v3.6.0-beta1302026-07-16
64Apple iPhone15,3Apple A16 Bionicx--LLAMA.CPP-1.0-5 GB--16.134 (1801)2026-07-06

How these numbers were produced

Each score comes from the benchmark built into the TokForge app, run on a physical phone that the owner chose to submit from. The prompt is fixed, the reply length is capped so every run generates the same amount of work (128 tokens on Android, 96 on iPhone), and the app runs several timed passes rather than one. The Method column records how those passes were combined: best3 is the best of three passes with the prefill figure taken from that same pass, median3 is the median of three, and warmup1_median2 is one warmup pass followed by the median of two. Rows submitted by app versions that predate the tag show a dash. Submissions are signed by the app before they are sent, which keeps invented numbers off the board, and they carry no personal data.

Each row is the best submitted score for one combination of phone chip, model, backend and quantization, so a device can appear more than once with different models. TTFT, peak memory, OS version and the method tag arrived with later app versions, so a dash means the run predates that field rather than a failed measurement. Nothing here is filtered for flattery: slow rows, old builds and odd model names all stay in.

To cite it: TokForge On-Device LLM Benchmark Dataset, snapshot date as shown above, https://tokforge.ai/leaderboard/. Machine-readable JSON for the same rows is at leaderboard.tokforge.ai/v1/board.

How to read this board

Decode t/s

How fast the phone writes its reply, in tokens per second. A token is roughly three quarters of a word. 15 or more feels like comfortable reading speed, 8 to 15 is usable, under 8 feels slow.

Prefill t/s

How fast the phone reads your prompt before it starts answering. Higher means less waiting before the first word appears. A dash means that run did not measure prefill.

TTFT

Time to first token: the pause between sending a message and seeing the first word of the answer, in milliseconds. Reported by newer app versions, so a dash just means an older submission.

Size classes

Models are grouped by total parameter count: Under 2B, 2-4B (2B up to 5B), 7-9B (5B up to 10B), and 14B+ (10B and up). MoE (mixture-of-experts) models get their own tab because only a fraction of their parameters is active per token, so they behave differently from dense models of the same total size. Models we cannot size yet appear under Other.

What counts as a run

Every score comes from the TokForge app's built-in benchmark on a real phone: a fixed prompt, a fixed reply length, and several timed passes. Each row shows the best submitted score for that phone chip, model, backend, and quantization. Scores are signed by the app before they are sent, which keeps drive-by fake numbers off the board and carries no personal data.

Get TokForge & submit your device Our reference benchmarks

Submissions come straight from the app's Forge/Benchmark screen (opt-in) and are anti-spam signed. Raw API: /v1/board · standalone dashboard: leaderboard.tokforge.ai. Decode tokens/sec is sustained generation speed; higher is faster.