Run Gemma 4 E2B on Android and iPhone

Google's Gemma 4 E2B on real Android phones: measured tokens per second, the 8 GB RAM floor, what the E in E2B means, and why the E4B build is not shipping.

2 billion effective parameters MNN 8 GB RAM floor Runs offline Free
Run this on your own phone

TokForge is free. Android is on Google Play open testing; iPhone and iPad are in open TestFlight beta. No account, and it works with the network off once a model is downloaded.

Get it on Google Play iPhone beta on TestFlight

Reading on a computer? See real measured speeds on the phone-speed leaderboard.

What Gemma 4 E2B is good at

Instruction following and a different answer style to the Qwen family, which is the main reason to keep both installed. The MNN bundle TokForge ships is an Omni export, so it carries vision and audio components alongside the text model. Note the disk footprint: at 3.6 GB it is larger on storage than Qwen3.5 4B despite having fewer effective parameters.

Builds TokForge ships

BuildFormatDownload sizeRAM floorSource repo
Gemma 4 E2BMNN3.6 GB8 GBdarkmaniac7/Gemma-4-E2B-it-MNN

Download sizes and RAM floors are the values the shipping app uses, taken from the catalogue in build v3.6.0-beta142. The RAM floor is a gate, not a suggestion: below it the model does not appear in the app at all.

What real phones record

These are submissions from the benchmark built into TokForge, not estimates and not numbers from a press release. Each row carries the handset, the chip, the app build and the date it was submitted. The table refreshes from the live board when you load this page; the version printed in the HTML is a snapshot taken on 2026-07-27.

DeviceSoCEngineDecode tok/sPrefill tok/sPeak RAMApp buildSubmitted
CPH2305SM8450MNN8.478.5-v3.6.0-beta1052026-07-02
BOLD N2mt6833-8.246.7-v3.6.0-beta552026-06-15
I2407SM7635MNN6.715.9-v3.6.0-beta1052026-06-26

Decode tokens per second is the speed you feel while an answer is being written. Prefill is how fast it reads what you sent. A dash means the submission did not carry that field. Peak RAM is what the process actually touched, and it is left blank where the reported figure is too small to be a real measurement of a model this size.

Open the full phone-speed leaderboard

Which phones can run it

TokForge offers it from 8 GB up, and keeps it in the recommended set at 12 GB. TokForge snaps your device to a RAM tier and hides anything above it, so the honest answer to "will it run on mine" is: if your phone reports 8 GB or more, the app will offer it to you. Whether it runs well is a separate question, and the table above is the place to look for a handset close to yours.

If you want the arithmetic rather than the rule of thumb, the phone RAM calculator works through weights, cache and operating system overhead for any model in the catalogue, and is explicit about how wide its error bars are.

Why there is no E4B build here

Gemma 4 ships in two sizes that share one set of weights, E2B and E4B, where the E stands for effective parameters: the larger model contains the smaller one and can drop down to it. TokForge ships E2B only. The E4B MNN exports we have tested still contain fused attention graphs that the app’s loader refuses, so rather than list a model that fails on load we leave it out of the catalogue until that is resolved. If you have seen E4B mentioned in connection with TokForge, that is the long-context work, not a shipping catalogue entry.

How to run it

  1. Install TokForge. Android is on Google Play open testing, iPhone and iPad are on TestFlight.
  2. Open the model manager. Gemma 4 E2B appears if your device clears the 8 GB floor.
  3. Download it once over Wi-Fi, then turn the network off if you like. It keeps working.

You are not limited to the catalogue. TokForge browses Hugging Face inside the app and imports GGUF and MNN files you already have, so any other build of this model is a search away.

Common questions

How much RAM does Gemma 4 E2B need on a phone?

TokForge will not offer Gemma 4 E2B to a device with less than 8 GB of RAM. That floor is the app’s own gate, not an estimate: below it the model is hidden rather than allowed to fail halfway through loading. Total RAM is not the whole story, because the operating system keeps a large share of it and a backgrounded app gets less again, so treat 8 GB as the minimum rather than the comfortable number. Handsets that have already recorded a run of it on the public board include CPH2305, BOLD N2, I2407.

How fast is Gemma 4 E2B on a phone?

That depends entirely on the handset, so we do not publish a single number for it. Every speed figure on this page is a submission from a real device through the benchmark built into the app, carrying the device, the SoC, the app build and the date. The live table above is the answer, and the full board is at /leaderboard/.

Can I run Gemma 4 E2B on an iPhone?

TokForge for iPhone and iPad is in open TestFlight beta, and it is not on the App Store yet. The model catalogue and the RAM floors are the same on both platforms. iOS is stricter about memory than Android is, so a model that is marginal on an Android phone with the same amount of RAM will be more marginal on iOS.

Does Gemma 4 E2B work offline?

Yes. The download needs a network. After that the model runs on the handset with the radios off, and nothing about the conversation leaves the device.

Related models