LiteRT-LM

litert-community/gemma-4-12B-it-litert-lm

Main Model Card: google/gemma-4-12B-it

This model card provides the Gemma 4 12B model in LiteRT-LM format ready for deployment on macOS, Linux and Windows, as well as web in a more limited capacity. The current gemma-4-12B-it.litertlm file supports text, vision and audio modalities as well as Multi-Token Prediction (MTP) for accelerated speculative decoding and lower latency inference.

Requirement: Running this model requires LiteRT-LM v0.17 or later.

Gemma is a family of lightweight, state-of-the-art open models from Google, built from the same research and technology used to create the Gemini models. This particular Gemma 4 model is medium sized, so it is ideal for desktop use cases. By running this model on device, users can have private access to Generative AI without requiring an internet connection.

These models are provided in the .litertlm format for use with the LiteRT-LM framework. LiteRT-LM is a specialized orchestration layer built directly on top of LiteRT, Google’s high-performance multi-platform runtime trusted by millions of Android and edge developers. LiteRT provides the foundational hardware acceleration via XNNPack for CPU and ML Drift for GPU. LiteRT-LM adds the specialized GenAI libraries and APIs, such as KV-cache management, prompt templating, and function calling. This integrated stack is the same technology powering the Google AI Edge Gallery showcase app.

Try Gemma 4 12B

Build with Gemma 4 12B and LiteRT-LM

Ready to integrate this into your product? Get started with LiteRT-LM documentation.

Gemma 4 12B Performance on LiteRT-LM

All benchmarks were taken using 1024 prefill tokens and 256 decode tokens with a context length of 4096 tokens via LiteRT-LM. The model can support up to 128k context length on devices with sufficent memory (please set the context length, max_num_tokens, to be smaller when encountering memory issue). Time-to-first-token does not include load time. Benchmarks were run with caches enabled and initialized. During the first run, the latency and memory usage may differ. Model size is the size of the file on disk.

Linux

Device                                      Backend Prefill (tokens/sec) Decode (tokens/sec) Time-to-first-token (sec) Model size (MB) GPU Memory (MB)
NVidia 4090 24GB GPU 3548 69 0.3 6883 ~7790

macOS

Device                                      Backend Prefill (tokens/sec) Decode (tokens/sec) Time-to-first-token (sec) Model size (MB) GPU Memory (MB)
Macbook M4 Pro 48GB GPU 297 29 3.48 6883 ~7870
Macbook Air M4 16GB GPU 114 15 9.13 6883 ~7900

Windows

Device                                      Backend Prefill (tokens/sec) Decode (tokens/sec) Time-to-first-token (sec) Model size (MB) GPU Memory (MB)
NVidia 5080 16GB GPU 391 50 2.5 6883 ~7300

Web

Device                                      Backend Prefill (tokens/sec) Decode (tokens/sec) Time-to-first-token (sec) Model size (MB) GPU Memory (MB) Peak CPU Memory (MB)
MacBook Pro M4 (M4 Max) GPU 388 26 3.5 5986 ~7700 ~1200
  • Web on LiteRT-LM uses a specially optimized model for Web because of its unique memory constraints. Currently the model is text-only.
  • Benchmarks taken in Chrome using a context length of 1280. The web model can support up to 195k context length on devices with sufficent memory.
Downloads last month
13,428
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 1 Ask for provider support

Model tree for litert-community/gemma-4-12B-it-litert-lm

Finetuned
(168)
this model

Spaces using litert-community/gemma-4-12B-it-litert-lm 3

Collections including litert-community/gemma-4-12B-it-litert-lm