Skip to content

Repository files navigation

Gemma 3 270M CPU Optimization Guide

This implementation is specifically configured for google/gemma-3-270m-it - the instruction-tuned 270M parameter model.

Model Specifications

  • Model: google/gemma-3-270m-it
  • Parameters: 270M
  • Architecture: Gemma 3 (instruction-tuned)
  • Context Length: Varies (check model card)
  • Precision: FP32 (for CPU inference)

Quick Start

# 1. Setup (installs dependencies)
chmod +x setup.sh quick_run.sh
./setup.sh

# 2. Configure environment (required for gated models like Gemma)
cp .env.example .env
# Edit .env and set HF_TOKEN=your_token (get one at https://huggingface.co/settings/tokens)

# 3. Run benchmark on Gemma 3 270M
./quick_run.sh

The script automatically uses google/gemma-3-270m-it. If .env exists, quick_run.sh loads it (so HF_TOKEN, OMP_NUM_THREADS, etc. are applied). See .env.example for all options.

Model-Specific Optimizations

Why 270M is Great for CPU Optimization

  1. Smaller = Faster: 270M parameters fit entirely in CPU cache hierarchy
  2. Memory Efficient: ~1.1GB in FP32, easily runs on 4GB+ RAM
  3. Low Latency: Can achieve 10-50ms per token on modern CPUs
  4. Instruction-Tuned: Better quality outputs than base model

Benchmark Results (8-core, 50 tokens, Gemma 3 270M)

Measured on Apple Silicon / 8-core CPU:

Stage Optimization Time (s) TPS Speedup vs baseline
Baseline HF generate() 1.90 26.3/s 1.00x
Stage 1 Manual loop + KV cache 1.35 37.1/s 1.41x
Stage 2 C++ MatVec (single-thread) 1.70 29.4/s 1.12x
Stage 3 OpenMP MatVec (8 threads) 1.19 42.1/s 1.60x
Stage 4 Fused MLP (C++) 1.18 42.3/s 1.61x
Stage 6 Static code path 1.14 43.8/s 1.67x

Best: Stage 6 (Static) — ~1.67x faster, ~44 tokens/s. Time saved: ~0.76s per 50-token generation.

Target: ~1.5–1.7x speedup on 8-core CPU (varies by hardware).

Using other models

The scripts support any HuggingFace causal LM that uses a compatible architecture (Gemma-style: embed_tokens, model.layers, input_layernorm / post_attention_layernorm, self_attn with q_proj/k_proj/v_proj/o_proj, mlp with gate_proj/up_proj/down_proj). Pass the model name with --model:

Inference (Stage 6, optimized):

python3 run_optimized.py --model "meta-llama/Llama-3.2-1B" "Your prompt"
python3 run_optimized.py --model "google/gemma-2-2b-it" --max-tokens 100 "Explain recursion"

Full benchmark (all stages, compare baseline → Stage 6):

python3 run_all_benchmarks.py --model "google/gemma-2-2b-it" --threads 8
python3 run_all_benchmarks.py --model "meta-llama/Llama-3.2-1B" --max-tokens 32 --num-prompts 5

Find best thread count for your model and CPU:

python3 find_best_params.py --model "google/gemma-2-2b-it" --runs 3

Stage 1 only (manual loop + KV cache, no C++):

# Edit stage1_manual_loop.py: change default model_name, or call with a wrapper that passes model name.
python3 -c "
from stage1_manual_loop import ManualTokenLoop
g = ManualTokenLoop('google/gemma-2-2b-it')
print(g.generate('Hello,', max_new_tokens=20))
"

Caveats for other models:

  • Stage 6 is written for Gemma (config: num_attention_heads, num_key_value_heads, head_dim, etc.). Other architectures (e.g. LLaMA, Qwen) may need small changes in stage6_static_path.py (e.g. layer attribute names or config keys) to run.
  • Stages 2–4 replace Linear/MLP layers; they are architecture-agnostic for standard nn.Linear and Gemma-style MLP.
  • Use HF_TOKEN (and accept license on the model page) for gated models.

Architecture Differences from Gemma 2

Gemma 3 270M uses:

  • Smaller hidden dimension (likely 1024 or 1536 vs 2304)
  • Fewer layers (likely 18-24 vs 42)
  • Fewer attention heads (likely 8-16 vs 18)
  • More efficient attention mechanism

This makes it ideal for CPU inference with our optimizations.

Running Benchmarks

Full Benchmark Suite

python3 run_all_benchmarks.py
# Or with 10 prompts and 8 threads:
python3 run_all_benchmarks.py --num-prompts 10 --threads 8

Output includes all stages with time, TPS, and speedup:

FINAL SUMMARY
======================================================================
Stage                Time (s)   TPS        Speedup    vs Prev
----------------------------------------------------------------------
Baseline (HF)           1.902     26.3/s     1.00x      1.00x
Stage 1 (Manual)        1.348     37.1/s     1.41x      1.41x
Stage 2 (C++ ST)        1.702     29.4/s     1.12x      0.79x
Stage 3 (OpenMP)        1.188     42.1/s     1.60x      1.43x
Stage 4 (Fused MLP)     1.181     42.3/s     1.61x      1.01x
Stage 6 (Static)        1.141     43.8/s     1.67x      1.04x
======================================================================
Best optimization: Stage 6 (Static)
Throughput: 43.8 tokens/s (at 50 tokens)

Find Optimal Thread Count

python3 find_best_params.py

This tests 1, 2, 4, 8, 16 threads and finds the sweet spot for your CPU.

Individual Stage Testing

# Test just Stage 1
python3 stage1_manual_loop.py

# Test just Stage 6
python3 stage6_static_path.py

Run inference with a prompt (all optimizations)

Use the optimized path (Stage 6) to generate from any prompt:

python3 run_optimized.py "Your prompt here"
python3 run_optimized.py --prompt "Explain quantum computing" --max-tokens 100 --temperature 0.7
python3 run_optimized.py   # default prompt: "The future of artificial intelligence is"

Options: --model, --max-tokens, --temperature, --top-k. Use --no-optimization to compare against baseline HuggingFace generate(). Requires HF_TOKEN in the environment for gated models (e.g. Gemma).

Quick run (10 prompts, 3 runs)

./quick_run.sh

Runs the full benchmark with 10 diverse prompts × 3 runs per stage and saves results to benchmark_results.json.

Memory Requirements

  • Model Weights: ~1.1GB (FP32)
  • KV Cache: ~50-200MB (depends on context length)
  • Activations: ~100-500MB
  • Total: ~2GB RAM minimum, 4GB recommended

CPU Optimization Tips for 270M

1. Thread Count

For Gemma 3 270M:

  • 1-4 cores: Use all cores (Stage 3)
  • 8 cores: Use 4-6 threads (diminishing returns)
  • 16+ cores: Use 6-8 threads (overhead increases)

2. CPU Governor

# Set performance mode
sudo cpupower frequency-set -g performance

3. Process Affinity

# Pin to specific cores
taskset -c 0-7 python3 run_all_benchmarks.py

4. Disable Turbo Boost (for consistent timing)

echo "1" | sudo tee /sys/devices/system/cpu/intel_pstate/no_turbo

Comparing with Other Models

Model Params Baseline Optimized Speedup
Gemma 3 270M 270M ~5 tok/s ~25 tok/s 5x
Gemma 2 2B 2B ~2 tok/s ~6 tok/s 3x
Gemma 2 7B 7B ~0.5 tok/s ~1.5 tok/s 3x

270M gets the best speedup because:

  • Entire model fits in L3 cache
  • Less memory bandwidth pressure
  • Lower computational overhead
  • Better parallelization efficiency

Production Deployment

For serving Gemma 3 270M in production:

from stage6_static_path import StaticGemmaForCausalLM
from transformers import AutoModelForCausalLM, AutoTokenizer

# Load once
model = AutoModelForCausalLM.from_pretrained(
    "google/gemma-3-270m-it",
    torch_dtype=torch.float32,
    trust_remote_code=True
)
static_model = StaticGemmaForCausalLM(model, max_seq_length=2048)
tokenizer = AutoTokenizer.from_pretrained(
    "google/gemma-3-270m-it",
    trust_remote_code=True
)

# Inference loop
def generate(prompt: str, max_tokens: int = 50):
    input_ids = tokenizer.encode(prompt, return_tensors="pt")
    output_ids = static_model.generate(
        input_ids, 
        max_new_tokens=max_tokens,
        temperature=0.7
    )
    return tokenizer.decode(output_ids, skip_special_tokens=True)

# Use it
response = generate("Explain quantum computing in simple terms:")
print(response)

Environment variables

Variable Purpose Example
HF_TOKEN Hugging Face API token (gated models) Get at huggingface.co/settings/tokens
TOKENIZERS_PARALLELISM Avoid fork warnings with OpenMP false
OMP_NUM_THREADS Thread count for C++ stages 8 (tune with find_best_params.py)
OMP_PROC_BIND, OMP_PLACES Pin threads to cores true, cores

Copy .env.example to .env, set your values (at least HF_TOKEN for Gemma), and run. Scripts that need these will load .env when present.

Troubleshooting

Issue: "trust_remote_code" error

Solution: Update transformers

pip install --upgrade transformers

Issue: Model not found / gated model / login required

Solution: Set HF_TOKEN in .env (copy from .env.example) and accept the model terms on the model card. Verify model name:

# Should be exactly:
google/gemma-3-270m-it

Issue: Slow on Stage 3

Solution: Check thread count

python3 find_best_params.py --runs 5

Issue: Out of memory

Solution: You likely have <2GB RAM. Try:

  • Close other applications
  • Reduce max_seq_length to 1024
  • Use swap space

Advanced: Quantization

For even better performance, combine with quantization:

from transformers import AutoModelForCausalLM, BitsAndBytesConfig

# 8-bit quantization (requires bitsandbytes)
quantization_config = BitsAndBytesConfig(load_in_8bit=True)

model = AutoModelForCausalLM.from_pretrained(
    "google/gemma-3-270m-it",
    quantization_config=quantization_config,
    device_map="cpu",
    trust_remote_code=True
)

# Apply our optimizations on top
# ... (use Stage 1 + Stage 6)

This can give you:

  • 2x memory reduction (550MB vs 1.1GB)
  • 1.5-2x additional speedup
  • Minimal quality loss

Benchmarking Tips

  1. Warm up: Run 1-2 iterations before timing
  2. Multiple runs: Average over 3-5 runs
  3. Fixed clock: Disable CPU frequency scaling
  4. Isolated CPU: Close background apps
  5. Consistent load: Same prompt length

Next Steps

  1. Run the quick benchmark: ./quick_run.sh
  2. Find optimal threads: python3 find_best_params.py
  3. Integrate into your application
  4. Consider quantization for extra speedup
  5. Profile with perf or vtune for micro-optimizations

Support


Note: Gemma 3 270M is instruction-tuned. Use proper prompt formatting:

<start_of_turn>user
Your question here<end_of_turn>
<start_of_turn>model

Check the model card for exact template format.

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages