This implementation is specifically configured for google/gemma-3-270m-it - the instruction-tuned 270M parameter model.
- Model: google/gemma-3-270m-it
- Parameters: 270M
- Architecture: Gemma 3 (instruction-tuned)
- Context Length: Varies (check model card)
- Precision: FP32 (for CPU inference)
# 1. Setup (installs dependencies)
chmod +x setup.sh quick_run.sh
./setup.sh
# 2. Configure environment (required for gated models like Gemma)
cp .env.example .env
# Edit .env and set HF_TOKEN=your_token (get one at https://huggingface.co/settings/tokens)
# 3. Run benchmark on Gemma 3 270M
./quick_run.shThe script automatically uses google/gemma-3-270m-it. If .env exists, quick_run.sh loads it (so HF_TOKEN, OMP_NUM_THREADS, etc. are applied). See .env.example for all options.
- Smaller = Faster: 270M parameters fit entirely in CPU cache hierarchy
- Memory Efficient: ~1.1GB in FP32, easily runs on 4GB+ RAM
- Low Latency: Can achieve 10-50ms per token on modern CPUs
- Instruction-Tuned: Better quality outputs than base model
Measured on Apple Silicon / 8-core CPU:
| Stage | Optimization | Time (s) | TPS | Speedup vs baseline |
|---|---|---|---|---|
| Baseline | HF generate() | 1.90 | 26.3/s | 1.00x |
| Stage 1 | Manual loop + KV cache | 1.35 | 37.1/s | 1.41x |
| Stage 2 | C++ MatVec (single-thread) | 1.70 | 29.4/s | 1.12x |
| Stage 3 | OpenMP MatVec (8 threads) | 1.19 | 42.1/s | 1.60x |
| Stage 4 | Fused MLP (C++) | 1.18 | 42.3/s | 1.61x |
| Stage 6 | Static code path | 1.14 | 43.8/s | 1.67x |
Best: Stage 6 (Static) — ~1.67x faster, ~44 tokens/s. Time saved: ~0.76s per 50-token generation.
Target: ~1.5–1.7x speedup on 8-core CPU (varies by hardware).
The scripts support any HuggingFace causal LM that uses a compatible architecture (Gemma-style: embed_tokens, model.layers, input_layernorm / post_attention_layernorm, self_attn with q_proj/k_proj/v_proj/o_proj, mlp with gate_proj/up_proj/down_proj). Pass the model name with --model:
Inference (Stage 6, optimized):
python3 run_optimized.py --model "meta-llama/Llama-3.2-1B" "Your prompt"
python3 run_optimized.py --model "google/gemma-2-2b-it" --max-tokens 100 "Explain recursion"Full benchmark (all stages, compare baseline → Stage 6):
python3 run_all_benchmarks.py --model "google/gemma-2-2b-it" --threads 8
python3 run_all_benchmarks.py --model "meta-llama/Llama-3.2-1B" --max-tokens 32 --num-prompts 5Find best thread count for your model and CPU:
python3 find_best_params.py --model "google/gemma-2-2b-it" --runs 3Stage 1 only (manual loop + KV cache, no C++):
# Edit stage1_manual_loop.py: change default model_name, or call with a wrapper that passes model name.
python3 -c "
from stage1_manual_loop import ManualTokenLoop
g = ManualTokenLoop('google/gemma-2-2b-it')
print(g.generate('Hello,', max_new_tokens=20))
"Caveats for other models:
- Stage 6 is written for Gemma (config:
num_attention_heads,num_key_value_heads,head_dim, etc.). Other architectures (e.g. LLaMA, Qwen) may need small changes instage6_static_path.py(e.g. layer attribute names or config keys) to run. - Stages 2–4 replace Linear/MLP layers; they are architecture-agnostic for standard
nn.Linearand Gemma-style MLP. - Use
HF_TOKEN(and accept license on the model page) for gated models.
Gemma 3 270M uses:
- Smaller hidden dimension (likely 1024 or 1536 vs 2304)
- Fewer layers (likely 18-24 vs 42)
- Fewer attention heads (likely 8-16 vs 18)
- More efficient attention mechanism
This makes it ideal for CPU inference with our optimizations.
python3 run_all_benchmarks.py
# Or with 10 prompts and 8 threads:
python3 run_all_benchmarks.py --num-prompts 10 --threads 8Output includes all stages with time, TPS, and speedup:
FINAL SUMMARY
======================================================================
Stage Time (s) TPS Speedup vs Prev
----------------------------------------------------------------------
Baseline (HF) 1.902 26.3/s 1.00x 1.00x
Stage 1 (Manual) 1.348 37.1/s 1.41x 1.41x
Stage 2 (C++ ST) 1.702 29.4/s 1.12x 0.79x
Stage 3 (OpenMP) 1.188 42.1/s 1.60x 1.43x
Stage 4 (Fused MLP) 1.181 42.3/s 1.61x 1.01x
Stage 6 (Static) 1.141 43.8/s 1.67x 1.04x
======================================================================
Best optimization: Stage 6 (Static)
Throughput: 43.8 tokens/s (at 50 tokens)
python3 find_best_params.pyThis tests 1, 2, 4, 8, 16 threads and finds the sweet spot for your CPU.
# Test just Stage 1
python3 stage1_manual_loop.py
# Test just Stage 6
python3 stage6_static_path.pyUse the optimized path (Stage 6) to generate from any prompt:
python3 run_optimized.py "Your prompt here"
python3 run_optimized.py --prompt "Explain quantum computing" --max-tokens 100 --temperature 0.7
python3 run_optimized.py # default prompt: "The future of artificial intelligence is"Options: --model, --max-tokens, --temperature, --top-k. Use --no-optimization to compare against baseline HuggingFace generate(). Requires HF_TOKEN in the environment for gated models (e.g. Gemma).
./quick_run.shRuns the full benchmark with 10 diverse prompts × 3 runs per stage and saves results to benchmark_results.json.
- Model Weights: ~1.1GB (FP32)
- KV Cache: ~50-200MB (depends on context length)
- Activations: ~100-500MB
- Total: ~2GB RAM minimum, 4GB recommended
For Gemma 3 270M:
- 1-4 cores: Use all cores (Stage 3)
- 8 cores: Use 4-6 threads (diminishing returns)
- 16+ cores: Use 6-8 threads (overhead increases)
# Set performance mode
sudo cpupower frequency-set -g performance# Pin to specific cores
taskset -c 0-7 python3 run_all_benchmarks.pyecho "1" | sudo tee /sys/devices/system/cpu/intel_pstate/no_turbo| Model | Params | Baseline | Optimized | Speedup |
|---|---|---|---|---|
| Gemma 3 270M | 270M | ~5 tok/s | ~25 tok/s | 5x |
| Gemma 2 2B | 2B | ~2 tok/s | ~6 tok/s | 3x |
| Gemma 2 7B | 7B | ~0.5 tok/s | ~1.5 tok/s | 3x |
270M gets the best speedup because:
- Entire model fits in L3 cache
- Less memory bandwidth pressure
- Lower computational overhead
- Better parallelization efficiency
For serving Gemma 3 270M in production:
from stage6_static_path import StaticGemmaForCausalLM
from transformers import AutoModelForCausalLM, AutoTokenizer
# Load once
model = AutoModelForCausalLM.from_pretrained(
"google/gemma-3-270m-it",
torch_dtype=torch.float32,
trust_remote_code=True
)
static_model = StaticGemmaForCausalLM(model, max_seq_length=2048)
tokenizer = AutoTokenizer.from_pretrained(
"google/gemma-3-270m-it",
trust_remote_code=True
)
# Inference loop
def generate(prompt: str, max_tokens: int = 50):
input_ids = tokenizer.encode(prompt, return_tensors="pt")
output_ids = static_model.generate(
input_ids,
max_new_tokens=max_tokens,
temperature=0.7
)
return tokenizer.decode(output_ids, skip_special_tokens=True)
# Use it
response = generate("Explain quantum computing in simple terms:")
print(response)| Variable | Purpose | Example |
|---|---|---|
HF_TOKEN |
Hugging Face API token (gated models) | Get at huggingface.co/settings/tokens |
TOKENIZERS_PARALLELISM |
Avoid fork warnings with OpenMP | false |
OMP_NUM_THREADS |
Thread count for C++ stages | 8 (tune with find_best_params.py) |
OMP_PROC_BIND, OMP_PLACES |
Pin threads to cores | true, cores |
Copy .env.example to .env, set your values (at least HF_TOKEN for Gemma), and run. Scripts that need these will load .env when present.
Solution: Update transformers
pip install --upgrade transformersSolution: Set HF_TOKEN in .env (copy from .env.example) and accept the model terms on the model card. Verify model name:
# Should be exactly:
google/gemma-3-270m-itSolution: Check thread count
python3 find_best_params.py --runs 5Solution: You likely have <2GB RAM. Try:
- Close other applications
- Reduce max_seq_length to 1024
- Use swap space
For even better performance, combine with quantization:
from transformers import AutoModelForCausalLM, BitsAndBytesConfig
# 8-bit quantization (requires bitsandbytes)
quantization_config = BitsAndBytesConfig(load_in_8bit=True)
model = AutoModelForCausalLM.from_pretrained(
"google/gemma-3-270m-it",
quantization_config=quantization_config,
device_map="cpu",
trust_remote_code=True
)
# Apply our optimizations on top
# ... (use Stage 1 + Stage 6)This can give you:
- 2x memory reduction (550MB vs 1.1GB)
- 1.5-2x additional speedup
- Minimal quality loss
- Warm up: Run 1-2 iterations before timing
- Multiple runs: Average over 3-5 runs
- Fixed clock: Disable CPU frequency scaling
- Isolated CPU: Close background apps
- Consistent load: Same prompt length
- Run the quick benchmark:
./quick_run.sh - Find optimal threads:
python3 find_best_params.py - Integrate into your application
- Consider quantization for extra speedup
- Profile with
perforvtunefor micro-optimizations
- Model Card: https://huggingface.co/google/gemma-3-270m-it
- Issues: Check model card discussions
- Performance: Expected 3-5x speedup on 8-core CPU
Note: Gemma 3 270M is instruction-tuned. Use proper prompt formatting:
<start_of_turn>user
Your question here<end_of_turn>
<start_of_turn>model
Check the model card for exact template format.