Linear probes map where information lives in a network and predict where to spend precision when you quantize it. Probe-guided mixed precision beats uniform allocation, and a per-task probe estimates how far each task compresses. BERT, RoBERTa, DistilBERT, and a vision model, in Colab notebooks.
nlp quantization bert model-compression linear-probing mechanistic-interpretability qwen linear-probes mixed-pre
-
Updated
Jul 22, 2026 - Jupyter Notebook