Phase-aware LLM activation steering and linear probing. A memory-efficient, practical implementation of Representation Engineering (RepE) for safety research.
-
Updated
Apr 1, 2026 - Python
Phase-aware LLM activation steering and linear probing. A memory-efficient, practical implementation of Representation Engineering (RepE) for safety research.
Do LLMs know when to say no? Extending Apple's ICLR 2025 instruction-following probes to agentic tool calling.
Layer-wise hidden-state probing for early detection of harmful intent in small instruction-tuned language models.
🏛️ Champollion cracked hieroglyphs in 1822. I applied the same logic to LLM internals. 95% accuracy, $0 cost, fully reproducible. Contributors welcome.
EECS E6895 final project measuring reward-gaming behavior in Gemma 2B with shell-game evals, LoRA SFT, and leakage-aware probes.
White-box detection of collusion in an untrusted monitor: model organisms, linear probes, and a control evaluation that prices what they buy.
Study 1 of the Hebbian Belief-State World Model: does a BDH core's plastic synapse state encode linearly readable beliefs? A preregistered negative result.
Open-source agent skills for Claude Code, Codex, GitHub Copilot, and other coding agents: GPT-5.6-style rigor, delegation, linear probes, and self-evolving workflows.
Inline safety probes on production LLM servers
Linear probes map where information lives in a network and predict where to spend precision when you quantize it. Probe-guided mixed precision beats uniform allocation, and a per-task probe estimates how far each task compresses. BERT, RoBERTa, DistilBERT, and a vision model, in Colab notebooks.
Preregistered AI-safety study of sandbagging model organisms: trigger type sets the sign of cross-capability alignment (task-local locks dismantle it, situational locks amplify it) and cue-sharing sets its size. All five predictions failed, four reversed.
Does a language model's self-explanation actually depend on the activation it explains? Pre-registered controls for introspective verbalization, building on Li et al. (arXiv:2511.08579). Apparatus and frozen pre-registration - no measurement yet.
Notebook-first layer-wise probes for physical variables in visual world models
Necessity tests for linear safety probes
Linear classifier probes (Alain & Bengio, 2016) — trains independent logistic regression classifiers on frozen GPT-2 representations to measure how sentiment information emerges across layers.
Maritime Intent Probe is a Phase 1 research programme on construct validity in neural probing. It introduces BC1 and uses a preregistered maritime routing counterexample to establish the identifiability requirements that motivate a crossed-design Phase 2 validation.
How fine-tuning breaks probe-based safety monitors
Research code for claim-level correctness probes on Llama activations.
Experiment testing whether a language model's expressed distress and its internal distress signal can be separated by a system prompt. On Gemma-3-12B-IT an affect-free prompt cut expressed distress by 83.5% of the natural separation while a final-token linear probe did not fall. Apart Research Digital Minds sprint.
A research tool for studying how deception emerges in multi-agent LLM systems and detecting it through activation analysis.
To associate your repository with the linear-probes topic, visit your repo's landing page and select "manage topics."