Skip to content
#

linear-probes

Here are 23 public repositories matching this topic...

Open-source agent skills for Claude Code, Codex, GitHub Copilot, and other coding agents: GPT-5.6-style rigor, delegation, linear probes, and self-evolving workflows.

  • Updated Jul 17, 2026
  • Shell

Linear probes map where information lives in a network and predict where to spend precision when you quantize it. Probe-guided mixed precision beats uniform allocation, and a per-task probe estimates how far each task compresses. BERT, RoBERTa, DistilBERT, and a vision model, in Colab notebooks.

  • Updated Jul 22, 2026
  • Jupyter Notebook

Preregistered AI-safety study of sandbagging model organisms: trigger type sets the sign of cross-capability alignment (task-local locks dismantle it, situational locks amplify it) and cue-sharing sets its size. All five predictions failed, four reversed.

  • Updated Sep 5, 2026
  • Python

Does a language model's self-explanation actually depend on the activation it explains? Pre-registered controls for introspective verbalization, building on Li et al. (arXiv:2511.08579). Apparatus and frozen pre-registration - no measurement yet.

  • Updated Aug 15, 2026
  • Python

Maritime Intent Probe is a Phase 1 research programme on construct validity in neural probing. It introduces BC1 and uses a preregistered maritime routing counterexample to establish the identifiability requirements that motivate a crossed-design Phase 2 validation.

  • Updated Aug 4, 2026
  • Python

Add this topic to your repo

To associate your repository with the linear-probes topic, visit your repo's landing page and select "manage topics."

Learn more