Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
43 commits
Select commit Hold shift + click to select a range
ddf5e10
chore(deps): add transformers, torch, datasets for inference
xzAscC Mar 20, 2026
9124e03
test(inference): add TDD tests for inference module
xzAscC Mar 20, 2026
50db867
feat(inference): implement MATH500 inference with Qwen3.5-0.8B
xzAscC Mar 20, 2026
6769b10
feat(inference): implement MATH500 inference with Qwen3.5-0.8B
xzAscC Mar 20, 2026
c699d23
feat(inference): implement MATH500 inference with Qwen3.5-0.8B
xzAscC Mar 20, 2026
0e95c2b
fix(inference): replace Any type with proper types, format tests
xzAscC Mar 20, 2026
2f83496
feat(eval): add LLM-as-Judge evaluation module
xzAscC Mar 21, 2026
e023022
feat: add reflection token diagnosis tool
xzAscC Mar 21, 2026
3f31e4f
feat(types): add ExtractVectorsConfig and SteeringVectorResult for st…
xzAscC Mar 21, 2026
6228103
feat(steering): implement classify_samples, find_reflection_token_pos…
xzAscC Mar 21, 2026
f9b5bd7
feat(steering): implement extract_batch_activations with batch proces…
xzAscC Mar 21, 2026
d39b410
feat(steering): implement classify_samples, find_reflection_token_pos…
xzAscC Mar 21, 2026
502b26e
feat(steering): implement extract_steering_vectors pipeline and extra…
xzAscC Mar 21, 2026
512b774
feat(steering): add complete steering inference infrastructure
xzAscC Mar 21, 2026
2ff6637
fix(steering): dtype mismatch and dataset split handling
xzAscC Mar 21, 2026
e4a0432
feat(linear-probe): implement linear probe module for reflection toke…
xzAscC Mar 22, 2026
2529c2e
feat(linear-probe): add run_linear_probe pipeline and update exports
xzAscC Mar 22, 2026
db27028
chore: configure local research file ignores
xzAscC Jul 16, 2026
280ee8b
chore: exclude research notebooks from Ruff
xzAscC Jul 16, 2026
ef01088
feat(utils): add shared model loading helpers
xzAscC Jul 16, 2026
437239d
refactor(inference): share batch processing helpers
xzAscC Jul 16, 2026
8923624
feat(prompts): centralize evaluation prompt templates
xzAscC Jul 16, 2026
f9b9ca7
refactor(linear-probe): reuse shared model utilities
xzAscC Jul 16, 2026
adb56ab
refactor(steering): reuse shared model loading
xzAscC Jul 16, 2026
9eafb76
refactor(judges): consolidate LLM evaluation classes
xzAscC Jul 16, 2026
ac7486a
feat(roscoe): add reasoning quality diagnosis
xzAscC Jul 16, 2026
ccd7469
test(datasets): mock Hugging Face loading
xzAscC Jul 16, 2026
a3d936c
style(cli): simplify command help text
xzAscC Jul 16, 2026
f17b6ec
feat(scripts): add cross-family judge tooling
xzAscC Jul 16, 2026
c75118e
chore(scripts): add experiment runner wrappers
xzAscC Jul 16, 2026
f970b82
build(deps): add Jupyter runtime
xzAscC Jul 16, 2026
4b73761
docs(readme): update paper citation
xzAscC Jul 16, 2026
a3cbafe
build(types): follow untyped dependency imports
xzAscC Jul 16, 2026
85b382e
refactor(models): type model lifecycle boundaries
xzAscC Jul 16, 2026
405c7c3
fix(inference): enforce limits and decode generated tokens
xzAscC Jul 16, 2026
1984cf5
fix(judges): decode only newly generated tokens
xzAscC Jul 16, 2026
751a94a
fix(steering): type hooks and decode generated tokens
xzAscC Jul 16, 2026
849c961
fix(steering-vectors): avoid duplicate OOM activations
xzAscC Jul 16, 2026
2f75e3d
refactor(datasets): validate untyped dataset rows
xzAscC Jul 16, 2026
adac210
refactor(evaluation): normalize result metadata types
xzAscC Jul 16, 2026
6bb9a3d
refactor(linear-probe): type sklearn boundaries
xzAscC Jul 16, 2026
998a67e
fix(scripts): prevent wrapper argument injection
xzAscC Jul 16, 2026
37b4170
docs(readme): document current research pipelines
xzAscC Jul 16, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
21 changes: 21 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -27,8 +27,29 @@ build/
htmlcov/
.pytest_cache/
.mypy_cache/
.ruff_cache/

# Jupyter
.ipynb_checkpoints/

# Environment
.env
.env.local

# Outputs
outputs/

# Local research data and generated artifacts
data/
figures/
src/probing_reflection/*.csv
src/probing_reflection/*.ipynb
src/probing_reflection/*.pdf
src/probing_reflection/*.png

# Local agent workflow state
.omo/
.sisyphus/

# Local TMLR repository
67d0fc7176c97e12c92d4271/
17 changes: 17 additions & 0 deletions .ignore
Original file line number Diff line number Diff line change
@@ -0,0 +1,17 @@
# Keep Git-ignored local research material searchable by OpenCode/ripgrep.
!.omo/
!.omo/**
!.sisyphus/
!.sisyphus/**
!67d0fc7176c97e12c92d4271/
!67d0fc7176c97e12c92d4271/**
!data/
!data/**
!figures/
!figures/**
!src/probing_reflection/.ipynb_checkpoints/
!src/probing_reflection/.ipynb_checkpoints/**
!src/probing_reflection/*.csv
!src/probing_reflection/*.ipynb
!src/probing_reflection/*.pdf
!src/probing_reflection/*.png
22 changes: 0 additions & 22 deletions .sisyphus/drafts/token-probe-experiment.md

This file was deleted.

151 changes: 115 additions & 36 deletions README.md
Original file line number Diff line number Diff line change
@@ -1,70 +1,149 @@
# ProbingReflection

Source code for the paper **"From Emergence to Control: Probing and Modulating Self-Reflection in Language Models"**
Official code for **[From Emergence to Control: Probing and Modulating
Self-Reflection in Language Models](https://arxiv.org/abs/2506.12217)**.

This repository studies where self-reflection appears in language-model
representations and how it can be measured and controlled. It includes
reproducible pipelines for inference, LLM-based evaluation, reflection-token
diagnosis, activation probing, steering-vector extraction, and steering
interventions.

## Features

- Run chain-of-thought inference on mathematical reasoning datasets.
- Evaluate generated answers with a bidirectional LLM-as-a-judge protocol.
- Detect and categorize self-reflection tokens in model outputs.
- Measure reasoning quality with ROSCOE-style faithfulness, coherence,
informativeness, repetition, and completeness scores.
- Train layer-wise linear probes to test whether reflective and
non-reflective token activations are linearly separable.
- Extract difference-in-means steering vectors from contrastive activations.
- Apply steering vectors during generation with configurable layers and
coefficients.

## Overview
## Installation

This project investigates self-reflection in Large Language Models (LLMs) through probing and steering techniques. We explore how reflection behaviors emerge and how they can be controlled through vector-based interventions.
The project requires Python 3.12 or newer and uses
[uv](https://docs.astral.sh/uv/) for dependency management.

## Key Concepts
```bash
git clone https://github.com/xzAscC/ProbingReflection.git
cd ProbingReflection
uv sync
```

- **Probing Vectors**: Techniques to detect and measure self-reflection patterns in model activations
- **Model Insertion**: Methods for injecting steering vectors to modulate reflection behavior
- **Reflection Analysis**: Frameworks for evaluating and understanding model self-reflection
Model-backed experiments require access to the configured Hugging Face models
and sufficient CPU/GPU memory. Some datasets, including GPQA, may require
accepting their access conditions on Hugging Face.

## Installation
## Command-Line Usage

List the available commands:

```bash
uv sync
uv run probing-reflection --help
```

## Development
### Inference

```bash
# Lint check
uv run ruff check src/ tests/
uv run probing-reflection inference \
--model Qwen/Qwen3.5-0.8B \
--dataset HuggingFaceH4/MATH-500 \
--batch-size 8 \
--max-new-tokens 256 \
--limit 100 \
--output outputs/math500.jsonl
```

# Format
uv run ruff format src/ tests/
### Answer evaluation

# Type check
uv run mypy src/
```bash
uv run probing-reflection evaluate outputs/math500.jsonl \
--model Qwen/Qwen3.5-27B \
--confidence-threshold 0.7 \
--output outputs/evaluation.json
```

# Run tests
uv run pytest
### Reflection diagnosis

```bash
uv run probing-reflection reflection-diagnose \
--input outputs/math500.jsonl \
--model Qwen/Qwen3.5-27B \
--output-dir outputs/reflection_diagnosis
```

## Project Structure
### Steering-vector extraction

```bash
uv run probing-reflection extract-vectors \
--input outputs/reflection_diagnosis/analyzed_samples.jsonl \
--model Qwen/Qwen2.5-0.5B \
--layers 8,12,16 \
--min-samples 10 \
--output outputs/steering_vectors.pt
```
.
├── src/probing_reflection/ # Source code
│ ├── __init__.py
│ └── py.typed
├── tests/ # Test files
├── docs/ # Documentation
│ └── design-docs/ # Design documents
├── AGENTS.md # AI agent instructions
├── ARCHITECTURE.md # System architecture
└── pyproject.toml # Project configuration

Additional wrappers for evaluation, diagnosis, linear probing, vector
extraction, and steering inference are available under
`scripts/probing_reflection/`. Use `scripts/run_experiments.py` to coordinate
multi-dataset steering experiments and `scripts/generate_report.py` to create
summary reports.

## Project Layout

```text
src/probing_reflection/
├── inference.py # Dataset inference pipeline
├── evaluation.py # Answer evaluation and reports
├── reflection_diagnosis.py # Reflection-token analysis
├── linear_probe.py # Layer-wise linear probing
├── steering_vectors.py # Steering-vector extraction
├── steering_inference.py # Steered generation
├── judges.py # Shared LLM judge implementations
├── roscoe_metrics.py # ROSCOE-style reasoning metrics
├── model_utils.py # Model loading and lifecycle helpers
├── batch_utils.py # Shared batching and decoding helpers
├── prompts.py # Prompt templates and reflection taxonomy
└── types.py # Typed configurations and result schemas

tests/ # Deterministic unit and integration tests
scripts/ # Experiment and reporting entry points
docs/ # Architecture and design documentation
```

## Research Workflow
Generated outputs, local datasets, figures, notebooks, and experiment state are
excluded from Git. The repository-level `.ignore` file keeps these local paths
searchable by OpenCode and ripgrep without uploading them to GitHub.

## Development

```bash
uv sync
uv run ruff check src/ tests/ scripts/
uv run ruff format --check src/ tests/ scripts/
uv run mypy src/
uv run pytest
```

This project follows an AI-assisted research workflow. See AGENTS.md for detailed instructions on how AI agents should work in this repository.
The current test suite contains 159 deterministic tests and mocks external
model and dataset boundaries.

## Citation

If you use this code, please cite:

```bibtex
@article{probing_reflection_2024,
title={From Emergence to Control: Probing and Modulating Self-Reflection in Language Models},
author={[Authors]},
year={2024}
@article{zhu2025emergence,
title={From emergence to control: Probing and modulating self-reflection in language models},
author={Zhu, Xudong and Jiang, Jiachen and Khalili, Mohammad Mahdi and Zhu, Zhihui},
journal={arXiv preprint arXiv:2506.12217},
year={2025}
}
```

## License

[Add your license here]
This project is released under the MIT License. See [LICENSE](LICENSE).
17 changes: 17 additions & 0 deletions pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -7,19 +7,35 @@ name = "probing_reflection"
version = "0.1.0"
requires-python = ">=3.12"
description = "ProbingReflection - Probing and Modulating Self-Reflection in Language Models"
dependencies = [
"jupyter>=1.1.1",
]

[project.scripts]
probing-reflection = "probing_reflection.__main__:main"

[dependency-groups]
dev = [
"ruff>=0.8.0",
"mypy>=1.13.0",
"pytest>=8.0.0",
"pyyaml>=6.0.3",
"transformers>=4.40.0",
"torch>=2.0.0",
"datasets>=2.14.0",
"tqdm>=4.65.0",
"types-tqdm>=4.65.0",
"accelerate>=0.26.0",
"bitsandbytes>=0.43.0",
"scikit-learn>=1.3.0",
"matplotlib>=3.7.0",
]

[tool.ruff]
target-version = "py312"
line-length = 100
src = ["src"]
extend-exclude = ["src/probing_reflection/**/*.ipynb"]

[tool.ruff.lint]
select = ["E", "F", "I", "UP", "B", "SIM", "N"]
Expand All @@ -32,6 +48,7 @@ python_version = "3.12"
strict = true
warn_return_any = true
warn_unused_ignores = true
follow_untyped_imports = true
packages = ["probing_reflection"]
mypy_path = "src"

Expand Down
Loading
Loading