Skip to content

Repository files navigation

TRAC: Teacher-Guided Token Reward with Adaptive Calibration for Robust Policy Optimization

ACL 2026 (Main, Long Paper)

Sitong Wu, Haoru Tan, Xichen Zhang, Bin Xia, Wenhu Zhang, Xiaojuan Qi, Bei Yu, Jiaya Jia

Paper Code Built on verl

TRAC method overview: overall framework and the three calibration coefficients

Introduction

TRAC (Teacher-guided token Reward with Adaptive Calibration) is a reinforcement learning method that turns a frozen teacher LLM into a source of dense, per-token reward for policy optimization — while explicitly guarding against the teacher's own mistakes.

Standard verifiable-reward RL (e.g. GRPO / PPO) only provides a single outcome reward per trajectory, which gives extremely sparse credit assignment over long reasoning chains. TRAC augments this with an On-Policy-Distillation (OPD) token reward: for every token y_t sampled by the student, the teacher tells us how much more (or less) likely it would have been:

$$ R^{\text{teacher}}_{t} = \log q_{\text{teacher}}(y_t \mid y_{<t}) - \log p_{\theta}(y_t \mid y_{<t}) $$

This is dense and cheap (one teacher forward pass), but a teacher is not an oracle — it is noisy on hard problems, may misrank trajectories, and can be uncertain token by token. Blindly trusting it hurts training. TRAC's key idea is Adaptive Calibration: it multiplies the raw teacher reward by a reliability weight λ ∈ [0, 1] that is estimated at three granularities,

$$ R^{\text{TRAC}}_{i,j,t} = \lambda_{i,j,t} \cdot R^{\text{teacher}}_{i,j,t}, \qquad \lambda_{i,j,t} = \lambda^{\text{expert}}_{i} \cdot \lambda^{\text{disc}}_{i,j} \cdot \lambda^{\text{conf}}_{i,j,t} $$

  • λ_expert (problem-level) — how competent the teacher is on this problem, proxied by its average entropy on the ground-truth solution. Low entropy → the teacher understands the problem → trust it more.
  • λ_disc (trajectory-level) — a binary consistency gate: does the teacher rank this trajectory correctly relative to its group (assign higher likelihood to correct than to incorrect solutions)? If not, the teacher signal for that trajectory is suppressed.
  • λ_conf (token-level) — the teacher's per-token certainty from its per-token entropy. Confident tokens keep their reward; uncertain tokens are attenuated.

Because every coefficient lies in [0, 1], calibration can only attenuate (never amplify) the teacher reward, keeping training robust. The calibrated token reward is combined with the usual group-relative outcome reward and plugged directly into GRPO / PPO.

The figure above details how each of the three coefficients is computed (subfigures (b)–(d)):

  • λ_expert (subfigure (b)) applies a linear rectification to the teacher's average entropy E* on the ground-truth solution: λ = 1 when E* ≤ ε*_min, decreasing linearly to 0 at E* ≥ ε*_max.
  • λ_disc (subfigure (c)) compares each trajectory's mean teacher log-prob P(y) against the group means of the correct (P+) and incorrect (P-) sets: a correct y keeps its reward only if P(y) > P-, an incorrect y only if P(y) < P+; otherwise the gate closes (λ = 0).
  • λ_conf (subfigure (d)) applies the same linear rectification to the teacher's per-token entropy E_t, so confident tokens keep their reward while uncertain tokens are attenuated.

Highlights

  • 🧭 Teacher-guided token rewards: dense per-token credit assignment from any off-the-shelf teacher LLM, no reward model training required.
  • 🛡️ Robust to a noisy teacher: multi-granularity calibration (λ_expert · λ_disc · λ_conf) down-weights unreliable teacher signals instead of trusting them blindly.
  • 🔬 Fully ablatable: each coefficient has its own switch, so you can enable / disable them independently.
  • 🚀 Drop-in on verl: implemented as a custom OPD advantage estimator on top of verl.

This repository implements TRAC (a.k.a. TC-OPD, Teacher-Calibrated OPD) on top of verl.

Installation

TRAC follows the standard verl setup. We recommend following the official verl installation guide.

# 1. Create the environment
conda create -n trac python=3.10 -y
conda activate trac

# 2. Clone the repository
git clone https://github.com/stonewst/TRAC.git
cd TRAC

# 3. Install the inference / training backends via the script provided by verl
#    If you need to run with Megatron:
bash scripts/install_vllm_sglang_mcore.sh
#    Or, if you simply need to run with FSDP:
USE_MEGATRON=0 bash scripts/install_vllm_sglang_mcore.sh

# 4. Install this package
pip install --no-deps -e .

Training

1. Prepare the data

The training data is a parquet file in the standard verl RL dataset format. A ready-to-use MATH-12K split is expected under data/MATH-12K/train.parquet. Each sample should also carry the ground-truth solution text under extra_info (default key solution), which is used to compute the problem-level coefficient λ_expert. You can preprocess your own dataset following the examples in examples/data_preprocess/.

2. Prepare the models

  • Policy (student) model: any HuggingFace causal LM (e.g. Qwen3-1.7B-Base).
  • Teacher model: any stronger causal LM (e.g. Qwen3-4B) used purely for inference to produce the per-token OPD reward and the calibration entropies.

3. Launch training

Edit the model / data paths and the calibration switches at the top of the example script, then run:

conda activate trac
export HYDRA_FULL_ERROR=1
ray start --head
bash examples/trac/example.sh   # an example script

Citation

If you find this work useful, please cite:

@inproceedings{wu2026trac,
  title={TRAC: Teacher-Guided Token Reward with Adaptive Calibration for Robust Policy Optimization},
  author={Wu, Sitong and Tan, Haoru and Zhang, Xichen and Xia, Bin and Zhang, Wenhu and Qi, Xiaojuan and Yu, Bei and Jia, Jiaya},
  booktitle={Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)},
  pages={47869--47884},
  year={2026}
}

Acknowledgement

This codebase is built on top of the excellent verl RL training library. We thank the verl team and community for their open-source contribution.

Contact

For questions about the paper or code, please contact Sitong Wu at stonewst@163.com, or open an issue on GitHub.

About

Official Code of ACL-2026 Paper "TRAC: Teacher-Guided Token Reward with Adaptive Calibration for Robust Policy Optimization"

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages