ACL 2026 (Main, Long Paper)
Sitong Wu, Haoru Tan, Xichen Zhang, Bin Xia, Wenhu Zhang, Xiaojuan Qi, Bei Yu, Jiaya Jia
TRAC (Teacher-guided token Reward with Adaptive Calibration) is a reinforcement learning method that turns a frozen teacher LLM into a source of dense, per-token reward for policy optimization — while explicitly guarding against the teacher's own mistakes.
Standard verifiable-reward RL (e.g. GRPO / PPO) only provides a single outcome reward per
trajectory, which gives extremely sparse credit assignment over long reasoning chains. TRAC
augments this with an On-Policy-Distillation (OPD) token reward: for every token y_t sampled
by the student, the teacher tells us how much more (or less) likely it would have been:
This is dense and cheap (one teacher forward pass), but a teacher is not an oracle — it is
noisy on hard problems, may misrank trajectories, and can be uncertain token by token. Blindly
trusting it hurts training. TRAC's key idea is Adaptive Calibration: it multiplies the raw
teacher reward by a reliability weight λ ∈ [0, 1] that is estimated at three granularities,
λ_expert(problem-level) — how competent the teacher is on this problem, proxied by its average entropy on the ground-truth solution. Low entropy → the teacher understands the problem → trust it more.λ_disc(trajectory-level) — a binary consistency gate: does the teacher rank this trajectory correctly relative to its group (assign higher likelihood to correct than to incorrect solutions)? If not, the teacher signal for that trajectory is suppressed.λ_conf(token-level) — the teacher's per-token certainty from its per-token entropy. Confident tokens keep their reward; uncertain tokens are attenuated.
Because every coefficient lies in [0, 1], calibration can only attenuate (never amplify)
the teacher reward, keeping training robust. The calibrated token reward is combined with the
usual group-relative outcome reward and plugged directly into GRPO / PPO.
The figure above details how each of the three coefficients is computed (subfigures (b)–(d)):
λ_expert(subfigure (b)) applies a linear rectification to the teacher's average entropyE*on the ground-truth solution:λ = 1whenE* ≤ ε*_min, decreasing linearly to0atE* ≥ ε*_max.λ_disc(subfigure (c)) compares each trajectory's mean teacher log-probP(y)against the group means of the correct (P+) and incorrect (P-) sets: a correctykeeps its reward only ifP(y) > P-, an incorrectyonly ifP(y) < P+; otherwise the gate closes (λ = 0).λ_conf(subfigure (d)) applies the same linear rectification to the teacher's per-token entropyE_t, so confident tokens keep their reward while uncertain tokens are attenuated.
Highlights
- 🧭 Teacher-guided token rewards: dense per-token credit assignment from any off-the-shelf teacher LLM, no reward model training required.
- 🛡️ Robust to a noisy teacher: multi-granularity calibration (
λ_expert · λ_disc · λ_conf) down-weights unreliable teacher signals instead of trusting them blindly. - 🔬 Fully ablatable: each coefficient has its own switch, so you can enable / disable them independently.
- 🚀 Drop-in on verl: implemented as a custom OPD advantage estimator on top of verl.
This repository implements TRAC (a.k.a. TC-OPD, Teacher-Calibrated OPD) on top of verl.
TRAC follows the standard verl setup. We recommend following the official verl installation guide.
# 1. Create the environment
conda create -n trac python=3.10 -y
conda activate trac
# 2. Clone the repository
git clone https://github.com/stonewst/TRAC.git
cd TRAC
# 3. Install the inference / training backends via the script provided by verl
# If you need to run with Megatron:
bash scripts/install_vllm_sglang_mcore.sh
# Or, if you simply need to run with FSDP:
USE_MEGATRON=0 bash scripts/install_vllm_sglang_mcore.sh
# 4. Install this package
pip install --no-deps -e .The training data is a parquet file in the standard verl RL dataset format. A ready-to-use
MATH-12K split is expected under data/MATH-12K/train.parquet. Each sample should also carry
the ground-truth solution text under extra_info (default key solution), which is used to
compute the problem-level coefficient λ_expert. You can preprocess your own dataset following
the examples in examples/data_preprocess/.
- Policy (student) model: any HuggingFace causal LM (e.g.
Qwen3-1.7B-Base). - Teacher model: any stronger causal LM (e.g.
Qwen3-4B) used purely for inference to produce the per-token OPD reward and the calibration entropies.
Edit the model / data paths and the calibration switches at the top of the example script, then run:
conda activate trac
export HYDRA_FULL_ERROR=1
ray start --head
bash examples/trac/example.sh # an example scriptIf you find this work useful, please cite:
@inproceedings{wu2026trac,
title={TRAC: Teacher-Guided Token Reward with Adaptive Calibration for Robust Policy Optimization},
author={Wu, Sitong and Tan, Haoru and Zhang, Xichen and Xia, Bin and Zhang, Wenhu and Qi, Xiaojuan and Yu, Bei and Jia, Jiaya},
booktitle={Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)},
pages={47869--47884},
year={2026}
}This codebase is built on top of the excellent verl RL training library. We thank the verl team and community for their open-source contribution.
For questions about the paper or code, please contact Sitong Wu at stonewst@163.com, or open an issue on GitHub.
