Learning at the Right Pace: Adaptive Data Scheduling improves LLM reinforcement learning with semantic-cluster and policy-boundary sampling.
-
Updated
Jul 5, 2026 - Python
Learning at the Right Pace: Adaptive Data Scheduling improves LLM reinforcement learning with semantic-cluster and policy-boundary sampling.
Your RL second brain: 34 source-cited topics from Q-learning to GRPO and agentic RL, kept fresh monthly. Obsidian vault + AI agent layer, Brainstein SSS+.
A novel RL algorithm that designs hierarchical token-level optimization objective for exploration-exploitation balanced policy optimization
Paired-intervention toolkit for testing when to switch from on-policy distillation to GRPO.
Controlled LLM post-training study of Agent robustness under reduced skill prompts.
金融工单 Agent 的数据工程与后训练实验:数据/意图/工具/政策四张注册表 → SFT / DPO / tool-use SFT 导出 → Qwen2.5-3B LoRA;held-out 政策评测 88/91,加推理时 guard 后 91/91。
按优化机制与反馈闭环导航 2017—2026 年关键方法、证据和失效边界|56 项记录
Machine-verifiable evidence chain for online LLM post-training (TRL GRPO + vLLM): trajectory contracts, fault injection F1-F8, guarded updates, paired gradient replay
To associate your repository with the llm-post-training topic, visit your repo's landing page and select "manage topics."