pritamqu / RRPO Star 10 Code Issues Pull requests [NeurIPS 2025] Self-alignment of Large Video Language Models with Refined Regularized Preference Optimization alignment video-understanding multimodal-large-language-models large-multimodal-models self-alignment Updated Nov 2, 2025 Python
NilayRaut / Self-Alignment-with-Instruction-Backtranslation Star 1 Code Issues Pull requests LLaMA-2-7B self-alignment via instruction backtranslation. 95% labeling reduction, 0.0622% trainable params (LoRA r=8), 3 HuggingFace artifacts published. lora synthetic-data peft huggingface llm instruction-tuning llama2 self-alignment Updated Mar 29, 2026 Jupyter Notebook
fledor / self-align-to-explain Star 1 Code Issues Pull requests Self-Align to Explain: Comparing Post-Training Methods for Counterfactual Generation post-training sft dpo counterfactual-explanations grpo self-alignment simpo gdpo Updated Aug 21, 2026 Python