Tag: Post-Training
All the articles with the tag "Post-Training".
-
Can RFT Learn From Failure? Negative Signals for Agent Fine-Tuning
· 16 min readA technical guide to learning from failed agent rollouts through step masking, unlikelihood, DPO, SimPO, and step-level credit assignment.
-
From GRPO Outcome Rewards to Token-Level Advantage
Updated: · 22 min readA practical framework for turning GRPO-style sequence rewards into token-level advantages, including GAE-style estimators, credit assignment routes, and multi-reward training design.
-
When Does Model Souping Work for LLMs?
Updated: · 21 min readA practical guide to when LLM weight averaging and model merging work, why they fail, how methods such as Task Arithmetic, TIES, DARE, and LoRA merging differ, and how to evaluate a merge before deployment.
-
Do Verifier Errors Grow Superlinearly with Horizon? A Six-Stage Experiment
Updated: · 17 min readA preregistered 224-artifact study confirms a 66% hybrid-verifier improvement, does not support the universal superlinear headline, and finds that task structure can reverse the apparent horizon effect.
-
From Tool Calls to Policy Updates: A Reproducible Agent RL Stack
· 45 min readA reproducible blueprint for Agent RL: environment snapshots, verifier contracts, credit assignment, three-policy semantics, asynchronous generation and training, partial rollouts, and runnable smoke tests.
-
A Repair Can Save a Trajectory Without Teaching the Agent
· 18 min readAn eight-seed counterfactual-replay experiment found that high-value one-action repairs reliably rescue failed agent trajectories—but a repair-value-dominant SFT selector did not beat simpler rules and recovered only 25% of matched-RL improvement.