Posts
All the articles I've posted.
-
How OSWorld Turns Computer Work into an Agent Benchmark
· 35 min readA technical guide to OSWorld's origins, 13 real tasks, Verified, 2.0 and 2.1, scoring contracts, and the experiments needed to distinguish local progress from complete computer work.
-
How Should We Repair Reasoning Traces Before Distillation?
· 23 min readA map of 26 research papers on reasoning distillation, trajectory correction, and student compatibility: where each method acts, what it costs to implement, and an experiment to test how much of an RL teacher trace to rewrite.
-
What Makes an SFT Example Worth Learning? Quality, Novelty, and Learnability
Updated: · 23 min readFrom curriculum learning and DAgger to LIMA, LESS, and on-policy distillation: a source-backed framework for separating SFT data quality, novelty, and learnability, with executable examples and a small negative result.
-
Can RFT Learn From Failure? Negative Signals for Agent Fine-Tuning
· 16 min readA technical guide to learning from failed agent rollouts through step masking, unlikelihood, DPO, SimPO, and step-level credit assignment.
-
Improving LLM Internationalization: Bridging the Gap in Tool Use and Agency
Updated: · 18 min readA practical multilingual agent playbook, updated with task-specific evidence from MASSIVE-Agents, GAIA-v2-LILT, and SEATauBench, with explicit limits on what transfers to current frontier models.
-
From GRPO Outcome Rewards to Token-Level Advantage
Updated: · 22 min readA practical framework for turning GRPO-style sequence rewards into token-level advantages, including GAE-style estimators, credit assignment routes, and multi-reward training design.