Tag: Evaluation
All the articles with the tag "Evaluation".
-
How OSWorld Turns Computer Work into an Agent Benchmark
· 35 min readA technical guide to OSWorld's origins, 13 real tasks, Verified, 2.0 and 2.1, scoring contracts, and the experiments needed to distinguish local progress from complete computer work.
-
How Should We Repair Reasoning Traces Before Distillation?
· 23 min readA map of 26 research papers on reasoning distillation, trajectory correction, and student compatibility: where each method acts, what it costs to implement, and an experiment to test how much of an RL teacher trace to rewrite.
-
What Makes an SFT Example Worth Learning? Quality, Novelty, and Learnability
Updated: · 23 min readFrom curriculum learning and DAgger to LIMA, LESS, and on-policy distillation: a source-backed framework for separating SFT data quality, novelty, and learnability, with executable examples and a small negative result.
-
Can RFT Learn From Failure? Negative Signals for Agent Fine-Tuning
· 16 min readA technical guide to learning from failed agent rollouts through step masking, unlikelihood, DPO, SimPO, and step-level credit assignment.
-
When Does Model Souping Work for LLMs?
Updated: · 21 min readA practical guide to when LLM weight averaging and model merging work, why they fail, how methods such as Task Arithmetic, TIES, DARE, and LoRA merging differ, and how to evaluate a merge before deployment.
-
Do LLM Agents Work Equally Well Across Languages?
· 11 min readAgent performance differs across languages, tasks, and localization settings. A visual review of 30 languages and 16 evidence sources identifies where Arabic, Thai, Japanese, and Hindi need targeted evaluation.