Tag: ML Engineering
All the articles with the tag "ML Engineering".
-
How OSWorld Turns Computer Work into an Agent Benchmark
· 35 min readA technical guide to OSWorld's origins, 13 real tasks, Verified, 2.0 and 2.1, scoring contracts, and the experiments needed to distinguish local progress from complete computer work.
-
How Should We Repair Reasoning Traces Before Distillation?
· 23 min readA map of 26 research papers on reasoning distillation, trajectory correction, and student compatibility: where each method acts, what it costs to implement, and an experiment to test how much of an RL teacher trace to rewrite.
-
What Makes an SFT Example Worth Learning? Quality, Novelty, and Learnability
Updated: · 23 min readFrom curriculum learning and DAgger to LIMA, LESS, and on-policy distillation: a source-backed framework for separating SFT data quality, novelty, and learnability, with executable examples and a small negative result.
-
Improving LLM Internationalization: Bridging the Gap in Tool Use and Agency
Updated: · 18 min readA practical multilingual agent playbook, updated with task-specific evidence from MASSIVE-Agents, GAIA-v2-LILT, and SEATauBench, with explicit limits on what transfers to current frontier models.
-
When Does Model Souping Work for LLMs?
Updated: · 21 min readA practical guide to when LLM weight averaging and model merging work, why they fail, how methods such as Task Arithmetic, TIES, DARE, and LoRA merging differ, and how to evaluate a merge before deployment.
-
What Makes an Agent Swarm Work? Simple Rules, Information Boundaries, and the Value of Verification
· 37 min readA theory-and-experiment study of agent swarm organization: operational roles, marginal-contribution incentives, selective information sharing, and why verification remains the limiting assumption.