Tag: Evaluation
All the articles with the tag "Evaluation".
-
When Does Model Souping Work for LLMs?
Updated: · 21 min readA practical guide to when LLM weight averaging and model merging work, why they fail, how methods such as Task Arithmetic, TIES, DARE, and LoRA merging differ, and how to evaluate a merge before deployment.
-
Do LLM Agents Work Equally Well Across Languages?
· 11 min readAgent performance differs across languages, tasks, and localization settings. A visual review of 30 languages and 16 evidence sources identifies where Arabic, Thai, Japanese, and Hindi need targeted evaluation.
-
What Makes an Agent Swarm Work? Simple Rules, Information Boundaries, and the Value of Verification
· 37 min readA theory-and-experiment study of agent swarm organization: operational roles, marginal-contribution incentives, selective information sharing, and why verification remains the limiting assumption.
-
Do Verifier Errors Grow Superlinearly with Horizon? A Six-Stage Experiment
Updated: · 17 min readA preregistered 224-artifact study confirms a 66% hybrid-verifier improvement, does not support the universal superlinear headline, and finds that task structure can reverse the apparent horizon effect.
-
From Tool Calls to Policy Updates: A Reproducible Agent RL Stack
· 45 min readA reproducible blueprint for Agent RL: environment snapshots, verifier contracts, credit assignment, three-policy semantics, asynchronous generation and training, partial rollouts, and runnable smoke tests.
-
A Repair Can Save a Trajectory Without Teaching the Agent
· 18 min readAn eight-seed counterfactual-replay experiment found that high-value one-action repairs reliably rescue failed agent trajectories—but a repair-value-dominant SFT selector did not beat simpler rules and recovered only 25% of matched-RL improvement.