Agent papers can make progress look like a single score. Split the system apart before accepting it.
Can the task be checked
If success cannot be judged automatically, scores are hard to compare. Look at how completion is defined and how often a person enters the evaluation.
What the model is allowed to do
List the tools, the step limit, and the writable surface. Two scores are comparable only when the action spaces are close.
Where the failures are
A table that reports only the average usually leaves the useful information outside the appendix.
Find the result on long tasks, permission errors, and retrieval failures. That is closer to the system you would actually build.
Discussion