Testing an agent when there is no right answer
An answer key only covers the questions that have one. How to evaluate an AI agent for scientific work when two good scientists would disagree, by scoring provenance, negatives, and the judge itself.
Carmen Kivisild