HarnessEval-W: Agentifying the Evaluation of Visual Worlds
This paper develops a new method for evaluating world models, called HarnessEval-W, which provides more detailed and justifiable results than existing benchmarks. Practitioners might care about HarnessEval-W because it can help them build more trustworthy world models that better align with human preferences.