OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models
This paper develops a standardized evaluation method for computer-using agents (CUAs) to ensure they fulfill task instructions, using vision-language models (VLMs) as judges. Practitioners can benefit from this work by using reliable and cost-effective reward signals for CUA training.