StateAct: Program State, before Pixels, for Long-Horizon Computer-Use Agents
This paper introduces a new approach to training computer-use agents by directly interacting with the underlying program state, rather than relying on visual perception. By doing so, agents can reason more effectively and make fewer mistakes, which can lead to significant improvements in performance.