StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling
This paper introduces StateM, a runtime system that improves the performance of long-horizon agents by organizing their execution around durable states, phase-local context, and other components. Practitioners might care because StateM can help agents achieve higher accuracy and efficiency, especially in complex tasks.