Two new benchmarks try to measure coding agents on realistic work

Tencent's WorkBuddy Bench evaluates coding agents across code, web, office and security domains, with task construction designed to resist training-data contamination. ICAE-Bench takes a different angle, scoring agents on building software from incomplete product intent rather than from a tidy specification. Both are early signals that the field is trying to move past single-file puzzle benchmarks toward something closer to how agents actually get used.

Read the source →

Research papersAgentic codingworkbuddy-benchicae-benchcoding-agentsbenchmarkscontamination

← All signals