Discover
4 results tagged “Research papers” · clear
Signals
-
Two new benchmarks try to measure coding agents on realistic work
Tencent's WorkBuddy Bench evaluates coding agents across code, web, office and security domains, with task construction designed to resist training-data contamination. ICAE-Bench t…
-
AREX proposes a recursively self-improving deep-research agent
AREX, a Hugging Face paper with 115 upvotes, describes an agent that recursively self-improves at deep research tasks. It's part of a run of work on agents that refine their own st…
-
This week's most-upvoted HF papers cluster around embodied AI
Four of the top Hugging Face daily papers this week are embodied and robotics-flavoured: Xiaomi-Robotics-1 (vision-language-action at scale), RynnBrain, HOMIE, and Apple-π. If you …
-
Paper: coder LLMs may already "know" what context to prune
A new paper (SWE-Pruner Pro) proposes that coding-focused LLMs already carry enough internal signal to identify which parts of a large context window are safe to prune, rather than…