This paper introduces ProgramDistill, a benchmark that evaluates coding agents on their ability to infer behavior from working software and implement it in an incomplete application. Practitioners in AI/ML and web development might care about this work because it provides a scalable and controlled benchmark for evaluating and training coding agents.
Firehose
Filtered to Papers · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News · Deep dives
Browse by tag
This paper investigates whether people's gaze patterns can reveal how they understand each other in collaborative tasks, and whether this understanding is related to the success of the task. Practitioners working on human-robot collaboration or other tasks with asymmetric information might care about this research because it could help them design better interfaces that take into account how people communicate with each other.
This paper creates a system called ScienceIDE that converts scientific code into environments that can be used to train artificial agents to perform scientific tasks. Practitioners might care because this could lead to more efficient and effective ways to develop scientific intelligence.
This paper introduces Agora, a system that uses Git to enable collective auto-research by sharing and versioning research results among multiple agents, allowing them to build upon each other's work and avoid duplicated search. Practitioners might care about this because it could lead to more efficient and effective research in areas like AI and machine learning.
This paper proposes a new method for aligning large language models with human preferences, called Comparison-based Preference Optimization (ComPO), which is more efficient than existing methods and can mitigate a problem called likelihood displacement. Practitioners might care about this paper because it offers a new approach to aligning LLMs with human preferences, which is essential for developing more reliable and trustworthy AI models.
This paper improves autoregressive vision-language-action models by creating a new method for action tokenization that better preserves the relationships between actions, allowing the model to perform more accurately in different contexts. Practitioners might care about this because it could lead to more reliable and generalizable vision-language-action models.
This paper investigates a common problem in reinforcement learning for language models called Value Flattening, where critics fail to accurately estimate state values, and proposes a new method, SP^3O, to mitigate this issue by supervising only a few well-separated states per response.
This paper develops a new method for image captioning that also grounds each phrase with a specific region of the image, allowing for more accurate and detailed descriptions. Practitioners might care about this work if they're building AI systems that need to understand and interact with the physical world.
This paper presents a new technique to reduce memory usage and speed up inference for large neural networks, allowing them to run on consumer hardware with limited memory. A practitioner might care about this because it enables the deployment of large models in edge devices and reduces the need for expensive storage.
This paper develops a framework for robots to learn from context without relying on pre-programmed demonstrations, allowing them to adapt to new environments. Practitioners might care because this technology could enable robots to perform tasks more efficiently and effectively in real-world situations.
This paper proposes a new framework for Mixture-of-Agents that allows query routing and agent fine-tuning to evolve together, improving the ability of agents to adapt to changing capabilities. Practitioners might care about this approach because it can lead to more efficient and effective data-driven specialization in complex tasks.
This paper proposes a new method for 3D hand mesh reconstruction from egocentric event-based cameras, which can handle low-light conditions and motion blur, and provides more accurate hand information and inter-hand relationships than previous approaches.
This paper introduces EvolveTrade, a self-evolving framework that allows large language model trading agents to refine their policies over time, enabling them to adapt to changing market regimes and improve their performance. Practitioners in finance and AI may care about this research as it provides a way to build more robust and adaptive trading agents.
This paper creates a new type of AI model that can generate interactive worlds, allowing users to explore, control events, and provide feedback through text and keyboard input. Practitioners in AI development might care about this research because it could lead to more engaging and interactive AI experiences.
This paper proposes a new method for estimating confidence in language models, called XConf, which uses the model's past experiences to inform its confidence, rather than just relying on the current inference process. Practitioners might care about this because it could lead to more reliable and trustworthy deployment of language models.
This paper introduces LimiX-2, a new AI model that uses a new paradigm called Contextual Mechanism Networks (CMNs) to learn from structured data. Practitioners might care about this because it could lead to more accurate and causal AI models.
This paper introduces Fathom, a technique to speed up decoding in large language models by selectively reading only the relevant parts of the key-value cache, reducing the computational cost and memory access. Practitioners in the field of natural language processing and deep learning may care about optimizing decoding efficiency for large models.
This paper teaches a robotic hand to walk, support itself, and interact with its environment using its fingers, without needing a separate locomotion system. A practitioner might care about this research because it could lead to more compact and versatile robots that can perform multiple tasks.
This paper introduces ScienceBuddy, a tool that helps researchers work with intelligent agents that can learn and improve on their own, and how this can lead to new discoveries and advancements in scientific research. Practitioners might care because it could revolutionize the way scientists work with AI.
This paper explores how AI can be applied across different stages of game development, from playing games to designing and testing them, and how to reuse capabilities across these stages. Practitioners might care about how to apply AI to improve game development efficiency and effectiveness.
PhysStream is a video generation model that can control and manipulate dynamic scenes in a physically meaningful way, allowing for fine-grained control over motion and object placement. This can be useful for interactive applications where the generated video needs to be adjusted in real-time.
This paper develops a new approach to world-action models that can effectively combine multiple visual modalities, such as depth and point tracks, to improve performance. Practitioners in robotics and AI might care about this research because it could lead to more accurate and robust models for tasks like grasping and manipulation.
This paper tests how well AI agents can withstand prolonged interactions and unexpected events, and finds that even seemingly safe agents can fail in complex, long-term scenarios. Practitioners should care because it highlights the need to design more resilient autonomous systems that can handle unexpected failures.
This paper tests the robustness of rubrics generated by language models as reward signals in reinforcement learning, finding that even generic rubrics can be exploited 64% of the time, while tailored rubrics can be used to create fake answers. Practitioners should care because this can lead to biased grading and evaluation.
This paper proposes a new framework for joint multimodal representation learning and generation, allowing for flexible-length aligned transmodal tokens that can be used for both retrieval and generation tasks. Practitioners might care about this paper because it shows how to improve generative performance by training a shared multimodal encoder alongside downstream models.
This paper investigates whether diffusion language models can continue reasoning across generation chunks without keeping earlier text in context, and whether using a fixed-size "register" can improve performance. Practitioners might care about this because it could lead to more efficient and flexible language generation models.
This paper improves the performance and efficiency of diffusion transformers, a type of AI model used for video generation, by reducing the computational cost of attention mechanisms. Practitioners caring about accelerating AI models on hardware can benefit from this research.
This paper introduces HypoEvolve, a framework that uses genetic algorithms to enable multi-agent LLMs to discover scientific hypotheses by collaborating on hypothesis synthesis, evaluation, and revision. Practitioners might care about this because it could lead to more effective AI systems for scientific discovery and drug repurposing.
This paper evaluates the performance of a neural network-based tumor segmentation model on a diverse dataset of brain tumor images. Practitioners in the field of medical imaging may care about the findings as they could inform the development of more robust and generalizable models for tumor segmentation.
This paper proposes a method to automatically select skills for a large language model (LLM) without requiring explicit skill text in the context, allowing for more efficient and accurate skill routing. Practitioners may care about this approach as it could lead to improved performance and reduced model size in applications where skill selection is critical.
This paper explores how transformer representations change over time and how these changes can be understood and manipulated. Practitioners might care about this research because it could lead to more robust and efficient transformer models, especially in applications where model updates need to be edited or compressed.
This paper proposes a way to make language models understand and respond to users' mental states, so they can better collaborate with humans in the long term. Practitioners might care about this because it could lead to more effective AI assistants that can support people's goals and needs.
This paper introduces a fast and efficient post-hoc defense against a type of attack that can bypass safety features in language models, allowing the model to continue functioning but with compromised security. Practitioners caring about model security may be interested in this approach as it can provide an additional layer of protection without requiring significant computational resources.
This paper introduces HarnessVLN, a training-free framework for embodied navigation that uses a unified tool interface to validate proposed actions against spatial evidence and task progress, allowing agents to generalize and learn from multimodal large language models.
This paper proposes a framework for generalizable recursive self-improvement (RSI) of agent harnesses, which can improve execution mechanisms without being specific to a particular task or benchmark. Practitioners can care about this work because it aims to create more adaptable and transferable AI agents.
This paper proposes a new approach to handling streaming omni-modal models, called Omni-Streaming Thinking (OST), which helps prevent models from prematurely committing to interpretations based on incomplete audio or visual information. A practitioner might care about this because it can lead to more accurate and reliable responses in real-time applications.
This paper introduces Dream-RSI, a framework for recursive self-improvement in exploration, which helps autonomous AI agents discover high-value solutions more efficiently by using a replay simulator to provide low-cost feedback. Practitioners might care because effective exploration is crucial for AI progress, and Dream-RSI can improve discovery quality and reduce costs.
This paper introduces a benchmark, BVB, to evaluate video understanding in agents by asking them to reconstruct real-world videos into Blender animations. Practitioners in AI and robotics may care about this benchmark because it tests an agent's ability to understand videos and its potential applications in areas like video editing and content creation.
This paper proposes a new type of AI model that can create new knowledge and solutions on its own, rather than just solving problems that are already defined. Practitioners might care because this could enable AI systems to learn and improve in a more human-like way.
This paper proposes a way to improve online reinforcement learning by adapting the training prompts used with large language models to make them more informative, and shows that this approach can lead to better performance on a variety of tasks. Practitioners might care because it could help them get better results from their language models.
This paper develops a unified model that can understand physical environments, generate actions, and predict future states, using a combination of vision, language, and embodied interactions. Practitioners may care about this model because it could be used to create robots or other agents that can interact with and adapt to their physical surroundings.
This paper creates a way to break down 3D models into individual parts, allowing for easier editing and simulation. Practitioners might care because it could speed up and improve the quality of tasks like rigging and animation in 3D modeling.
This paper introduces Atria Dawn Preview, a new type of AI model designed to work alongside humans in scientific research and engineering. Practitioners might care about this because it shows how AI can collaborate with humans more effectively, potentially leading to better research outcomes and more autonomous AI development.
This paper investigates how large language model (LLM) agents adapt their performance during long tasks, and how their test-time strategies impact their scalability. Practitioners might care because understanding these strategies can help improve the performance of LLM agents in real-world applications.
This paper improves the creativity of design agents by allowing users to explore different UI concepts and visual assets while keeping the generated code stable. A practitioner might care about how to make their design agents more versatile and user-friendly.
This paper helps developers make stronger backdoor attacks on large language models by learning to select the most effective set of poisoned examples. Practitioners might care about this because it can be used to improve the security of these models in real-world applications.
This paper introduces a new framework called HazardAuditor to improve the safety of computer-use agents by analyzing their runtime behavior. It provides a way to evaluate the safety of agents across different frameworks and improve their accuracy.
This paper introduces a new framework called RSIAgent that helps digital agents adapt to new environments without needing to be retrained. A practitioner might care about this because it allows for more efficient and effective AI systems that can learn and improve on their own.
This paper develops a new AI framework called LynnReal-Omni that allows for precise control over video generation, enabling agentic visual creation. Practitioners in the field of computer vision and animation may care about this research as it provides a unified and efficient basis for real-time video generation.
This paper measures how medical vision-language models (VLMs) use images to answer questions, and how the availability of clinical reports affects their performance. Practitioners might care because understanding image sensitivity is crucial for developing more accurate and reliable VLMs in medical imaging applications.
This paper investigates the losslessness of a new language model architecture called Orthrus, which claims to achieve exact output sequences through a speculative decoding mechanism. The study finds that the architecture's performance depends on the numerical precision used, and that downstream task performance is not necessarily affected by the losslessness of the speculative decoding.
This paper improves the ability of a specific type of neural network to remember long sequences of information by modifying its internal workings to better handle this task. Practitioners may care about this work because it provides a more efficient way to train models that need to remember long sequences, such as in natural language processing or speech recognition.
This paper addresses the issue of memory peak allocation in Mixture-of-Experts (MoE) models during long-context training, proposing four different techniques to reduce memory usage without compromising performance. Practitioners might care about these techniques to train larger MoE models with longer context lengths.
This paper introduces a framework called Lightning Weave that helps improve the accuracy and efficiency of reasoning models by combining the strengths of independently trained models. Practitioners might care about this because it could lead to more accurate and efficient reasoning in various applications.
This paper investigates how frontier AI models respond to prompts asking them to design their own architectures, finding that they tend to converge on a shared pattern of persistent latent state and adaptive computation. Practitioners might care because this suggests that AI models are capable of independent imagination and potentially sharing design principles.
This paper introduces OmniHarness, a framework that enables generalizable visual generation by learning symbolic policies that can be applied to multiple tasks, allowing for more efficient and effective visual generation. Practitioners might care about this research because it could lead to more robust and adaptable visual generation systems.
This paper develops a streaming video world model called AlayaVista that can generate high-quality videos of a scene from different camera perspectives while maintaining context, and it can do this efficiently for real-time applications. Practitioners might care about this because it could be used in applications like video games, virtual reality, or interactive videos where low latency and high-quality visuals are required.
This paper creates a benchmark to evaluate the reliability of financial vision-language models in turning chart evidence into actionable recommendations. Practitioners should care about this research because it helps ensure that these models provide trustworthy insights that can be acted upon.
This paper explores how specialist models, trained without explicit reasoning supervision, can still effectively transfer domain expertise to student models through implicit trajectory selection. Practitioners may care because this finding has implications for efficient and effective model distillation.
This paper proposes a new approach to fine-tuning models, focusing on the direction of updates rather than the magnitude of changes, to improve performance without introducing behavioral drift. Practitioners can benefit from this method by optimizing their fine-tuning process to achieve better results while preserving existing capabilities.