This paper introduces Agora, a system that uses Git to enable collective auto-research by sharing and versioning research results among multiple agents, allowing them to build upon each other's work and avoid duplicated search. Practitioners might care about this because it could lead to more efficient and effective research in areas like AI and machine learning.
Firehose
Filtered to Papers, tagged “continual learning” · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News · Deep dives
Browse by tag
This paper develops a framework for robots to learn from context without relying on pre-programmed demonstrations, allowing them to adapt to new environments. Practitioners might care because this technology could enable robots to perform tasks more efficiently and effectively in real-world situations.
This paper proposes a new framework for Mixture-of-Agents that allows query routing and agent fine-tuning to evolve together, improving the ability of agents to adapt to changing capabilities. Practitioners might care about this approach because it can lead to more efficient and effective data-driven specialization in complex tasks.
This paper proposes a new method for 3D hand mesh reconstruction from egocentric event-based cameras, which can handle low-light conditions and motion blur, and provides more accurate hand information and inter-hand relationships than previous approaches.
This paper introduces EvolveTrade, a self-evolving framework that allows large language model trading agents to refine their policies over time, enabling them to adapt to changing market regimes and improve their performance. Practitioners in finance and AI may care about this research as it provides a way to build more robust and adaptive trading agents.
This paper proposes a new method for estimating confidence in language models, called XConf, which uses the model's past experiences to inform its confidence, rather than just relying on the current inference process. Practitioners might care about this because it could lead to more reliable and trustworthy deployment of language models.
This paper introduces ScienceBuddy, a tool that helps researchers work with intelligent agents that can learn and improve on their own, and how this can lead to new discoveries and advancements in scientific research. Practitioners might care because it could revolutionize the way scientists work with AI.
This paper explores how AI can be applied across different stages of game development, from playing games to designing and testing them, and how to reuse capabilities across these stages. Practitioners might care about how to apply AI to improve game development efficiency and effectiveness.
This paper develops a new approach to world-action models that can effectively combine multiple visual modalities, such as depth and point tracks, to improve performance. Practitioners in robotics and AI might care about this research because it could lead to more accurate and robust models for tasks like grasping and manipulation.
This paper tests how well AI agents can withstand prolonged interactions and unexpected events, and finds that even seemingly safe agents can fail in complex, long-term scenarios. Practitioners should care because it highlights the need to design more resilient autonomous systems that can handle unexpected failures.
This paper proposes a new framework for joint multimodal representation learning and generation, allowing for flexible-length aligned transmodal tokens that can be used for both retrieval and generation tasks. Practitioners might care about this paper because it shows how to improve generative performance by training a shared multimodal encoder alongside downstream models.
This paper investigates whether diffusion language models can continue reasoning across generation chunks without keeping earlier text in context, and whether using a fixed-size "register" can improve performance. Practitioners might care about this because it could lead to more efficient and flexible language generation models.
This paper introduces HypoEvolve, a framework that uses genetic algorithms to enable multi-agent LLMs to discover scientific hypotheses by collaborating on hypothesis synthesis, evaluation, and revision. Practitioners might care about this because it could lead to more effective AI systems for scientific discovery and drug repurposing.
This paper proposes a method to automatically select skills for a large language model (LLM) without requiring explicit skill text in the context, allowing for more efficient and accurate skill routing. Practitioners may care about this approach as it could lead to improved performance and reduced model size in applications where skill selection is critical.
This paper proposes a way to make language models understand and respond to users' mental states, so they can better collaborate with humans in the long term. Practitioners might care about this because it could lead to more effective AI assistants that can support people's goals and needs.
This paper introduces a fast and efficient post-hoc defense against a type of attack that can bypass safety features in language models, allowing the model to continue functioning but with compromised security. Practitioners caring about model security may be interested in this approach as it can provide an additional layer of protection without requiring significant computational resources.
This paper proposes a framework for generalizable recursive self-improvement (RSI) of agent harnesses, which can improve execution mechanisms without being specific to a particular task or benchmark. Practitioners can care about this work because it aims to create more adaptable and transferable AI agents.
This paper proposes a new approach to handling streaming omni-modal models, called Omni-Streaming Thinking (OST), which helps prevent models from prematurely committing to interpretations based on incomplete audio or visual information. A practitioner might care about this because it can lead to more accurate and reliable responses in real-time applications.
This paper proposes a new type of AI model that can create new knowledge and solutions on its own, rather than just solving problems that are already defined. Practitioners might care because this could enable AI systems to learn and improve in a more human-like way.
This paper creates a way to break down 3D models into individual parts, allowing for easier editing and simulation. Practitioners might care because it could speed up and improve the quality of tasks like rigging and animation in 3D modeling.
This paper introduces Atria Dawn Preview, a new type of AI model designed to work alongside humans in scientific research and engineering. Practitioners might care about this because it shows how AI can collaborate with humans more effectively, potentially leading to better research outcomes and more autonomous AI development.
This paper investigates how large language model (LLM) agents adapt their performance during long tasks, and how their test-time strategies impact their scalability. Practitioners might care because understanding these strategies can help improve the performance of LLM agents in real-world applications.
This paper helps developers make stronger backdoor attacks on large language models by learning to select the most effective set of poisoned examples. Practitioners might care about this because it can be used to improve the security of these models in real-world applications.
This paper introduces a new framework called HazardAuditor to improve the safety of computer-use agents by analyzing their runtime behavior. It provides a way to evaluate the safety of agents across different frameworks and improve their accuracy.
This paper improves the ability of a specific type of neural network to remember long sequences of information by modifying its internal workings to better handle this task. Practitioners may care about this work because it provides a more efficient way to train models that need to remember long sequences, such as in natural language processing or speech recognition.
This paper introduces OmniHarness, a framework that enables generalizable visual generation by learning symbolic policies that can be applied to multiple tasks, allowing for more efficient and effective visual generation. Practitioners might care about this research because it could lead to more robust and adaptable visual generation systems.