This paper creates a new dataset and benchmark for detecting landmines in images taken by drones or ground vehicles, and tests how well different AI detectors can handle variations in conditions. Practitioners who work on drone or ground vehicle safety systems might care about this research because it could help them build more reliable systems that can detect landmines in different environments.
Firehose
Filtered to tagged “continual learning” · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News · Deep dives
Browse by tag
This paper introduces ODEWorld, a new approach to modeling the physical world by learning a continuous latent velocity field that operates in physical time, allowing for more efficient and realistic predictions of future events. Practitioners in robotics and computer vision may care about ODEWorld's ability to provide rich planning-oriented information and high-quality image reconstruction.
This paper proposes a method to improve long-CoT reasoning in large language models by addressing the issue of unequal token contributions to the final outcome. It shows that current methods, such as GRPO, assign too much credit to highly sensitive tokens and proposes a new method, CSCR, that reduces credit for these tokens to improve performance.
This paper develops a method to improve the performance of Vision-Language-Action models by adapting their steering strategy at test time, allowing them to generalize better to new tasks and domains. Practitioners can benefit from this approach by improving the robustness of their VLA models in real-world applications.
This paper proposes a new approach to memory-augmentation in large language model agents, allowing them to actively reconstruct and adapt past experiences to fit the current context, rather than simply replaying them. Practitioners might care because this approach can improve the robustness and intrinsic reasoning capabilities of agents in complex scenarios.
This paper introduces Chimera, a hybrid visual diffusion transformer that efficiently processes text, image, and video tokens to generate high-resolution images, videos, and multimodal context. Practitioners might care about this paper because it provides a scalable solution for large-scale visual generation tasks.
This paper proposes a new approach to improve vision-language models for visual retrieval, which can handle long visual contexts and large numbers of distractors. Practitioners might care because it can lead to better performance on image and video benchmarks.
This paper introduces a new method for training computer-use agents, called Echoverse, which generates evolving environments that mimic real-world applications. By using these environments, agents can learn more effectively and improve their performance on real-world tasks.
This paper introduces Qwen-UI-Agent, a type of artificial intelligence system that can perform tasks on various devices, such as smartphones and computers, and improve its abilities on its own. Practitioners might care about this research because it aims to create more practical and autonomous AI systems that can be used in real-world scenarios.
This paper develops a method to scale up language models by increasing their memory capacity, allowing for better performance and more efficient use of parameters. Practitioners may care about this research as it could lead to more powerful and efficient language models for applications like language translation and text generation.
This paper introduces a new type of memory system for large language model (LLM) based multi-agent systems that tracks which agents can be trusted and under what conditions. Practitioners might care because it can help improve the reliability and coordination of these systems.
This paper develops a new method for improving reasoning language models, called β-OPSD, which combines policy optimization and self-distillation to improve stability and performance. Practitioners might care about this method because it provides a more efficient and effective way to improve language model reasoning abilities.
This paper introduces a new type of AI model called memory foundation models, which allows the model to learn and retain information internally, rather than relying on external memory modules. This could be useful for practitioners who want to build more efficient and flexible AI agents.
This paper studies how lossy verification schemes can improve the efficiency of speculative decoding in large language models, but may also degrade generation quality. Practitioners may care about understanding the trade-offs between speed and quality when using these schemes.
This paper introduces Explorative Modeling, a new approach to training generative models that allows for end-to-end generation by exploring multiple candidate matches between model generations and data. This can lead to improved performance and efficiency in various applications.
This paper investigates how Large Language Model (LLM) agents can use a file system to store and organize their memories, and whether this approach improves their performance. Practitioners might care because it shows that using a file system as memory can be beneficial for LLM agents, but there are limitations to this approach.
This paper introduces TurboVLA, a new vision-language-action model that reduces computation and memory overhead by directly exchanging information between visual observations and language instructions, allowing for faster and more efficient robotic manipulation. Practitioners might care about this approach for building more efficient and effective VLA models.
This paper investigates whether AI agents can conduct open-ended AI research and provides early evidence that they can perform the engineering aspects but struggle with critical parts of the research lifecycle, such as making progress on research questions and judgment about publishable research.
This paper introduces SkillRise, a framework that enables large language model agents to learn skills across related tasks, allowing them to reuse solution patterns and improve performance on multiple tasks. Practitioners can use SkillRise to train more efficient LLM agents that can adapt to new tasks and improve their performance over time.
This paper creates a benchmark to test the security capabilities of AI agents in a real-world setting, specifically incident response, and finds that current agents struggle to detect and remediate silent intrusions and produce verified plans.
This paper introduces a new approach to agentic speech recognition that uses a memory to help correct mistakes and improve accuracy. By limiting the corrections made, the system can avoid over-correcting and improve performance on challenging tasks.
This paper improves autoregressive video distillation methods by aligning the initialization and distribution matching stages, focusing on matching the target distribution's mode coverage rather than just visual quality. Practitioners can benefit from this approach to generate higher-quality videos with better diversity and coverage.
This paper teaches language models to synthesize complete software programs from scratch, which is a challenging task. Practitioners might care because this can improve the models' performance on software engineering tasks.
Intelligence on the Edge: Liquid AI's Ramin Hasani on the Search for Device-Native Foundation Models
This episode features Ramin Hassani, CEO of Liquid AI, discussing the company's journey from biologically inspired neural networks at MIT to developing device-native foundation models. He makes a technically grounded case for efficient, har…
In this episode, Thomas Ahle discusses the development of thermodynamic computing chips and the challenges of chip design automation using AI agents. He explains how his team built an open-source Verilog simulator with AI collaboration to o…