This paper investigates how text conditioning affects visual generation and proposes ways to improve it, leading to better performance on various benchmarks. Practitioners might care about the findings to develop more effective text-to-image models.
Firehose
Filtered to Papers · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News · Deep dives
Browse by tag
This paper introduces a benchmark for schema-guided document extraction, which is a crucial task in enterprise workflows, and evaluates various models' performance on this task, including their accuracy, grounding, and cost-effectiveness.
This paper develops a new method to evaluate and verify image editing consistency across multiple references, addressing a challenge in reinforcement learning for multi-reference editing. Practitioners may care about this approach as it enables more accurate and reliable reinforcement learning for image editing tasks.
This paper creates a new dataset and benchmark for detecting landmines in images taken by drones or ground vehicles, and tests how well different AI detectors can handle variations in conditions. Practitioners who work on drone or ground vehicle safety systems might care about this research because it could help them build more reliable systems that can detect landmines in different environments.
This paper proposes a new method for low-light imaging that combines RGB and Near-Infrared (NIR) images in 3D space to improve image quality without requiring clean RGB data. Practitioners might care because it could lead to more robust and reliable low-light imaging systems.
This paper proposes a method to combine reinforcement learning with verifiable rewards and on-policy distillation to improve performance on complex tasks, and shows that this method can lead to more stable training and better results.
This paper introduces a system called EMBL AI Librarian that helps AI agents find relevant life-science papers and evidence by providing a natural language interface. Practitioners in life sciences and AI development may care about this paper because it shows how a better knowledge retrieval system can improve the performance of AI agents in various tasks.
This paper introduces a framework to audit system prompts in AI applications, examining how developers design and use these prompts to govern the behaviors of foundation models. Practitioners should care because the lack of transparency and accountability in system prompts can erode trust in AI systems.
This paper proposes a new regularization technique for world models to improve planning performance by aligning latent samples with quantiles of a Gaussian distribution, which helps control heavy-tailed deviations. Practitioners caring about efficient planning in complex environments may find this approach useful.
This paper explores the limitations of current safeguards for Large Language Models (LLMs) in preventing misuse, and proposes a new approach that combines capability release with evidence about downstream use to improve safety. Practitioners caring about the responsible development and deployment of LLMs might care about this research as it addresses a key challenge in ensuring the safe and trustworthy use of these models.
This paper shows that large language models struggle with commonsense reasoning due to a bias towards explicit conditions, which can be misled by irrelevant information, and that this issue can be improved by adjusting the task framing or using lightweight prompting. Practitioners caring about the reliability of language models in real-world applications might want to consider this when using them for tasks that require critical thinking.
This paper introduces ODEWorld, a new approach to modeling the physical world by learning a continuous latent velocity field that operates in physical time, allowing for more efficient and realistic predictions of future events. Practitioners in robotics and computer vision may care about ODEWorld's ability to provide rich planning-oriented information and high-quality image reconstruction.
This paper proposes a new method for robots in a swarm to predict the same future state from local observations and limited messages, and shows that it can outperform a simpler approach with less training data. Practitioners might care because it could lead to more efficient and effective collective decision-making in swarms of robots.
This paper proposes a method to improve long-CoT reasoning in large language models by addressing the issue of unequal token contributions to the final outcome. It shows that current methods, such as GRPO, assign too much credit to highly sensitive tokens and proposes a new method, CSCR, that reduces credit for these tokens to improve performance.
This paper develops a method to improve the performance of Vision-Language-Action models by adapting their steering strategy at test time, allowing them to generalize better to new tasks and domains. Practitioners can benefit from this approach by improving the robustness of their VLA models in real-world applications.
This paper introduces a new research paradigm for emotional dialogue that focuses on sustaining users' emotional capacities and capabilities, rather than just providing relief. Practitioners in human-computer interaction, dialogue systems, and affective computing may care about this approach as it has implications for designing more effective and long-lasting support systems.
This paper develops a framework, SpatialCLI, to help vision-language models (VLMs) better understand and use visual tools to make better decisions. By training VLMs to reason with spatial tools and then internalize those capabilities, SpatialCLI can improve the performance of VLMs in tasks that require visual reasoning.
This paper proposes a new approach to memory-augmentation in large language model agents, allowing them to actively reconstruct and adapt past experiences to fit the current context, rather than simply replaying them. Practitioners might care because this approach can improve the robustness and intrinsic reasoning capabilities of agents in complex scenarios.
This paper introduces PhiZero, a world model that uses physical language to predict how the physical world evolves, allowing for more explicit reasoning and potentially more realistic simulations. Practitioners might care because this approach could lead to more realistic and interactive world modeling in applications like robotics, video games, and virtual reality.
This paper introduces a new dataset called ACE-Data-0, which provides a comprehensive and synchronized record of human behavior in real-world environments, capturing various aspects of embodied intelligence such as perception, action, and interaction. Practitioners in the field of embodied AI and machine learning can use this dataset to develop more sophisticated models that can learn from human demonstrations and perform tasks that involve complex manipulation, locomotion, and interaction.
This paper introduces Chimera, a hybrid visual diffusion transformer that efficiently processes text, image, and video tokens to generate high-resolution images, videos, and multimodal context. Practitioners might care about this paper because it provides a scalable solution for large-scale visual generation tasks.
This paper proposes a new approach to improve vision-language models for visual retrieval, which can handle long visual contexts and large numbers of distractors. Practitioners might care because it can lead to better performance on image and video benchmarks.
This paper trains an AI model to improve itself in the process of building AI, with a focus on machine learning engineering, and shows promising results in various benchmarks. Practitioners may care about this research as it could lead to more efficient and autonomous AI development.
This paper creates a system to help scientists and AI agents find and verify specific information in chemistry literature by converting research papers into claims, which are then linked together to form a network of evidence. Practitioners might care because this system could improve the efficiency and accuracy of literature synthesis in chemistry research.
This paper introduces a new method for training computer-use agents, called Echoverse, which generates evolving environments that mimic real-world applications. By using these environments, agents can learn more effectively and improve their performance on real-world tasks.
This paper proposes a new framework for search agents in reinforcement learning, called Harness-G, which addresses a problem called retrieval-equivalence collapse where the same query string can be generated multiple times but the retrieved information becomes increasingly similar. This can make it difficult for the agent to learn effective retrieval strategies.
This paper proposes a new method for training large language models in open-ended domains, using evolving contexts as in-training supervision to capture task preferences. Practitioners may care about this approach because it can lead to better performance on open-ended tasks.
This paper compares the performance of different retrieval-augmented generation (RAG) paradigms at varying corpus sizes, finding that BM25 outperforms others at larger scales, but not at smaller ones, and that lexical retrieval is the strongest scalable default.
This paper proposes a new approach to agentic visual reasoning, which helps large language models (LLMs) perform better on complex tasks by using tools more efficiently. Practitioners might care about this research because it aims to improve the performance of LLMs on challenging problems.
This paper introduces a new task called multi-reference image-grounded video captioning, where models must describe video content while referencing multiple images. Practitioners might care because this can improve the accuracy and faithfulness of video captions in real-world applications.
This paper introduces Qwen-UI-Agent, a type of artificial intelligence system that can perform tasks on various devices, such as smartphones and computers, and improve its abilities on its own. Practitioners might care about this research because it aims to create more practical and autonomous AI systems that can be used in real-world scenarios.
This paper develops a method to scale up language models by increasing their memory capacity, allowing for better performance and more efficient use of parameters. Practitioners may care about this research as it could lead to more powerful and efficient language models for applications like language translation and text generation.
This paper explores using large language models to improve execution costs in algorithmic trading by breaking down a large order into smaller ones, and finds that these models can outperform human traders and other approaches in certain situations.
This paper introduces a benchmark to evaluate the performance of models in editing multi-person images, focusing on anatomical and geometric accuracy. Practitioners in the field of computer vision and image editing might care about this research as it aims to improve the quality of human-like images with multiple people.
This paper proposes a new method for evaluating role-playing agents (RPAs) that simulates user experiences in interactive role-playing to provide more accurate and personalized assessments of RPA capabilities. Practitioners might care about this because it helps them evaluate RPAs more effectively, leading to better systems that meet user needs.
This paper develops a new approach to multimodal question answering that ensures the provenance of the reasoning process, making it more trustworthy and transparent. Practitioners might care about this because it helps to identify and mitigate common pitfalls in multimodal question answering models.
This paper introduces ShadowDancer, a method for teaching video world models to perform any-action control by learning unified dynamics representations from a video and its shadow. Practitioners might care because it can improve action transfer and long-term control in complex environments.
This paper introduces a new type of memory system for large language model (LLM) based multi-agent systems that tracks which agents can be trusted and under what conditions. Practitioners might care because it can help improve the reliability and coordination of these systems.
This paper develops a technique to identify and mitigate demographic bias in large language models by selectively pruning specific neurons in neural networks, without significantly impacting the model's overall performance. Practitioners in AI and NLP may care about this method as it could help create more fair and transparent language models.
This paper investigates how sparse mixture-of-experts language models route tokens to multiple experts and how this routing affects their performance. Practitioners may care because understanding how to optimize these models can lead to better language understanding and generation.
This paper develops a new method for improving reasoning language models, called β-OPSD, which combines policy optimization and self-distillation to improve stability and performance. Practitioners might care about this method because it provides a more efficient and effective way to improve language model reasoning abilities.
This paper develops a new framework for world modeling that takes into account the mental state of agents, which is essential for predicting human decisions. Practitioners caring about human decision-making and planning might find this research useful.
This paper investigates how AI-assisted coding assistants can better understand and respond to users' ambiguous coding requests by leveraging their past experiences. A practitioner might care about developing more effective coding assistants that can reduce the need for repeated clarification.
OVEarth-Bench: Evaluating Category Breadth and Query Diversity for Open-Vocabulary Earth Observation
This paper introduces a new benchmark (OVEarth-Bench) to evaluate open-vocabulary Earth observation models, focusing on both category breadth and query diversity. Practitioners in this field can benefit from understanding the importance of developing more realistic and diverse benchmarks for reliable model evaluation.
This paper introduces VideoCoCo, a system that uses executable Blender code to generate physically consistent videos from text prompts, and shows promising results in improving video generation quality. Practitioners might care about this research if they want to generate realistic videos that are also physically plausible.
This paper introduces a new type of AI model called memory foundation models, which allows the model to learn and retain information internally, rather than relying on external memory modules. This could be useful for practitioners who want to build more efficient and flexible AI agents.
This paper studies how lossy verification schemes can improve the efficiency of speculative decoding in large language models, but may also degrade generation quality. Practitioners may care about understanding the trade-offs between speed and quality when using these schemes.
This paper introduces Explorative Modeling, a new approach to training generative models that allows for end-to-end generation by exploring multiple candidate matches between model generations and data. This can lead to improved performance and efficiency in various applications.
This paper investigates how Large Language Model (LLM) agents can use a file system to store and organize their memories, and whether this approach improves their performance. Practitioners might care because it shows that using a file system as memory can be beneficial for LLM agents, but there are limitations to this approach.
This paper evaluates how multimodal large language models use intermediate visual states during reasoning and finds that these visual states are not as crucial as previously thought, but can still impact model performance under certain conditions. Practitioners might care because understanding how these models use visual states can help improve their performance and reliability.
This paper evaluates whether vision-language models can act through a physical body and how they can make decisions about what actions to take, without being hindered by issues like balance and motor control. Practitioners in AI and robotics might care because understanding how models interact with their physical bodies can help improve their ability to navigate and interact with the world.
This paper introduces TurboVLA, a new vision-language-action model that reduces computation and memory overhead by directly exchanging information between visual observations and language instructions, allowing for faster and more efficient robotic manipulation. Practitioners might care about this approach for building more efficient and effective VLA models.
This paper investigates whether AI agents can conduct open-ended AI research and provides early evidence that they can perform the engineering aspects but struggle with critical parts of the research lifecycle, such as making progress on research questions and judgment about publishable research.
This paper introduces a benchmark to evaluate large language model (LLM) agents on long-horizon office-suite tasks, considering their cost-effectiveness and quality. Practitioners can care about this research because it aims to ensure LLM agents can assist users efficiently and effectively.
This paper introduces SkillRise, a framework that enables large language model agents to learn skills across related tasks, allowing them to reuse solution patterns and improve performance on multiple tasks. Practitioners can use SkillRise to train more efficient LLM agents that can adapt to new tasks and improve their performance over time.
This paper creates a benchmark to test the security capabilities of AI agents in a real-world setting, specifically incident response, and finds that current agents struggle to detect and remediate silent intrusions and produce verified plans.
This paper proposes a new model that can generate realistic and interactive game environments while respecting the game's underlying mechanics, such as rules for health and skill activation. Practitioners who work on game generation or virtual worlds may care about this research because it can lead to more realistic and engaging game experiences.
This paper introduces a new approach to agentic speech recognition that uses a memory to help correct mistakes and improve accuracy. By limiting the corrections made, the system can avoid over-correcting and improve performance on challenging tasks.
This paper improves autoregressive video distillation methods by aligning the initialization and distribution matching stages, focusing on matching the target distribution's mode coverage rather than just visual quality. Practitioners can benefit from this approach to generate higher-quality videos with better diversity and coverage.
This paper teaches language models to synthesize complete software programs from scratch, which is a challenging task. Practitioners might care because this can improve the models' performance on software engineering tasks.