Papers

Filtered to transformer-based models · clear filter

Browse by term

continual learning 64reinforcement learning 34large language models 17benchmarking 12vision-language models 10generative models 8language models 8video generation 7multimodal models 6natural language processing 6robotics 6world models 6benchmarks 5diffusion models 5on-policy distillation 5policy optimization 5scalability 5self-distillation 5vision-language-action models 5autoregressive models 4computer vision 4diffusion transformers 4LLMs 4multimodal large language models 4verifiable rewards 4attention mechanisms 3embodied intelligence 3image editing 3long-term memory 3multimodal learning 3

Matching papers

Scaling Native Multimodal Pre-Training From Scratch

21 upvotes · 24 JUL 2026 · Haoyuan Wu, Aoqi Wu, Hai Wang et al.

This paper investigates how to scale large language models to also understand and interact with the physical world by training them on multiple types of data from scratch, allowing them to reason about both text and images. Practitioners might care because this could lead to more robust and versatile AI systems that can handle a wider range of tasks.

DiffGI: Differentiable Geometry Images for High-Fidelity Thin-Shell 3D Generation

11 upvotes · 15 JUL 2026 · Eungjune Shim, Hansol Lee, Eunjung Ju

This paper develops a new method for generating high-fidelity 3D images of thin-shell objects, like garments, by learning a continuous surface representation. Practitioners caring about 3D generation and object modeling might care about this approach because it achieves better results with fewer computational resources.