DynaFLIP: Rethinking Robotics Perception via Tri-Modal-Dynamics Guided Representation Paper β’ 2605.30350 β’ Published May 28 β’ 13
Contrastive Distribution Matching for Amortized Sequential Monte Carlo in Discrete Diffusion Paper β’ 2605.23346 β’ Published May 22
optimize_anything: A Universal API for Optimizing any Text Parameter Paper β’ 2605.19633 β’ Published May 19 β’ 6
MultiGen: Level-Design for Editable Multiplayer Worlds in Diffusion Game Engines Paper β’ 2603.06679 β’ Published Mar 30 β’ 7
AVO: Agentic Variation Operators for Autonomous Evolutionary Search Paper β’ 2603.24517 β’ Published Mar 25 β’ 11
V-Co: A Closer Look at Visual Representation Alignment via Co-Denoising Paper β’ 2603.16792 β’ Published Mar 17 β’ 3
SE-Bench: Benchmarking Self-Evolution with Knowledge Internalization Paper β’ 2602.04811 β’ Published Feb 4 β’ 2
Motion 3-to-4: 3D Motion Reconstruction for 4D Synthesis Paper β’ 2601.14253 β’ Published Jan 20 β’ 10
V-DPM: 4D Video Reconstruction with Dynamic Point Maps Paper β’ 2601.09499 β’ Published Jan 14 β’ 11
UM-Text: A Unified Multimodal Model for Image Understanding Paper β’ 2601.08321 β’ Published Jan 13 β’ 21
ResTok: Learning Hierarchical Residuals in 1D Visual Tokenizers for Autoregressive Image Generation Paper β’ 2601.03955 β’ Published Jan 7 β’ 3
FlowBlending: Stage-Aware Multi-Model Sampling for Fast and High-Fidelity Video Generation Paper β’ 2512.24724 β’ Published Dec 31, 2025 β’ 9
Dream2Flow: Bridging Video Generation and Open-World Manipulation with 3D Object Flow Paper β’ 2512.24766 β’ Published Dec 31, 2025 β’ 9
What matters for Representation Alignment: Global Information or Spatial Structure? Paper β’ 2512.10794 β’ Published Dec 11, 2025 β’ 11
ThreadWeaver: Adaptive Threading for Efficient Parallel Reasoning in Language Models Paper β’ 2512.07843 β’ Published Nov 24, 2025 β’ 22