Discrete Diffusion Models: A Unified Framework from Tokenization to Generation Paper • 2607.13431 • Published 5 days ago • 16
X-Lens: Real-Time Metric Depth Estimation with Heterogeneous Cameras Paper • 2607.12993 • Published 6 days ago • 131
Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable Paper • 2607.13285 • Published 6 days ago • 198
Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation Paper • 2607.11886 • Published 7 days ago • 80
The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning Paper • 2606.29526 • Published 22 days ago • 168
Drop-Then-Recovery: How Redundant Are Vision-Language-Action Models? Paper • 2606.27755 • Published 24 days ago • 6
Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models Paper • 2606.03988 • Published Jun 3 • 126
Crafter: A Multi-Agent Harness for Editable Scientific Figure Generation from Diverse Inputs Paper • 2605.30611 • Published May 28 • 252
OmniVerifier-M1: Multimodal Meta-Verifier with Explicit Structured Recalibration Paper • 2605.28805 • Published May 27 • 11
GenEvolve: Self-Evolving Image Generation Agents via Tool-Orchestrated Visual Experience Distillation Paper • 2605.21605 • Published May 20 • 15
Perception or Prejudice: Can MLLMs Go Beyond First Impressions of Personality? Paper • 2605.22109 • Published May 21 • 171
Multi-Objective and Mixed-Reward Reinforcement Learning via Reward-Decorrelated Policy Optimization Paper • 2605.13641 • Published May 13 • 51
ShotStream: Streaming Multi-Shot Video Generation for Interactive Storytelling Paper • 2603.25746 • Published Mar 26 • 155
SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models Paper • 2603.16859 • Published Mar 17 • 248