TransFG: A Transformer Architecture for Fine-grained Recognition Paper • 2103.07976 • Published Dec 1, 2021
Vibe Spaces for Creatively Connecting and Expressing Visual Concepts Paper • 2512.14884 • Published Dec 16, 2025 • 2
Pillar-0: A New Frontier for Radiology Foundation Models Paper • 2511.17803 • Published Nov 21, 2025 • 25
GRAID: Enhancing Spatial Reasoning of VLMs Through High-Fidelity Data Generation Paper • 2510.22118 • Published Oct 25, 2025
Transformers Discover Molecular Structure Without Graph Priors Paper • 2510.02259 • Published Oct 2, 2025 • 8
TARDIS STRIDE: A Spatio-Temporal Road Image Dataset for Exploration and Autonomy Paper • 2506.11302 • Published Jun 12, 2025 • 3
AlphaOne: Reasoning Models Thinking Slow and Fast at Test Time Paper • 2505.24863 • Published May 30, 2025 • 98
LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning Paper • 2406.11815 • Published Jun 17, 2024 • 1
Making Your First Choice: To Address Cold Start Problem in Vision Active Learning Paper • 2210.02442 • Published Oct 5, 2022 • 1
Delving into Masked Autoencoders for Multi-Label Thorax Disease Classification Paper • 2210.12843 • Published Oct 23, 2022
Masked Autoencoders Enable Efficient Knowledge Distillers Paper • 2208.12256 • Published Aug 25, 2022 • 1
Sequential Modeling Enables Scalable Learning for Large Vision Models Paper • 2312.00785 • Published Dec 1, 2023 • 1
Discovering Failure Modes of Text-guided Diffusion Models via Adversarial Search Paper • 2306.00974 • Published Jun 1, 2023