Direct Preference Optimization: Your Language Model is Secretly a Reward Model Paper • 2305.18290 • Published May 29, 2023 • 68
My models: daily driver rotation Collection A rotating list of models I created and currently use as daily drivers. From my many models, these are the ones I’m actively using. • 7 items • Updated 3 days ago • 10
Ultralytics YOLO26: Unified Real-Time End-to-End Vision Models Paper • 2606.03748 • Published Jun 2 • 19
SafeDiffusion-R1: Online Reward Steering for Safe Diffusion Post-Training Paper • 2605.18719 • Published May 18 • 7
view article Article Chitos: From Detection to Proof — An Autonomous Security AI That Actually Exploits FINAL-Bench • 21 days ago • 19
Refusal in Language Models Is Mediated by a Single Direction Paper • 2406.11717 • Published Jun 17, 2024 • 15
Safety Tax: Safety Alignment Makes Your Large Reasoning Models Less Reasonable Paper • 2503.00555 • Published Mar 1, 2025 • 3
Gemma 4 — DECKARD HERETIC, Multimodal & Speculators Collection Gemma 4 abliterated/quantized — DECKARD HERETIC 31B, SuperGemma4-26B multimodal, 26B-A4B MoE, plus EAGLE3/DFlash drafters. • 14 items • Updated 18 days ago • 8