deqing/convergent-llama-300M-muon-6digit-addition_6digit_custom6 1B • Updated about 4 hours ago • 709
Value-Aware Stochastic KV Cache Eviction for Reasoning Models Paper • 2606.03928 • Published 2 days ago • 8
Value-Aware Stochastic KV Cache Eviction for Reasoning Models Paper • 2606.03928 • Published 2 days ago • 8
deqing/convergent-llama-300M-muon-6digit-addition_6digit_llama Text Generation • 0.3B • Updated 1 day ago • 1.3k • 1
deqing/convergent-llama-300M-muon-6digit-addition_6digit_custom3 Text Generation • 0.2B • Updated 3 days ago • 1.92k • 1
deqing/convergent-llama-300M-muon-base15-addition_base15 Text Generation • 0.2B • Updated 5 days ago • 853
deqing/convergent-llama-300M-muon-6digit-addition_6digit_llama Text Generation • 0.3B • Updated 1 day ago • 1.3k • 1
deqing/convergent-llama-300M-muon-base12-addition_base12 Text Generation • 0.2B • Updated 5 days ago • 920
deqing/convergent-llama-300M-muon-4digit-addition_4digit_custom3_right2left 0.2B • Updated 6 days ago • 120
deqing/convergent-llama-300M-muon-4digit-addition_4digit_custom3_right2left 0.2B • Updated 6 days ago • 120
deqing/convergent-llama-300M-muon-base15-addition_base15 Text Generation • 0.2B • Updated 5 days ago • 853
deqing/convergent-llama-300M-muon-base12-addition_base12 Text Generation • 0.2B • Updated 5 days ago • 920