penfever/a3-rl-SankalpKJ_swesmith-oracle-filtered-40-8B Reinforcement Learning • 8B • Updated Jun 11 • 261
penfever/kimi-k2-swesmith_with_plain_docker-sandboxes-maxeps-32k Text Generation • 308k • Updated Apr 13 • 9
penfever/GLM-4_6-gemini25flash-stackexchange-overflow-32ep-512k-fixeps Text Generation • 308k • Updated Apr 12 • 10
penfever/rl_rl-conf_24GP_base-yaml_mode-path_r2eg-nl2b-stac-bugs-fixt_trai-data_exp_rpt_soft-v2-45 8B • Updated Feb 24 • 2
penfever/rl_rl-conf_24GP_base_noth-yaml_mode-path_r2eg-nl2b-stac-bugs_trai-data_exp_rpt_stac-bash-110 8B • Updated Feb 24 • 3
penfever/rl_rl-conf_20GP_base-yaml_mode-path_r2eg-nl2b-stac-bugs-fixt_trai-data_exp_rpt_stac-pyte-v2-25 8B • Updated Feb 24 • 2
penfever/rl_rl-conf_20GP_base-yaml_mode-path_r2eg-nl2b-stac-bugs-fixt_trai-data_exp_rpt_code-v2-25 8B • Updated Feb 24 • 3