genpark-reinforcement-learning-math-reasoning-synthesizer-skill

源码

RL mathematical reasoning synthesizer & step-by-step theorem prover (DeepSeek-R1 style)

⭐ 8开发工具Python仓库 ↗更新于 2026-08-30收录于 2026-10-04

安装配置

查看源码仓库 →

相关服务器