When LLM Agents Meet Reinforcement Learning
Section: Base Framework · Tsinghua University (THUDM)
Entry
Appears in 5 awesome lists
LLM post-training framework for RL Scaling from THUDM. Supports SFT and RL training with multi-turn compilation feedback, powering projects like TritonForge for automated GPU kernel generation. Apache 2.0 licensed.
Section: Base Framework · Tsinghua University (THUDM)
Section: Agentic RL
Section: Training and Fine-tuning · an LLM post-training framework for RL Scaling
Section: 7. Training & Fine-tuning Ecosystem · LLM post-training framework for RL Scaling from THUDM. Supports SFT and RL training with multi-turn compilation feedback, powering projects like TritonForge for automated GPU kernel generation. Apache 2.0 licensed.
Section: Industry Strength Reinforcement Learning · slime is an LLM post-training framework for RL Scaling.
Build, Evaluate, and Optimize AI Systems. Includes evals, RAG, agents, fine-tuning, synthetic data generation, dataset management, MCP, and more.
an easy-to-use, high-performance open-source RLHF framework built on Ray, vLLM, ZeRO-3 and HuggingFace Transformers, designed to make RLHF training simple and accessible
Scalable open-source RL infrastructure for post-training foundation models via reinforcement learning. Features M2Flow paradigm for embodied AI and agentic workflows with real-world robotics integrations. Apache 2.0 licensed.
a Python library for using and training embedding and reranker models for applications like retrieval augmented generation, semantic search, and more
Open-source system for automatic censorship and robustness suppression removal in language-model outputs.
Agentic RL Training at Scale from Prime Intellect. Framework for large-scale reinforcement learning capable of scaling to 1000+ GPUs with fully asynchronous RL, FSDP2 training, and vLLM inference. Apache 2.0 licensed.
Volcano Engine Reinforcement Learning for LLMs with PPO, GRPO, REINFORCE++, DAPO (EuroSys 2025).