-
Efficient RL Training for LLMs with Experience Replay
Paper • 2604.08706 • Published • 23 -
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO
Paper • 2605.30789 • Published • 26 -
Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning
Paper • 2607.18722 • Published • 36
Chintu Kumar
chang2394
AI & ML interests
None yet
Recent Activity
updated a collection 17 days ago
Off policy/entropy updated a collection about 2 months ago
Off policy/entropy updated a collection 8 months ago
Tool useOrganizations
None yet