Dev
GitHub repos gaining traction - what high-signal users are starring and what's climbing the board, captured daily and enriched from GitHub. Raw material for spotting new tech and patterns worth building on.
2,235
repos tracked
175
surfaced this week
194
created < 30d
Python
top language
8 repos
-
Manipulating RLHF corpus to reduce sycophancy outcomes
-
Implement a reasoning LLM in PyTorch from scratch, step by step
-
RewardBench: the first evaluation tool for reward models.
-
Textbook on reinforcement learning from human feedback
-
An Easy-to-use, Scalable and High-performance Agentic RL Framework based on Ray (PPO & DAPO & REINFORCE++ & VLM & TIS & vLLM & Ray & Async RL)
-
RLHF中文手册 - 详细解析RLHF全流程优化阶段,涵盖指令调优、奖励模型训练,以及拒绝采样、强化学习和直接对齐算法等关键技术。
-
🚀 An open-source, hands-on curriculum bridging the gap from basic RL concepts to LLM alignment, RLVR, and advanced Agentic systems.
-
A curated collection of papers and resources on On-Policy Distillation for Large Language Models.