Dev
GitHub repos gaining traction - what high-signal users are starring and what's climbing the board, captured daily and enriched from GitHub. Raw material for spotting new tech and patterns worth building on.
2,235
repos tracked
175
surfaced this week
194
created < 30d
Python
top language
24 repos
-
Our library for RL environments + evals
-
Homework assignments for Evaluating and Improving AI Agents
-
Evaluation harness and tooling for measuring model compliance with the OpenAI Model Spec.
-
cotrust.ai
-
Our library for RL environments + evals
-
Public results and task definitions for FrontierHarness Eval
-
Code, Build and Evaluate agents - excellent Model and Skills/MCP/ACP/A2A Support
-
Give your coding agent the power to write and run agent evals.
-
Our library for RL environments + evals
-
Automatic evals for LLMs
-
Workspace manager for coding agents. Interactively solve and develop Harbor tasks.
-
Same model, different wrapper: a from-scratch benchmark comparing coding-agent harnesses (codex, pi, opencode, cursor, devin) and open models on correctness, speed, and token cost
-
Open-source test harness for AI agents. Stress-test production agents with adversarial multi-turn scenarios in CI
-
Framework for evaluating and improving agents
-
Realistic examples of building evals and optimizing agents with Harbor
-
A curated, non-BS library of the best resources for building and evaluating AI agents — papers, blogs, talks, tools, benchmarks. Maintained by BenchFlow.
-
Skills that guide AI coding agents to help you build product-specific AI evals.
-
Run Inspect AI evals in the cloud
-
Readymade evaluators for agent trajectories
-
Readymade evaluators for your LLM apps
-
A benchmark for evaluating AI agents on realistic business workflows