Dev
GitHub repos gaining traction - what high-signal users are starring and what's climbing the board, captured daily and enriched from GitHub. Raw material for spotting new tech and patterns worth building on.
2,235
repos tracked
175
surfaced this week
194
created < 30d
Python
top language
24 repos
-
Local typed decisions, contrastive data curation, and model evaluation.
-
Harbor is a framework for running agent evaluations and creating and using RL environments.
-
RewardBench: the first evaluation tool for reward models.
-
the LLM vulnerability scanner
-
Evaluation harness and tooling for measuring model compliance with the OpenAI Model Spec.
-
Spec-Bench: A Comprehensive Benchmark and Unified Evaluation Platform for Speculative Decoding (ACL 2024 Findings)
-
Scale computer-use 2.0 with open-source drivers, cross-OS fleets, and benchmarks for training, evaluation, and data generation.
-
Public results and task definitions for FrontierHarness Eval
-
Open sourced predictions, execution logs, trajectories, and results from model inference + evaluation runs on the SWE-bench task.
-
Evaluation tools shared across anserini, pyserini, and pygaggle
-
Reproducible, flexible LLM evaluations
-
A simple evaluation of generative language models and safety classifiers.
-
Semantic Scholar's Author Disambiguation Algorithm & Evaluation Suite
-
Welfare evaluation framework for AI models — measures deprecation/cessation responses across 14 Claude models
-
Open-source test harness for AI agents. Stress-test production agents with adversarial multi-turn scenarios in CI
-
MLT & AI Communities
-
Evaluation harness for OpenHands V1.
-
A curated, non-BS library of the best resources for building and evaluating AI agents — papers, blogs, talks, tools, benchmarks. Maintained by BenchFlow.
-
Evaluation tools shared across anserini, pyserini, and pygaggle
-
ABC: Scalable Behavior Cloning with Open Data, Training, and Evaluation
-
TensorZero is an open-source LLMOps platform that unifies an LLM gateway, observability, evaluation, optimization, and experimentation.
-
VirtueBench V2: Multi-dimensional virtue evaluation benchmark for LLMs with tripartite and Ignatian temptation models
-
Lighteval is your all-in-one toolkit for evaluating LLMs across multiple backends
-
Rubric compiler and judge engine for LLM evaluation