Dev
GitHub repos gaining traction - what high-signal users are starring and what's climbing the board, captured daily and enriched from GitHub. Raw material for spotting new tech and patterns worth building on.
2,235
repos tracked
175
surfaced this week
194
created < 30d
Python
top language
53 repos
-
Local typed decisions, contrastive data curation, and model evaluation.
-
Harbor is a framework for running agent evaluations and creating and using RL environments.
-
Homework assignments for Evaluating and Improving AI Agents
-
Terminal-Bench-Science: Evaluating AI agents on research workflows across scientific domains
-
RewardBench: the first evaluation tool for reward models.
-
the LLM vulnerability scanner
-
Evaluation harness and tooling for measuring model compliance with the OpenAI Model Spec.
-
Spec-Bench: A Comprehensive Benchmark and Unified Evaluation Platform for Speculative Decoding (ACL 2024 Findings)
-
Scale computer-use 2.0 with open-source drivers, cross-OS fleets, and benchmarks for training, evaluation, and data generation.
-
Public results and task definitions for FrontierHarness Eval
-
Code, Build and Evaluate agents - excellent Model and Skills/MCP/ACP/A2A Support
-
Open sourced predictions, execution logs, trajectories, and results from model inference + evaluation runs on the SWE-bench task.
-
A benchmark for evaluating AI agents on realistic business workflows
-
Ridiculously fast symbolic expressions
-
Evaluation tools shared across anserini, pyserini, and pygaggle
-
ExploitGym is a large-scale, realistic benchmark built from real-world vulnerabilities designed to evaluate AI agents' ability to develop exploits.
-
Wrap Antigravity, ChatGPT Codex, Claude Code, Grok Build as an OpenAI/Gemini/Claude/Codex compatible API service, allowing you to enjoy the free Gemini 3.1 Pro, GPT 5.6 Series, Grok 4.5, Claude model through API
-
Reproducible, flexible LLM evaluations
-
A simple evaluation of generative language models and safety classifiers.
-
Semantic Scholar's Author Disambiguation Algorithm & Evaluation Suite
-
Starter Kit for Auto-Judge developers
-
Welfare evaluation framework for AI models — measures deprecation/cessation responses across 14 Claude models
-
Open-source test harness for AI agents. Stress-test production agents with adversarial multi-turn scenarios in CI
-
A benchmark built to evaluate and improve agent capabilities for supporting legal work.
-
MLT & AI Communities
-
Evaluation harness for OpenHands V1.
-
[EMNLP 2026] A counterfactual chart generation pipeline for evaluating chart reasoning in vision-language models.
-
Framework for evaluating and improving agents
-
SWE-Together: Evaluating Coding Agents in Interactive User Sessions
-
Testbed to evaluate dense supervision methods.
-
An extensive math library for JavaScript and Node.js
-
DeepSpec: a full-stack codebase for training and evaluating speculative decoding algorithms
-
A curated, non-BS library of the best resources for building and evaluating AI agents — papers, blogs, talks, tools, benchmarks. Maintained by BenchFlow.
-
Improving and Evaluating Hand-Object Interaction Detection
-
Benchmark environment for evaluating vision-language models (VLMs) on popular video games!
-
Vim-fork focused on extensibility and usability
-
Open source code for ICLR 2026 Paper: Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions
-
[NeurIPS 2025] The first web-based benchmark and platform to evaluate visual reasoning and interaction capabilities of MLLM powered agents through diverse and dynamic CAPTCHA puzzles.
-
Rojo enables Roblox developers to use professional-grade software engineering tools
-
Evaluation tools shared across anserini, pyserini, and pygaggle
-
ABC: Scalable Behavior Cloning with Open Data, Training, and Evaluation
-
A public-domain dataset of prompts and scenarios for evaluating compliance with the OpenAI Model Spec.
-
Readymade evaluators for agent trajectories
-
Readymade evaluators for your LLM apps
-
TensorZero is an open-source LLMOps platform that unifies an LLM gateway, observability, evaluation, optimization, and experimentation.
-
[ICML '26] Code repo for the paper entitled "Convex Dataset Valuation for Post-Training" at ICML 2026.
-
VirtueBench V2: Multi-dimensional virtue evaluation benchmark for LLMs with tripartite and Ignatian temptation models
-
A benchmark for evaluating AI agents on realistic business workflows
-
Lighteval is your all-in-one toolkit for evaluating LLMs across multiple backends