Dev
GitHub repos gaining traction - what high-signal users are starring and what's climbing the board, captured daily and enriched from GitHub. Raw material for spotting new tech and patterns worth building on.
2,235
repos tracked
175
surfaced this week
194
created < 30d
Python
top language
83 repos
-
Our library for RL environments + evals
-
Pyserini is a Python toolkit for reproducible information retrieval research with sparse and dense representations.
-
The largest collection of PyTorch image encoders / backbones. Including train, eval, inference, export scripts, and pretrained weights -- ResNet, ResNeXT, EfficientNet, NFNet, Vision Transformer (ViT), MobileNetV4, MobileNet-V3 & V2, RegNet, DPN, CSPNet, Swin Transformer, MaxViT, CoAtNet, ConvNeXt, and more
-
Local typed decisions, contrastive data curation, and model evaluation.
-
Anserini is a Lucene toolkit for reproducible information retrieval research
-
Harbor is a framework for running agent evaluations and creating and using RL environments.
-
Homework assignments for Evaluating and Improving AI Agents
-
Terminal-Bench-Science: Evaluating AI agents on research workflows across scientific domains
-
RewardBench: the first evaluation tool for reward models.
-
the LLM vulnerability scanner
-
Evaluation harness and tooling for measuring model compliance with the OpenAI Model Spec.
-
Spec-Bench: A Comprehensive Benchmark and Unified Evaluation Platform for Speculative Decoding (ACL 2024 Findings)
-
cotrust.ai
-
Our library for RL environments + evals
-
Scale computer-use 2.0 with open-source drivers, cross-OS fleets, and benchmarks for training, evaluation, and data generation.
-
Proofs for Erlang/Elixir programs through LEAN (via automatic translation of Core Erlang)
-
The ultimate RAG for your monorepo. Query, understand, and edit multi-language codebases with the power of AI and knowledge graphs
-
Implementation for AutoIndex: Learning Representation Programs for Retrieval
-
Public results and task definitions for FrontierHarness Eval
-
RankLLM is a Python toolkit for reproducible information retrieval research using rerankers, with a focus on listwise reranking.
-
Code, Build and Evaluate agents - excellent Model and Skills/MCP/ACP/A2A Support
-
Open sourced predictions, execution logs, trajectories, and results from model inference + evaluation runs on the SWE-bench task.
-
A benchmark for evaluating AI agents on realistic business workflows
-
Give your coding agent the power to write and run agent evals.
-
Ridiculously fast symbolic expressions
-
~95% on SimpleQA (e.g. Qwen3.6-27B on a 3090). Supports all local and cloud LLMs (llama.cpp, Ollama, Google, ...). 10+ search engines - arXiv, PubMed, your private documents. Everything Local & Encrypted.
-
A toy eval suite for tracing generalization dynamics of LM pre-training
-
Evaluation tools shared across anserini, pyserini, and pygaggle
-
ExploitGym is a large-scale, realistic benchmark built from real-world vulnerabilities designed to evaluate AI agents' ability to develop exploits.
-
Our library for RL environments + evals
-
Reproducible, flexible LLM evaluations
-
Automatic evals for LLMs
-
A simple evaluation of generative language models and safety classifiers.
-
Semantic Scholar's Author Disambiguation Algorithm & Evaluation Suite
-
Workspace manager for coding agents. Interactively solve and develop Harbor tasks.
-
Same model, different wrapper: a from-scratch benchmark comparing coding-agent harnesses (codex, pi, opencode, cursor, devin) and open models on correctness, speed, and token cost
-
Starter Kit for Auto-Judge developers
-
Welfare evaluation framework for AI models — measures deprecation/cessation responses across 14 Claude models
-
Open-source test harness for AI agents. Stress-test production agents with adversarial multi-turn scenarios in CI
-
A benchmark built to evaluate and improve agent capabilities for supporting legal work.
-
MLT & AI Communities
-
Evaluation harness for OpenHands V1.
-
[EMNLP 2026] A counterfactual chart generation pipeline for evaluating chart reasoning in vision-language models.
-
Framework for evaluating and improving agents
-
Realistic examples of building evals and optimizing agents with Harbor
-
Official eval scripts for JobBench