Dev
GitHub repos gaining traction - what high-signal users are starring and what's climbing the board, captured daily and enriched from GitHub. Raw material for spotting new tech and patterns worth building on.
1,506
repos tracked
261
surfaced this week
215
created < 30d
Python
top language
27 repos
-
Easily benchmark a Julia package over its commit history
-
ExploitGym is a large-scale, realistic benchmark built from real-world vulnerabilities designed to evaluate AI agents' ability to develop exploits.
-
FrontierSWE is an ultra long-horizon coding agent benchmark that tests implementation, performance eng and ML research
-
Easily benchmark a Julia package over its commit history
-
Same model, different wrapper: a from-scratch benchmark comparing coding-agent harnesses (codex, pi, opencode, cursor, devin) and open models on correctness, speed, and token cost
-
Solver benchmarking tools
-
A benchmark built to evaluate and improve agent capabilities for supporting legal work.
-
Evaluation harness for OpenHands V1.
-
OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks
-
Benchmarking Goal-Oriented Software Engineering
-
FlowWM stochastic world modeling via flow matching in DINOv3 feature space, with the FuturePerception (Waymo) benchmark.
-
A curated, non-BS library of the best resources for building and evaluating AI agents — papers, blogs, talks, tools, benchmarks. Maintained by BenchFlow.
-
τ-Bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
-
Benchmark environment for evaluating vision-language models (VLMs) on popular video games!
-
★ 123 prinz-ai/prinzbenchprinzbench is a private benchmark that ranks LLMs based on their ability to conduct legal research and analysis and locate obscure publicly available information online.
-
Agent memory for LLMs: 30 runnable Jupyter notebooks covering conversation buffers, vector stores, knowledge graphs, episodic and semantic memory, MemGPT, Mem0, Letta, Zep, Graphiti, LoCoMo benchmarks, and production patterns.
-
Benchmarking Chat Assistants on Long-Term Interactive Memory (ICLR 2025)
-
Official Repo: AutoResearchBench: Benchmarking AI Agents on Complex Scientific Literature Discovery
-
[NeurIPS 2025] The first web-based benchmark and platform to evaluate visual reasoning and interaction capabilities of MLLM powered agents through diverse and dynamic CAPTCHA puzzles.
-
Tools for various benchmarking scenarios of Weaviate's Query Agent.
-
Landing page + leaderboard for SWE-Bench benchmark
-
VirtueBench V2: Multi-dimensional virtue evaluation benchmark for LLMs with tripartite and Ignatian temptation models
-
A benchmark for evaluating AI agents on realistic business workflows
-
Scale computer-use 2.0 with open-source drivers, cross-OS fleets, and benchmarks for training, evaluation, and data generation.
-
Submission pipeline and results store for the lean-eval benchmark (https://github.com/leanprover/lean-eval)
-
Code for "Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly" (CVPR 2026)