Dev
GitHub repos gaining traction - what high-signal users are starring and what's climbing the board, captured daily and enriched from GitHub. Raw material for spotting new tech and patterns worth building on.
2,235
repos tracked
175
surfaced this week
194
created < 30d
Python
top language
44 repos
-
Raw results of solver benchmark campaigns run with bodono/solver_benchmarks
-
Solver benchmarking tools
-
Benchmarks scikit-learn-compatible machine learning libraries
-
Physical Analysis: embodied AI benchmarks and aggregate results
-
Open Source Continuous Inference Benchmark Research Platform — Kimi K3 2.8T, MiniMax M3, DeepSeekv4, GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72 & soon™ TPUv6e/v7/Trainium2/3 | 开源持续推理基准研究平台 — Kimi K2.7-Code、MiniMax M3、DeepSeekv4、GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72,即将推出™ TPUv6e/v7/Trainium2/3
-
A curated collection of papers, research blogs, open-source tools, benchmarks, and community demos for robot-use agents, including demos powered by GPT-6 Astra. —— Explore our searchable website.
-
Spec-Bench: A Comprehensive Benchmark and Unified Evaluation Platform for Speculative Decoding (ACL 2024 Findings)
-
OpenMMLab Detection Toolbox and Benchmark
-
Autonomous research agents that co-develop benchmark tracked repos
-
Scale computer-use 2.0 with open-source drivers, cross-OS fleets, and benchmarks for training, evaluation, and data generation.
-
AM-Bench: a modular simulation suite and benchmark for aerial manipulation policy learning (CoRL 2026)
-
Offline LLM-judged arena for embedding models
-
Public results and task definitions for FrontierHarness Eval
-
Official Code Repository for the ICLR 2026 Paper "Human Behavior Atlas: Benchmarking Unified Psychological and Social Behavior Understanding" & ICML 2026 Paper "Omnisapiens: A foundation model for social behavior processing via heterogeneity-aware relative policy optimization"
-
CommerceAgentBench: Benchmarking Long-Horizon Agents in High-Fidelity, Stateful, and Reproducible Replicas of Real Online Services
-
An MCP server that lets LLM agents play Civilization VI.
-
The agent benchmark that scores the full stack — harness, config, and model — not just the LLM. Trace-based scoring, reliability metrics, configuration diagnostics.
-
A benchmark for evaluating AI agents on realistic business workflows
-
ExploitGym is a large-scale, realistic benchmark built from real-world vulnerabilities designed to evaluate AI agents' ability to develop exploits.
-
FrontierSWE is an ultra long-horizon coding agent benchmark that tests implementation, performance eng and ML research
-
Easily benchmark a Julia package over its commit history
-
Easily benchmark a Julia package over its commit history
-
Same model, different wrapper: a from-scratch benchmark comparing coding-agent harnesses (codex, pi, opencode, cursor, devin) and open models on correctness, speed, and token cost
-
A benchmark built to evaluate and improve agent capabilities for supporting legal work.
-
Evaluation harness for OpenHands V1.
-
OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks
-
Benchmarking Goal-Oriented Software Engineering
-
FlowWM stochastic world modeling via flow matching in DINOv3 feature space, with the FuturePerception (Waymo) benchmark.
-
A curated, non-BS library of the best resources for building and evaluating AI agents — papers, blogs, talks, tools, benchmarks. Maintained by BenchFlow.
-
τ-Bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
-
Benchmark environment for evaluating vision-language models (VLMs) on popular video games!
-
Project page for Planning with the Views: ViewAgent + ViewSuite, a 6-DoF view planning benchmark for VLM agents
-
★ 132 prinz-ai/prinzbenchprinzbench is a private benchmark that ranks LLMs based on their ability to conduct legal research and analysis and locate obscure publicly available information online.
-
Agent memory for LLMs: 30 runnable Jupyter notebooks covering conversation buffers, vector stores, knowledge graphs, episodic and semantic memory, MemGPT, Mem0, Letta, Zep, Graphiti, LoCoMo benchmarks, and production patterns.
-
Benchmarking Chat Assistants on Long-Term Interactive Memory (ICLR 2025)
-
Official Repo: AutoResearchBench: Benchmarking AI Agents on Complex Scientific Literature Discovery
-
[NeurIPS 2025] The first web-based benchmark and platform to evaluate visual reasoning and interaction capabilities of MLLM powered agents through diverse and dynamic CAPTCHA puzzles.
-
Tools for various benchmarking scenarios of Weaviate's Query Agent.
-
Landing page + leaderboard for SWE-Bench benchmark
-
VirtueBench V2: Multi-dimensional virtue evaluation benchmark for LLMs with tripartite and Ignatian temptation models
-
A benchmark for evaluating AI agents on realistic business workflows
-
Submission pipeline and results store for the lean-eval benchmark (https://github.com/leanprover/lean-eval)
-
Code for "Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly" (CVPR 2026)