Dev
GitHub repos gaining traction - what high-signal users are starring and what's climbing the board, captured daily and enriched from GitHub. Raw material for spotting new tech and patterns worth building on.
2,235
repos tracked
175
surfaced this week
194
created < 30d
Python
top language
19 repos
-
A high-throughput and memory-efficient inference and serving engine for LLMs
-
SGLang is a high-performance serving framework for large language models and multimodal models.
-
Official Problem Sets / Reference Kernels for the GPU MODE Leaderboard!
-
Open Source Continuous Inference Benchmark Research Platform — Kimi K3 2.8T, MiniMax M3, DeepSeekv4, GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72 & soon™ TPUv6e/v7/Trainium2/3 | 开源持续推理基准研究平台 — Kimi K2.7-Code、MiniMax M3、DeepSeekv4、GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72,即将推出™ TPUv6e/v7/Trainium2/3
-
DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm
-
FlashInfer: Kernel Library for LLM Serving
-
A retargetable MLIR-based machine learning compiler and runtime toolkit.
-
CUDA Craftax-Classic: 7x faster RL training than JAX
-
A flyweight in situ visualization and analysis runtime for multi-physics HPC simulations
-
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
-
PyTorch native quantization for training and inference
-
High-Performance Cross-Platform Monte Carlo Renderer Based on LuisaCompute
-
UCCL is an efficient communication library for GPUs, covering collectives, P2P (e.g., KV cache transfer, RL weight transfer), and EP (e.g., GPU-driven)
-
The open-source AI voice studio. Clone, dictate, create.
-
Build configuration for PyCUDA
-
Fork of Apache MXNet 2.0 — runs legacy MXNet code on CUDA 13 / Blackwell (sm_120) GPUs and native Apple Silicon CPU.
-
Throwaway staging: validate CUDA 13.3 prebuilt build leg