Dev
GitHub repos gaining traction - what high-signal users are starring and what's climbing the board, captured daily and enriched from GitHub. Raw material for spotting new tech and patterns worth building on.
2,235
repos tracked
175
surfaced this week
194
created < 30d
Python
top language
48 repos
-
An open-source execution engine for AI-SQL and LLM-powered dataflow
-
👻 Ghostty is a fast, feature-rich, and cross-platform terminal emulator that uses platform-native UI and GPU acceleration.
-
Open Machine Learning Compiler Framework
-
Official Problem Sets / Reference Kernels for the GPU MODE Leaderboard!
-
JAX in JavaScript – ML library for the web, running on WebGPU & Wasm
-
Official CLI and Python SDK for Prime Intellect - access GPU compute, remote sandboxes, RL environments, and distributed training infrastructure for AI development at scale.
-
Open ABI and FFI for Machine Learning Systems
-
Open Source Continuous Inference Benchmark Research Platform — Kimi K3 2.8T, MiniMax M3, DeepSeekv4, GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72 & soon™ TPUv6e/v7/Trainium2/3 | 开源持续推理基准研究平台 — Kimi K2.7-Code、MiniMax M3、DeepSeekv4、GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72,即将推出™ TPUv6e/v7/Trainium2/3
-
Winner 🏆 (Agent-only) MLSys 2026 - FlashInfer AI Kernel Generation Contest for the DeepSeek Sparse Attention (DSA) track with an average speedup of 34.93x
-
Composable transformations of Python+NumPy programs: differentiate, vectorize, JIT to GPU/TPU, and more
-
Composable transformations of Python+NumPy programs: differentiate, vectorize, JIT to GPU/TPU, and more
-
FlashInfer: Kernel Library for LLM Serving
-
RTX 6000 Pro Wiki — Running Large LLMs (Qwen3.5-397B, Kimi-K2.5, GLM-5) on PCIe GPUs without NVLink
-
DeepGEMM: clean and efficient BLAS kernel library on GPU
-
If you live in the terminal, kitty is made for you! Cross-platform, fast, feature-rich, GPU based.
-
UI components for Omarchy system style
-
The fastai deep learning library
-
Numerical differential equation solvers in JAX. Autodifferentiable and GPU-capable. https://docs.kidger.site/diffrax/
-
Powerful web graphics runtime built on WebGL, WebGPU, WebXR and glTF
-
Codex skill for GPU MODE YouTube thumbnails
-
NVSentinel detects and remediates GPU faults on Kubernetes nodes
-
GPU-optimized version of the MuJoCo physics simulator, designed for NVIDIA hardware.
-
The AI Compute Platform for frontier teams. SkyPilot turns fragmented AI compute into one AI supercomputer, so frontier AI teams build custom intelligence faster.
-
WebGPU Samples
-
AI-powered video frame extraction tool that automatically identifies and extracts high-quality frames containing people, with intelligent pose categorization (standing/sitting/squatting), head orientation detection, and shot type classification. Features GPU acceleration, resumable processing, and extensive configuration options.
-
⚡️A native, local-first alternative to Logitech Options+, written in Rust 🦀 — remap buttons, DPI, and SmartShift over HID++. No account, no telemetry.
-
OCR with VLMs running on a single GPU with high throughput
-
Tensors and Dynamic neural networks in Python with strong GPU acceleration
-
kernelboard is the webapp for https://www.gpumode.com
-
👻 Ghostty is a fast, feature-rich, and cross-platform terminal emulator that uses platform-native UI and GPU acceleration.
-
Tensors and Dynamic neural networks in Python with strong GPU acceleration
-
Command-line utility for monitoring GPU hardware.
-
Write a fast kernel and see how you compare against the best humans and AI on gpumode.com
-
Apple Silicon (Metal) backend for Triton: write standard @triton.jit kernels on your Mac GPU. The same source runs bit-identical on NVIDIA and AMD, so you develop kernel logic locally and rent a datacenter GPU only for the perf pass.
-
Machine Learning Engineering Open Book
-
DiRe-RAPIDS + EVoC integration (Igor Rivin): GPU Phase-2 hybrid clustering, intrinsic dimension from the kNN graph, out-of-sample predict
-
Code at the speed of thought – Zed is a high-performance, multiplayer code editor from the creators of Atom and Tree-sitter.
-
High-Performance Cross-Platform Monte Carlo Renderer Based on LuisaCompute
-
A Python DSL to write Nvidia PTX for Hopper and Blackwell in JAX and PyTorch
-
UCCL is an efficient communication library for GPUs, covering collectives, P2P (e.g., KV cache transfer, RL weight transfer), and EP (e.g., GPU-driven)
-
AI agents running research on single-GPU nanochat training automatically
-
A tutorial on modern GPU programming for machine learning systems
-
Fork of Apache MXNet 2.0 — runs legacy MXNet code on CUDA 13 / Blackwell (sm_120) GPUs and native Apple Silicon CPU.
-
Rust & GPUI Native Agent CLIs manager for macOS. Ghostty Terminals + Codex App Features/UX = Ghostex! Embedded browser & IDE. Tons of useful features.
-
Production-Grade Autoresearch. Ideal for agent harness engineering, prompt engineering, ML model development, GPU kernels, and other optimizable code.