Dev
GitHub repos gaining traction - what high-signal users are starring and what's climbing the board, captured daily and enriched from GitHub. Raw material for spotting new tech and patterns worth building on.
2,235
repos tracked
175
surfaced this week
194
created < 30d
Python
top language
56 repos
-
Proxy Policy Steering: inference-time adaptation for robotics foundation models
-
Batched single-token choice inference for open language models, compatible with TypeSafe
-
A high-throughput and memory-efficient inference and serving engine for LLMs
-
LLM inference in C/C++
-
Tree Decision Diagrams for Boolean functions, model counting, and probabilistic inference in Rust
-
An open-source execution engine for AI-SQL and LLM-powered dataflow
-
SGLang is a high-performance serving framework for large language models and multimodal models.
-
Programmable chat templates for LLM training and inference.
-
The largest collection of PyTorch image encoders / backbones. Including train, eval, inference, export scripts, and pretrained weights -- ResNet, ResNeXT, EfficientNet, NFNet, Vision Transformer (ViT), MobileNetV4, MobileNet-V3 & V2, RegNet, DPN, CSPNet, Swin Transformer, MaxViT, CoAtNet, ConvNeXt, and more
-
Implement a reasoning LLM in PyTorch from scratch, step by step
-
A high-throughput and memory-efficient inference and serving engine for LLMs
-
Open Source Continuous Inference Benchmark Research Platform — Kimi K3 2.8T, MiniMax M3, DeepSeekv4, GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72 & soon™ TPUv6e/v7/Trainium2/3 | 开源持续推理基准研究平台 — Kimi K2.7-Code、MiniMax M3、DeepSeekv4、GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72,即将推出™ TPUv6e/v7/Trainium2/3
-
DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm
-
A high-throughput and memory-efficient inference and serving engine for LLMs
-
Accurate, large-scale, and extensible simulator for LLM inference Systems
-
FlashInfer: Kernel Library for LLM Serving
-
RTX 6000 Pro Wiki — Running Large LLMs (Qwen3.5-397B, Kimi-K2.5, GLM-5) on PCIe GPUs without NVLink
-
Port of OpenAI's Whisper model in C/C++
-
Turn your PC, Mac, or Linux box into an AI server. LLM inference, chat UI, voice, agents, workflows, RAG, and image generation.
-
Use Hugging Face with JavaScript
-
Compile natural-language function descriptions into reusable local PAW programs.
-
FreeToken brings datacenter-scale model serving to your desktop. Run massive models locally, fast and efficiently.
-
Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
-
Open sourced predictions, execution logs, trajectories, and results from model inference + evaluation runs on the SWE-bench task.
-
Continual learning infra for self-improving agents
-
Embarrassingly parallel fasttext batch inference in Rust.
-
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
-
Sub-real-time MiniMax H3 on Hopper: 13.506 s for a 14.375 s 768p video with audio on 8xH100. Four runtime patches removing 9.72 s of non-model overhead from FastVideo.
-
Native LLM inference server for Apple Silicon. OpenAI + Anthropic API compatible. No Python. Includes MLX Core macOS app with chat, agent mode, and tool calling.
-
dInfer: An Efficient Inference Framework for Diffusion Language Models
-
A Simple Transport Layer For Teleoperation And Inference
-
ggml speech-to-text inference for 16+ model families
-
A Datacenter Scale Distributed Inference Serving Framework
-
LLM inference in C/C++
-
LLM inference in C/C++
-
Machine Learning Engineering Open Book
-
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
-
1D dilated causal convolutions with extreme caching for 5µs inference on an FPGA
-
PyTorch native quantization for training and inference
-
Quantization, kernels, runtime and inference engine for mobiles, wearables, smart home and robots.
-
Home for "How To Scale Your Model", a short blog-style textbook about scaling LLMs on TPUs
-
Diffusion model(SD,Flux,Wan,Qwen Image,Z-Image,...) inference in pure C/C++
-
The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider.
-
Noumena's internal version of your favorite coding tui but improved and optimized for our inference stack
-
Office inference code for World Tracing (object/scene/dynamic). Live demos: https://haoz19.github.io/world-tracing-page/
-
Arctic Training and Inference Platform
-
The new data-free filesystem!
-
Independent reliability monitoring for AI inference providers: Fireworks, Together, and Baseten
-
high-performance inference and serving library for interactive autoregressive video and world models
-
LLM inference in C/C++