Dev
GitHub repos gaining traction - what high-signal users are starring and what's climbing the board, captured daily and enriched from GitHub. Raw material for spotting new tech and patterns worth building on.
1,506
repos tracked
261
surfaced this week
215
created < 30d
Python
top language
36 repos
-
ggml speech-to-text inference for 16+ model families
-
SGLang is a high-performance serving framework for large language models and multimodal models.
-
A Datacenter Scale Distributed Inference Serving Framework
-
LLM inference in C/C++
-
LLM inference in C/C++
-
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
-
Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
-
A high-throughput and memory-efficient inference and serving engine for LLMs
-
Machine Learning Engineering Open Book
-
LLM inference in C/C++
-
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
-
Use Hugging Face with JavaScript
-
1D dilated causal convolutions with extreme caching for 5µs inference on an FPGA
-
Programmable chat templates for LLM training and inference.
-
PyTorch native quantization and sparsity for training and inference
-
The largest collection of PyTorch image encoders / backbones. Including train, eval, inference, export scripts, and pretrained weights -- ResNet, ResNeXT, EfficientNet, NFNet, Vision Transformer (ViT), MobileNetV4, MobileNet-V3 & V2, RegNet, DPN, CSPNet, Swin Transformer, MaxViT, CoAtNet, ConvNeXt, and more
-
Tiny AI for tiny devices
-
Home for "How To Scale Your Model", a short blog-style textbook about scaling LLMs on TPUs
-
Diffusion model(SD,Flux,Wan,Qwen Image,Z-Image,...) inference in pure C/C++
-
Port of OpenAI's Whisper model in C/C++
-
The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider.
-
Noumena's internal version of your favorite coding tui but improved and optimized for our inference stack
-
Office inference code for World Tracing (object/scene/dynamic). Live demos: https://haoz19.github.io/world-tracing-page/
-
DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm
-
Arctic Training and Inference Platform
-
The new data-free filesystem!
-
Independent reliability monitoring for AI inference providers: Fireworks, Together, and Baseten
-
high-performance inference and serving library for interactive autoregressive video and world models
-
LLM inference in C/C++
-
LLM inference in C/C++
-
LLM inference in C/C++
-
Implement a reasoning LLM in PyTorch from scratch, step by step
-
LLM inference in C/C++
-
KVarN is a native vLLM KV-cache quantization backend for your agents: 3-5x more context, throughput above FP16, and FP16-level accuracy. Calibration-free, one flag.
-
Fast LLM inference with Elixir and Bumblebee
-
FastCrest Tether: the OSS edge-to-cloud AI deploy CLI. Optimize, verify, deploy across Jetson, RTX, Apple Silicon, AMD. Hybrid edge-cloud inference with parity certs.