Awesome local LLM
Section: Inference engines · official inference framework for 1-bit LLMs
Entry
Appears in 5 awesome lists
Official inference framework for 1-bit LLMs (BitNet b1.58). Enables running large models on CPU with minimal memory footprint. Features custom kernels for ternary weight quantization and efficient matmul operations. MIT licensed.
Section: Inference engines · official inference framework for 1-bit LLMs
Section: Inference and hardware · Inference framework for supported native low-bit BitNet models.
Section: 3. Inference Engines & Serving · Official inference framework for 1-bit LLMs (BitNet b1.58). Enables running large models on CPU with minimal memory footprint. Features custom kernels for ternary weight quantization and efficient matmul operations. MIT licensed.
Section: Developer tools · Official inference framework for 1-bit LLMs, by Microsoft. #opensource
Section: Other · Official inference framework for 1-bit LLMs
Ollama is a tool for running large language models locally, offering easy setup for macOS, Windows, Linux, and Docker, along with a library of models and quickstart guides for customization and integration github | github profile
State-of-the-art serving engine with PagedAttention and continuous batching. Currently the fastest production-grade LLM server.
Pure C/C++ inference engine with GGUF format support. The gold standard for CPU/GPU/Apple Silicon on-device running. Includes llama-server for OpenAI-compatible API. Now at 100K+ stars.
Production-grade platform for running any open-source LLMs as OpenAI-compatible API endpoints. Supports 50+ models with built-in streaming, batching, and auto-acceleration. Apache 2.0 licensed.
(MPL-2.0) allows specifying JSON schemas using regular expressions or Pydantic models for constrained decoding. Its high-performance runtime accelerates JSON decoding.
🟢🟠 — Apache-licensed AI and MCP gateway with multi-provider routing, virtual-key access controls, budgets, rate limits, MCP aggregation, OAuth, automatic fallbacks, and load balancing. (Maxim) — note: model-provider credentials and proxied request/response data are sensitive; restrict and…;…
Python-free Rust inference server with OpenAI API compatibility. Supports GGUF and SafeTensors formats with hot model swap, auto-discovery, and single binary deployment for zero-dependency inference. Apache 2.0 licensed.
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference…