Skip to content

Entry

vLLM

Appears in 11 awesome lists

State-of-the-art serving engine with PagedAttention and continuous batching. Currently the fastest production-grade LLM server.

Open github.comvllm-project/vllm

Found in these lists

Awesome Ai Agents 2026

Section: Local LLM Runners · High-throughput serving. PagedAttention. Production-grade.

ActiveScore 74

Awesome Data Analysis

Section: Tools · High-throughput and memory-efficient inference library for LLMs.

FreshScore 80

Awesome LLM Resources

Section: 推理 Inference · A high-throughput and memory-efficient inference and serving engine for LLMs.

FreshScore 87

Awesome LLMOps

Section: Large Model Serving · A high-throughput and memory-efficient inference and serving engine for LLMs.

ActiveScore 75

Awesome local LLM

Section: Inference engines · a high-throughput and memory-efficient inference and serving engine for LLMs

FreshScore 87

Awesome Model Quantization

Section: Inference and hardware · LLM serving with supported low-bit kernels and quantized KV caches.

FreshScore 83

awesome-nlp

Section: Efficient and Small Language Models · PagedAttention-based high-throughput LM serving.

FreshScore 90

Awesome Open Source AI

Section: 3. Inference Engines & Serving · State-of-the-art serving engine with PagedAttention and continuous batching. Currently the fastest production-grade LLM server.

FreshScore 89

Awesome Production Machine Learning

Section: Deployment and Serving · vLLM is a high-throughput and memory-efficient inference and serving engine for LLMs.

FreshScore 92

Awesome Python

Section: AI and Agents · A high-throughput and memory-efficient inference and serving engine for LLMs.

FreshScore 94

awesome-python

Section: Other · A high-throughput and memory-efficient inference and serving engine for LLMs

FreshScore 81

LangChain

Langchain integrates various providers like Anthropic, AWS, and OpenAI, and offers tools for components such as LLMs, chat models, and data analysis, supporting functionalities from Alpha Vantage to YouTube github | docs

In 20 listsDetails

LiteLLM

Unified proxy and SDK that routes to 100+ LLM providers behind a single OpenAI-compatible interface, with a Router handling retry/fallback across deployments, per-project cost and rate-limit tracking, and OTEL callback integrations. The right infrastructure layer when your harness needs provider…

In 16 listsDetails

Opik

Comet's open-source AI observability and evaluation platform: deep tracing of LLM calls, conversation logging, and agent activity, plus built-in eval metrics, prompt versioning, guardrails, and the Opik Agent Optimizer. Worth including because it unifies observability, verification, and…

In 16 listsDetails

LlamaIndex

(MIT) provides modules for structured outputs at different levels of abstraction, including output parsers for text completion endpoints, Pydantic programs for mapping prompts to structured outputs using function calling or output parsing, and pre-defined Pydantic programs for specific output types.

In 14 listsDetails

LocalAI

robot: The free, Open Source alternative to OpenAI, Claude and others. Self-hosted and local-first. Drop-in replacement for OpenAI, running on consumer-grade hardware. No GPU required. Runs gguf, transformers, diffusers and many more models architectures. Features: Generate Text, Audio, Video,…

In 14 listsDetails

Colossal-AI

(from Hpcaitech) - A Unified Deep Learning System for Large-Scale Parallel Training (1D, 2D, 2.5D, 3D and sequence parallelism, and ZeRO protocol).

In 14 listsDetails

Mem0

Mem0 is an intelligent memory layer for Large Language Models that enhances personalized AI experiences by retaining and utilizing contextual information across various applications. github | website | docs | discord | twitter | github profile | linkedin

In 13 listsDetails

Haystack

Open-source AI orchestration framework for building context-engineered, production-ready LLM applications. Design modular pipelines and agent workflows with explicit control over retrieval, routing, memory, and generation. Built for scalable agents, RAG, multimodal applications, semantic search,…

In 13 listsDetails