Awesome Ai Agents 2026
Section: Local LLM Runners · High-throughput serving. PagedAttention. Production-grade.
Entry
Appears in 11 awesome lists
State-of-the-art serving engine with PagedAttention and continuous batching. Currently the fastest production-grade LLM server.
Section: Local LLM Runners · High-throughput serving. PagedAttention. Production-grade.
Section: Tools · High-throughput and memory-efficient inference library for LLMs.
Section: 推理 Inference · A high-throughput and memory-efficient inference and serving engine for LLMs.
Section: Large Model Serving · A high-throughput and memory-efficient inference and serving engine for LLMs.
Section: Inference engines · a high-throughput and memory-efficient inference and serving engine for LLMs
Section: Inference and hardware · LLM serving with supported low-bit kernels and quantized KV caches.
Section: Efficient and Small Language Models · PagedAttention-based high-throughput LM serving.
Section: 3. Inference Engines & Serving · State-of-the-art serving engine with PagedAttention and continuous batching. Currently the fastest production-grade LLM server.
Section: Deployment and Serving · vLLM is a high-throughput and memory-efficient inference and serving engine for LLMs.
Section: AI and Agents · A high-throughput and memory-efficient inference and serving engine for LLMs.
Section: Other · A high-throughput and memory-efficient inference and serving engine for LLMs
Langchain integrates various providers like Anthropic, AWS, and OpenAI, and offers tools for components such as LLMs, chat models, and data analysis, supporting functionalities from Alpha Vantage to YouTube github | docs
Unified proxy and SDK that routes to 100+ LLM providers behind a single OpenAI-compatible interface, with a Router handling retry/fallback across deployments, per-project cost and rate-limit tracking, and OTEL callback integrations. The right infrastructure layer when your harness needs provider…
Comet's open-source AI observability and evaluation platform: deep tracing of LLM calls, conversation logging, and agent activity, plus built-in eval metrics, prompt versioning, guardrails, and the Opik Agent Optimizer. Worth including because it unifies observability, verification, and…
(MIT) provides modules for structured outputs at different levels of abstraction, including output parsers for text completion endpoints, Pydantic programs for mapping prompts to structured outputs using function calling or output parsing, and pre-defined Pydantic programs for specific output types.
robot: The free, Open Source alternative to OpenAI, Claude and others. Self-hosted and local-first. Drop-in replacement for OpenAI, running on consumer-grade hardware. No GPU required. Runs gguf, transformers, diffusers and many more models architectures. Features: Generate Text, Audio, Video,…
(from Hpcaitech) - A Unified Deep Learning System for Large-Scale Parallel Training (1D, 2D, 2.5D, 3D and sequence parallelism, and ZeRO protocol).
Mem0 is an intelligent memory layer for Large Language Models that enhances personalized AI experiences by retaining and utilizing contextual information across various applications. github | website | docs | discord | twitter | github profile | linkedin
Open-source AI orchestration framework for building context-engineered, production-ready LLM applications. Design modular pipelines and agent workflows with explicit control over retrieval, routing, memory, and generation. Built for scalable agents, RAG, multimodal applications, semantic search,…