Awesome Ai Agents 2026
Section: Tracing and Monitoring · OSS AI observability. Traces, evals, embeddings.
Entry
Appears in 10 awesome lists
Open-source AI observability & evaluation platform (Arize) — OpenTelemetry-native tracing for agents, LLM-as-judge evals, versioned datasets & experiments for prompt regression testing, prompt management with version control and replay, plus an MCP endpoint so Claude Code/Cursor can query traces…
Section: Tracing and Monitoring · OSS AI observability. Traces, evals, embeddings.
Section: Prompts · AI Observability & Evaluation
Section: Tools · AI observability platform. Tracing, datasets, experiments, and playground for troubleshooting and evaluating LLM apps.
Section: Observability & Tracing · Self-hostable trace UI and eval runtime for agent workflows. Lets harness engineers audit and replay every reasoning step and tool call offline, without sending data to a third-party cloud.
Section: LLMOps · ML observability for LLMs, vision, language, and tabular models.
Section: 8. MLOps / LLMOps & Production · AI observability & evaluation platform.
Section: Evaluation and Monitoring · Phoenix is an open-source AI observability platform designed for experimentation, evaluation, and troubleshooting.
Section: Eval & Observability · Open-source AI observability & evaluation platform (Arize) — OpenTelemetry-native tracing for agents, LLM-as-judge evals, versioned datasets & experiments for prompt regression testing, prompt management with version control and replay, plus an MCP endpoint so Claude Code/Cursor can query traces…
Section: Testing · Open source tool for testing changes in AI agent or application
Section: Other · AI Observability & Evaluation
Langchain integrates various providers like Anthropic, AWS, and OpenAI, and offers tools for components such as LLMs, chat models, and data analysis, supporting functionalities from Alpha Vantage to YouTube github | docs
(MIT) provides modules for structured outputs at different levels of abstraction, including output parsers for text completion endpoints, Pydantic programs for mapping prompts to structured outputs using function calling or output parsing, and pre-defined Pydantic programs for specific output types.
robot: The free, Open Source alternative to OpenAI, Claude and others. Self-hosted and local-first. Drop-in replacement for OpenAI, running on consumer-grade hardware. No GPU required. Runs gguf, transformers, diffusers and many more models architectures. Features: Generate Text, Audio, Video,…
February 2026 release making human oversight a native workflow primitive: suspend execution at critical decision points, expose review-and-edit UI mid-flow, and route subsequent execution based on human action (approve/reject/escalate). Demonstrates how HITL transitions from bolt-on approval gates…
Mem0 is an intelligent memory layer for Large Language Models that enhances personalized AI experiences by retaining and utilizing contextual information across various applications. github | website | docs | discord | twitter | github profile | linkedin
Microsoft's multi-agent conversation framework with a complete AgentChat layer covering agent loop, tool integration, termination conditions, and human-in-the-loop. The most comprehensive open-source reference for large-scale multi-agent harness design.