Skip to content

Entry

Phoenix

Appears in 10 awesome lists

Open-source AI observability & evaluation platform (Arize) — OpenTelemetry-native tracing for agents, LLM-as-judge evals, versioned datasets & experiments for prompt regression testing, prompt management with version control and replay, plus an MCP endpoint so Claude Code/Cursor can query traces…

Open github.comarize-ai/phoenix

Found in these lists

Awesome Ai Agents 2026

Section: Tracing and Monitoring · OSS AI observability. Traces, evals, embeddings.

ActiveScore 74

awesome-ChatGPT-repositories

Section: Prompts · AI Observability & Evaluation

FreshScore 87

Awesome Data Analysis

Section: Tools · AI observability platform. Tracing, datasets, experiments, and playground for troubleshooting and evaluating LLM apps.

FreshScore 80

Awesome Harness Engineering

Section: Observability & Tracing · Self-hostable trace UI and eval runtime for agent workflows. Lets harness engineers audit and replay every reasoning step and tool call offline, without sending data to a third-party cloud.

FreshScore 88

Awesome LLMOps

Section: LLMOps · ML observability for LLMs, vision, language, and tabular models.

ActiveScore 75

Awesome Open Source AI

Section: 8. MLOps / LLMOps & Production · AI observability & evaluation platform.

FreshScore 89

Awesome Production Machine Learning

Section: Evaluation and Monitoring · Phoenix is an open-source AI observability platform designed for experimentation, evaluation, and troubleshooting.

FreshScore 92

Awesome Prompts

Section: Eval & Observability · Open-source AI observability & evaluation platform (Arize) — OpenTelemetry-native tracing for agents, LLM-as-judge evals, versioned datasets & experiments for prompt regression testing, prompt management with version control and replay, plus an MCP endpoint so Claude Code/Cursor can query traces…

FreshScore 90

Awesome AI Agents: Tools, Resources, and Projects

Section: Testing · Open source tool for testing changes in AI agent or application

SlowScore 68

awesome-python

Section: Other · AI Observability & Evaluation

FreshScore 81

LangChain

Langchain integrates various providers like Anthropic, AWS, and OpenAI, and offers tools for components such as LLMs, chat models, and data analysis, supporting functionalities from Alpha Vantage to YouTube github | docs

In 20 listsDetails

LlamaIndex

(MIT) provides modules for structured outputs at different levels of abstraction, including output parsers for text completion endpoints, Pydantic programs for mapping prompts to structured outputs using function calling or output parsing, and pre-defined Pydantic programs for specific output types.

In 14 listsDetails

LocalAI

robot: The free, Open Source alternative to OpenAI, Claude and others. Self-hosted and local-first. Drop-in replacement for OpenAI, running on consumer-grade hardware. No GPU required. Runs gguf, transformers, diffusers and many more models architectures. Features: Generate Text, Audio, Video,…

In 14 listsDetails

Dify

February 2026 release making human oversight a native workflow primitive: suspend execution at critical decision points, expose review-and-edit UI mid-flow, and route subsequent execution based on human action (approve/reject/escalate). Demonstrates how HITL transitions from bolt-on approval gates…

In 14 listsDetails

Mem0

Mem0 is an intelligent memory layer for Large Language Models that enhances personalized AI experiences by retaining and utilizing contextual information across various applications. github | website | docs | discord | twitter | github profile | linkedin

In 13 listsDetails

AutoGen

Microsoft's multi-agent conversation framework with a complete AgentChat layer covering agent loop, tool integration, termination conditions, and human-in-the-loop. The most comprehensive open-source reference for large-scale multi-agent harness design.

In 14 listsDetails

Whisper

Whisper is a general-purpose speech recognition model that can be run locally offline. It can transcribe audio from and to multiple languages.

In 11 listsDetails

promptfoo

Test your prompts, models, RAGs. Evaluate and compare LLM outputs, catch regressions, and improve prompt quality. LLM evals for OpenAI/Azure GPT, Anthropic Claude, VertexAI Gemini, Ollama, Local & private models like Mistral/Mixtral/Llama with CI/CD

In 11 listsDetails