Skip to content

Entry

VLMEvalKit

Appears in 4 awesome lists

Open-source evaluation toolkit for large multi-modality models (LMMs). Supports 220+ LMMs and 80+ benchmarks including MMMU, MathVista, and ChartQA. Powers the OpenVLM Leaderboard. Apache 2.0 licensed.

Open github.comopen-compass/vlmevalkit

Found in these lists

awesome-ChatGPT-repositories

Section: NLP · Open-source evaluation toolkit of large vision-language models (LVLMs), support GPT-4v, Gemini, QwenVLPlus, 30+ HF models, 15+ benchmarks

FreshScore 87

Awesome LLM Resources

Section: 评估 Evaluation · Open-source evaluation toolkit of large vision-language models (LVLMs), support ~100 VLMs, 40+ benchmarks.

FreshScore 87

Awesome Open Source AI

Section: 9. Evaluation, Benchmarks & Datasets · Open-source evaluation toolkit for large multi-modality models (LMMs). Supports 220+ LMMs and 80+ benchmarks including MMMU, MathVista, and ChartQA. Powers the OpenVLM Leaderboard. Apache 2.0 licensed.

FreshScore 89

Awesome Production Machine Learning

Section: Evaluation and Monitoring · VLMEvalKit is an open-source evaluation toolkit of large vision-language models (LVLMs).

FreshScore 92

Opik

Comet's open-source AI observability and evaluation platform: deep tracing of LLM calls, conversation logging, and agent activity, plus built-in eval metrics, prompt versioning, guardrails, and the Opik Agent Optimizer. Worth including because it unifies observability, verification, and…

In 16 listsDetails

transformers

(formerly known as pytorch-transformers and pytorch-pretrained-bert) provides state-of-the-art general-purpose architectures (BERT, GPT-2, RoBERTa, XLM, DistilBert, XLNet, CTRL...) for Natural Language Understanding (NLU) and Natural Language Generation (NLG) with over 32+ pretrained models in…

In 14 listsDetails

Mem0

Mem0 is an intelligent memory layer for Large Language Models that enhances personalized AI experiences by retaining and utilizing contextual information across various applications. github | website | docs | discord | twitter | github profile | linkedin

In 13 listsDetails

Haystack

Open-source AI orchestration framework for building context-engineered, production-ready LLM applications. Design modular pipelines and agent workflows with explicit control over retrieval, routing, memory, and generation. Built for scalable agents, RAG, multimodal applications, semantic search,…

In 13 listsDetails

Semantic Kernel

Semantic Kernel is an SDK that integrates Large Language Models (LLMs) like OpenAI, Azure OpenAI, and Hugging Face with conventional programming languages like C#, Python, and Java. Semantic Kernel achieves this by allowing you to define plugins that can be chained together in just a few lines of…

In 11 listsDetails

OpenLLM

Production-grade platform for running any open-source LLMs as OpenAI-compatible API endpoints. Supports 50+ models with built-in streaming, batching, and auto-acceleration. Apache 2.0 licensed.

In 10 listsDetails

Markstream

Multi-framework streaming Markdown renderer for AI chat interfaces, with incomplete Markdown handling, Mermaid, KaTeX, Shiki/Monaco code blocks, SSR, and packages for Vue, React, Svelte, and Angular.

In 9 listsDetails

OpenAI Agents SDK

Lightweight multi-agent framework built around handoffs and guardrails; the production successor to Swarm. Complements LangGraph for harnesses where delegation patterns are simpler than full graph orchestration.

In 9 listsDetails