Awesome LLM Resources
Section: 评估 Evaluation · A framework for few-shot evaluation of language models.
Entry
Appears in 7 awesome lists
Language Model Evaluation Harness is a framework to test generative language models on a large number of different evaluation tasks.
Section: 评估 Evaluation · A framework for few-shot evaluation of language models.
Section: Evaluation and Benchmarks · unified framework for LM benchmark evaluation.
Section: 9. Evaluation, Benchmarks & Datasets · De-facto standard for generative model evaluation.
Section: Evaluation and Monitoring · Language Model Evaluation Harness is a framework to test generative language models on a large number of different evaluation tasks.
Section: Tools & Libraries · EleutherAI's unified LLM evaluation framework
Section: LLM 评测与治理 (LLM Evaluation & Harness)
Section: Other · A framework for few-shot evaluation of language models.
The most complete open-source LLM/agent eval framework: 20+ built-in metrics (hallucination, answer relevancy, RAGAs, tool correctness), pytest integration, and a CI-friendly runner. Removes the need to hand-roll eval infrastructure when you need structured, repeatable agent quality gates.
Framework for evaluating LLMs and LLM systems with an open-source registry of 100+ community-contributed benchmarks. MIT licensed.
The largest hub of ready-to-use NLP datasets for ML models with fast, easy-to-use and efficient data manipulation tools.
RAG AutoML tool for automatically finding optimal RAG pipelines. Evaluates and optimizes retrieval-augmented generation with AutoML-style automation for your own data and use-case. Apache 2.0 licensed.
OpenCompass is an LLM evaluation platform, supporting a wide range of models (Llama3, Mistral, InternLM2,GPT-4,LLaMa2, Qwen,GLM, Claude, etc) over 100+ datasets.
Open-source evaluation toolkit for large multi-modality models (LMMs). Supports 220+ LMMs and 80+ benchmarks including MMMU, MathVista, and ChartQA. Powers the OpenVLM Leaderboard. Apache 2.0 licensed.
Evaluate machine learning models (huggingface). pycm - Multi-class confusion matrix. pandas_ml - Confusion matrix. Plotting learning curve: link. yellowbrick - Learning curve. pyroc - Receiver Operating Characteristic (ROC) curves.
Open corpus of the instructions and tool schemas shipping AI agents are actually sent: 119 artifacts from 43 products, 44 of them recorded off the wire with the command that reproduces each, every file marked captured or reported. AGPL-3.0.