Awesome-LLM-Eval: a curated list of tools, datasets/benchmark, demos, leaderboard, papers, docs and models, mainly for Evaluation on Large Language Models and exploring the boundaries and limits of Generative AI. Use it to build a structured path from fundamentals to hands-on practice.
Table of Contents
Snapshot 2026-08-04 16:17:00 UTC · version 1
Research document
Table of Contents
Awesome-LLM-Eval: a curated list of tools, datasets/benchmark, demos, leaderboard, papers, docs and models, mainly for Evaluation on Large Language Models and exploring the boundaries and limits of Generative AI. Use it to build a structured path from fundamentals to hands-on practice.
Editorial note: curated source snapshot published by Collider.club under the MIT License. Source attribution is preserved in the front matter.
Source snapshot
Awesome-LLM-Eval: a curated list of tools, datasets/benchmark, demos, leaderboard, papers, docs and models, mainly for Evaluation on Large Language Models and exploring the boundaries and limits of Generative AI.
The is the official project of our survey: Beyond Benchmark: LLMs Evaluation with an Anthropomorphic and Value-oriented Roadmap.
NOTE: As we cannot update the arXiv paper in real time, please refer to this repo for the latest updates and the paper may be updated later. We also welcome any pull request or issues to help us improve this work. Your contributions will be acknowledged in acknowledgements.
If you find our survey useful, please kindly cite our paper:
@misc{wang2025llmevalroadmap,
title={Beyond Benchmark: LLMs Evaluation with an Anthropomorphic and Value-oriented Roadmap},
author={Jun Wang and Ninglun Gu and Kailai Zhang and Zijiao Zhang and Yelun Bao and Jin Yang and Xu Yin and Liwei Liu and Yihuan Liu and Pengyong Li and Gary G. Yen and Junchi Yan},
year={2025},
eprint={2508.18646},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2508.18646},
}
Table of Contents
- News
- Tools
- Datasets / Benchmark
- Demos
- Leaderboards
- Papers
- LLM-List
- LLMOps
- Frameworks for Training
- Courses
- Others
- Other Awesome Lists
- Licenses
- Citation
News
- [2025/08/20] We added the Anthropomorphic-Taxonomy section.
- [2024/04/26] We added the Inference-Speed section.
- [2024/02/26] We added the Coding-Evaluation section.
- [2024/02/08] We added the lighteval tool from Huggingface.
- [2024/01/15] We added CRUXEval, DebugBench, OpenFinData, and LAiW.
- [2023/12/20] We added the RAG-Evaluation section.
- [2023/11/15] We added Instruction-Following-Evaluation and LLMBar for evaluating the instruction following capabilities of LLMs.
- [2023/10/20] We added SuperCLUE-Agent for LLM agent evaluation.
- [2023/09/25] We added ColossalEval from Colossal-AI.
- [2023/09/22] We added the LeaderboardFinder chapter.
- [2023/09/20] We added DeepEval, FinEval, and SuperCLUE-Safety from CLUEbenchmark.
- [2023/09/18] We added OpenCompass from Shanghai AI Lab.
- [2023/08/03] We added new Chinese LLMs: Baichuan and Qwen.
- [2023/06/28] We added AlpacaEval and multiple tools.
- [2023/04/26] We released the V0.1 evaluation list with multiple benchmarks.
Anthropomorphic-Taxonomy
Typical Intelligence Quotient (IQ)-General Intelligence evaluation benchmarks
| Name | Year | Task Type | Institution | Evaluation Focus | Datasets | Url |
|---|---|---|---|---|---|---|
| MMLU-Pro | 2024 | Multi-Choice Knowledge | TIGER-AI-Lab | Subtle Reasoning, Fewer Noise | MMLU-Pro | link |
| DyVal | 2024 | Dynamic Evaluation | Microsoft | Data Pollution, Complexity Control | DyVal | link |
| PertEval | 2024 | General | USTC | Knowledge capacity | PertEval | link |
| LV-Eval | 2024 | Long Text QA | Infinigence-AI | Length Variability, Factuality | 11 Subsets | link |
| LLM-Uncertainty-Bench | 2024 | NLP Tasks | Tencent | Uncertainty Quantification | 5 NLP Tasks | link |
| CommonGen-Eval | 2024 | Generation | AI2 | Common Sense | CommonGen-lite | link |
| MathBench | 2024 | Math | Shanghai AI Lab | Theoretical and practical problem-solving | Various | link |
| AIME | 2024 | Math | MAA | American Invitational Mathematics Examination | Various | link |
| FrontierMath | 2024 | Math | Epoch AI | Original, challenging mathematics problems | Various | link |
| FELM | 2023 | Factuality | HKUST | Factuality | 847 Questions | link |
| Just-Eval-Instruct | 2023 | General | AI2 Mosaic | Helpfulness, Explainability | Various | link |
| MLAgentBench | 2023 | ML Research | snap-stanford | End-to-End ML Tasks | 15 Tasks | link |
| UltraEval | 2023 | General | OpenBMB | Lightweight, Flexible, Fast | Various | link |
| FMTI | 2023 | Transparency | Stanford | Model Transparency | 100 Metrics | link |
| BAMBOO | 2023 | Long Text | RUCAIBox | Long Text Modeling | 10 Datasets | link |
| TRACE | 2023 | Continuous Learning | Fudan University | Continuous Learning | 8 Datasets | link |
| ColossalEval | 2023 | General | Colossal-AI | Unified Evaluation | Various | link |
| LLMEval² | 2023 | General | AlibabaResearch | Wide and Deep Evaluation | 2,553 Samples | link |
| BigBench | 2023 | General | knowledge, language, reasoning | Various | link | |
| LucyEval | 2023 | General | Oracle | Maturity Assessment | Various | link |
| Zhujiu | 2023 | General | IACAS | Comprehensive Evaluation | 51 Tasks | link |
| ChatEval | 2023 | Chat | THU-NLP | Human-like Evaluation | Various | link |
| FlagEval | 2023 | General | THU | Subjective and Objective Scoring | Various | link |
| AlpacaEval | 2023 | General | tatsu-lab | Automatic Evaluation | Various | link |
| GPQA | 2023 | General | NYU | Graduate-Level Google-Proof QA | Various | link |
| MuSR | 2023 | Reasoning | Zayne Sprague | Narrative-Based Reasoning | 756 | link |
| FreshQA | 2023 | Knowledge | FreshLLMs | Current World Knowledge | 599 | link |
| AGIEval | 2023 | General | Microsoft | Human-Centric Reasoning | NA | link |
| SummEdits | 2023 | General | Salesforce | Inconsistency Detection | 6,348 | link |
| ScienceQA | 2022 | Reasoning | UCLA | Science Reasoning | 21,208 | link |
| e-CARE | 2022 | Reasoning | HIT | Explainable Causality | 21,000 | link |
| BigBench Hard | 2022 | Reasoning | BigBench | Challenging Subtasks | 6,500 | link |
| PlanBench | 2022 | Reasoning | ASU | Action Planning | 11,113 | link |
| MGSM | 2022 | Math | Grade-school math problems in 10 languages | Various | link | |
| MATH | 2021 | Math | UC Berkeley | Mathematical Problem Solving | Various | link |
| GSM8K | 2021 | Math | OpenAI | Diverse grade school math word problems | Various | link |
| SVAMP | 2021 | Math | Microsoft | Arithmetic Reasoning | 1,000 | link |
| SpartQA | 2021 | Reasoning | MSU | Textual Spatial QA | 510 | link |
| MLSUM | 2020 | General | Thomas Scialom | News Summarization | 535,062 | link |
| Natural Questions | 2019 | Language, Reasoning | Search-Based QA | 300,000 | link | |
| ANLI | 2019 | Language, Reasoning | Facebook AI | Adversarial Reasoning | 169,265 | link |
| BoolQ | 2019 | Language, Reasoning | Binary QA | 16,000 | link | |
| SuperGLUE | 2019 | Language, Reasoning | NYU | Advanced GLUE Tasks | NA | link |
| DROP | 2019 | Language, Reasoning | UCI NLP | Paragraph-Level Reasoning | 96,000 | link |
| HellaSwag | 2019 | Language, Reasoning | AI2 | Commonsense Inference | 59,950 | link |
| Winogrande | 2019 | Language, Reasoning | AI2 | Pronoun Disambiguation | 44,000 | link |
| PIQA | 2019 | Language, Reasoning | AI2 | Physical Interaction QA | 18,000 | link |
| HotpotQA | 2018 | Language, Reasoning | HotpotQA | Explainable QA | 113,000 | link |
| GLUE | 2018 | Language, Reasoning | NYU | Foundational NLU Tasks | NA | link |
| OpenBookQA | 2018 | Language, Reasoning | AI2 | Open Book Exams | 12,000 | link |
| SQuAD2.0 | 2018 | Language, Reasoning | Stanford University | Unanswerable Questions | 150,000 | link |
| ARC | 2018 | Language, Reasoning | AI2 | AI2 Reasoning Challenge | 7,787 | link |
| SWAG | 2018 | Language, Reasoning | AI2 | Adversarial Commonsense | 113,000 | link |
| CommonsenseQA | 2018 | Language, Reasoning | AI2 | Commonsense Reasoning | 12,102 | link |
| RACE | 2017 | Language, Reasoning | CMU | Exam-Style QA | 100,000 | link |
| SciQ | 2017 | Language, Reasoning | AI2 | Crowd-Sourced Science | 13,700 | link |
| TriviaQA | 2017 | Language, Reasoning | AI2 | Distant Supervision | 650,000 | link |
| MultiNLI | 2017 | Language, Reasoning | NYU | Cross-Genre Entailment | 433,000 | link |
| SQuAD | 2016 | Language, Reasoning | Stanford University | Wikipedia-Based QA | 100,000 | link |
| LAMBADA | 2016 | Language, Reasoning | CIMEC | Discourse Context | 12,684 | link |
| MS MARCO | 2016 | Language, Reasoning | Microsoft | Search-Based QA | 1,112,939 | link |
Typical Professional Quotient (PQ)-Professional Expertise evaluation benchmarks
| Domain | Name | Institution | Scope of Tasks | Unique Contributions | Url |
|---|---|---|---|---|---|
| BLURB | Mindrank AI | Six diverse NLP tasks, thirteen datasets | A macro-average score across all tasks | link | |
| Seismometer | Epic | Using local data and workflows | patient demographics, clinical interventions, and outcomes | link | |
| Healthcare | Medbench | OpenMEDLab | Emphasizes scientific rigor and fairness | 40,041 questions from medical exams and reports | link |
| GenMedicalEval | E | 16 majors, 3 training stages, 6 clinical scenarios | Open-ended metrics and automated assessment models | link | |
| PsyEval | SJTU | Six subtasks covering three dimensions | Customized benchmark for mental health LLMs | link | |
| Fin-Eva | Ant Group | Wealth management, insurance, investment research | Both industrial and academic financial evaluations | link | |
| Finance | FinEval | SUFE-AIFLM-Lab | Multiple-choice QA on finance, economics, accounting | Focuses on high-quality evaluation questions | link |
| OpenFinData | Shanghai AI Lab | Multi-scenario financial tasks | First comprehensive finance evaluation dataset | link | |
| FinBen | FinAI | 35 datasets across 23 financial tasks | Inductive reasoning, quantitative reasoning | link | |
| LAiW | Sichuan University | 13 fundamental legal NLP tasks | Divides legal NLP capabilities into three major abilities | link | |
| Legal | LawBench | Nanjing University | Legal entity recognition, reading comprehension | Real-world tasks, "abstention rate" metric | link |
| LegalBench | Stanford University | 162 tasks covering six types of legal reasoning | Enables interdisciplinary conversations | link | |
| LexEval | Tsinghua University | Legal cognitive abilities to organize different tasks | Larger legal evaluation dataset, examining the ethical issues | link | |
| SPEC5G | Purdue University | Security-related text classification and summarization | 5G protocol analysis automation | link | |
| Telecom | TeleQnA | Huawei(Paris) | General telecom inquiries | Proficiency in telecom-related questions | link |
| OpsEval | Tsinghua University | Wired network ops, 5G, database ops | Focus on AIOps, evaluates proficiency | link | |
| TelBench | SK Telecom | Math modeling, open-ended QA, code generation | Holistic evaluation in telecom | link | |
| TelecomGPT | UAE | Telecom Math Modeling, Open QnA and Code Tasks | Holistic evaluation in telecom | link | |
| Linguistic | Queen's University | Multiple language-centric tasks | zero-shot evaluation | link | |
| TelcoLM | Orange | Multiple-choice questionnaires | Domain-specific data (800M tokens, 80K instructions) | link | |
| ORAN-Bench-13K | GMU | Multiple-choice questions | Open Radio Access Networks (O-RAN) | link | |
| Open-Telco Benchmarks | GSMA | Multiple language-centric tasks | zero-shot evaluation | link | |
| FullStackBench | ByteDance | Code writing, debugging, code review | Featuring the most recent Stack Overflow QA | link | |
| Coding | StackEval | Prosus AI | 11 real-world scenarios, 16 languages | Evaluation across diverse & practical coding environments | link |
| CodeBenchGen | Various Institutions | Execution-based code generation tasks | Benchmarks scaling with the size and complexity | link | |
| HumanEval | University of Washington | Rigorous testing | Stricter protocol for assessing correctness of generated code | link | |
| APPS | University of California | Coding challenges from competitive platforms | Checking problem-solving of generated code on test cases | link | |
| MBPP | Google Research | Programming problems sourced from various origins | Diverse programming tasks | link | |
| ClassEval | Tsinghua University | Class-level code generation | Manually crafted, object-oriented programming concepts | link | |
| CoderEval | Peking University | Pragmatic code generation | Proficiency to generate functional code patches for described issues | link | |
| MultiPL-E | Princeton University | Neural code generation | Benchmarking neural code generation models | link | |
| CodeXGLUE | Microsoft | Code intelligence | Wide tasks covering: code-code, text-code, code-text and text-text | link | |
| EvoCodeBench | Peking University | Evolving code generation benchmark | Aligned with real-world code repositories, evolving over time | link |
Typical Emotional Quotient (EQ)-Alignment Ability evaluation benchmarks
| Name | Year | Task Type | Institution | Category | Datasets | Url |
|---|---|---|---|---|---|---|
| DiffAware | 2025 | Bias | Stanford | General Bias | 8 datasets | link |
| CASE-Bench | 2025 | Safety | Cambridge | Context-Aware Safety | CASE-Bench | link |
| Fairness | 2025 | Fairness | PSU | Distributive Fairness | - | - |
| HarmBench | 2024 | Safety | UIUC | Adversarial Behaviors | 510 | link |
| SimpleQA | 2024 | Safety | OpenAI | Factuality | 4,326 | link |
| AgentHarm | 2024 | Safety | BEIS | Malicious Agent Tasks | 110 | link |
| StrongReject | 2024 | Safety | dsbowen | Attack Resistance | n/a | link |
| LLMBar | 2024 | Instruction | Princeton | Instruction Following | 419 Instances | link |
| AIR-Bench | 2024 | Safety | Stanford | Regulatory Alignment | 5,694 | link |
| TrustLLM | 2024 | General | TrustLLM | Trustworthiness | 30+ | link |
| RewardBench | 2024 | Alignment | AIAI | Human preference | RewardBench | link |
| EQ-Bench | 2024 | Emotion | Paech | Emotional intelligence | 171 Questions | link |
| Forbidden | 2023 | Safety | CISPA | Jailbreak Detection | 15,140 | link |
| MaliciousInstruct | 2023 | Safety | Princeton | Malicious Intentions | 100 | link |
| SycophancyEval | 2023 | Safety | Anthropic | Opinion Alignment | n/a | link |
| DecodingTrust | 2023 | Safety | UIUC | Trustworthiness | 243,877 | link |
| AdvBench | 2023 | Safety | CMU | Adversarial Attacks | 1,000 | link |
| XSTest | 2023 | Safety | Bocconi | Safety Overreach | 450 | link |
| OpinionQA | 2023 | Safety | tatsu-lab | Demographic Alignment | 1,498 | link |
| SafetyBench | 2023 | Safety | THU | Content Safety | 11,435 | link |
| HarmfulQA | 2023 | Safety | declare-lab | Harmful Topics | 1,960 | link |
| QHarm | 2023 | Safety | vinid | Safety Sampling | 100 | link |
| BeaverTails | 2023 | Safety | PKU | Red Teaming | 334,000 | link |
| DoNotAnswer | 2023 | Safety | Libr-AI | Safety Mechanisms | 939 | link |
| AlignBench | 2023 | Alignment | THUDM | Alignment, Reliability | Various | link |
| IFEval | 2023 | Instruction | Instruction Following | 500 Prompts | link | |
| ToxiGen | 2022 | Safety | Microsoft | Toxicity Detection | 274,000 | link |
| HHH | 2022 | Safety | Anthropic | Human Preferences | 44,849 | link |
| RedTeam | 2022 | Safety | Anthropic | Red Teaming | 38,921 | link |
| BOLD | 2021 | Bias | Amazon | Bias in Generation | 23,679 | link |
| BBQ | 2021 | Bias | NYU | Social Bias | 58,492 | link |
| StereoSet | 2020 | Bias | McGill | Stereotype Detection | 4,229 | link |
| ETHICS | 2020 | Ethics | Berkeley | Moral Judgement | 134,400 | link |
| ToxicityPrompt | 2020 | Safety | AllenAI | Toxicity Assessment | 99,442 | link |
| CrowS-Pairs | 2020 | Bias | NYU | Stereotype Measurement | 1,508 | link |
| SEAT | 2019 | Bias | Princeton | Encoder Bias | n/a | link |
| WinoGender | 2018 | Bias | UMass | Gender Bias | 720 | link |
Tools
| Name | Organization | Website | Description |
|---|---|---|---|
| prometheus-eval | prometheus-eval | prometheus-eval | PROMETHEUS Open Evaluation Dedicated Language Model, which is more powerful than its predecessor. It can closely imitate the judgments of humans and GPT-4. Additionally, it can handle both direct evaluation and pairwise ranking formats, and can be used with user-defined evaluation criteria. On four direct evaluation benchmarks and four pairwise ranking benchmarks, PROMETHEUS 2 achieves the highest correlation and consistency with human evaluators and proprietary language models among all tested open-source evaluation language models (2024-05-04). |
| athina-evals | athina-ai | athina-ai | Athina-ai is an open-source library that provides plug-and-play preset evaluations and a modular, extensible framework for writing and running evaluations. It helps engineers systematically improve the reliability and performance of their large language models through evaluation-driven development. Athina-ai offers a system for evaluation-driven development, overcoming the limitations of traditional workflows, enabling rapid experimentation, and providing customizable evaluators with consistent metrics. |
| LeaderboardFinder | Huggingface | LeaderboardFinder | LeaderboardFinder helps you find suitable leaderboards for specific scenarios, a leaderboard of leaderboards (2024-04-02). |
| LightEval | Huggingface | lighteval | LightEval is a lightweight framework developed by Hugging Face for evaluating large language models (LLMs). Originally designed as an internal tool for assessing Hugging Face's recently released LLM data processing library datatrove and LLM training library nanotron, it is now open-sourced for community use and improvement. Key features of LightEval include: (1) lightweight design, making it easy to use and integrate; (2) an evaluation suite supporting multiple tasks and models; (3) compatibility with evaluation on CPUs or GPUs, and integration with Hugging Face's acceleration library (Accelerate) and frameworks like Nanotron; (4) support for distributed evaluation, which is particularly useful for evaluating large models; (5) applicability to all benchmarks on the Open LLM Leaderboard; and (6) customizability, allowing users to add new metrics and tasks to meet specific evaluation needs (2024-02-08). |
| LLM Comparator | LLM Comparator | A visual analytical tool for comparing and evaluating large language models (LLMs). Compared to traditional human evaluation methods, this tool offers a scalable automated approach to comparative evaluation. It leverages another LLM as an evaluator to demonstrate quality differences between models and provide reasons for these differences. Through interactive tables and summary visualizations, the LLM Comparator helps users understand why models perform well or poorly in specific contexts, as well as the qualitative differences between model responses. Developed in collaboration with Google researchers and engineers, this tool has been widely used internally at Google, attracting over 400 users and evaluating more than 1,000 experiments within three months (2024-02-16). | |
| Arthur Bench | Arthur-AI | Arthur Bench | Arthur Bench is an open-source evaluation tool designed to compare and analyze the performance of large language models (LLMs). It supports various evaluation tasks, including question answering, summarization, translation, and code generation, and provides detailed reports on LLM performance across these tasks. Key features and advantages of Arthur Bench include: (1) model comparison, enabling the evaluation of different suppliers, versions, and training datasets of LLMs; (2) prompt and hyperparameter evaluation, assessing the impact of different prompts on LLM performance and testing the control of model behavior through various hyperparameter settings; (3) task definition and model selection, allowing users to define specific evaluation tasks and select evaluation targets from a range of supported LLM models; (4) parameter configuration, enabling users to adjust prompts and hyperparameters to finely control LLM behavior; (5) automated evaluation workflows, simplifying the execution of evaluation tasks; and (6) application scenarios such as model selection and validation, budget and privacy optimization, and the transformation of academic benchmarks into real-world performance evaluations. Additionally, it offers comprehensive scoring metrics, supports both local and cloud versions, and encourages community collaboration and project development (2023-10-06). |
| llm-benchmarker-suite | FormulaMonks | llm-benchmarker-suite | This open-source initiative aims to address fragmentation and ambiguity in LLM benchmarking. The suite provides a structured methodology, a collection of diverse benchmarks, and toolkits to streamline the assessment of LLM performance. By offering a common platform, this project seeks to promote collaboration, transparency, and high-quality research in NLP. |
| autoevals | braintrust | autoevals | AutoEvals is an AI model output evaluation tool that leverages best practices to quickly and easily assess AI model outputs. It integrates multiple automatic evaluation methods, supports customizable evaluation prompts and custom scorers, and simplifies the evaluation process of model outputs. Autoevals incorporates model-graded evaluation for various subjective tasks, including fact-checking, safety, and more. Many of these evaluations are adapted from OpenAI's excellent evals project but are implemented in a flexible way to allow users to tweak prompts and debug outputs. |
| EVAL | OPENAI | EVAL | EVAL is a tool developed by OpenAI for evaluating large language models (LLMs). It can test the performance and generalization capabilities of models across different tasks and datasets. |
| lm-evaluation-harness | EleutherAI | lm-evaluation-harness | lm-evaluation-harness is a tool developed by EleutherAI for evaluating large language models (LLMs). It can test the performance and generalization capabilities of models across different tasks and datasets. |
| lm-evaluation | AI21Labs | lm-evaluation | Evaluations and reproducing the results from the Jurassic-1 Technical Paper, with current support for running tasks via both the AI21 Studio API and OpenAI's GPT-3 API. |
| OpenCompass | Shanghai AI Lab | OpenCompass | OpenCompass is a one-stop platform for evaluating large models. Its main features include: open-source and reproducible evaluation schemes; comprehensive capability dimensions covering five major areas with over 50 datasets and approximately 300,000 questions to assess model capabilities; support for over 20 Hugging Face and API models; distributed and efficient evaluation with one-line task splitting and distributed evaluation, enabling full evaluation of trillion-parameter models within hours; diverse evaluation paradigms supporting zero-shot, few-shot, and chain-of-thought evaluations with standard or conversational prompt templates to easily elicit peak model performance. |
| Large language model evaluation and workflow framework from Phase AI | wgryc | phasellm | A framework provided by Phase AI for evaluating and managing LLMs, helping users select appropriate models, datasets, and metrics, as well as visualize and analyze results. |
| Evaluation benchmark for LLM | FreedomIntelligence | LLMZoo | LLMZoo is an evaluation benchmark for LLMs developed by FreedomIntelligence, featuring multiple domain and task datasets, metrics, and pre-trained models with results. |
| Holistic Evaluation of Language Models (HELM) | Stanford | HELM | HELM is a comprehensive evaluation method for LLMs proposed by the Stanford research team, considering multiple aspects such as model language ability, knowledge, reasoning, fairness, and safety. |
| A lightweight evaluation tool for question-answering | Langchain | auto-evaluator | auto-evaluator is a lightweight tool developed by Langchain for evaluating question-answering systems. It can automatically generate questions and answers and calculate metrics such as model accuracy, recall, and F1 score. |
| PandaLM | WeOpenML | PandaLM | PandaLM is an LLM assessment tool developed by WeOpenML for automated and reproducible evaluation. It allows users to select appropriate datasets, metrics, and models based on their needs and preferences, and generates reports and charts. |
| FlagEval | Tsinghua University | FlagEval | FlagEval is an evaluation platform for LLMs developed by Tsinghua University, offering multiple tasks and datasets, as well as online testing, leaderboards, and analysis functions. |
| AlpacaEval | tatsu-lab | alpaca_eval | AlpacaEval is an evaluation tool for LLMs developed by tatsu-lab, capable of testing models across various languages, domains, and tasks, and providing explainability, robustness, and credibility metrics. |
| Prompt flow | Microsoft | promptflow | A set of development tools designed by Microsoft to simplify the end-to-end development cycle of AI applications based on LLMs, from conception, prototyping, testing, and evaluation to production deployment and monitoring. It makes prompt engineering easier and enables the development of product-level LLM applications. |
| DeepEval | mr-gpt | DeepEval | DeepEval is a simple-to-use, open-source LLM evaluation framework. Similar to Pytest but specialized for unit testing LLM outputs, it incorporates the latest research to evaluate LLM outputs based on metrics such as G-Eval, hallucination, answer relevance, RAGAS, etc., utilizing LLMs and various other NLP models that run locally on your machine for evaluation. |
| CONNER | Tencent AI Lab | CONNER | CONNER is a comprehensive large model knowledge evaluation framework designed to systematically and automatically assess the information generated from six critical perspectives: factuality, relevance, coherence, informativeness, usefulness, and validity. |
Datasets or Benchmarks
General
| Name | Organization | Website | Description |
|---|---|---|---|
| MMLU-Pro | TIGER-AI-Lab | MMLU-Pro | MMLU-Pro is an improved version of the MMLU dataset. MMLU has long been a reference for multiple-choice knowledge datasets. However, recent studies have shown that it contains noise (some questions are unanswerable) and is too easy (due to the evolution of model capabilities and increased contamination). MMLU-Pro provides ten options instead of four, requires reasoning on more questions, and has undergone expert review to reduce noise. It is of higher quality and more challenging than the original. MMLU-Pro reduces the impact of prompt variations on model performance, a common issue with its predecessor MMLU. Research indicates that models using "Chain of Thought" reasoning perform better on this new benchmark, suggesting that MMLU-Pro is better suited for evaluating the subtle reasoning abilities of AI. (2024-05-20) |
| TrustLLM Benchmark | TrustLLM | TrustLLM | TrustLLM is a benchmark for evaluating the trustworthiness of large language models. It covers six dimensions of trustworthiness and includes over 30 datasets to comprehensively assess the functional capabilities of LLMs, ranging from simple classification tasks to complex generative tasks. Each dataset presents unique challenges and has benchmarked 16 mainstream LLMs (including commercial and open-source models). |
| DyVal | Microsoft | DyVal | Concerns have been raised about the potential data contamination in the vast training corpora of LLMs. Additionally, the static nature and fixed complexity of current benchmarks may not adequately measure the evolving capabilities of LLMs. DyVal is a general and flexible protocol for dynamically evaluating LLMs. Leveraging the advantages of directed acyclic graphs, DyVal dynamically generates evaluation samples with controllable complexity. It has created challenging evaluation sets for reasoning tasks such as mathematics, logical reasoning, and algorithmic problems. Various LLMs, from Flan-T5-large to GPT-3.5-Turbo and GPT-4, have been evaluated. Experiments show that LLMs perform worse on DyVal-generated samples of different complexities, highlighting the importance of dynamic evaluation. The authors also analyze failure cases and results of different prompting methods. Furthermore, DyVal-generated samples not only serve as evaluation sets but also aid in fine-tuning to enhance LLM performance on existing benchmarks. (2024-04-20) |
| RewardBench | AIAI | RewardBench | RewardBench is an evaluation benchmark for language model reward models, assessing the strengths and weaknesses of various models. It reveals that existing models still exhibit significant shortcomings in reasoning and instruction following. It includes a Leaderboard, Code, and Dataset (2024-03-20). |
| LV-Eval | Infinigence-AI | LVEval | LV-Eval is a long-text evaluation benchmark featuring five length tiers (16k, 32k, 64k, 128k, and 256k), with a maximum text test length of 256k. The average text length of LV-Eval is 102,380 characters, with a minimum/maximum text length of 11,896/387,406 characters. LV-Eval primarily consists of two types of evaluation tasks: single-hop QA and multi-hop QA, encompassing 11 sub-datasets in Chinese and English. During its design, LV-Eval introduced three key technologies: Confusion Facts Insertion (CFI) to enhance challenge, Keyword and Phrase Replacement (KPR) to reduce information leakage, and Answer Keywords (AK) based evaluation metrics (combining answer keywords and word blacklists) to improve the objectivity of evaluation results (2024-02-06). |
| LLM-Uncertainty-Bench | Tencent | LLM-Uncertainty-Bench | A new benchmark method for LLMs has been introduced, incorporating uncertainty quantification. Based on nine LLMs tested across five representative NLP tasks, it was found that: I) More accurate LLMs may exhibit lower certainty; II) Larger-scale LLMs may display greater uncertainty than smaller models; III) Instruction fine-tuning tends to increase the uncertainty of LLMs. These findings underscore the importance of including uncertainty in LLM evaluations (2024-01-22). |
| Psychometrics Eval | Microsoft Research Asia | Psychometrics Eval | Microsoft Research Asia has proposed a generalized evaluation method for AI based on psychometrics, aiming to address limitations in traditional evaluation methods concerning predictive power, information volume, and test tool quality. This approach draws on psychometric theories to identify key psychological constructs of AI, design targeted tests, and apply Item Response Theory for precise scoring. It also introduces concepts of reliability and validity to ensure evaluation reliability and accuracy. This framework extends psychometric methods to assess AI performance in handling unknown complex tasks but also faces open questions such as distinguishing between AI "individuals" and "populations," addressing prompt sensitivity, and evaluating differences between human and AI constructs (2023-10-19). |
| CommonGen-Eval | AllenAI | CommonGen-Eval | A study using the CommonGen-lite dataset to evaluate LLMs, employing GPT-4 for assessment and comparing the performance of different models, with results listed on the leaderboard (2024-01-04). |
| felm | HKUST | felm | FELM is a meta-benchmark for evaluating the factual assessment of large language models. The benchmark comprises 847 questions spanning five distinct domains: world knowledge, science/technology, writing/recommendation, reasoning, and mathematics. Prompts corresponding to each domain are gathered from various sources, including standard datasets like TruthfulQA, online platforms like GitHub repositories, ChatGPT-generated prompts, or those drafted by authors. For each response, fine-grained annotation at the segment level is employed, including reference links, identified error types, and reasons behind these errors as provided by annotators (2023-10-03). |
| just-eval | AI2 Mosaic | just-eval | A GPT-based evaluation tool for multi-faceted and explainable assessment of LLMs, capable of evaluating aspects such as helpfulness, clarity, factuality, depth, and engagement (2023-12-05). |
| EQ-Bench | EQ-Bench | EQ-Bench | A benchmark for evaluating the emotional intelligence of language models, featuring 171 questions (compared to 60 in v1) and a new scoring system that better distinguishes performance differences among models (2023-12-20). |
| CRUXEval | MIT CSAIL | CRUXEval | CRUXEval is a benchmark for evaluating code reasoning, understanding, and execution. It includes 800 Python functions and their input-output pairs, testing input prediction and output prediction tasks. Many models that perform well on HumanEval underperform on CRUXEval, highlighting the need for improved code reasoning capabilities. The best model, GPT-4 with chain-of-thought (CoT), achieved pass@1 rates of 75% and 81% for input prediction and output prediction, respectively. The benchmark exposes gaps between open-source and closed-source models. GPT-4 failed to fully pass CRUXEval, providing insights into its limitations and directions for improvement (2024-01-05). |
| MLAgentBench | snap-stanford | MLAgentBench | MLAgentBench is a suite of end-to-end machine learning (ML) research tasks for benchmarking AI research agents. These agents aim to autonomously develop or improve an ML model based on a given dataset and ML task description. Each task represents an interactive environment that directly reflects what human researchers encounter. Agents can read available files, run multiple experiments on compute clusters, and analyze results to achieve the specified research objectives. Specifically, it includes 15 diverse ML engineering tasks that can be accomplished by attempting different ML methods, data processing, architectures, and training processes (2023-10-05). |
| AlignBench | THUDM | AlignBench | AlignBench is a comprehensive and multi-dimensional benchmark for evaluating the alignment performance of Chinese large language models. It constructs a human-in-the-loop data creation process to ensure dynamic data updates. AlignBench employs a multi-dimensional, rule-based model evaluation method (LLM-as-Judge) and combines chain-of-thought (CoT) to generate multi-dimensional analyses and final comprehensive scores for model responses, enhancing the reliability and explainability of evaluations (2023-12-01). |
| UltraEval | OpenBMB | UltraEval | UltraEval is an open-source foundational model capability evaluation framework offering a lightweight and easy-to-use evaluation system that supports mainstream large model performance assessments. Its key features include: (1) a lightweight and user-friendly evaluation framework with intuitive design, minimal dependencies, easy deployment, and good scalability for various evaluation scenarios; (2) flexible and diverse evaluation methods with unified prompt templates and rich evaluation metrics, supporting customization; (3) efficient and rapid inference deployment supporting multiple model deployment solutions, including torch and vLLM, and enabling multi-instance deployment to accelerate the evaluation process; (4) a transparent and open leaderboard with publicly accessible, traceable, and reproducible evaluation results driven by the community to ensure transparency; and (5) official and authoritative evaluation data using widely recognized official datasets to guarantee evaluation fairness and standardization, ensuring result comparability and reproducibility (2023-11-24). |
| IFEval | google-research | Instruction Following Eval | Following natural language instructions is a core capability of large language models. However, the evaluation of this capability lacks standardization: human evaluation is expensive, slow, and lacks objective reproducibility, while automated evaluation based on LLMs may be biased by the evaluator LLM's capabilities or limitations. To address these issues, researchers at Google introduced Instruction Following Evaluation (IFEval), a simple and reproducible benchmark focusing on a set of "verifiable instructions," such as "write over 400 words" and "mention the AI keyword at least 3 times." IFEval identifies 25 such verifiable instructions and constructs approximately 500 prompts, each containing one or more verifiable instructions (2023-11-15). |
| LLMBar | princeton-nlp | LLMBar | LLMBar is a challenging meta-evaluation benchmark designed to test the ability of LLM evaluators to identify instruction-following outputs. It contains 419 instances, each consisting of an instruction and two outputs: one faithfully and correctly following the instruction, and the other deviating from it. Each instance also includes a gold label indicating which output is objectively better (2023-10-29). |
| HalluQA | Fudan, Shanghai AI Lab | HalluQA | HalluQA is a Chinese LLM hallucination evaluation benchmark, featuring 450 data points including 175 misleading entries, 69 hard misleading entries, and 206 knowledge-based entries. Each question has an average of 2.8 correct and incorrect answers annotated. To enhance the usability of HalluQA, the authors designed a GPT-4-based evaluation method. Specifically, hallucination criteria and correct answers are input as instructions to GPT-4, which evaluates whether the model's response contains hallucinations. |
| FMTI | Stanford | FMTI | The Foundation Model Transparency Index (FMTI) evaluates the transparency of developers in model training and deployment across 100 indicators, including data, computational resources, and labor. Evaluations of flagship models from 10 companies reveal an average transparency score of only 37/100, indicating significant room for improvement. |
| ColossalEval | Colossal-AI | ColossalEval | A project by Colossal-AI offering a unified evaluation workflow for assessing language models on public datasets or custom datasets using traditional metrics and GPT-assisted evaluations. |
| LLMEval²-WideDeep | Alibaba Research | LLMEval² | Constructed as the largest and most diverse English evaluation benchmark for LLM evaluators, featuring 15 tasks, 8 capabilities, and 2,553 samples. Experimental results indicate that a wider network (involving many reviewers) with two layers (one round of discussion) performs best, improving the Kappa correlation coefficient from 0.28 to 0.34. WideDeep is also utilized to assist in evaluating Chinese LLMs, accelerating the evaluation process by 4.6 times and reducing costs by 60%. |
| Aviary | Ray Project | Aviary | Enables interaction with various large language models (LLMs) in one place. Direct comparison of different model outputs, ranking by quality, and obtaining cost and latency estimates are supported. It particularly supports models hosted on Hugging Face and in many cases, also supports DeepSpeed inference acceleration. |
| Do-Not-Answer | Libr-AI | Do-Not-Answer | An open-source dataset designed to evaluate the safety mechanisms of LLMs at a low cost. It consists of prompts that responsible language models should not respond to. In addition to human annotations, it implements model-based evaluation, where a BERT-like evaluator fine-tuned 600 million times achieves results comparable to humans and GPT-4. |
| LucyEval | Oracle | LucyEval | Chinese LLM maturity evaluation—LucyEval can objectively test various aspects of model capabilities, identify model shortcomings, and help designers and engineers more accurately adjust and train models, aiding LLMs in advancing toward greater intelligence. |
| Zhujiu | Institute of Automation, CAS | Zhujiu | Covers seven capability dimensions and 51 tasks; employs three complementary evaluation methods; offers comprehensive Chinese benchmarking with English evaluation capabilities. |
| ChatEval | THU-NLP | ChatEval | ChatEval aims to simplify the human evaluation process of generated text. Given different text fragments, roles (played by master's students) in ChatEval can autonomously discuss nuances and differences, providing judgments based on their designated roles. |
| FlagEval | Zhiyuan/Tsinghua | FlagEval | Produced by Zhiyuan, combining subjective and objective scoring to offer LLM score rankings. |
| InfoQ Comprehensive LLM Evaluation | InfoQ | InfoQ Evaluation | Chinese-oriented ranking: ChatGPT > Wenxin Yiyang > Claude > Xinghuo. |
This HTML preview is truncated for page performance. The canonical Markdown file contains the complete snapshot.
Why MDRSS assigned this score
- evidence comes from multiple domains
- some evidence URLs look like primary-source hosts
Evidence (4)
concept:prompting-and-agent-evaluationorg:collider-club Discussion 0
Sign in to join the discussion.