Table of Contents

Snapshot 2026-08-04 16:17:00 UTC · version 1

published
C
Collider.club487 cards · 9.8/10 MDRSS

Awesome-LLM-Eval: a curated list of tools, datasets/benchmark, demos, leaderboard, papers, docs and models, mainly for Evaluation on Large Language Models and exploring the boundaries and limits of Generative AI. Use it to build a structured path from fundamentals to hands-on practice.

MARKDOWN SNAPSHOT

Loading…

Direct .mdRaw + metadata0 commentsMDRSS 9.8/10
INDEXABLE MARKDOWN SNAPSHOT

Research document

Open canonical .md

Table of Contents

Awesome-LLM-Eval: a curated list of tools, datasets/benchmark, demos, leaderboard, papers, docs and models, mainly for Evaluation on Large Language Models and exploring the boundaries and limits of Generative AI. Use it to build a structured path from fundamentals to hands-on practice.

Editorial note: curated source snapshot published by Collider.club under the MIT License. Source attribution is preserved in the front matter.

Source snapshot

Awesome LLM Eval

English | 中文

Awesome-LLM-Eval: a curated list of tools, datasets/benchmark, demos, leaderboard, papers, docs and models, mainly for Evaluation on Large Language Models and exploring the boundaries and limits of Generative AI.

The is the official project of our survey: Beyond Benchmark: LLMs Evaluation with an Anthropomorphic and Value-oriented Roadmap.

NOTE: As we cannot update the arXiv paper in real time, please refer to this repo for the latest updates and the paper may be updated later. We also welcome any pull request or issues to help us improve this work. Your contributions will be acknowledged in acknowledgements.

If you find our survey useful, please kindly cite our paper:

@misc{wang2025llmevalroadmap,
      title={Beyond Benchmark: LLMs Evaluation with an Anthropomorphic and Value-oriented Roadmap}, 
      author={Jun Wang and Ninglun Gu and Kailai Zhang and Zijiao Zhang and Yelun Bao and Jin Yang and Xu Yin and Liwei Liu and Yihuan Liu and Pengyong Li and Gary G. Yen and Junchi Yan},
      year={2025},
      eprint={2508.18646},
      archivePrefix={arXiv},
      primaryClass={cs.AI},
      url={https://arxiv.org/abs/2508.18646}, 
}

Table of Contents

News

Anthropomorphic-Taxonomy

Typical Intelligence Quotient (IQ)-General Intelligence evaluation benchmarks

Name Year Task Type Institution Evaluation Focus Datasets Url
MMLU-Pro 2024 Multi-Choice Knowledge TIGER-AI-Lab Subtle Reasoning, Fewer Noise MMLU-Pro link
DyVal 2024 Dynamic Evaluation Microsoft Data Pollution, Complexity Control DyVal link
PertEval 2024 General USTC Knowledge capacity PertEval link
LV-Eval 2024 Long Text QA Infinigence-AI Length Variability, Factuality 11 Subsets link
LLM-Uncertainty-Bench 2024 NLP Tasks Tencent Uncertainty Quantification 5 NLP Tasks link
CommonGen-Eval 2024 Generation AI2 Common Sense CommonGen-lite link
MathBench 2024 Math Shanghai AI Lab Theoretical and practical problem-solving Various link
AIME 2024 Math MAA American Invitational Mathematics Examination Various link
FrontierMath 2024 Math Epoch AI Original, challenging mathematics problems Various link
FELM 2023 Factuality HKUST Factuality 847 Questions link
Just-Eval-Instruct 2023 General AI2 Mosaic Helpfulness, Explainability Various link
MLAgentBench 2023 ML Research snap-stanford End-to-End ML Tasks 15 Tasks link
UltraEval 2023 General OpenBMB Lightweight, Flexible, Fast Various link
FMTI 2023 Transparency Stanford Model Transparency 100 Metrics link
BAMBOO 2023 Long Text RUCAIBox Long Text Modeling 10 Datasets link
TRACE 2023 Continuous Learning Fudan University Continuous Learning 8 Datasets link
ColossalEval 2023 General Colossal-AI Unified Evaluation Various link
LLMEval² 2023 General AlibabaResearch Wide and Deep Evaluation 2,553 Samples link
BigBench 2023 General Google knowledge, language, reasoning Various link
LucyEval 2023 General Oracle Maturity Assessment Various link
Zhujiu 2023 General IACAS Comprehensive Evaluation 51 Tasks link
ChatEval 2023 Chat THU-NLP Human-like Evaluation Various link
FlagEval 2023 General THU Subjective and Objective Scoring Various link
AlpacaEval 2023 General tatsu-lab Automatic Evaluation Various link
GPQA 2023 General NYU Graduate-Level Google-Proof QA Various link
MuSR 2023 Reasoning Zayne Sprague Narrative-Based Reasoning 756 link
FreshQA 2023 Knowledge FreshLLMs Current World Knowledge 599 link
AGIEval 2023 General Microsoft Human-Centric Reasoning NA link
SummEdits 2023 General Salesforce Inconsistency Detection 6,348 link
ScienceQA 2022 Reasoning UCLA Science Reasoning 21,208 link
e-CARE 2022 Reasoning HIT Explainable Causality 21,000 link
BigBench Hard 2022 Reasoning BigBench Challenging Subtasks 6,500 link
PlanBench 2022 Reasoning ASU Action Planning 11,113 link
MGSM 2022 Math Google Grade-school math problems in 10 languages Various link
MATH 2021 Math UC Berkeley Mathematical Problem Solving Various link
GSM8K 2021 Math OpenAI Diverse grade school math word problems Various link
SVAMP 2021 Math Microsoft Arithmetic Reasoning 1,000 link
SpartQA 2021 Reasoning MSU Textual Spatial QA 510 link
MLSUM 2020 General Thomas Scialom News Summarization 535,062 link
Natural Questions 2019 Language, Reasoning Google Search-Based QA 300,000 link
ANLI 2019 Language, Reasoning Facebook AI Adversarial Reasoning 169,265 link
BoolQ 2019 Language, Reasoning Google Binary QA 16,000 link
SuperGLUE 2019 Language, Reasoning NYU Advanced GLUE Tasks NA link
DROP 2019 Language, Reasoning UCI NLP Paragraph-Level Reasoning 96,000 link
HellaSwag 2019 Language, Reasoning AI2 Commonsense Inference 59,950 link
Winogrande 2019 Language, Reasoning AI2 Pronoun Disambiguation 44,000 link
PIQA 2019 Language, Reasoning AI2 Physical Interaction QA 18,000 link
HotpotQA 2018 Language, Reasoning HotpotQA Explainable QA 113,000 link
GLUE 2018 Language, Reasoning NYU Foundational NLU Tasks NA link
OpenBookQA 2018 Language, Reasoning AI2 Open Book Exams 12,000 link
SQuAD2.0 2018 Language, Reasoning Stanford University Unanswerable Questions 150,000 link
ARC 2018 Language, Reasoning AI2 AI2 Reasoning Challenge 7,787 link
SWAG 2018 Language, Reasoning AI2 Adversarial Commonsense 113,000 link
CommonsenseQA 2018 Language, Reasoning AI2 Commonsense Reasoning 12,102 link
RACE 2017 Language, Reasoning CMU Exam-Style QA 100,000 link
SciQ 2017 Language, Reasoning AI2 Crowd-Sourced Science 13,700 link
TriviaQA 2017 Language, Reasoning AI2 Distant Supervision 650,000 link
MultiNLI 2017 Language, Reasoning NYU Cross-Genre Entailment 433,000 link
SQuAD 2016 Language, Reasoning Stanford University Wikipedia-Based QA 100,000 link
LAMBADA 2016 Language, Reasoning CIMEC Discourse Context 12,684 link
MS MARCO 2016 Language, Reasoning Microsoft Search-Based QA 1,112,939 link

Typical Professional Quotient (PQ)-Professional Expertise evaluation benchmarks

Domain Name Institution Scope of Tasks Unique Contributions Url
BLURB Mindrank AI Six diverse NLP tasks, thirteen datasets A macro-average score across all tasks link
Seismometer Epic Using local data and workflows patient demographics, clinical interventions, and outcomes link
Healthcare Medbench OpenMEDLab Emphasizes scientific rigor and fairness 40,041 questions from medical exams and reports link
GenMedicalEval E 16 majors, 3 training stages, 6 clinical scenarios Open-ended metrics and automated assessment models link
PsyEval SJTU Six subtasks covering three dimensions Customized benchmark for mental health LLMs link
Fin-Eva Ant Group Wealth management, insurance, investment research Both industrial and academic financial evaluations link
Finance FinEval SUFE-AIFLM-Lab Multiple-choice QA on finance, economics, accounting Focuses on high-quality evaluation questions link
OpenFinData Shanghai AI Lab Multi-scenario financial tasks First comprehensive finance evaluation dataset link
FinBen FinAI 35 datasets across 23 financial tasks Inductive reasoning, quantitative reasoning link
LAiW Sichuan University 13 fundamental legal NLP tasks Divides legal NLP capabilities into three major abilities link
Legal LawBench Nanjing University Legal entity recognition, reading comprehension Real-world tasks, "abstention rate" metric link
LegalBench Stanford University 162 tasks covering six types of legal reasoning Enables interdisciplinary conversations link
LexEval Tsinghua University Legal cognitive abilities to organize different tasks Larger legal evaluation dataset, examining the ethical issues link
SPEC5G Purdue University Security-related text classification and summarization 5G protocol analysis automation link
Telecom TeleQnA Huawei(Paris) General telecom inquiries Proficiency in telecom-related questions link
OpsEval Tsinghua University Wired network ops, 5G, database ops Focus on AIOps, evaluates proficiency link
TelBench SK Telecom Math modeling, open-ended QA, code generation Holistic evaluation in telecom link
TelecomGPT UAE Telecom Math Modeling, Open QnA and Code Tasks Holistic evaluation in telecom link
Linguistic Queen's University Multiple language-centric tasks zero-shot evaluation link
TelcoLM Orange Multiple-choice questionnaires Domain-specific data (800M tokens, 80K instructions) link
ORAN-Bench-13K GMU Multiple-choice questions Open Radio Access Networks (O-RAN) link
Open-Telco Benchmarks GSMA Multiple language-centric tasks zero-shot evaluation link
FullStackBench ByteDance Code writing, debugging, code review Featuring the most recent Stack Overflow QA link
Coding StackEval Prosus AI 11 real-world scenarios, 16 languages Evaluation across diverse & practical coding environments link
CodeBenchGen Various Institutions Execution-based code generation tasks Benchmarks scaling with the size and complexity link
HumanEval University of Washington Rigorous testing Stricter protocol for assessing correctness of generated code link
APPS University of California Coding challenges from competitive platforms Checking problem-solving of generated code on test cases link
MBPP Google Research Programming problems sourced from various origins Diverse programming tasks link
ClassEval Tsinghua University Class-level code generation Manually crafted, object-oriented programming concepts link
CoderEval Peking University Pragmatic code generation Proficiency to generate functional code patches for described issues link
MultiPL-E Princeton University Neural code generation Benchmarking neural code generation models link
CodeXGLUE Microsoft Code intelligence Wide tasks covering: code-code, text-code, code-text and text-text link
EvoCodeBench Peking University Evolving code generation benchmark Aligned with real-world code repositories, evolving over time link

Typical Emotional Quotient (EQ)-Alignment Ability evaluation benchmarks

Name Year Task Type Institution Category Datasets Url
DiffAware 2025 Bias Stanford General Bias 8 datasets link
CASE-Bench 2025 Safety Cambridge Context-Aware Safety CASE-Bench link
Fairness 2025 Fairness PSU Distributive Fairness - -
HarmBench 2024 Safety UIUC Adversarial Behaviors 510 link
SimpleQA 2024 Safety OpenAI Factuality 4,326 link
AgentHarm 2024 Safety BEIS Malicious Agent Tasks 110 link
StrongReject 2024 Safety dsbowen Attack Resistance n/a link
LLMBar 2024 Instruction Princeton Instruction Following 419 Instances link
AIR-Bench 2024 Safety Stanford Regulatory Alignment 5,694 link
TrustLLM 2024 General TrustLLM Trustworthiness 30+ link
RewardBench 2024 Alignment AIAI Human preference RewardBench link
EQ-Bench 2024 Emotion Paech Emotional intelligence 171 Questions link
Forbidden 2023 Safety CISPA Jailbreak Detection 15,140 link
MaliciousInstruct 2023 Safety Princeton Malicious Intentions 100 link
SycophancyEval 2023 Safety Anthropic Opinion Alignment n/a link
DecodingTrust 2023 Safety UIUC Trustworthiness 243,877 link
AdvBench 2023 Safety CMU Adversarial Attacks 1,000 link
XSTest 2023 Safety Bocconi Safety Overreach 450 link
OpinionQA 2023 Safety tatsu-lab Demographic Alignment 1,498 link
SafetyBench 2023 Safety THU Content Safety 11,435 link
HarmfulQA 2023 Safety declare-lab Harmful Topics 1,960 link
QHarm 2023 Safety vinid Safety Sampling 100 link
BeaverTails 2023 Safety PKU Red Teaming 334,000 link
DoNotAnswer 2023 Safety Libr-AI Safety Mechanisms 939 link
AlignBench 2023 Alignment THUDM Alignment, Reliability Various link
IFEval 2023 Instruction Google Instruction Following 500 Prompts link
ToxiGen 2022 Safety Microsoft Toxicity Detection 274,000 link
HHH 2022 Safety Anthropic Human Preferences 44,849 link
RedTeam 2022 Safety Anthropic Red Teaming 38,921 link
BOLD 2021 Bias Amazon Bias in Generation 23,679 link
BBQ 2021 Bias NYU Social Bias 58,492 link
StereoSet 2020 Bias McGill Stereotype Detection 4,229 link
ETHICS 2020 Ethics Berkeley Moral Judgement 134,400 link
ToxicityPrompt 2020 Safety AllenAI Toxicity Assessment 99,442 link
CrowS-Pairs 2020 Bias NYU Stereotype Measurement 1,508 link
SEAT 2019 Bias Princeton Encoder Bias n/a link
WinoGender 2018 Bias UMass Gender Bias 720 link

Tools

Name Organization Website Description
prometheus-eval prometheus-eval prometheus-eval PROMETHEUS Open Evaluation Dedicated Language Model, which is more powerful than its predecessor. It can closely imitate the judgments of humans and GPT-4. Additionally, it can handle both direct evaluation and pairwise ranking formats, and can be used with user-defined evaluation criteria. On four direct evaluation benchmarks and four pairwise ranking benchmarks, PROMETHEUS 2 achieves the highest correlation and consistency with human evaluators and proprietary language models among all tested open-source evaluation language models (2024-05-04).
athina-evals athina-ai athina-ai Athina-ai is an open-source library that provides plug-and-play preset evaluations and a modular, extensible framework for writing and running evaluations. It helps engineers systematically improve the reliability and performance of their large language models through evaluation-driven development. Athina-ai offers a system for evaluation-driven development, overcoming the limitations of traditional workflows, enabling rapid experimentation, and providing customizable evaluators with consistent metrics.
LeaderboardFinder Huggingface LeaderboardFinder LeaderboardFinder helps you find suitable leaderboards for specific scenarios, a leaderboard of leaderboards (2024-04-02).
LightEval Huggingface lighteval LightEval is a lightweight framework developed by Hugging Face for evaluating large language models (LLMs). Originally designed as an internal tool for assessing Hugging Face's recently released LLM data processing library datatrove and LLM training library nanotron, it is now open-sourced for community use and improvement. Key features of LightEval include: (1) lightweight design, making it easy to use and integrate; (2) an evaluation suite supporting multiple tasks and models; (3) compatibility with evaluation on CPUs or GPUs, and integration with Hugging Face's acceleration library (Accelerate) and frameworks like Nanotron; (4) support for distributed evaluation, which is particularly useful for evaluating large models; (5) applicability to all benchmarks on the Open LLM Leaderboard; and (6) customizability, allowing users to add new metrics and tasks to meet specific evaluation needs (2024-02-08).
LLM Comparator Google LLM Comparator A visual analytical tool for comparing and evaluating large language models (LLMs). Compared to traditional human evaluation methods, this tool offers a scalable automated approach to comparative evaluation. It leverages another LLM as an evaluator to demonstrate quality differences between models and provide reasons for these differences. Through interactive tables and summary visualizations, the LLM Comparator helps users understand why models perform well or poorly in specific contexts, as well as the qualitative differences between model responses. Developed in collaboration with Google researchers and engineers, this tool has been widely used internally at Google, attracting over 400 users and evaluating more than 1,000 experiments within three months (2024-02-16).
Arthur Bench Arthur-AI Arthur Bench Arthur Bench is an open-source evaluation tool designed to compare and analyze the performance of large language models (LLMs). It supports various evaluation tasks, including question answering, summarization, translation, and code generation, and provides detailed reports on LLM performance across these tasks. Key features and advantages of Arthur Bench include: (1) model comparison, enabling the evaluation of different suppliers, versions, and training datasets of LLMs; (2) prompt and hyperparameter evaluation, assessing the impact of different prompts on LLM performance and testing the control of model behavior through various hyperparameter settings; (3) task definition and model selection, allowing users to define specific evaluation tasks and select evaluation targets from a range of supported LLM models; (4) parameter configuration, enabling users to adjust prompts and hyperparameters to finely control LLM behavior; (5) automated evaluation workflows, simplifying the execution of evaluation tasks; and (6) application scenarios such as model selection and validation, budget and privacy optimization, and the transformation of academic benchmarks into real-world performance evaluations. Additionally, it offers comprehensive scoring metrics, supports both local and cloud versions, and encourages community collaboration and project development (2023-10-06).
llm-benchmarker-suite FormulaMonks llm-benchmarker-suite This open-source initiative aims to address fragmentation and ambiguity in LLM benchmarking. The suite provides a structured methodology, a collection of diverse benchmarks, and toolkits to streamline the assessment of LLM performance. By offering a common platform, this project seeks to promote collaboration, transparency, and high-quality research in NLP.
autoevals braintrust autoevals AutoEvals is an AI model output evaluation tool that leverages best practices to quickly and easily assess AI model outputs. It integrates multiple automatic evaluation methods, supports customizable evaluation prompts and custom scorers, and simplifies the evaluation process of model outputs. Autoevals incorporates model-graded evaluation for various subjective tasks, including fact-checking, safety, and more. Many of these evaluations are adapted from OpenAI's excellent evals project but are implemented in a flexible way to allow users to tweak prompts and debug outputs.
EVAL OPENAI EVAL EVAL is a tool developed by OpenAI for evaluating large language models (LLMs). It can test the performance and generalization capabilities of models across different tasks and datasets.
lm-evaluation-harness EleutherAI lm-evaluation-harness lm-evaluation-harness is a tool developed by EleutherAI for evaluating large language models (LLMs). It can test the performance and generalization capabilities of models across different tasks and datasets.
lm-evaluation AI21Labs lm-evaluation Evaluations and reproducing the results from the Jurassic-1 Technical Paper, with current support for running tasks via both the AI21 Studio API and OpenAI's GPT-3 API.
OpenCompass Shanghai AI Lab OpenCompass OpenCompass is a one-stop platform for evaluating large models. Its main features include: open-source and reproducible evaluation schemes; comprehensive capability dimensions covering five major areas with over 50 datasets and approximately 300,000 questions to assess model capabilities; support for over 20 Hugging Face and API models; distributed and efficient evaluation with one-line task splitting and distributed evaluation, enabling full evaluation of trillion-parameter models within hours; diverse evaluation paradigms supporting zero-shot, few-shot, and chain-of-thought evaluations with standard or conversational prompt templates to easily elicit peak model performance.
Large language model evaluation and workflow framework from Phase AI wgryc phasellm A framework provided by Phase AI for evaluating and managing LLMs, helping users select appropriate models, datasets, and metrics, as well as visualize and analyze results.
Evaluation benchmark for LLM FreedomIntelligence LLMZoo LLMZoo is an evaluation benchmark for LLMs developed by FreedomIntelligence, featuring multiple domain and task datasets, metrics, and pre-trained models with results.
Holistic Evaluation of Language Models (HELM) Stanford HELM HELM is a comprehensive evaluation method for LLMs proposed by the Stanford research team, considering multiple aspects such as model language ability, knowledge, reasoning, fairness, and safety.
A lightweight evaluation tool for question-answering Langchain auto-evaluator auto-evaluator is a lightweight tool developed by Langchain for evaluating question-answering systems. It can automatically generate questions and answers and calculate metrics such as model accuracy, recall, and F1 score.
PandaLM WeOpenML PandaLM PandaLM is an LLM assessment tool developed by WeOpenML for automated and reproducible evaluation. It allows users to select appropriate datasets, metrics, and models based on their needs and preferences, and generates reports and charts.
FlagEval Tsinghua University FlagEval FlagEval is an evaluation platform for LLMs developed by Tsinghua University, offering multiple tasks and datasets, as well as online testing, leaderboards, and analysis functions.
AlpacaEval tatsu-lab alpaca_eval AlpacaEval is an evaluation tool for LLMs developed by tatsu-lab, capable of testing models across various languages, domains, and tasks, and providing explainability, robustness, and credibility metrics.
Prompt flow Microsoft promptflow A set of development tools designed by Microsoft to simplify the end-to-end development cycle of AI applications based on LLMs, from conception, prototyping, testing, and evaluation to production deployment and monitoring. It makes prompt engineering easier and enables the development of product-level LLM applications.
DeepEval mr-gpt DeepEval DeepEval is a simple-to-use, open-source LLM evaluation framework. Similar to Pytest but specialized for unit testing LLM outputs, it incorporates the latest research to evaluate LLM outputs based on metrics such as G-Eval, hallucination, answer relevance, RAGAS, etc., utilizing LLMs and various other NLP models that run locally on your machine for evaluation.
CONNER Tencent AI Lab CONNER CONNER is a comprehensive large model knowledge evaluation framework designed to systematically and automatically assess the information generated from six critical perspectives: factuality, relevance, coherence, informativeness, usefulness, and validity.

Datasets or Benchmarks

General

Name Organization Website Description
MMLU-Pro TIGER-AI-Lab MMLU-Pro MMLU-Pro is an improved version of the MMLU dataset. MMLU has long been a reference for multiple-choice knowledge datasets. However, recent studies have shown that it contains noise (some questions are unanswerable) and is too easy (due to the evolution of model capabilities and increased contamination). MMLU-Pro provides ten options instead of four, requires reasoning on more questions, and has undergone expert review to reduce noise. It is of higher quality and more challenging than the original. MMLU-Pro reduces the impact of prompt variations on model performance, a common issue with its predecessor MMLU. Research indicates that models using "Chain of Thought" reasoning perform better on this new benchmark, suggesting that MMLU-Pro is better suited for evaluating the subtle reasoning abilities of AI. (2024-05-20)
TrustLLM Benchmark TrustLLM TrustLLM TrustLLM is a benchmark for evaluating the trustworthiness of large language models. It covers six dimensions of trustworthiness and includes over 30 datasets to comprehensively assess the functional capabilities of LLMs, ranging from simple classification tasks to complex generative tasks. Each dataset presents unique challenges and has benchmarked 16 mainstream LLMs (including commercial and open-source models).
DyVal Microsoft DyVal Concerns have been raised about the potential data contamination in the vast training corpora of LLMs. Additionally, the static nature and fixed complexity of current benchmarks may not adequately measure the evolving capabilities of LLMs. DyVal is a general and flexible protocol for dynamically evaluating LLMs. Leveraging the advantages of directed acyclic graphs, DyVal dynamically generates evaluation samples with controllable complexity. It has created challenging evaluation sets for reasoning tasks such as mathematics, logical reasoning, and algorithmic problems. Various LLMs, from Flan-T5-large to GPT-3.5-Turbo and GPT-4, have been evaluated. Experiments show that LLMs perform worse on DyVal-generated samples of different complexities, highlighting the importance of dynamic evaluation. The authors also analyze failure cases and results of different prompting methods. Furthermore, DyVal-generated samples not only serve as evaluation sets but also aid in fine-tuning to enhance LLM performance on existing benchmarks. (2024-04-20)
RewardBench AIAI RewardBench RewardBench is an evaluation benchmark for language model reward models, assessing the strengths and weaknesses of various models. It reveals that existing models still exhibit significant shortcomings in reasoning and instruction following. It includes a Leaderboard, Code, and Dataset (2024-03-20).
LV-Eval Infinigence-AI LVEval LV-Eval is a long-text evaluation benchmark featuring five length tiers (16k, 32k, 64k, 128k, and 256k), with a maximum text test length of 256k. The average text length of LV-Eval is 102,380 characters, with a minimum/maximum text length of 11,896/387,406 characters. LV-Eval primarily consists of two types of evaluation tasks: single-hop QA and multi-hop QA, encompassing 11 sub-datasets in Chinese and English. During its design, LV-Eval introduced three key technologies: Confusion Facts Insertion (CFI) to enhance challenge, Keyword and Phrase Replacement (KPR) to reduce information leakage, and Answer Keywords (AK) based evaluation metrics (combining answer keywords and word blacklists) to improve the objectivity of evaluation results (2024-02-06).
LLM-Uncertainty-Bench Tencent LLM-Uncertainty-Bench A new benchmark method for LLMs has been introduced, incorporating uncertainty quantification. Based on nine LLMs tested across five representative NLP tasks, it was found that: I) More accurate LLMs may exhibit lower certainty; II) Larger-scale LLMs may display greater uncertainty than smaller models; III) Instruction fine-tuning tends to increase the uncertainty of LLMs. These findings underscore the importance of including uncertainty in LLM evaluations (2024-01-22).
Psychometrics Eval Microsoft Research Asia Psychometrics Eval Microsoft Research Asia has proposed a generalized evaluation method for AI based on psychometrics, aiming to address limitations in traditional evaluation methods concerning predictive power, information volume, and test tool quality. This approach draws on psychometric theories to identify key psychological constructs of AI, design targeted tests, and apply Item Response Theory for precise scoring. It also introduces concepts of reliability and validity to ensure evaluation reliability and accuracy. This framework extends psychometric methods to assess AI performance in handling unknown complex tasks but also faces open questions such as distinguishing between AI "individuals" and "populations," addressing prompt sensitivity, and evaluating differences between human and AI constructs (2023-10-19).
CommonGen-Eval AllenAI CommonGen-Eval A study using the CommonGen-lite dataset to evaluate LLMs, employing GPT-4 for assessment and comparing the performance of different models, with results listed on the leaderboard (2024-01-04).
felm HKUST felm FELM is a meta-benchmark for evaluating the factual assessment of large language models. The benchmark comprises 847 questions spanning five distinct domains: world knowledge, science/technology, writing/recommendation, reasoning, and mathematics. Prompts corresponding to each domain are gathered from various sources, including standard datasets like TruthfulQA, online platforms like GitHub repositories, ChatGPT-generated prompts, or those drafted by authors. For each response, fine-grained annotation at the segment level is employed, including reference links, identified error types, and reasons behind these errors as provided by annotators (2023-10-03).
just-eval AI2 Mosaic just-eval A GPT-based evaluation tool for multi-faceted and explainable assessment of LLMs, capable of evaluating aspects such as helpfulness, clarity, factuality, depth, and engagement (2023-12-05).
EQ-Bench EQ-Bench EQ-Bench A benchmark for evaluating the emotional intelligence of language models, featuring 171 questions (compared to 60 in v1) and a new scoring system that better distinguishes performance differences among models (2023-12-20).
CRUXEval MIT CSAIL CRUXEval CRUXEval is a benchmark for evaluating code reasoning, understanding, and execution. It includes 800 Python functions and their input-output pairs, testing input prediction and output prediction tasks. Many models that perform well on HumanEval underperform on CRUXEval, highlighting the need for improved code reasoning capabilities. The best model, GPT-4 with chain-of-thought (CoT), achieved pass@1 rates of 75% and 81% for input prediction and output prediction, respectively. The benchmark exposes gaps between open-source and closed-source models. GPT-4 failed to fully pass CRUXEval, providing insights into its limitations and directions for improvement (2024-01-05).
MLAgentBench snap-stanford MLAgentBench MLAgentBench is a suite of end-to-end machine learning (ML) research tasks for benchmarking AI research agents. These agents aim to autonomously develop or improve an ML model based on a given dataset and ML task description. Each task represents an interactive environment that directly reflects what human researchers encounter. Agents can read available files, run multiple experiments on compute clusters, and analyze results to achieve the specified research objectives. Specifically, it includes 15 diverse ML engineering tasks that can be accomplished by attempting different ML methods, data processing, architectures, and training processes (2023-10-05).
AlignBench THUDM AlignBench AlignBench is a comprehensive and multi-dimensional benchmark for evaluating the alignment performance of Chinese large language models. It constructs a human-in-the-loop data creation process to ensure dynamic data updates. AlignBench employs a multi-dimensional, rule-based model evaluation method (LLM-as-Judge) and combines chain-of-thought (CoT) to generate multi-dimensional analyses and final comprehensive scores for model responses, enhancing the reliability and explainability of evaluations (2023-12-01).
UltraEval OpenBMB UltraEval UltraEval is an open-source foundational model capability evaluation framework offering a lightweight and easy-to-use evaluation system that supports mainstream large model performance assessments. Its key features include: (1) a lightweight and user-friendly evaluation framework with intuitive design, minimal dependencies, easy deployment, and good scalability for various evaluation scenarios; (2) flexible and diverse evaluation methods with unified prompt templates and rich evaluation metrics, supporting customization; (3) efficient and rapid inference deployment supporting multiple model deployment solutions, including torch and vLLM, and enabling multi-instance deployment to accelerate the evaluation process; (4) a transparent and open leaderboard with publicly accessible, traceable, and reproducible evaluation results driven by the community to ensure transparency; and (5) official and authoritative evaluation data using widely recognized official datasets to guarantee evaluation fairness and standardization, ensuring result comparability and reproducibility (2023-11-24).
IFEval google-research Instruction Following Eval Following natural language instructions is a core capability of large language models. However, the evaluation of this capability lacks standardization: human evaluation is expensive, slow, and lacks objective reproducibility, while automated evaluation based on LLMs may be biased by the evaluator LLM's capabilities or limitations. To address these issues, researchers at Google introduced Instruction Following Evaluation (IFEval), a simple and reproducible benchmark focusing on a set of "verifiable instructions," such as "write over 400 words" and "mention the AI keyword at least 3 times." IFEval identifies 25 such verifiable instructions and constructs approximately 500 prompts, each containing one or more verifiable instructions (2023-11-15).
LLMBar princeton-nlp LLMBar LLMBar is a challenging meta-evaluation benchmark designed to test the ability of LLM evaluators to identify instruction-following outputs. It contains 419 instances, each consisting of an instruction and two outputs: one faithfully and correctly following the instruction, and the other deviating from it. Each instance also includes a gold label indicating which output is objectively better (2023-10-29).
HalluQA Fudan, Shanghai AI Lab HalluQA HalluQA is a Chinese LLM hallucination evaluation benchmark, featuring 450 data points including 175 misleading entries, 69 hard misleading entries, and 206 knowledge-based entries. Each question has an average of 2.8 correct and incorrect answers annotated. To enhance the usability of HalluQA, the authors designed a GPT-4-based evaluation method. Specifically, hallucination criteria and correct answers are input as instructions to GPT-4, which evaluates whether the model's response contains hallucinations.
FMTI Stanford FMTI The Foundation Model Transparency Index (FMTI) evaluates the transparency of developers in model training and deployment across 100 indicators, including data, computational resources, and labor. Evaluations of flagship models from 10 companies reveal an average transparency score of only 37/100, indicating significant room for improvement.
ColossalEval Colossal-AI ColossalEval A project by Colossal-AI offering a unified evaluation workflow for assessing language models on public datasets or custom datasets using traditional metrics and GPT-assisted evaluations.
LLMEval²-WideDeep Alibaba Research LLMEval² Constructed as the largest and most diverse English evaluation benchmark for LLM evaluators, featuring 15 tasks, 8 capabilities, and 2,553 samples. Experimental results indicate that a wider network (involving many reviewers) with two layers (one round of discussion) performs best, improving the Kappa correlation coefficient from 0.28 to 0.34. WideDeep is also utilized to assist in evaluating Chinese LLMs, accelerating the evaluation process by 4.6 times and reducing costs by 60%.
Aviary Ray Project Aviary Enables interaction with various large language models (LLMs) in one place. Direct comparison of different model outputs, ranking by quality, and obtaining cost and latency estimates are supported. It particularly supports models hosted on Hugging Face and in many cases, also supports DeepSpeed inference acceleration.
Do-Not-Answer Libr-AI Do-Not-Answer An open-source dataset designed to evaluate the safety mechanisms of LLMs at a low cost. It consists of prompts that responsible language models should not respond to. In addition to human annotations, it implements model-based evaluation, where a BERT-like evaluator fine-tuned 600 million times achieves results comparable to humans and GPT-4.
LucyEval Oracle LucyEval Chinese LLM maturity evaluation—LucyEval can objectively test various aspects of model capabilities, identify model shortcomings, and help designers and engineers more accurately adjust and train models, aiding LLMs in advancing toward greater intelligence.
Zhujiu Institute of Automation, CAS Zhujiu Covers seven capability dimensions and 51 tasks; employs three complementary evaluation methods; offers comprehensive Chinese benchmarking with English evaluation capabilities.
ChatEval THU-NLP ChatEval ChatEval aims to simplify the human evaluation process of generated text. Given different text fragments, roles (played by master's students) in ChatEval can autonomously discuss nuances and differences, providing judgments based on their designated roles.
FlagEval Zhiyuan/Tsinghua FlagEval Produced by Zhiyuan, combining subjective and objective scoring to offer LLM score rankings.
InfoQ Comprehensive LLM Evaluation InfoQ InfoQ Evaluation Chinese-oriented ranking: ChatGPT > Wenxin Yiyang > Claude > Xinghuo.

This HTML preview is truncated for page performance. The canonical Markdown file contains the complete snapshot.

MARKDOWN METRICS
14830words
38headings
634links
2code blocks
MDRSS ASSESSMENT
Scam / risk5/100low
Evidence100/100high confidence
Why MDRSS assigned this score
  • evidence comes from multiple domains
  • some evidence URLs look like primary-source hosts
Evidence (4)
concept:prompting-and-agent-evaluationorg:collider-club

Discussion 0

Sign in to join the discussion.