AlpacaEval
AlpacaEval, along with MT-Bench, is one of the best LLM evaluations for understanding the relative ranking of LLMs compared to their peers. While not perfect, it provides an automated comparison.
AlpacaEval is an automated tool for evaluating instruction-following language models against the AlpacaFarm dataset. It stands out for its human-validated, high-quality assessments that are both cost-effective and rapid.
Imagine you're playing fetch with your dog. You throw the ball in different directions and see how well your dog catches it. AlpacaEval is like that, but instead of throwing a ball, you give your language model (like a fancy talking dog) different tasks like writing stories or summarizing things. By seeing how well it completes these tasks, you get a sense of how good it is at understanding and following instructions.
The evaluator is specifically designed for chat-based large language models (LLMs) and features a leaderboard to benchmark model performance.
AlpacaEval calculates win-rates for models across a variety of tasks, including traditional NLP and instruction-tuning datasets, providing a comprehensive measure of model capabilities.
AlpacaEval is a single-turn benchmark, which means it evaluates models based on their responses to single-turn prompts. It has been used to assess models like OpenAI GPT-4, Mistral Mixtral, Anthropic Claude 2, and others.
Current Leaderboard
As of January 15, 2024, the current leaderboard is led by GPT-4 Turbo, mirroring the human preference results of MT-Bench.
Model Name | Win Rate | License |
|---|---|---|
GPT-4 Turbo | 50.00% | Proprietary |
Yi 34B Chat | 29.66% | Open Source |
GPT-4 | 23.58% | Proprietary |
GPT-4 0314 | 22.07% | Proprietary |
Mistral Medium | 21.86% | Proprietary |
Mixtral 8x7B v0.1 | 18.26% | Open Source |
Claude 2 | 17.19% | Proprietary |
Claude | 16.99% | Proprietary |
Gemini Pro | 16.85% | Proprietary |
Tulu 2+DPO 70B | 15.98% | Open Source |
GPT-4 0613 | 15.76% | Proprietary |
Claude 2.1 | 15.73% | Proprietary |
Mistral 7B v0.2 | 14.72% | Open Source |
GPT 3.5 Turbo 0613 | 14.13% | Proprietary |
LLaMA2 Chat 70B | 13.87% | Open Source |
Cohere Command | 12.90% | Proprietary |
Vicuna 33B v1.3 | 12.71% | Open Source |
OpenHermes-2.5-Mistral | 10.34% | Open Source |
GPT 3.5 Turbo 0301 | 9.62% | Proprietary |
GPT 3.5 Turbo 1106 | 9.18% | Proprietary |
