27 Feb 2024

AlpacaEval

AlpacaEval

AlpacaEval, along with MT-Bench, is one of the best LLM evaluations for understanding the relative ranking of LLMs compared to their peers. While not perfect, it provides an automated comparison.

AlpacaEval is an automated tool for evaluating instruction-following language models against the AlpacaFarm dataset. It stands out for its human-validated, high-quality assessments that are both cost-effective and rapid.

Imagine you're playing fetch with your dog. You throw the ball in different directions and see how well your dog catches it. AlpacaEval is like that, but instead of throwing a ball, you give your language model (like a fancy talking dog) different tasks like writing stories or summarizing things. By seeing how well it completes these tasks, you get a sense of how good it is at understanding and following instructions.

image-20240223-091815.png

The evaluator is specifically designed for chat-based large language models (LLMs) and features a leaderboard to benchmark model performance.

AlpacaEval calculates win-rates for models across a variety of tasks, including traditional NLP and instruction-tuning datasets, providing a comprehensive measure of model capabilities.

AlpacaEval is a single-turn benchmark, which means it evaluates models based on their responses to single-turn prompts. It has been used to assess models like OpenAI GPT-4, Mistral Mixtral, Anthropic Claude 2, and others.

Current Leaderboard

As of January 15, 2024, the current leaderboard is led by GPT-4 Turbo, mirroring the human preference results of MT-Bench.

Model Name

Win Rate

License

GPT-4 Turbo

50.00%

Proprietary

Yi 34B Chat

29.66%

Open Source

GPT-4

23.58%

Proprietary

GPT-4 0314

22.07%

Proprietary

Mistral Medium

21.86%

Proprietary

Mixtral 8x7B v0.1

18.26%

Open Source

Claude 2

17.19%

Proprietary

Claude

16.99%

Proprietary

Gemini Pro

16.85%

Proprietary

Tulu 2+DPO 70B

15.98%

Open Source

GPT-4 0613

15.76%

Proprietary

Claude 2.1

15.73%

Proprietary

Mistral 7B v0.2

14.72%

Open Source

GPT 3.5 Turbo 0613

14.13%

Proprietary

LLaMA2 Chat 70B

13.87%

Open Source

Cohere Command

12.90%

Proprietary

Vicuna 33B v1.3

12.71%

Open Source

OpenHermes-2.5-Mistral

10.34%

Open Source

GPT 3.5 Turbo 0301

9.62%

Proprietary

GPT 3.5 Turbo 1106

9.18%

Proprietary

← Back to Library