Introduction
MT-Bench, short for Multi-Turn Bench, is another tool designed to evaluate the capabilities of large language models (LLMs) like me, but with a different twist compared to AlpacaEval. Imagine you're having a conversation with a friend. You throw different topics at them, ask questions, and even give instructions. You want to see how well they follow your train of thought, understand your intentions, and respond in a natural and engaging way. MT-Bench is like that friend, but for your LLM. Instead of you, it throws questions and challenges at the LLM, asking it to follow instructions, answer open-ended questions, and engage in multi-turn dialogues. It then assesses how well the LLM performs in these conversational scenarios.
MT-Bench unique Features
Focuses on conversation: Unlike AlpacaEval's individual tasks, MT-Bench tests how LLMs handle the flow and coherence of a conversation.
Human feedback: The evaluation is based on human judgment, not just pre-defined metrics, making it more nuanced and subjective.
Open-ended questions: It doesn't just ask for simple factual answers, but challenges the LLM to think creatively and respond in a natural way.
In simpler terms:
AlpacaEval tests your LLM's ability to follow specific instructions.
MT-Bench tests your LLM's ability to have a good conversation.
Key features of the MT Bench Leaderboard
Challenging Multi-Turn Benchmark
The MT-Bench incorporates challenging follow-up questions as part of its design, ensuring that models demonstrate a deep understanding of the task at hand.
Three Metrics
The leaderboard uses three metrics for evaluation: Chatbot Arena Elo, based on 42K anonymous votes from Chatbot Arena using the Elo rating system; MT-Bench score, based on a challenging multi-turn benchmark and GPT-4 grading; and MMLU, a widely adopted benchmark.
Regular Updates
The leaderboard is updated regularly, providing a constantly evolving view of the latest LLM performance.
Triangulating relative model performance with MT-Bench and AlpacaEval provides the best benchmark for general performance from both human preference and LLM-as-judge perspectives. While performance on individual use cases may vary between models, these two benchmarks offer the most reliable standard.
MT-Bench Leaderboard
(February 2024)
Model | Arena Elo rating | MT-bench (score) | MMLU | License |
|---|---|---|---|---|
1253 | 9.32 | — | Proprietary | |
1252 | 9.32 | — | Proprietary | |
1224 | 9.18 | — | Proprietary | |
1190 | 8.96 | 86.4 | Proprietary | |
1160 | 9.18 | — | Proprietary | |
1150 | 8.61 | 75.3 | Proprietary | |
1149 | 7.9 | 77 | Proprietary | |
1131 | 8.06 | 78.5 | Proprietary | |
1123 | 8.3 | 70.6 | Apache 2.0 | |
1120 | — | 71.8 | Proprietary | |
1119 | 8.18 | — | Proprietary | |
1116 | 8.39 | — | Proprietary | |
1110 | 7.85 | 73.4 | Proprietary | |
1110 | 7.89 | — | AI2 ImpACT | |
1110 | — | 73.5 | Yi License | |
1111 | — | 71.8 | Proprietary | |
1105 | 7.94 | 70 | Proprietary |
As of February 2, the leaderboard now includes the gpt-4-0125-preview and Gemini Pro via Bard Assistant. Previously, on January 10, the experimental Mixtral Medium model was added, surpassing the performance of all Anthropic models.
The MT Bench Leaderboard, updated regularly, ranks Large Language Models (LLMs) like GPT-3.5-turbo, Vicuna-33B, WizardLM-30B, WizardLM-13B, Guanaco-33B, and Vicuna-13B based on their task performance, providing a dynamic snapshot of the field's top performers as of January 2024.