https://gorilla.cs.berkeley.edu/leaderboard.html
Leaderboard Try it Out! Blog Gorilla
UC Berkeley Logo Berkeley Function-Calling Leaderboard
Leaderboard
This live leaderboard evaluates the LLM's ability to call functions
(aka tools) accurately. This leaderboard consists of real-world data
and will be updated periodically. For more information on the
evaluation dataset and methodology, please refer to our blog post and
code release.
Abstract Syntax Tree (AST) Evaluation Evaluation by Executing APIs
Rank Overall Model Organization License Simple Multiple Parallel Parallel Simple Multiple Parallel Parallel Relevance
Acc Function Functions Functions Multiple Function Functions Functions Multiple Detection
1 83.80 GPT-4-0125-Preview OpenAI Proprietary 82.18 90.00 90.00 91.00 54.12 70.00 76.00 55.00 87.50
2 83.55 GPT-4-1106-Preview OpenAI Proprietary 81.64 89.50 92.00 92.00 53.53 62.00 72.00 50.00 88.75
3 83.55 OpenFunctions-v2 Gorilla LLM Apache 2.0 88.73 89.50 79.50 78.00 78.82 74.00 76.00 60.00 71.67
4 81.63 GPT-3.5-Turbo OpenAI Proprietary 81.27 88.00 87.50 88.00 74.12 74.00 70.00 47.50 68.33
5 79.46 Mistral-medium Mistral AI Proprietary 80.18 84.50 71.00 68.00 75.88 72.00 62.00 47.50 90.00
6 75.78 Claude-2.1 Anthropic Proprietary 85.64 83.00 72.00 56.50 61.18 48.00 60.00 45.00 78.33
7 59.52 Mistral-tiny Mistral AI Proprietary 59.27 59.50 53.50 41.50 58.24 64.00 42.00 40.00 77.08
8 59.22 Claude-instant Anthropic Proprietary 68.73 59.00 53.00 39.50 51.76 52.00 50.00 37.50 61.67
9 55.80 Mistral-large Mistral AI Proprietary 71.82 90.50 4.00 0.00 61.76 66.00 0.00 5.00 84.58
10 54.46 Nexusflow-Raven-v2 Nexusflow Apache 2.0 76.55 83.50 39.50 34.00 45.88 78.00 68.00 45.00 0.00
11 53.95 Firefunction-v1 Fireworks-ai Apache 2.0 73.19 87.00 4.00 0.00 48.24 64.00 0.00 5.00 81.25
12 53.86 Mistral-small Mistral AI Proprietary 46.55 68.00 48.50 58.00 14.12 30.00 40.00 37.50 89.58
13 53.49 GPT-4-0613 OpenAI Proprietary 74.55 86.00 4.00 0.00 37.65 50.00 0.00 0.00 87.08
14 43.19 Deepseek-v1.5 Deepseek Deepseek 48.36 61.00 35.00 43.50 5.29 2.00 0.00 7.50 66.25
License
15 33.61 OpenFunctions-v0 Gorilla LLM Apache 2.0 60.00 56.00 1.00 2.50 39.41 62.00 0.00 0.00 4.58
16 24.76 Glaive-v1 Glaive cc-by-sa-4.0 34.55 26.00 2.00 0.00 21.18 0.00 34.00 2.50 46.25
Wagon Wheel
The following chart shows the comparison of the models based on a few
metrics. You can select and deselect which models to compare. More
information on each metric can be found in the blog.
Function Calling Demo
In this demo for function calling, you can enter a prompt and a
function and see the output. There will be two outputs (and two
output boxes accordingly): one in the actual code format (the top
one) and the other in the OpenAI compatible format (the bottom one).
Note that the OpenAI compatible format output is only available if
the actual code output has valid syntax and can be parsed. We also
provide you a few examples to try out and get a sense of the input
format and the output.
Example 1 Example 2 Example 3
Model: [Gorilla OpenFunctions-v2]
Temperature: [0.7 ] 0.7
[ ]
[ ]
[ ]
[ ]
[ ]
[ ]
[ ]
[ ] [ ]
[ ] [ ]
[ ] [ ] Submit
Output will be shown here:
OpenAI compatible format output here:
Report Issue
Contact Us
[ ] [ ] [ ]
[ ]
[ ]
[ ]
[ ]
[ ]
[ ] Submit
Citation
@misc{berkeley-function-calling-leaderboard,
title={Berkeley Function Calling Leaderboard},
author={Fanjia Yan and Huanzhi Mao and Charlie Cheng-Jie Ji and
Tianjun Zhang and Shishir G. Patil and Ion Stoica and Joseph E.
Gonzalez},
howpublished={\url{https://gorilla.cs.berkeley.edu/blogs/
8_berkeley_function_calling_leaderboard.html}},
year={2024},
}