https://gorilla.cs.berkeley.edu/leaderboard.html Leaderboard Try it Out! Blog Gorilla UC Berkeley Logo Berkeley Function-Calling Leaderboard Leaderboard This live leaderboard evaluates the LLM's ability to call functions (aka tools) accurately. This leaderboard consists of real-world data and will be updated periodically. For more information on the evaluation dataset and methodology, please refer to our blog post and code release. Abstract Syntax Tree (AST) Evaluation Evaluation by Executing APIs Rank Overall Model Organization License Simple Multiple Parallel Parallel Simple Multiple Parallel Parallel Relevance Acc Function Functions Functions Multiple Function Functions Functions Multiple Detection 1 83.80 GPT-4-0125-Preview OpenAI Proprietary 82.18 90.00 90.00 91.00 54.12 70.00 76.00 55.00 87.50 2 83.55 GPT-4-1106-Preview OpenAI Proprietary 81.64 89.50 92.00 92.00 53.53 62.00 72.00 50.00 88.75 3 83.55 OpenFunctions-v2 Gorilla LLM Apache 2.0 88.73 89.50 79.50 78.00 78.82 74.00 76.00 60.00 71.67 4 81.63 GPT-3.5-Turbo OpenAI Proprietary 81.27 88.00 87.50 88.00 74.12 74.00 70.00 47.50 68.33 5 79.46 Mistral-medium Mistral AI Proprietary 80.18 84.50 71.00 68.00 75.88 72.00 62.00 47.50 90.00 6 75.78 Claude-2.1 Anthropic Proprietary 85.64 83.00 72.00 56.50 61.18 48.00 60.00 45.00 78.33 7 59.52 Mistral-tiny Mistral AI Proprietary 59.27 59.50 53.50 41.50 58.24 64.00 42.00 40.00 77.08 8 59.22 Claude-instant Anthropic Proprietary 68.73 59.00 53.00 39.50 51.76 52.00 50.00 37.50 61.67 9 55.80 Mistral-large Mistral AI Proprietary 71.82 90.50 4.00 0.00 61.76 66.00 0.00 5.00 84.58 10 54.46 Nexusflow-Raven-v2 Nexusflow Apache 2.0 76.55 83.50 39.50 34.00 45.88 78.00 68.00 45.00 0.00 11 53.95 Firefunction-v1 Fireworks-ai Apache 2.0 73.19 87.00 4.00 0.00 48.24 64.00 0.00 5.00 81.25 12 53.86 Mistral-small Mistral AI Proprietary 46.55 68.00 48.50 58.00 14.12 30.00 40.00 37.50 89.58 13 53.49 GPT-4-0613 OpenAI Proprietary 74.55 86.00 4.00 0.00 37.65 50.00 0.00 0.00 87.08 14 43.19 Deepseek-v1.5 Deepseek Deepseek 48.36 61.00 35.00 43.50 5.29 2.00 0.00 7.50 66.25 License 15 33.61 OpenFunctions-v0 Gorilla LLM Apache 2.0 60.00 56.00 1.00 2.50 39.41 62.00 0.00 0.00 4.58 16 24.76 Glaive-v1 Glaive cc-by-sa-4.0 34.55 26.00 2.00 0.00 21.18 0.00 34.00 2.50 46.25 Wagon Wheel The following chart shows the comparison of the models based on a few metrics. You can select and deselect which models to compare. More information on each metric can be found in the blog. Function Calling Demo In this demo for function calling, you can enter a prompt and a function and see the output. There will be two outputs (and two output boxes accordingly): one in the actual code format (the top one) and the other in the OpenAI compatible format (the bottom one). Note that the OpenAI compatible format output is only available if the actual code output has valid syntax and can be parsed. We also provide you a few examples to try out and get a sense of the input format and the output. Example 1 Example 2 Example 3 Model: [Gorilla OpenFunctions-v2] Temperature: [0.7 ] 0.7 [ ] [ ] [ ] [ ] [ ] [ ] [ ] [ ] [ ] [ ] [ ] [ ] [ ] Submit Output will be shown here: OpenAI compatible format output here: Report Issue Contact Us [ ] [ ] [ ] [ ] [ ] [ ] [ ] [ ] [ ] Submit Citation @misc{berkeley-function-calling-leaderboard, title={Berkeley Function Calling Leaderboard}, author={Fanjia Yan and Huanzhi Mao and Charlie Cheng-Jie Ji and Tianjun Zhang and Shishir G. Patil and Ion Stoica and Joseph E. Gonzalez}, howpublished={\url{https://gorilla.cs.berkeley.edu/blogs/ 8_berkeley_function_calling_leaderboard.html}}, year={2024}, }