https://github.com/FranxYao/chain-of-thought-hub Skip to content Toggle navigation Sign up * Product + Actions Automate any workflow + Packages Host and manage packages + Security Find and fix vulnerabilities + Codespaces Instant dev environments + Copilot Write better code with AI + Code review Manage code changes + Issues Plan and track work + Discussions Collaborate outside of code Explore + All features + Documentation + GitHub Skills + Blog * Solutions For + Enterprise + Teams + Startups + Education By Solution + CI/CD & Automation + DevOps + DevSecOps Case Studies + Customer Stories + Resources * Open Source + GitHub Sponsors Fund open source developers + The ReadME Project GitHub community articles Repositories + Topics + Trending + Collections * Pricing [ ] * # In this repository All GitHub | Jump to | * No suggested jump to results * # In this repository All GitHub | Jump to | * # In this user All GitHub | Jump to | * # In this repository All GitHub | Jump to | Sign in Sign up {{ message }} FranxYao / chain-of-thought-hub Public * Notifications * Fork 40 * Star 750 Benchmarking large language models' complex reasoning ability with chain-of-thought prompting 750 stars 40 forks Star Notifications * Code * Issues 2 * Pull requests 1 * Actions * Projects 0 * Security * Insights More * Code * Issues * Pull requests * Actions * Projects * Security * Insights FranxYao/chain-of-thought-hub This commit does not belong to any branch on this repository, and may belong to a fork outside of the repository. main Switch branches/tags [ ] Branches Tags Could not load branches Nothing to show {{ refName }} default View all branches Could not load tags Nothing to show {{ refName }} default View all tags Name already in use A tag already exists with the provided branch name. Many Git commands accept both tag and branch names, so creating this branch may cause unexpected behavior. Are you sure you want to create this branch? Cancel Create 5 branches 0 tags Code * Local * Codespaces * Clone HTTPS GitHub CLI [https://github.com/F] Use Git or checkout with SVN using the web URL. [gh repo clone FranxY] Work fast with our official CLI. Learn more about the CLI. * Open with GitHub Desktop * Download ZIP Sign In Required Please sign in to use Codespaces. Launching GitHub Desktop If nothing happens, download GitHub Desktop and try again. Launching GitHub Desktop If nothing happens, download GitHub Desktop and try again. Launching Xcode If nothing happens, download Xcode and try again. Launching Visual Studio Code Your codespace will open once ready. There was a problem preparing your codespace, please try again. Latest commit Yao Fu update ... 409ff3f May 29, 2023 update 409ff3f Git stats * 146 commits Files Permalink Failed to load latest commit information. Type Name Latest commit message Commit time BBH add complexity-based prompting files May 23, 2023 17:34 MATH update flan t5 May 23, 2023 21:56 MMLU update flan t5 May 23, 2023 21:56 gsm8k update flan t5 May 23, 2023 21:56 notebooks update TheoremQA, alpaca and vicuna May 27, 2023 12:46 research update research directory May 23, 2023 17:40 resources Merge branch 'main' of https://github.com/FranxYao/ chain-of-thought-hub... May 25, 2023 13:17 .gitignore update flan t5 May 23, 2023 21:56 readme.md update May 28, 2023 17:07 View code [ ] Chain-of-Thought Hub: Measuring LLMs' Reasoning Performance Results - Overall Run FAQ References Pretraining/ Continue Training Finetuning Reinforcement Learning Prompting Under Development readme.md Chain-of-Thought Hub: Measuring LLMs' Reasoning Performance Title "A fantasy graph illustrating a chain of stars in a dark night with blue sky, digital art, super resolution". Midjourney V5 --------------------------------------------------------------------- By Yao Fu, Litu Ou, Mingyu Chen, Yuhao Wan, Hao Peng, Tushar Khot, Wenhu Chen From University of Edinburgh, University of Washington, Allen Institute for AI, University of Waterloo Recently, there are a lot of progress in LLMs. Many claim that a small model less than 10B can achieve comparable performance to GPT-3.5. Really? In a casual conversation, the distinction between GPT-3.5 and GPT-4 can be subtle. The difference comes out when *the complexity of the task reaches a sufficient threshold* -- GPT-4 is more reliable, creative, and able to handle much more nuanced instructions than GPT-3.5. -- GPT-4 release blog The key differentiator is whether a model can do complex tasks, like the old saying: "chit-chat is cheap, show me the reasoning." This is why we compile a list of complex reasoning tasks including math (GSM8K), science (MATH, TheoremQA), symbolic (BBH), knowledge (MMLU, C-Eval), coding (HumanEval) to measure the models' performance on challenging tasks. For more detailed discussion about complex reasoning, see article Towards Complex Reasoning: the Polaris of Large Language Models [UPDATE 20230527]: Call for contribution! If you are interested in fill in a missing number in our table, feel free to send a PR (especially for smaller models like Vicuna), much appreciated! [UPDATE 20230527]: Add TheoremQA, add Vicuna, Alpaca, InstructCodeT5. Yet still many numbers missing ... Results - Overall Model Param. Type GSM8K MATH MMLU BBH HumanEval C-Eval TheoremQA gpt-4 ? RLHF 92.0 42.5 86.4 - 67.0 68.7* 43.4 claude-v1.3 ? RLHF 81.8* - 74.8* 67.3* - 54.2* 24.9 PaLM-2 ? Base 80.7 34.3 78.3 78.1 - - 31.8 gpt-3.5-turbo ? RLHF 74.9* - 67.3* 70.1* 48.1 54.4* 30.2 claude-instant ? RLHF 70.8* - - 66.9* - 45.9* 23.6 text-davinci-003 ? RLHF - - 64.6 70.7 - - 22.8 code-davinci-002 ? Base 66.6 19.1 64.5 73.7 47.0 - - text-davinci-002 ? SIFT 55.4 - 60.0 67.2 - - 16.6 Minerva 540B SIFT 58.8 33.6 - - - - - Flan-PaLM 540B SIFT - - 70.9 66.3 - - - Flan-U-PaLM 540B SIFT - - 69.8 64.9 - - - PaLM 540B Base 56.9 8.8 62.9 62.0 26.2 - - LLaMA 65B Base 50.9 10.6 63.4 - 23.7 38.8* - PaLM 64B Base 52.4 4.4 49.0 42.3 - - - LLaMA 33B Base 35.6 7.1 57.8 - 21.7 - - InstructCodeT5+ 16B SIFT - - - - 35.0 - 11.6 StarCoder 15B Base 8.4 15.1 33.9 - 33.6 - 12.2 Vicuna 13B SIFT - - - - - - 12.9 LLaMA 13B Base 17.8 3.9 46.9 - 15.8 - - Flan-T5 11B SIFT 16.1* - 48.6 41.4 - - - Alpaca 7B SIFT - - - - - - 13.5 LLaMA 7B Base 11.0 2.9 35.1 - 10.5 - - Flan-T5 3B SIFT 13.5* - 45.5 35.2 - - - Base means the pretrained checkpoint. SIFT means the checkpoint after supervised instruction finetuning. RLHF means the checkpoint after Reinforcement Learning from Human Feedback. Numbers marked with an asterisk * are from our own run, otherwise from multiple sources which we explain below. What's different than HeLM and other evaluation? * HeLM uses answer-only prompting, we use chain-of-thought promoting * HeLM evaluates everything. We only focus on complex reasoning, the key differentiator of LLMs' capability. How the models are ranked * If we know model scale, we rank it by scale. * If we do not know model scale, we rank it by GSM8K, the classical benchmark measuring chain-of-thought math reasoning performance. + This is definitely not the only metric, but a good interpretation is "how good the model can do math while maintaining other generic abilities" -- which is also very hard. + GPT-4 is already pretrained on GSM8k training split, others may not. So for GPT-4, its perf. on GSM8k is in-distribution generalization, while for others are ood. generalization. Yet even for in-dist. FlanT5 is also trained on GSM8k, still shows perf. difference. * Generally it is very hard to rigiously compare model perf. due to multiple factors (whether trained on the corresponding training split, whether trained on code, whether optimize prompt .etc). View our results as approximate reference. Source of numbers * GPT-4 from its website and Bubeck et al Mar 2023. Note that the version that Bubeck uses is GPT-4 Early which is supposedly to be more powerful than GPT-4 Launch (OpenAI paid a lot of alignment tax to make GPT-4 safer). * *-davinci-00* and *PaLM are from the Flan-PaLM paper appendix. * LLaMA from LLaMA paper (TODO: test LLaMA on BBH). Note that the prompt of LLaMA used in these tasks are not released so reproduction may have varied numbers, see this twitter thread for more discussions. * PaLM-2 from their tech report. * Claude is from our own test script, see below about how to run it. * The HumanEval results for LLaMA models, PaLM and StartCoder are from HuggingFace report. Code-davinci-002's performance on HumanEval is from CodeT5+ paper * C-Eval is from their website * TheoremQA is from their github Current results: * GPT-4 clearly outperforms all other models on GSM8K and MMLU. * **The 65B LLaMA is very close to text/code-davinci-002, which means that based on it, if SFT and RLHF are done correctly, it is very likely that we could reproduce ChatGPT based on the 65B LLaMA** * Claude is the only model family that is comparable to GPT family. * On GSM8K, gpt-3.5-turbo improves over text-davinci-003. This confirms OpenAI's Jan 30 2023 release notes "improved mathematical capabilities." * On MMLU, gpt-3.5-turbo is slightly better than text-davinci-003. But this level of margin is NOT SIGNIFICANT * Also remember that gpt-3.5-turbo is 10 times cheaper than text-davinci-003 * Also be careful that GPT-4/ 3.5's performance on GSM8K is not true few-shot -- in GPT-4 report they said that they mixed a portion of GSM8K training set to train the model * LLaMA performance on MMLU is from their paper and probably not CoT but AO. Generally on MMLU, AO is better than CoT but just slightly better. So the LLaMA numbers on MMLU might be slightly overestimated. Why choosing the above tasks? * We mostly care about complex reasoning. + Other abilities of LLMs such as summarization or translation are not considered here as they are rather standard and probably not challenging enough. + We consider + MMLU: high school and college knowledge + GSM8K: elementary school math. -- Performance improvements on this dataset directly translate to daily math abilities when interacting with LLMs + MATH (Hard!): very hard math and natural science. All current models struggle. + BBH: a collection of 27 hard reasoning problems + HumanEval: a classical dataset for evaluating coding capability. + C-Eval: a collection of 52 disciplines of knowledge test in Chinese + TheoremQA (Hard!): a question-answering dataset driven by STEM theorems Run MMLU cd MMLU mkdir outputs API_KEY= python run_mmlu_gpt_3.5_turbo.py --api_key=${API_KEY} python run_mmlu_claude.py --api_key=${API_KEY} --engine=claude-v1.3 GSM8k cd gsm8k mkdir outputs # run gpt-3.5 # codex_gsm8k_complex.ipynb -- code-davinci-002 + complex prompt # gpt3.5turbo_gsm8k_complex.ipynb -- gpt-3.5-turbo + complex prompt # run claude python run_gsm8k_claude.py\ --anthropic_key=${API_KEY}\ --prompt_file=lib_prompt/prompt_original.txt\ --engine=claude-v1.3\ --output_file=outputs/gsm8k_claude_v1.3_original_test.txt # run FlanT5 # flan_t5_11b_gsm8k.ipynb BBH cd BBH mkdir outputs # then run jupyter notebook to see an example penguins dataset cd penguins # gpt3.5trubo_penguins_original.ipynb # Or run the script for all datasets API_KEY= TASK= python run_bbh_gpt_3.5_turbo.py --api_key=${API_KEY} --task=${TASK} # task=all by default python run_bbh_claude_v1.3.py --api_key=${API_KEY} --model_index=claude-v1.3 --task=${TASK} # task=all by default FAQ * What are the prompts used in the complexity-based prompting paper? + See research/complexity_based_prompting/ * I want to try some open-sourced model + See gsm8k/flan_t5_11b_gsm8k.ipynb for a place to start * There are some prompts that have wrong answer + Yes, but we keep it as they are used in the original papers + Generally the model can be robust under prompt perturbation: even if sometimes there are errors in the prompt, as long as the format of the prompt is about the corresponding task, the model tend to only look at the format, ignore the prompt error, and make its own prediction. + See https://arxiv.org/abs/2202.12837 and https://arxiv.org/ abs/2212.10001 about more analysis how the model can ignore errors in the prompt References We first discuss the recipe of building models of strong reasoning abilities, which is the same as generic LLM recipe: pretraining, finetuning, reinforcement learning. Then we discuss prompting methods for releasing the reasoning power of large language models. Pretraining/ Continue Training * Lewkowycz et. al. 2022. Minerva: Solving Quantitative Reasoning Problems with Language Models * Taylor et. al. 2022. Galactica: A Large Language Model for Science Finetuning * Chung et. al. 2022. Scaling Instruction-Finetuned Language Models * Li et. al. 2022. Competition-Level Code Generation with AlphaCode * Fu et. al. 2023. Specializing Smaller Language Models towards Multi-Step Reasoning Reinforcement Learning * Uesato et. al. 2022. Solving math word problems with process- and outcome-based feedback * Le et. al. 2022. CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning Prompting * Wei et. al. 2022. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models * Suzgun et. al. 2022. Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them * Fu et. al. 2023. Complexity-Based Prompting for Multi-Step Reasoning More detailed discussions Under Development TODOs Literature Prompting Tricks Detailed Results About Benchmarking large language models' complex reasoning ability with chain-of-thought prompting Resources Readme Stars 750 stars Watchers 26 watching Forks 40 forks Report repository Releases No releases published Packages 0 No packages published Contributors 6 * @Yuhao-Wan * @FranxYao * @Spehhhhh * @Leonard907 * @wenhuchen * @LauJames Languages * Jupyter Notebook 98.4% * Python 1.6% Footer (c) 2023 GitHub, Inc. Footer navigation * Terms * Privacy * Security * Status * Docs * Contact GitHub * Pricing * API * Training * Blog * About You can't perform that action at this time. You signed in with another tab or window. Reload to refresh your session. You signed out in another tab or window. Reload to refresh your session.