https://github.com/microsoft/Llama-2-Onnx Skip to content Toggle navigation Sign up * Product + Actions Automate any workflow + Packages Host and manage packages + Security Find and fix vulnerabilities + Codespaces Instant dev environments + Copilot Write better code with AI + Code review Manage code changes + Issues Plan and track work + Discussions Collaborate outside of code Explore + All features + Documentation + GitHub Skills + Blog * Solutions For + Enterprise + Teams + Startups + Education By Solution + CI/CD & Automation + DevOps + DevSecOps Resources + Customer Stories + White papers, Ebooks, Webinars + Partners * Open Source + GitHub Sponsors Fund open source developers + The ReadME Project GitHub community articles Repositories + Topics + Trending + Collections * Pricing Search or jump to... Search code, repositories, users, issues, pull requests... Search [ ] Clear Search syntax tips Provide feedback We read every piece of feedback, and take your input very seriously. [ ] [ ] Include my email address so I can be contacted Cancel Submit feedback Saved searches Use saved searches to filter your results more quickly Name [ ] Query [ ] To see all available qualifiers, see our documentation. Cancel Create saved search Sign in Sign up You signed in with another tab or window. Reload to refresh your session. You signed out in another tab or window. Reload to refresh your session. You switched accounts on another tab or window. Reload to refresh your session. {{ message }} microsoft / Llama-2-Onnx Public * Notifications * Fork 14 * Star 151 License View license 151 stars 14 forks Activity Star Notifications * Code * Issues 3 * Pull requests 1 * Actions * Projects 0 * Security * Insights More * Code * Issues * Pull requests * Actions * Projects * Security * Insights microsoft/Llama-2-Onnx This commit does not belong to any branch on this repository, and may belong to a fork outside of the repository. main Switch branches/tags [ ] Branches Tags Could not load branches Nothing to show {{ refName }} default View all branches Could not load tags Nothing to show {{ refName }} default View all tags Name already in use A tag already exists with the provided branch name. Many Git commands accept both tag and branch names, so creating this branch may cause unexpected behavior. Are you sure you want to create this branch? Cancel Create 7 branches 0 tags Code * Local * Codespaces * Clone HTTPS GitHub CLI [https://github.com/m] Use Git or checkout with SVN using the web URL. [gh repo clone micros] Work fast with our official CLI. Learn more about the CLI. * Open with GitHub Desktop * Download ZIP Sign In Required Please sign in to use Codespaces. Launching GitHub Desktop If nothing happens, download GitHub Desktop and try again. Launching GitHub Desktop If nothing happens, download GitHub Desktop and try again. Launching Xcode If nothing happens, download Xcode and try again. Launching Visual Studio Code Your codespace will open once ready. There was a problem preparing your codespace, please try again. Latest commit @JoshuaElsdon JoshuaElsdon Merge pull request #17 from adarshxs/main ... a20f9bc Aug 9, 2023 Merge pull request #17 from adarshxs/main Update Example.md Readme had old versions of the argument names. a20f9bc Git stats * 30 commits Files Permalink Failed to load latest commit information. Type Name Latest commit message Commit time 13B_FT_float16 @ 1717145 Added information files July 17, 2023 23:47 13B_FT_float32 @ cc96b86 Added information files July 17, 2023 23:47 13B_float16 @ 07d9ff3 Added information files July 17, 2023 23:47 13B_float32 @ 21c147f Added information files July 17, 2023 23:47 7B_FT_float16 @ f860aae Added information files July 17, 2023 23:47 7B_FT_float32 @ b0c0799 Added information files July 17, 2023 23:47 7B_float16 @ 0a4e52d Added information files July 17, 2023 23:47 7B_float32 @ e630a4f Added information files July 17, 2023 23:47 ChatApp Added Chat App example, updated README.md and other documents accordi... July 31, 2023 22:05 Images Added Chat App example, updated README.md and other documents accordi... July 31, 2023 22:05 MinimumExample Update Example.md August 8, 2023 10:43 .gitignore added initial example file. July 17, 2023 18:23 .gitmodules Added information files July 17, 2023 23:47 LICENSE Added information files July 17, 2023 23:47 MODEL-CARD-META-LLAMA-2.md Added information files July 17, 2023 23:47 README.md Added Chat App example, updated README.md and other documents accordi... July 31, 2023 22:05 RESPONSIBLE-USE-GUIDE-META-LLAMA-2.pdf Added information files July 17, 2023 23:47 SECURITY.md Microsoft mandatory file July 18, 2023 15:04 USE-POLICY-META-LLAMA-2.md Added information files July 17, 2023 23:47 tokenizer.model Renamed sub-modules. Added tokenizer.model July 17, 2023 19:31 View code [ ] Llama 2 Powered By ONNX Before You Start Cloning This Repository And The Submodules What is Llama 2? What Is The Structure Of Llama 2? FAQ Is There A Simple Code Example Running Llama 2 With ONNX? Is There A More Complete Code Example Running Llama 2 With ONNX? How Do I Use The Fine-tuned Models? Why Is The First Inference Session Slow? Why Is FP16 ONNX Slower Than ONNX FP32 On My Device? How Do I Get Better Inference Speed? What Parameters Should I Test With? How Can I Develop With Llama 2 Responsibly? README.md Llama 2 Powered By ONNX This is an optimized version of the Llama 2 model, available from Meta under the Llama Community License Agreement found on this repository. Microsoft permits you to use, modify, redistribute and create derivatives of Microsoft's contributions to the optimized version subject to the restrictions and disclaimers of warranty and liability in the Llama Community License agreement. Before You Start The sub-modules that contain the ONNX files in this repository are access controlled. To get access permissions to the Llama 2 model, please fill out the Llama 2 access request form. If allowable, you will receive GitHub access in the next 48 hours, but usually much sooner. Cloning This Repository And The Submodules Chose from the following sub-modules: * 7B_FT_float16 * 7B_FT_float32 * 7B_float16 * 7B_float32 * 13B_FT_float16 * 13B_FT_float32 * 13B_float16 * 13B_float32 git clone https://github.com/microsoft/Llama-2-Onnx.git cd Llama-2-Onnx git submodule init git submodule update You can repeate the init command with a different submodule name to initialize multiple submodules. Be careful, the contained files are very large! (7B Float16 models are about 10GB) What is Llama 2? Llama 2 is a collection of pretrained and fine-tuned generative text models. To learn more about Llama 2, review the Llama 2 model card. What Is The Structure Of Llama 2? Llama 2 model consists of a stack of decoder layers. Each decoder layer (or transformer block) is constructed from one self-attention layer and one feed-forward multi-layer perceptron. Llama models use different projection sizes compared with classic transformers in the feed-forward layer, for instance, both Llama 1 and Llama 2 projection use 2.7x hidden size rather than the standard 4x hidden size. A key difference between Llama 1 and Llama 2 is the architectural change of attention layer, in which Llama 2 takes advantage of Grouped Query Attention (GQA) mechanism to improve efficiency. Llama 2 Model FAQ Is There A Simple Code Example Running Llama 2 With ONNX? There are two examples provided in this repository. There is a minimum working example shown in Llama-2-Onnx/MinimumExample. This is simply a command line program that will complete some text with the chosen version of Llama 2. Given the following input: python MinimumExample/Example_ONNX_LlamaV2.py --onnx_file 7B_FT_float16/ONNX/LlamaV2_7B_FT_float16.onnx --embedding_file 7B_FT_float16/embeddings.pth --tokenizer_path tokenizer.model --prompt "What is the lightest element?" Output: The lightest element is hydrogen. Hydrogen is the lightest element on the periodic table, with an atomic mass of 1.00794 u (unified atomic mass units). Is There A More Complete Code Example Running Llama 2 With ONNX? There is a more complete chat bot interface that is available in Llama-2-Onnx/ChatApp. This is a python program based on the popular Gradio web interface. It will allow you to interact with the chosen version of Llama 2 in a chat bot interface. An example interaction can be seen here: Chat App How Do I Use The Fine-tuned Models? The fine-tuned models were trained for dialogue applications. To get the expected features and performance for them, a specific formatting needs to be followed, including the INST tag, BOS and EOS tokens, and the whitespaces and breaklines in between (we recommend calling strip() on inputs to avoid double-spaces). This enables models in chat mode as well as additional safeguards to reduce potentially undesirable output. Why Is The First Inference Session Slow? ONNX runtime execution provider might need to generate JIT binaries for the underlying hardware, typically the binary is cache and will be loaded directly in the subsequent runs to reduce the overhead. Why Is FP16 ONNX Slower Than ONNX FP32 On My Device? It is possible that your device does not support native FP16 math, therefore weights will be cast to FP32 at runtime. Using the FP32 version of the model will avoid the cast overhead. How Do I Get Better Inference Speed? It is recommended that inputs/outputs are put on target device to avoid expensive data copies, please refer to the following document for details. I/O Binding | onnxruntime What Parameters Should I Test With? Users can perform temperature and top-p sampling using the model's output logits. Please refer to Meta's guidance for the best parameters combination; an example is located here. How Can I Develop With Llama 2 Responsibly? In order to help developers innovate responsibly, Meta encourages you to review the Responsible Use Guide for the Llama 2 models. Microsoft encourages you to learn more about its Responsible AI approach, including many publicly available resources and tools for developers. About No description, website, or topics provided. Resources Readme License View license Code of conduct Code of conduct Security policy Security policy Activity Stars 151 stars Watchers 154 watching Forks 14 forks Report repository Releases No releases published Packages 0 No packages published Contributors 5 * @JoshuaElsdon * @vriveras * @microsoftopensource * @adarshxs * @tammanygrantmsft Languages * Python 82.3% * CSS 17.6% * JavaScript 0.1% Footer (c) 2023 GitHub, Inc. Footer navigation * Terms * Privacy * Security * Status * Docs * Contact GitHub * Pricing * API * Training * Blog * About You can't perform that action at this time.