https://huggingface.co/NousResearch/Yarn-Mistral-7b-128k Hugging Face's logo Hugging Face [ ] * Models * Datasets * Spaces * Docs * Solutions * Pricing * * ----------------------------------------------------------------- * Log In * Sign Up [ngyQKdhayk] NousResearch / Yarn-Mistral-7b-128k like 368 Text Generation Transformers PyTorch emozilla/yarn-train-tokenized-16k-mistral English mistral custom_code Inference Endpoints text-generation-inference arxiv: 2309.00071 License: apache-2.0 Model card Files Files and versions Community 11 Train Deploy Use in Transformers Edit model card * Model Card: Nous-Yarn-Mistral-7b-128k + Model Description + Benchmarks + Collaborators Model Card: Nous-Yarn-Mistral-7b-128k Preprint (arXiv) GitHub yarn Model Description Nous-Yarn-Mistral-7b-128k is a state-of-the-art language model for long context, further pretrained on long context data for 1500 steps using the YaRN extension method. It is an extension of Mistral-7B-v0.1 and supports a 128k token context window. To use, pass trust_remote_code=True when loading the model, for example model = AutoModelForCausalLM.from_pretrained("NousResearch/Yarn-Mistral-7b-128k", use_flash_attention_2=True, torch_dtype=torch.bfloat16, device_map="auto", trust_remote_code=True) In addition you will need to use the latest version of transformers (until 4.35 comes out) pip install git+https://github.com/huggingface/transformers Benchmarks Long context benchmarks: Model Context 8k 16k 32k 64k 128k Window PPL PPL PPL PPL PPL Mistral-7B-v0.1 8k 2.96 - - - - Yarn-Mistral-7b-64k 64k 3.04 2.65 2.44 2.20 - Yarn-Mistral-7b-128k 128k 3.08 2.68 2.47 2.24 2.19 Short context benchmarks showing that quality degradation is minimal: Model Context Window ARC-c Hellaswag MMLU Truthful QA Mistral-7B-v0.1 8k 59.98 83.31 64.16 42.15 Yarn-Mistral-7b-64k 64k 59.38 81.21 61.32 42.50 Yarn-Mistral-7b-128k 128k 58.87 80.58 60.64 42.46 Collaborators * bloc97: Methods, paper and evals * @theemozilla: Methods, paper, model training, and evals * @EnricoShippole: Model training * honglu2875: Paper and evals The authors would like to thank LAION AI for their support of compute for this model. It was trained on the JUWELS supercomputer. Downloads last month 22,398 Inference API Text Generation Examples Compute Model state unknown JSON Output Maximize Dataset used to train NousResearch/Yarn-Mistral-7b-128k emozilla/yarn-train-tokenized-16k-mistral Viewer * Updated Oct 11 * 654 * 7 Spaces using NousResearch/Yarn-Mistral-7b-128k 23 Collection including NousResearch/Yarn-Mistral-7b-128k YaRN Collection Extension of Llama 2 to 128k context windows * 13 items * Updated 9 days ago * 9 Company (c) Hugging Face TOS Privacy About Jobs Website Models Datasets Spaces Pricing Docs