https://www.sbert.net/index.html Logo Sentence-Transformers [Twitter_Lo] [ ] Overview * Installation * Quickstart * Pretrained Models * Pretrained Cross-Encoders * Publications * Hugging Face Usage * Computing Sentence Embeddings * Semantic Textual Similarity * Embedding Quantization * Semantic Search * Retrieve & Re-Rank * Clustering * Paraphrase Mining * Translated Sentence Mining * Cross-Encoders * Image Search Training * Training Overview * Loss Overview * Matryoshka Embeddings * Adaptive Layers * Multilingual-Models * Model Distillation * Cross-Encoders * Augmented SBERT * Training Datasets Training Examples * Semantic Textual Similarity * Natural Language Inference * Paraphrase Data * Quora Duplicate Questions * MS MARCO Unsupervised Learning * Unsupervised Learning * Domain Adaptation Package Reference * SentenceTransformer * util * quantization * Models * Losses * Evaluation * Datasets * cross_encoder Sentence-Transformers * >> * SentenceTransformers Documentation * Edit on GitHub --------------------------------------------------------------------- SentenceTransformers DocumentationP SentenceTransformers is a Python framework for state-of-the-art sentence, text and image embeddings. The initial work is described in our paper Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. You can use this framework to compute sentence / text embeddings for more than 100 languages. These embeddings can then be compared e.g. with cosine-similarity to find sentences with a similar meaning. This can be useful for semantic textual similarity, semantic search, or paraphrase mining. The framework is based on PyTorch and Transformers and offers a large collection of pre-trained models tuned for various tasks. Further, it is easy to fine-tune your own models. InstallationP You can install it using pip: pip install -U sentence-transformers We recommend Python 3.8 or higher, and at least PyTorch 1.11.0. See installation for further installation options, especially if you want to use a GPU. UsageP The usage is as simple as: from sentence_transformers import SentenceTransformer model = SentenceTransformer("all-MiniLM-L6-v2") # Our sentences to encode sentences = [ "This framework generates embeddings for each input sentence", "Sentences are passed as a list of string.", "The quick brown fox jumps over the lazy dog." ] # Sentences are encoded by calling model.encode() embeddings = model.encode(sentences) # Print the embeddings for sentence, embedding in zip(sentences, embeddings): print("Sentence:", sentence) print("Embedding:", embedding) print("") PerformanceP Our models are evaluated extensively and achieve state-of-the-art performance on various tasks. Further, the code is tuned to provide the highest possible speed. Have a look at Pre-Trained Models for an overview of available models and the respective performance on different tasks. ContactP Contact person: Tom Aarsen, tom.aarsen@huggingface.co Don't hesitate to open an issue on the repository if something is broken (and it shouldn't be) or if you have further questions. This repository contains experimental software and is published for the sole purpose of giving additional background details on the respective publication. Citing & AuthorsP If you find this repository helpful, feel free to cite our publication Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks: @inproceedings{reimers-2019-sentence-bert, title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks", author = "Reimers, Nils and Gurevych, Iryna", booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing", month = "11", year = "2019", publisher = "Association for Computational Linguistics", url = "https://arxiv.org/abs/1908.10084", } If you use one of the multilingual models, feel free to cite our publication Making Monolingual Sentence Embeddings Multilingual using Knowledge Distillation: @inproceedings{reimers-2020-multilingual-sentence-bert, title = "Making Monolingual Sentence Embeddings Multilingual using Knowledge Distillation", author = "Reimers, Nils and Gurevych, Iryna", booktitle = "Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing", month = "11", year = "2020", publisher = "Association for Computational Linguistics", url = "https://arxiv.org/abs/2004.09813", } If you use the code for data augmentation, feel free to cite our publication Augmented SBERT: Data Augmentation Method for Improving Bi-Encoders for Pairwise Sentence Scoring Tasks: @inproceedings{thakur-2020-AugSBERT, title = "Augmented {SBERT}: Data Augmentation Method for Improving Bi-Encoders for Pairwise Sentence Scoring Tasks", author = "Thakur, Nandan and Reimers, Nils and Daxenberger, Johannes and Gurevych, Iryna", booktitle = "Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies", month = jun, year = "2021", address = "Online", publisher = "Association for Computational Linguistics", url = "https://www.aclweb.org/anthology/2021.naacl-main.28", pages = "296--310", } Overview * Installation + Install SentenceTransformers + Install PyTorch with CUDA support * Quickstart + Comparing Sentence Similarities + Pre-Trained Models + Training your own Embeddings * Pretrained Models + Model Overview + Semantic Search + Multi-Lingual Models + Image & Text-Models + Other Models * Pretrained Cross-Encoders + MS MARCO + SQuAD (QNLI) + STSbenchmark + Quora Duplicate Questions + NLI * Publications * Hugging Face + The Hugging Face Hub + Using Hugging Face models + Sharing your models + Sharing your embeddings + Additional resources Usage * Computing Sentence Embeddings + Prompt Templates + Input Sequence Length + Storing & Loading Embeddings + Multi-Process / Multi-GPU Encoding + Sentence Embeddings with Transformers * Semantic Textual Similarity * Embedding Quantization + Binary Quantization + Scalar (int8) Quantization + Additional extensions + Demo + Try it yourself * Semantic Search + Background + Symmetric vs. Asymmetric Semantic Search + Python + util.semantic_search + Speed Optimization + Elasticsearch + Approximate Nearest Neighbor + Retrieve & Re-Rank + Examples * Retrieve & Re-Rank + Retrieve & Re-Rank Pipeline + Retrieval: Bi-Encoder + Re-Ranker: Cross-Encoder + Example Scripts + Pre-trained Bi-Encoders (Retrieval) + Pre-trained Cross-Encoders (Re-Ranker) * Clustering + k-Means + Agglomerative Clustering + Fast Clustering + Topic Modeling * Paraphrase Mining * Translated Sentence Mining + Marging Based Mining + Examples * Cross-Encoders + Bi-Encoder vs. Cross-Encoder + When to use Cross- / Bi-Encoders? + Cross-Encoders Usage + Combining Bi- and Cross-Encoders + Training Cross-Encoders * Image Search + Installation + Usage + Examples Training * Training Overview + Network Architecture + Creating Networks from Scratch + Training Data + Loss Functions + Evaluators + Loading Custom SentenceTransformer Models + Multitask Training + Adding Special Tokens + Best Transformer Model * Loss Overview + Loss modifiers + Distillation + Commonly used Loss Functions * Matryoshka Embeddings + Use Cases + Results + Training + Inference + Code Examples * Adaptive Layers + Use Cases + Results + Training + Inference + Code Examples * Multilingual-Models + Available Pre-trained Models + Usage + Performance + Extend your own models + Training + Data Format + Loading Training Datasets + Sources for Training Data + Evaluation + Citation * Model Distillation + Knowledge Distillation + Speed - Performance Trade-Off + Dimensionality Reduction + Quantization * Cross-Encoders + Examples + Training CrossEncoders * Augmented SBERT + Motivation + Extend to your own datasets + Methodology + Scenario 1: Limited or small annotated datasets (few labeled sentence-pairs) + Scenario 2: No annotated datasets (Only unlabeled sentence-pairs) + Training + Citation * Training Datasets + Datasets on the Hugging Face Hub Training Examples * Semantic Textual Similarity + Training data + Loss Function * Natural Language Inference + Data + SoftmaxLoss + MultipleNegativesRankingLoss * Paraphrase Data + Datasets + Training + Pre-Trained Models + Work in Progress * Quora Duplicate Questions + Pretrained Models + Dataset + Usage + Training + MultipleNegativesRankingLoss * MS MARCO + Bi-Encoder + Cross-Encoder + Cross-Encoder Knowledge Distillation Unsupervised Learning * Unsupervised Learning + TSDAE + SimCSE + CT + CT (In-Batch Negative Sampling) + Masked Language Model (MLM) + GenQ + GPL + Performance Comparison * Domain Adaptation + Domain Adaptation vs. Unsupervised Learning + Adaptive Pre-Training + GPL: Generative Pseudo-Labeling Package Reference * SentenceTransformer * util * quantization * Models * Losses * Evaluation * Datasets * cross_encoder Next --------------------------------------------------------------------- (c) Copyright 2024, Nils Reimers * Contact Built with Sphinx using a theme provided by Read the Docs.