https://www.sbert.net/index.html
Logo Sentence-Transformers
[Twitter_Lo]
[ ]
Overview
* Installation
* Quickstart
* Pretrained Models
* Pretrained Cross-Encoders
* Publications
* Hugging Face
Usage
* Computing Sentence Embeddings
* Semantic Textual Similarity
* Embedding Quantization
* Semantic Search
* Retrieve & Re-Rank
* Clustering
* Paraphrase Mining
* Translated Sentence Mining
* Cross-Encoders
* Image Search
Training
* Training Overview
* Loss Overview
* Matryoshka Embeddings
* Adaptive Layers
* Multilingual-Models
* Model Distillation
* Cross-Encoders
* Augmented SBERT
* Training Datasets
Training Examples
* Semantic Textual Similarity
* Natural Language Inference
* Paraphrase Data
* Quora Duplicate Questions
* MS MARCO
Unsupervised Learning
* Unsupervised Learning
* Domain Adaptation
Package Reference
* SentenceTransformer
* util
* quantization
* Models
* Losses
* Evaluation
* Datasets
* cross_encoder
Sentence-Transformers
* >>
* SentenceTransformers Documentation
* Edit on GitHub
---------------------------------------------------------------------
SentenceTransformers DocumentationP
SentenceTransformers is a Python framework for state-of-the-art
sentence, text and image embeddings. The initial work is described in
our paper Sentence-BERT: Sentence Embeddings using Siamese
BERT-Networks.
You can use this framework to compute sentence / text embeddings for
more than 100 languages. These embeddings can then be compared e.g.
with cosine-similarity to find sentences with a similar meaning. This
can be useful for semantic textual similarity, semantic search, or
paraphrase mining.
The framework is based on PyTorch and Transformers and offers a large
collection of pre-trained models tuned for various tasks. Further, it
is easy to fine-tune your own models.
InstallationP
You can install it using pip:
pip install -U sentence-transformers
We recommend Python 3.8 or higher, and at least PyTorch 1.11.0. See
installation for further installation options, especially if you want
to use a GPU.
UsageP
The usage is as simple as:
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("all-MiniLM-L6-v2")
# Our sentences to encode
sentences = [
"This framework generates embeddings for each input sentence",
"Sentences are passed as a list of string.",
"The quick brown fox jumps over the lazy dog."
]
# Sentences are encoded by calling model.encode()
embeddings = model.encode(sentences)
# Print the embeddings
for sentence, embedding in zip(sentences, embeddings):
print("Sentence:", sentence)
print("Embedding:", embedding)
print("")
PerformanceP
Our models are evaluated extensively and achieve state-of-the-art
performance on various tasks. Further, the code is tuned to provide
the highest possible speed. Have a look at Pre-Trained Models for an
overview of available models and the respective performance on
different tasks.
ContactP
Contact person: Tom Aarsen, tom.aarsen@huggingface.co
Don't hesitate to open an issue on the repository if something is
broken (and it shouldn't be) or if you have further questions.
This repository contains experimental software and is published for
the sole purpose of giving additional background details on the
respective publication.
Citing & AuthorsP
If you find this repository helpful, feel free to cite our
publication Sentence-BERT: Sentence Embeddings using Siamese
BERT-Networks:
@inproceedings{reimers-2019-sentence-bert,
title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
author = "Reimers, Nils and Gurevych, Iryna",
booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
month = "11",
year = "2019",
publisher = "Association for Computational Linguistics",
url = "https://arxiv.org/abs/1908.10084",
}
If you use one of the multilingual models, feel free to cite our
publication Making Monolingual Sentence Embeddings Multilingual using
Knowledge Distillation:
@inproceedings{reimers-2020-multilingual-sentence-bert,
title = "Making Monolingual Sentence Embeddings Multilingual using Knowledge Distillation",
author = "Reimers, Nils and Gurevych, Iryna",
booktitle = "Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing",
month = "11",
year = "2020",
publisher = "Association for Computational Linguistics",
url = "https://arxiv.org/abs/2004.09813",
}
If you use the code for data augmentation, feel free to cite our
publication Augmented SBERT: Data Augmentation Method for Improving
Bi-Encoders for Pairwise Sentence Scoring Tasks:
@inproceedings{thakur-2020-AugSBERT,
title = "Augmented {SBERT}: Data Augmentation Method for Improving Bi-Encoders for Pairwise Sentence Scoring Tasks",
author = "Thakur, Nandan and Reimers, Nils and Daxenberger, Johannes and Gurevych, Iryna",
booktitle = "Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies",
month = jun,
year = "2021",
address = "Online",
publisher = "Association for Computational Linguistics",
url = "https://www.aclweb.org/anthology/2021.naacl-main.28",
pages = "296--310",
}
Overview
* Installation
+ Install SentenceTransformers
+ Install PyTorch with CUDA support
* Quickstart
+ Comparing Sentence Similarities
+ Pre-Trained Models
+ Training your own Embeddings
* Pretrained Models
+ Model Overview
+ Semantic Search
+ Multi-Lingual Models
+ Image & Text-Models
+ Other Models
* Pretrained Cross-Encoders
+ MS MARCO
+ SQuAD (QNLI)
+ STSbenchmark
+ Quora Duplicate Questions
+ NLI
* Publications
* Hugging Face
+ The Hugging Face Hub
+ Using Hugging Face models
+ Sharing your models
+ Sharing your embeddings
+ Additional resources
Usage
* Computing Sentence Embeddings
+ Prompt Templates
+ Input Sequence Length
+ Storing & Loading Embeddings
+ Multi-Process / Multi-GPU Encoding
+ Sentence Embeddings with Transformers
* Semantic Textual Similarity
* Embedding Quantization
+ Binary Quantization
+ Scalar (int8) Quantization
+ Additional extensions
+ Demo
+ Try it yourself
* Semantic Search
+ Background
+ Symmetric vs. Asymmetric Semantic Search
+ Python
+ util.semantic_search
+ Speed Optimization
+ Elasticsearch
+ Approximate Nearest Neighbor
+ Retrieve & Re-Rank
+ Examples
* Retrieve & Re-Rank
+ Retrieve & Re-Rank Pipeline
+ Retrieval: Bi-Encoder
+ Re-Ranker: Cross-Encoder
+ Example Scripts
+ Pre-trained Bi-Encoders (Retrieval)
+ Pre-trained Cross-Encoders (Re-Ranker)
* Clustering
+ k-Means
+ Agglomerative Clustering
+ Fast Clustering
+ Topic Modeling
* Paraphrase Mining
* Translated Sentence Mining
+ Marging Based Mining
+ Examples
* Cross-Encoders
+ Bi-Encoder vs. Cross-Encoder
+ When to use Cross- / Bi-Encoders?
+ Cross-Encoders Usage
+ Combining Bi- and Cross-Encoders
+ Training Cross-Encoders
* Image Search
+ Installation
+ Usage
+ Examples
Training
* Training Overview
+ Network Architecture
+ Creating Networks from Scratch
+ Training Data
+ Loss Functions
+ Evaluators
+ Loading Custom SentenceTransformer Models
+ Multitask Training
+ Adding Special Tokens
+ Best Transformer Model
* Loss Overview
+ Loss modifiers
+ Distillation
+ Commonly used Loss Functions
* Matryoshka Embeddings
+ Use Cases
+ Results
+ Training
+ Inference
+ Code Examples
* Adaptive Layers
+ Use Cases
+ Results
+ Training
+ Inference
+ Code Examples
* Multilingual-Models
+ Available Pre-trained Models
+ Usage
+ Performance
+ Extend your own models
+ Training
+ Data Format
+ Loading Training Datasets
+ Sources for Training Data
+ Evaluation
+ Citation
* Model Distillation
+ Knowledge Distillation
+ Speed - Performance Trade-Off
+ Dimensionality Reduction
+ Quantization
* Cross-Encoders
+ Examples
+ Training CrossEncoders
* Augmented SBERT
+ Motivation
+ Extend to your own datasets
+ Methodology
+ Scenario 1: Limited or small annotated datasets (few labeled
sentence-pairs)
+ Scenario 2: No annotated datasets (Only unlabeled
sentence-pairs)
+ Training
+ Citation
* Training Datasets
+ Datasets on the Hugging Face Hub
Training Examples
* Semantic Textual Similarity
+ Training data
+ Loss Function
* Natural Language Inference
+ Data
+ SoftmaxLoss
+ MultipleNegativesRankingLoss
* Paraphrase Data
+ Datasets
+ Training
+ Pre-Trained Models
+ Work in Progress
* Quora Duplicate Questions
+ Pretrained Models
+ Dataset
+ Usage
+ Training
+ MultipleNegativesRankingLoss
* MS MARCO
+ Bi-Encoder
+ Cross-Encoder
+ Cross-Encoder Knowledge Distillation
Unsupervised Learning
* Unsupervised Learning
+ TSDAE
+ SimCSE
+ CT
+ CT (In-Batch Negative Sampling)
+ Masked Language Model (MLM)
+ GenQ
+ GPL
+ Performance Comparison
* Domain Adaptation
+ Domain Adaptation vs. Unsupervised Learning
+ Adaptive Pre-Training
+ GPL: Generative Pseudo-Labeling
Package Reference
* SentenceTransformer
* util
* quantization
* Models
* Losses
* Evaluation
* Datasets
* cross_encoder
Next
---------------------------------------------------------------------
(c) Copyright 2024, Nils Reimers * Contact
Built with Sphinx using a theme provided by Read the Docs.