[HN Gopher] Lightweight Safety Classification Using Pruned Langu...
       ___________________________________________________________________
        
       Lightweight Safety Classification Using Pruned Language Models
        
       Layer Enhanced Classification (LEC) is a novel technique that
       outperforms current industry leaders like GPT-4o, LlamaGuards 1 and
       8B, and deBERTa v3 Prompt Injection v2 for content safety and
       prompt injection tasks.  We prove that the intermediate hidden
       layers in transformers are robust feature extractors for text
       classification.  On content safety, LEC models achieved a 0.96 F1
       score vs GPT-4o's 0.82 and Llama Guard 8B's 0.71.The LEC models
       were able to outperform the other models with only 15 training
       examples for binary classification and 50 examples for multi-class
       classification across 66 categories.  On prompt injection,LEC
       models achieved a 0.98 F1 score vs GPT-4o's 0.92 and deBERTa v3
       Prompt Injection v2's 0.73. LEC models were able to outperform
       deBERTa with only 5 training examples and GPT-4o with only 55
       training examples.  Read the full paper and our approach here:
       https://arxiv.org/abs/2412.13435
        
       Author : sandijean90
       Score  : 16 points
       Date   : 2024-12-19 17:49 UTC (5 hours ago)
        
 (HTM) web link (arxiv.org)
 (TXT) w3m dump (arxiv.org)
        
       | bberenberg wrote:
       | Are these models available for us to try out?
        
       | gdiamos wrote:
       | This is really easy to set up - and is much more accurate than
       | asking the LLM to predict True/False.
       | 
       | Just feed the outputs of an embedding API into logistic
       | regression, e.g. from sklearn.                 import voyageai
       | vo = voyageai.Client()       # This will automatically use the
       | environment variable VOYAGE_API_KEY.       # Alternatively, you
       | can use vo = voyageai.Client(api_key="<your secret key>")
       | texts = ["Sample text 1", "Sample text 2"]            result =
       | vo.embed(texts, model="voyage-2", input_type="document")
       | print(result.embeddings[0])
       | 
       | https://scikit-learn.org/1.5/modules/generated/sklearn.linea...
        
       | flatpepsi17 wrote:
       | I'd pay good money for a local LLM with no "content safety" at
       | all.
        
       ___________________________________________________________________
       (page generated 2024-12-19 23:02 UTC)