[HN Gopher] Punica: Serving multiple LoRA finetuned LLM as one
       ___________________________________________________________________
        
       Punica: Serving multiple LoRA finetuned LLM as one
        
       Author : abcdabcd987
       Score  : 25 points
       Date   : 2023-11-08 20:42 UTC (2 hours ago)
        
 (HTM) web link (github.com)
 (TXT) w3m dump (github.com)
        
       | junrushao1994 wrote:
       | This is great! Have you guys considered integrating with one of
       | the existing systems?
        
       | yyding wrote:
       | Good job! I observed that you implemented many cuda kernels by
       | yourselves. Just wondering your consideration or trade-off
       | between implementating the kernels via pure CUDA code vs.
       | implementing based on compiler like TVM/Triton.
        
       ___________________________________________________________________
       (page generated 2023-11-08 23:01 UTC)