[HN Gopher] Punica: Serving multiple LoRA finetuned LLM as one
___________________________________________________________________
Punica: Serving multiple LoRA finetuned LLM as one
Author : abcdabcd987
Score : 25 points
Date : 2023-11-08 20:42 UTC (2 hours ago)
(HTM) web link (github.com)
(TXT) w3m dump (github.com)
| junrushao1994 wrote:
| This is great! Have you guys considered integrating with one of
| the existing systems?
| yyding wrote:
| Good job! I observed that you implemented many cuda kernels by
| yourselves. Just wondering your consideration or trade-off
| between implementating the kernels via pure CUDA code vs.
| implementing based on compiler like TVM/Triton.
___________________________________________________________________
(page generated 2023-11-08 23:01 UTC)