[HN Gopher] LLama.cpp now has a web interface
       ___________________________________________________________________
        
       LLama.cpp now has a web interface
        
       Author : xal
       Score  : 267 points
       Date   : 2023-07-05 17:33 UTC (5 hours ago)
        
 (HTM) web link (github.com)
 (TXT) w3m dump (github.com)
        
       | wiihack wrote:
       | Nice simple interface! And great to see a CEO keeping up with
       | recent technological developments :)
        
       | creekcreek wrote:
       | [dead]
        
       | vdfs wrote:
       | Good to see a CEO is still hacking around!
       | 
       | Shopify did change my life in a way i could never imaging, it
       | will always hold a special place in my heart
        
         | Xenoamorphous wrote:
         | A billionare co-founder at that!
        
           | swyx wrote:
           | actually did tobi even have any other cofounders? i just
           | realized i only ever hear about him and harvey finkelstein
           | 
           | edit: google says scott lake was actually the founding CEO.
           | TIL! https://www.linkedin.com/mwlite/profile/in/scottlake?ori
           | gina...
        
             | jonny_eh wrote:
             | > harvey finkelstein
             | 
             | Harley Finkelstein, he's a nice guy. Not a co-founder but
             | very important to the history and running of Shopify.
        
             | Xenoamorphous wrote:
             | Just going by the Google snippet I got when searching for
             | "shopify ceo". Not sure if the downvotes I got are because
             | somone thinks that the co-founding aspect is wrong or
             | because they think that him being a billionaire is
             | irrelevant, but IMO it makes the fact he's opening PRs even
             | cooler.
        
         | samstave wrote:
         | Can you give a detailed account on specifics?
        
           | [deleted]
        
       | isoprophlex wrote:
       | > I tried to match the spirit of llama.cpp and used minimalistic
       | js dependencies and went with the ozempic css style of ggml.ai.
       | 
       | Ozempic? The anti-diabetes drug? That's either a glorious typo or
       | an interesting new adjective...
        
         | jahewson wrote:
         | I did a double-take when I saw that too. It's now prescribed as
         | a weight-loss drug and is very much in the zeitgeist so yeah I
         | think it's a new adjective. Personally I'm going to stick with
         | "light weight".
        
           | isoprophlex wrote:
           | The more I think about it, the more I'm loving it
        
         | xal wrote:
         | Nat coined this usage
         | https://twitter.com/natfriedman/status/1668656170645749761
        
       | eclectic29 wrote:
       | Can someone shed light on how does a CEO with 3 kids get time to
       | hack on something like this? Some might even argue that all the
       | time spent doing this would've been better spent on CEO
       | activities, but thankfully this is HN and people have hobbies, so
       | that's that, but this can be a very time consuming hobby.
        
         | xal wrote:
         | I've tried to eliminate language around being "too busy" from
         | my vocabulary and attempt to replace it with "can't
         | prioritize". That's sometimes a bit awkward, but really trains
         | better habits. Sometimes I prioritize hobbies if I feel I need
         | it, and doing this seemed fun and useful besides!
        
           | eclectic29 wrote:
           | Thanks for your helpful response.
        
             | vdfs wrote:
             | He have this in his Twitter bio: "CEO by day, Dad in
             | evening, hacker at night"
        
       | mliker wrote:
       | very cool to see the CEO of Shopify putting together this
       | interface.
        
       | mteam88 wrote:
       | Are these comments about the CEO bots?
        
         | mrtranscendence wrote:
         | Were you aware that the author is Shopify's CEO? How
         | deliciously absurd!
        
         | generalizations wrote:
         | I think this has been seen before, where it turned out to be a
         | bunch of employees. But that time it was a product launch, not
         | a CEO's pr.
        
         | isanjay wrote:
         | Dedo something fishy. But accounts are created ages ago
        
           | Kiro wrote:
           | Not really. I also think it's remarkable and would have
           | posted something similar if it hadn't already been pointed
           | out.
        
         | Cyph0n wrote:
         | People are just trying to point out who the author is for those
         | not familiar with him.
        
         | mliker wrote:
         | I think it's just impressive to folks that the CEO of a public
         | company is still coding and putting out useful PRs.
        
       | steren wrote:
       | And the PR is from Shopify's CEO!
        
         | vdfs wrote:
         | OP is Tobi
        
       | [deleted]
        
       | seydor wrote:
       | didnt people run llamacpp with oobabooga et al?
        
       | zoklet-enjoyer wrote:
       | Wow. Every comment pointing out the author but not the content
        
       | ynniv wrote:
       | i'm importing from js cdns instead of adding them here
       | 
       | FWIW this seems counter to llama.cpp's philosophy.
        
         | mhh__ wrote:
         | The philosophy in practice seems to be dirty deeds done dirt
         | cheap rather than genuine simplicity (it only looks simple
         | compared to python ecosystem)
        
         | xal wrote:
         | I agree, I ended up getting rid of it. Only one dependency is
         | downloaded (via bash script) and everything is baked into the
         | binary now.
        
       | franchiser wrote:
       | [dead]
        
       | mk_stjames wrote:
       | I'm always wondering about the, I don't even know what to call
       | this, etiquette? of proposing PR's to projects like these that
       | add a feature or a demo or whatnot to the main branch of a very
       | focused project by adding something that is very different in
       | interface, language, set and setting etc.
       | 
       | So in this case, Tobi made this awesome little web interface that
       | uses minimal HTML and JS as to stay in line with llama.cpp's
       | stripped-down-ness. But it is still a completely different mode
       | of operation, it's a 'new venue' essentially.
       | 
       | What if GG didn't want such a thing? When is something like this
       | better for a separately maintained repo and not a main merge? How
       | do you know when it is OK to submit a PR to add something like
       | this without overstepping (or is it always?)
       | 
       | I see this with a few projects on github that really 'blow up'
       | and everyone starts working on. They get a million PR's from
       | people hacking things on it in their domain of knowledge,
       | expanding the complexity (and potentially difficulty to maintain
       | quality). Sometimes it gets weird feeling watching from the
       | outside at least (I'm not a maintainer on any public FOSS).
       | 
       | Just curious what others think because those are my thoughts that
       | came to mind when I saw this.
        
         | LawnGnome wrote:
         | I generally think it's fine to do this sort of thing for your
         | own benefit and open a PR as long as you're really 100% fine
         | with "no, I'm not interested in merging this" being the answer.
         | 
         | Where the problems tend to arise (in my experience, at least)
         | is when people hack on something expecting that it will be
         | merged, get invested in it, and then get upset when the
         | maintainer(s) aren't interested.
         | 
         | Checking in before starting to work on something is important
         | if your goal is to have it merged, not just to do the work. The
         | problem is that a lot of people start in the first category,
         | but then move into the second category as they get invested in
         | their project.
        
         | [deleted]
        
         | ggerganov wrote:
         | My POV is that llama.cpp is primarily a playground for adding
         | new features to the core ggml library and in the long run an
         | interface for efficient LLM inference. The purpose of the
         | examples in the repo is to demonstrate ways of how to use the
         | ggml library and the LLM interface. The examples are decoupled
         | from the primary code - i.e. you can delete all of them and the
         | project will continue to function and build properly. So we can
         | afford to expand them more freely as long as people find them
         | useful and there is enough help for maintaining them. Still, we
         | try to keep the 3rd party dependencies to a minimum so that the
         | build process is simple and accessible
         | 
         | There was a similar "dilemma" about the GPU support - initially
         | I didn't envision adding GPU support to the core library as I
         | thought that things will become very entangled and hard to
         | maintain. But eventually, we found a way to extend the library
         | with different GPU backends in a relatively well decoupled way.
         | So now, we have various developers maintaining and contributing
         | to the backends in a nice independent way. Each backend can be
         | deleted and you will still be able to build the project and use
         | it.
         | 
         | So I guess we are optimizing for how easy it is to delete
         | things :)
         | 
         | Note that the project is still pretty much a "big hack" - it
         | supports just LLaMA models and derivatives, therefore it is
         | easy atm. The more "general purpose" it becomes, the more
         | difficult things become to design and maintain. This is the
         | main challenge I'm thinking how to solve, but for sure keeping
         | stuff minimalistic and small is a great help so far
         | 
         | > What if GG didn't want such a thing? When is something like
         | this better for a separately maintained repo and not a main
         | merge? How do you know when it is OK to submit a PR to add
         | something like this without overstepping (or is it always?)
         | 
         | I try to explain my vision for the project in the issues and
         | the discussion. I think most of the developers are very well
         | aligned with it and can already tell what is a good addition or
         | not
        
           | mk_stjames wrote:
           | Thanks for replying to me directly! I'm finding it
           | fascinating to follow this project. Good luck with your
           | company Georgi.
        
           | vitaminka wrote:
           | i'm curious, what's is the approach for maintainable and
           | decoupled various gpu backends?
        
             | ggerganov wrote:
             | It was designed in #915 (read just the OP and the linked
             | PRs at the end) and the implementation pretty much follows
             | it closely, at least for the Metal backend. The CUDA and
             | OpenCL backends are currently slightly coupled in ggml as
             | they started developing before #915, but I think we'll
             | resolve this eventually.
             | 
             | #915 -
             | https://github.com/ggerganov/llama.cpp/discussions/915
        
               | vitaminka wrote:
               | interesting decoupling method, ty :)
        
           | aidenn0 wrote:
           | Thank you for the ggml library, by the way. It let me play
           | around with whisper in a sane manner. To run the CUDA torch
           | versions, I needed to shut down X to free enough GPU memory
           | for the medium model, and the small model might require me to
           | quit firefox. With ggml, I can use cublas and run even the
           | large model with a _huge_ speedup compared to CPU only torch.
        
         | xal wrote:
         | GG would just say no and that's that. No hard feelings, that's
         | what makes open source so great.
        
         | grepLeigh wrote:
         | Some tips/tricks/tidbits:
         | 
         | * Open a draft PR early in the process with a Request for
         | Comment [RFC] tag. Explain your goal/approach in words, then
         | follow up with code.
         | 
         | * Be succinct.
         | 
         | * Provide minimal viable examples and build more complex
         | concepts from these.
         | 
         | * Accept feedback with grace, and execute promptly.
         | 
         | * Don't take personal offense if your work isn't merged, or
         | even responded to.
         | 
         | * Single-maintainer open-source looks very different than
         | consortium & working group FOSS.
        
         | renewiltord wrote:
         | You just saw it happen in OP. Just do as others do and don't
         | sweat this stuff. The principle is code sharing and an offer of
         | a thing you've done.
         | 
         | Don't overthink it.
         | 
         | This is fantastic. I love the way he handles his project. Just
         | great for adoption and contribution.
        
         | version_five wrote:
         | A nice thing about llama.cpp is that it's well organized to
         | accept a feature like this without really disturbing any other
         | part or potentially stepping on someone's toes. There is the
         | core repo and then this is on examples/server (as are various
         | other "example" features). This organization feel like it would
         | make it much easier to accept a pr like this than if doing the
         | same thing required wider changes.
        
       ___________________________________________________________________
       (page generated 2023-07-05 23:01 UTC)