https://dstack.ai/ [ ] [ ] dstack 0.15.1 is here! Kubernetes integration preview. Learn more. logo dstack Orchestrate GPU workloads effortlessly on any cloud ( ) Sign in Get started * Home * Docs * Examples * Changelog * Discord * GitHub logo dstack dstackai/dstack * [ ] Home * [ ] Docs Docs + [ ] Getting started Getting started o Introduction o Installation o Quickstart + [ ] Concepts Concepts o Dev environments o Tasks o Services + [ ] Reference Reference o CLI o [ ] API API # Python API # REST API o .dstack.yml o server/config.yml o profiles.yml * [ ] Examples Examples + Examples + [ ] LLMs LLMs o Mixtral 8x7B + [ ] Fine-tuning Fine-tuning o QLoRA + [ ] Deployment Deployment o Text Generation Inference o vLLM o Text Embedding Interface o SDXL o Infinity + [ ] RAG RAG o Llama Index * [ ] Changelog Changelog + Changelog + [ ] Archive Archive o 2024 o 2023 * Discord * GitHub Orchestrate GPU workloads effortlessly on any cloud dstack is an open-source engine for running GPU workloads. It simplifies development, training, and deployment of gen AI models on any cloud. Get started Documentation Development Experiment interactively in your IDE, terminal, or Jupyter notebooks before submitting long tasks or deploying models. With dstack, a single command provisions the necessary cloud resources, code, and environment for your dev setup. Learn more [dstack-dev-environment] [dstack-task] Training With dstack, running tasks such as training or fine-tuning scripts, or any other batch jobs, is incredibly easy. Simply provide the commands, ports, and choose the Python version or a Docker image. dstack will handle the execution on configured cloud GPU provider(s) with the necessary resources. Learn more Deployment With dstack, deploying models or any other web apps is straightforward. Just provide commands, port, and select Python version or Docker image. dstack handles deployment on configured cloud GPU provider(s), giving you a public HTTPS endpoint. Learn more [dstack-service-openai] Featured examples Mixtral 8x7B Deploy Mixtral 8x7B as a service using vLLM, an open-source serving library. LLMs Text Embeddings Inference Deploy text embeddings models using Services and TEI, an open-source text embeddings toolkit by Hugging Face. Deployment Llama Index Use Llama Index and Weaviate to enhance the capabilities of LLMs with the context of your data. RAG QLoRA Fine-tune Llama 2 on a custom dataset, with QLoRA and your own script, using Tasks. Fine-tuning Text Generation Inference Deploy LLMs using Services and TGI, an open-source serving framework by Hugging Face. Deployment vLLM Deploy LLMs with Services and vLLM, an open-source serving library. Deployment Browse all examples Get started in a minute Open-source Use with your own cloud accounts or on-premises. [aws-] Amazon Web Services # Azure # Google Cloud Platform # Lambda # TensorDock # Vast.ai # Kubernetes Install open-source 100% open-source dstack Cloud Access GPUs from a wide range of providers, ensuring the best price and availability. [aws-] Amazon Web Services # Azure # Google Cloud Platform # Lambda # TensorDock # Vast.ai Sign in with GitHub Pay per compute only Community Need help, have questions, or simply want to stay updated? Join Discord FAQ What is the difference between the open-source version and dstack Cloud? The open-source version allows you to run workloads using your own cloud accounts. It can be utilized via the CLI or API and enables the configuration of multiple projects and users. dstack Cloud is a fully managed service that enables you to run workloads across multiple cloud providers, guaranteeing optimal GPU pricing and availability. You don't need individual accounts with each provider - dstack Cloud manages everything for you. Back to top (c) 2024 dstack GmbH Community Discord GitHub Twitter Resources Docs Examples Changelog Legal Terms Privacy