https://gpuopen.com/learn/pytorch-windows-amd-llm-guide/ AMD Logo GPUOpen Search CtrlK Cancel AMD FidelityFX(tm) AMD FidelityFX SDK v2 Super Resolution 4 (FSR 4) More AMD FidelityFX Super Resolution 3 (FSR 3) Super Resolution 2 (FSR 2) Super Resolution 1 (FSR 1) AMD FidelityFX SDK v1 Blur Breadcrumbs library Brixelizer/GI Ambient Occlusion (CACAO) Contrast Adaptive Sharpening (CAS) Denoiser Depth of Field (DoF) Lens HDR Mapper (LPM) Parallel Sort Downsampler (SPD) Screen Space Reflections (SSSR) Variable Shading TressFX Developer testimonials Cauldron Framework FidelityFX Naming Guidelines Hybrid Shadows Hybrid Stochastic Reflections Tools What Tools Do We Have? Radeon Developer Tool Suite Radeon(tm) GPU Detective Radeon(tm) Raytracing Analyzer Radeon(tm) GPU Profiler Radeon(tm) GPU Analyzer Radeon(tm) Memory Visualizer Radeon(tm) Developer Panel GPU Reshape Compressonator Frame Latency Meter OCAT SDKs What SDKs Do We Have? AMD Radeon Anti-Lag 2 SDK AMD Driver/Hardware SDKs AMD GPU Services AMD Device Library eXtra Advanced Media Framework Streaming SDK GPU Performance API Content Creation Radeon(tm) ProRender Suite Radeon(tm) ProRender SDK GPUOpen MaterialX Library Radeon(tm) Rays Vulkan(r) Memory Allocator Direct3D(r)12 Memory Allocator HIP Ray Tracing Orochi Capsaicin Framework (GI-1.0) Render Pipeline Shaders Brotli-G SDK Dense Geometry Format SDK Platform Support Unreal Engine FSR 4 Unreal Engine 5 plugin Unreal Engine Performance Guide AMD Schola (Unreal NPCs) Unity Unity CPU Profiling Guide Vulkan(r) Radeon(tm) Vulkan(r) Drivers Version Table Developing Vulkan(r) Applications DirectX(r)12 DirectX(r)12 Ultimate Developing DirectX(r)12 Applications Docs/Research Meet all our blogs Getting Started Getting Started with our Software Getting Started with Development How to Become a Graphics Programmer General Developer Tech Articles Popular articles Integrating Anti-Lag 2 SDK Matrix Compendium Mesh Shaders Work Graphs Crash Course in Deep Learning (Graphics) Our Publications Advanced Rendering Research Group AMD Lab Notes (HPC) AMD RDNA(tm) Performance Guide AMD GPU Architecture Machine-readable ISAs AMD Ryzen(tm) Performance Guide CPU Performance and Optimization Software Manuals Presentations Samples News/Events Latest Developer News Recent Software Releases Hot new articles AMD FSR 4 is now available Introducing AMD FidelityFX SDK v2.0 Use WMMA intrinsics on AMD RDNA 4 Tool Suite updated for Radeon RX 9060 XT Generative AI model for GI effects Neural Networks for Geometric Representation Generative AI on AMD Radeon GPUs Events All Events GDC Digital Dragons Other events Watch Our Videos * AMD FidelityFX(tm) AMD FidelityFX SDK v2 Super Resolution 4 (FSR 4) More AMD FidelityFX Super Resolution 3 (FSR 3) Super Resolution 2 (FSR 2) Super Resolution 1 (FSR 1) AMD FidelityFX SDK v1 Blur Breadcrumbs library Brixelizer/GI Ambient Occlusion (CACAO) Contrast Adaptive Sharpening (CAS) Denoiser Depth of Field (DoF) Lens HDR Mapper (LPM) Parallel Sort Downsampler (SPD) Screen Space Reflections (SSSR) Variable Shading TressFX Developer testimonials Cauldron Framework FidelityFX Naming Guidelines Hybrid Shadows Hybrid Stochastic Reflections * Tools What Tools Do We Have? Radeon Developer Tool Suite Radeon(tm) GPU Detective Radeon(tm) Raytracing Analyzer Radeon(tm) GPU Profiler Radeon(tm) GPU Analyzer Radeon(tm) Memory Visualizer Radeon(tm) Developer Panel GPU Reshape Compressonator Frame Latency Meter OCAT * SDKs What SDKs Do We Have? AMD Radeon Anti-Lag 2 SDK AMD Driver/Hardware SDKs AMD GPU Services AMD Device Library eXtra Advanced Media Framework Streaming SDK GPU Performance API Content Creation Radeon(tm) ProRender Suite Radeon(tm) ProRender SDK GPUOpen MaterialX Library Radeon(tm) Rays Vulkan(r) Memory Allocator Direct3D(r)12 Memory Allocator HIP Ray Tracing Orochi Capsaicin Framework (GI-1.0) Render Pipeline Shaders Brotli-G SDK Dense Geometry Format SDK * Platform Support Unreal Engine FSR 4 Unreal Engine 5 plugin Unreal Engine Performance Guide AMD Schola (Unreal NPCs) Unity Unity CPU Profiling Guide Vulkan(r) Radeon(tm) Vulkan(r) Drivers Version Table Developing Vulkan(r) Applications DirectX(r)12 DirectX(r)12 Ultimate Developing DirectX(r)12 Applications * Docs/Research Meet all our blogs Getting Started Getting Started with our Software Getting Started with Development How to Become a Graphics Programmer General Developer Tech Articles Popular articles Integrating Anti-Lag 2 SDK Matrix Compendium Mesh Shaders Work Graphs Crash Course in Deep Learning (Graphics) Our Publications Advanced Rendering Research Group AMD Lab Notes (HPC) AMD RDNA(tm) Performance Guide AMD GPU Architecture Machine-readable ISAs AMD Ryzen(tm) Performance Guide CPU Performance and Optimization Software Manuals Presentations Samples * News/Events Latest Developer News Recent Software Releases Hot new articles AMD FSR 4 is now available Introducing AMD FidelityFX SDK v2.0 Use WMMA intrinsics on AMD RDNA 4 Tool Suite updated for Radeon RX 9060 XT Generative AI model for GI effects Neural Networks for Geometric Representation Generative AI on AMD Radeon GPUs Events All Events GDC Digital Dragons Other events Watch Our Videos 1. Home 2. >> Learn 3. >> A beginner's guide to deploying LLMs with AMD on Windows using PyTorch Share on Bluesky Share on Mastadon Share on LinkedIn Share on Twitter /X Share on Reddit Share on Facebook Share on Whatsapp Share via Email A beginner's guide to deploying LLMs with AMD on Windows using PyTorch Originally posted: September 24, 2025 Warren Eng's avatar Warren Eng Sheen Lam's avatar Sheen Lam Alexander Blake-Davies's avatar Alexander Blake-Davies If you're interested in deploying advanced AI models on your local hardware, leveraging a modern AMD GPU or APU can provide an efficient and scalable solution. You don't need dedicated AI infrastructure to experiment with Large Language Model (LLMs); a capable Microsoft(r) Windows(r) PC with PyTorch installed and equipped with a recent AMD graphics card is all you need. PyTorch for AMD on Windows and Linux is now available as a public preview. You can now use native PyTorch for AI inference on AMD Radeon(tm) RX 7000 and 9000 series GPUs and select AMD Ryzen(tm) AI 300 and AI Max APUs, enabling seamless AI workload execution on AMD hardware in Windows without any need for workarounds or dual-boot configurations. If you are new and just getting started with AMD ROCm(tm), be sure to check out our getting started guides here. This guide is designed for developers seeking to set up, configure, and execute LLMs locally on a Windows PC using PyTorch with an AMD GPU or APU. No previous experience with PyTorch or deep learning frameworks is needed. What you'll need (the prerequisites) * The currently supported AMD platforms and hardware for PyTorch on Windows are listed here: AMD Radeon(tm) AI AMD Radeon(tm) RX AMD Radeon(tm) PRO AMD Ryzen(tm) AI PRO R9700 7900 XTX W7900 Max+ 395 AMD Radeon(tm) RX AMD Radeon(tm) PRO AMD Ryzen(tm) AI 9070 XT W7900 Dual Slot Max 390 AMD Radeon(tm) RX AMD Ryzen(tm) AI 9070 Max 385 AMD Radeon(tm) RX AMD Ryzen(tm) AI 9 9060 XT HX 375 AMD Ryzen(tm) AI 9 HX 370 AMD Ryzen(tm) AI 9 365 * OS Supported: Microsoft(r) Windows(r) 11 * AMD Software: PyTorch on Windows Preview Edition 25.20.01.14 driver * Python 3.12 - Python Release Python 3.12.0 | Python.org + During the Python installation, make sure to check the box that says "Add Python to PATH" Part 1: Setting up your workspace Step 1: Open the Command Prompt First, we need to open the Command Prompt * Click the Start Menu, type cmd, and press Enter. A black terminal window will pop up. Terminal Window Step 2: Create and activate a virtual environment A "virtual environment" is like a clean, empty sandbox for a Python project. In your Command Prompt, type the following command and press Enter. This creates a new folder named llm-pyt that will house our project. python -m venv llm-pyt Next, we need to "activate" this environment. Think of this as stepping inside the sandbox. llm-pyt\Scripts\activate You'll know it worked because you'll see (llm-pyt) appear at the beginning of your command line prompt. llm-pyt prompt Step 3: Install PyTorch and other essential libraries Now we'll install the software libraries that do the heavy lifting. The most important one is PyTorch, an open-source framework for building and running AI models. We need a special version of PyTorch built to work with AMD's ROCm technology. We will also install Transformers and Accelerate, two libraries from Hugging Face that make it incredibly easy to download and run state-of-the-art AI models. Run the following command in your activated Command Prompt. This command tells Python's package installer (pip) to download and install PyTorch for ROCm, along with the other necessary tools. Terminal window pip install --no-cache-dir https://repo.radeon.com/rocm/windows/rocm-rel-6.4.4/torch-2.8.0a0%2Bgitfc14c65-cp312-cp312-win_amd64.whl pip install --no-cache-dir https://repo.radeon.com/rocm/windows/rocm-rel-6.4.4/torchaudio-2.6.0a0%2B1a8f621-cp312-cp312-win_amd64.whl pip install --no-cache-dir https://repo.radeon.com/rocm/windows/rocm-rel-6.4.4/torchvision-0.24.0a0%2Bc85f008-cp312-cp312-win_amd64.whl pip install transformers accelerate Part 2: Putting your new LLM setup to the test The moment of truth. Let's give our new setup a task: running a small but powerful language model called Llama 3.2 1B. Step 1: Launch the interactive Python session Make sure your Command Prompt still has the (llm-pyt) environment active. If you closed it, just re-open cmd and run llm-pyt\Scripts\ activate. Now, start Python: Terminal window python Step 2: Run the language model Copy the entire code block below. Paste it into your Python terminal (where you see the >>>) and press Enter. The first time you do this, it will download the model (which is a few gigabytes), so it may take several minutes. Subsequent runs will be much faster. import torch from transformers import pipeline model_id = "unsloth/Llama-3.2-1B-Instruct" pipe = pipeline( "text-generation", model=model_id, dtype=torch.float16, device_map="auto" ) pipe("The key to life is") You should see an output similar to this: [{'generated_text': 'The key to life is not to get what you want, but to give what you have. The best way to make life more meaningful is to practice gratitude, and to cultivate a sense of contentment with what you have. If you want to make life more interesting, you must be willing to take risks, and to embrace the unknown. The best way to avoid disappointment is to be patient and persistent, and to trust in the process. By following these principles, you can live a more fulfilling life, and make the most of the time you have.'}] You can return to your command prompt by typing exit() and pressing Enter. exit() Level Up: Create an interactive AI chatbot Running a single prompt is fun, but a real conversation is better. In this section, we'll create an interactive chat loop that "remembers" the conversation, allowing you to have a back-and-forth with the AI. Step 1: Create the chatbot script 1. Open a new file in your text editor. 2. Copy and paste the chatbot code below. import torch from transformers import pipeline print("Loading chat model...") model_id = "unsloth/Llama-3.2-1B-Instruct" pipe = pipeline( "text-generation", model=model_id, dtype=torch.float16, device_map="auto", ) # This list will store our conversation history messages = [] print("\nChatbot ready! Type 'quit' or 'exit' to end the conversation.") print("-" * 20) while True: # Get input from the user user_input = input("You: ") # Check if the user wants to exit if user_input.lower() in ["quit", "exit"]: print("Chat session ended.") break # Add the user's message to the conversation history messages.append({"role": "user", "content": user_input}) # Generate the AI's response using the full conversation history outputs = pipe(messages, max_new_tokens=500, do_sample=True, temperature=0.7) # The pipeline returns the full conversation. The last message is the new one. assistant_response = outputs[0]['generated_text'][-1]['content'] # Add the AI's response to our history messages.append({"role": "assistant", "content": assistant_response}) # Print just the AI's new response print(f"AI: {assistant_response}") 3. Save this new file as run_chat.py in the same user folder. Step 2: Run your chatbot In your Command Prompt, run the new script: Terminal window python run_chat.py The terminal will now prompt you with You:. Type a question and press Enter. The AI will respond, and you can ask follow-up questions. The chatbot will remember the context of the conversation. Results Note: When you run the LLM, you will see a warning message like this: UserWarning: 1Torch was not compiled with memory efficient attention. (Triggered internally at C:\develop\pytorch-test\aten\src\ATen\native\transformers\hip\sdp_utils.cpp:726.) Don't worry, this is expected, and your code is working correctly! What it means in simple terms: PyTorch 2.0+ introduced a feature called "Memory-Efficient Attention" to speed things up. The current version of PyTorch for AMD on Windows doesn't include this specific optimization out-of-the-box. When PyTorch can't find it, it prints this warning and automatically falls back to the standard, reliable method. Summary By following this blog, you should be able to get started with running transformer-based LLMs with PyTorch on Windows using AMD consumer graphics hardware. You can learn more about our road to AMD ROCm on Radeon for Windows and Linux in the blog from Andrej Zdravkovic, SVP and AMD Chief Software Officer, here. View endnotes PyTorch, the PyTorch logo and any related marks are trademarks of The Linux Foundation. Windows is a trademark of the Microsoft group of companies. Warren Eng's avatar Warren Eng Warren Eng is a Product Marketing Manager at AMD. During his time here, he has done technical marketing for consumer graphics, product marketing for workstation and datacenter products, managed a team of software marketing experts responsible for features like FSR 4, and currently owns product marketing for our AMD Ryzen processors for gamers and enthusiasts. On his free time he likes to travel and explore with his family, eat good food or relax on the couch watching Star Trek. Sheen Lam's avatar Sheen Lam Sheen Lam is a member of technical staff at AMD, working on ROCm on Radeon. Prior to that, he worked on Cloud and Virtualization products. Alexander Blake-Davies's avatar Alexander Blake-Davies Alexander Blake-Davies is a Senior Software Product Marketing Specialist for AMD Developer Programs. Related news and technical articles Accelerating Generative AI on AMD Radeon(tm) GPUs Accelerating Generative AI on AMD Radeon(tm) GPUs Discover AMD-optimized ONNX models on Hugging Face for AMD Ryzen(tm) AI APUs and Radeon(tm) GPUs and incredible performance with the AMD Radeon RX 9000 Series' advanced AI accelerators. Crash Course in Deep Learning (for Computer Graphics) Crash Course in Deep Learning (for Computer Graphics) If you're a graphics dev looking to understand more about deep learning, this blog introduces the basic principles in a graphics dev context. Jacobi Solver with HIP and OpenMP offloading Jacobi Solver with HIP and OpenMP offloading In this blog, we explore GPU offloading using HIP and OpenMP target directives and discuss their relative merits in terms of implementation efforts and performance. Creating a PyTorch/TensorFlow Code Environment on AMD GPUs Creating a PyTorch/TensorFlow Code Environment on AMD GPUs The machine learning ecosystem is quickly exploding and this article is designed to assist data scientists/ML practitioners get their machine learning environments up and running on AMD GPUs. GPU-aware MPI with ROCm GPU-aware MPI with ROCm MPI is the de facto standard for inter-process communication in High-Performance Computing. This post will guide you through the process of setting up an MPI application that supports execution on GPU clusters. Finite Difference Method - Laplacian part 3 Finite Difference Method - Laplacian part 3 In this third part, we cover additional optimizations to fine tune the performance of the kernel, and introduce temporary files, register pressure, and occupancy. AMD ROCm(tm) installation AMD ROCm(tm) installation Installation of the AMD ROCm(tm) software package can be challenging. This introductory material shows how to install ROCm on a workstation with an AMD GPU card that supports the AMD GFX9 architecture. Introducing AMD lab notes - new programming tutorials for HPC and ML Introducing AMD lab notes - new programming tutorials for HPC and ML In this blog series, we share the lessons learned from tuning a wide range of scientific applications, libraries, and frameworks for AMD GPUs. AMD GPUOpen Privacy Trademarks Supply Chain Transparency Terms & Conditions Cookie Policy Cookie Settings (c)2025 Advanced Micro Devices, Inc.