Run any AI model in the cloud with RunPod templates — 14 custom AI templates, no local GPU required
Reference Guide

Run Any AI Model in the Cloud: My Complete RunPod Template Collection

August 2026 · 9 min read · Cloud GPU · ComfyUI · WAN 2.2 · LTX 2.3 · Flux · LoRA Training

This one's a little different. Instead of a single workflow or installer, this is the master list — every RunPod template I've built and maintain, all in one place, with the tutorial video for each and the GPU I'd actually rent to run it.

If you've ever wanted to try WAN 2.2, LTX 2.3, Flux, Qwen, Hunyuan 3D, or LoRA training but your GPU couldn't handle it, this is the page to bookmark. I keep it updated as I add templates, so it should stay the go-to reference rather than something that goes stale in a month.

Disclosure: the template links below are my referral links. If you sign up through one, I earn RunPod credits at no cost to you — and you get the same signup bonus you'd get otherwise. That's it. I've been building on RunPod for years and would recommend it regardless, but you should know how the links work. See our affiliate disclosure for more.

The signup bonus (and how to actually get it)

RunPod gives both of us bonus credits when you sign up through a referral link and load your first $10.

Being straight with you about that number: the random bonus is weighted, and most people land near the bottom of the range. Treat it as a credit match with a small chance of something much better, not a lottery ticket. Even $5 is a few hours on a 4090.

Two things that will cost you the bonus if you miss them:
  1. Sign up with Google SSO. RunPod's referral program requires it. If you create the account with an email and password instead, the referral doesn't register — no bonus for you, no credit for me. This is the single most common way people miss out.
  2. Credits expire after 90 days and are non-transferable. Don't claim them and sit on them.

The bonus unlocks once you've loaded $10. Use any template link below to sign up.

What cloud GPU actually solves

I build low VRAM workflows for a reason — most people don't have a 4090 sitting in their case. But there's a ceiling to what quantization can fix, and if you've hit it, cloud GPU is the answer:

Getting started — full setup walkthrough

Clicking any template link takes you straight to the deploy page with my template preloaded. Here's the whole flow.

Step 1 — Pick your GPU

You'll see the full GPU list with live hourly pricing. A few worth knowing (prices as of July 2026):

RunPod GPU selection screen showing hourly pricing and availability for each card
The GPU picker. Note the RTX PRO 4000 at $0.57/hr — cheaper than the 4090 sitting right next to it.
GPUVRAMApprox. priceBest for
RTX 4000 Ada20 GB$0.28/hrImage generation — hard to beat on value
RTX PRO 400032 GB$0.57/hrMost ComfyUI templates — more VRAM than a 4090 for less
RTX 409024 GB$0.69/hrThe reliable default for video and 3D
RTX 509032 GB$0.99/hrLoRA training and the heaviest LLMs
L40S48 GB$0.99/hrLonger video clips and higher resolutions

The RTX PRO 4000 is the one to watch — 32 GB at $0.57/hr gives you more VRAM than a 4090 at a lower price. For most of my ComfyUI templates that's the best value on the board.

Watch the availability tag on each card (High / Medium / Low). A cheap GPU with Low availability may not actually deploy.

Step 2 — Confirm the template

The pod template should already show my template. If it shows something else, hit Change template and search for it. Leave GPU count at 1 and On-Demand selected.

RunPod configure deployment screen with the Local Lab LTXV template preloaded
The template is already filled in — here it is the Local Lab LTXV Video Generation pod. Leave GPU count at 1 and On-Demand selected.

Step 3 — Set your storage

This is the step people get wrong, so read this bit carefully.

Container disk is temporary — it's wiped when the pod stops. 200 GB is plenty for any of my templates.

You'll see a red warning: "No volume disk or network volume configured. ALL data will be lost on Pod restart." That's not a bug, it's the default. If you're doing a one-off session, it's fine — just download your outputs before you stop the pod.

If you're coming back regularly, add a network volume instead. At $0.07/GB/month under 1 TB (dropping to $0.05/GB/month above it) it's cheap, it survives pod termination, it mounts to any pod in that region, and your models, LoRAs and outputs persist between sessions. Note this is not the same as the volume disk next to it — that one is tied to the pod's lifecycle and is deleted with it. Skipping this is the single biggest source of wasted money on RunPod — people re-download 40 GB of models every session and pay GPU rates while they wait.

Then hit Deploy On-Demand.

RunPod storage configuration showing container disk, network volume pricing and the data loss warning
Storage configuration. The red warning in the pod summary is the default state, not an error — but it does mean everything is wiped on restart unless you add a volume.

Step 4 — Connect

Once the pod boots, open the Connect tab. You'll see several ports:

Wait for the status to read Ready before clicking. If it says Initializing, the service is still starting — give it a minute rather than assuming it's broken.

RunPod Connect tab listing ports 7860, 8188 and 8888 with ready status
The Connect tab. Port 8188 is ComfyUI. Note port 7860 still shows Initializing here while the others are Ready — that is normal, just wait.

Step 5 — You're in

ComfyUI opens in your browser at the pod's proxy address. Everything's preinstalled. Load a workflow and go.

ComfyUI running in a browser at a RunPod proxy address with the template browser open
ComfyUI running at the pod proxy address, fully preinstalled. Load a workflow and start generating.
When you're done, stop the pod. An idle pod bills exactly like a working one. This is the mistake everyone makes once.

All my RunPod templates

Video generation

WAN Video 2.2

The workhorse for local-quality AI video without local hardware. Text to video and image to video.

RTX 4090 24GB or RTX PRO 4000 32GB · L40S 48GB for longer clips

WAN SCAIL 2

Character animation and motion transfer — drive a character with a reference video.

RTX 4090 24GB or higher

LTXV Video Generation

LTX 2.3 in ComfyUI. Fast video generation with native audio and outpainting support.

RTX 4090 24GB or RTX PRO 4000 32GB

WAN2GP

A simpler UI for video generation, including LTX-2 audio to video and lip sync work.

RTX 4090 24GB

Image generation

Krea 2 Image Generation

Krea 2 Turbo in ComfyUI. One of the strongest open weight image models available right now.

RTX 4000 Ada 20GB or RTX 4090 24GB

Flux Klein 9B ComfyUI

Flux.2 Klein 9B, including the face swap workflows.

RTX 4090 24GB

Z-Image Turbo ComfyUI

Fast image generation and face swapping with Z-Image Turbo.

RTX 4000 Ada 20GB or RTX 4090 24GB

Qwen-Image

Qwen's image generator — fast, realistic, strong prompt adherence.

RTX 4090 24GB

3D model generation

Trellis 2 ComfyUI

Microsoft's Trellis 2 — image to 3D model, and the best free 3D generator I've tested.

RTX 4090 24GB

Hunyuan 3D 2.1 ComfyUI

Next generation image to 3D with full PBR texture generation.

RTX 4090 24GB

LoRA & model training

AI Toolkit

Train LoRAs for image and video models through a clean UI. This is what I use for consistent character work.

RTX 5090 32GB

FluxGym

Flux LoRA training with a Kohya backend. No local VRAM required.

RTX 4090 24GB

Audio & text to speech

VibeVoice TTS

Microsoft's realistic TTS with voice cloning. Genuinely competitive with the paid services.

RTX 4000 Ada 20GB

LLMs

Ollama & OpenWebUI

Run any open weight LLM with a ChatGPT-style interface. Great for models too large for local hardware.

RTX 5090 32GB

Tips that will save you money

Prefer to run it locally?

Cloud GPU is the answer when your hardware genuinely can't keep up — but plenty of these models run fine on a consumer card with the right setup. Our free ComfyUI workflows are tuned for low VRAM, and the one-click Windows installers handle the models, nodes and CUDA versions for you.

FAQ

Do I need my own GPU to use these templates? No. Everything runs in your browser on a rented cloud GPU — Mac, laptop and integrated graphics all work fine.

How much does it cost to try a model? An RTX 4000 Ada runs about $0.28/hr and an RTX PRO 4000 about $0.57/hr, so testing a workflow for an hour costs well under a dollar.

Will I lose my models when the pod stops? Yes, unless you attach a network volume. Container disk is wiped on restart — the volume at $0.07/GB/month is what makes your models and outputs persist.

Which GPU should I pick? The RTX 4090 is the reliable default. The RTX PRO 4000 gives you 32 GB for less, and the RTX 4000 Ada is the value pick for image generation.