The trainer I would use for a FLUX LoRA today is AI Toolkit, Ostris's open-source trainer: it gets commits almost daily, the latest on 2026-09-27, and its FLUX.1 [dev] example config is written for a 24 GB card, so a 48 GB RTX A6000 trains it with room to spare. The steps below take a folder of captioned images to a LoRA file on your own computer: the Bare Metal template, AI Toolkit installed from its README, the example config with five lines changed, and a copy step before you stop.
AI Toolkit or FluxGym#
The honest answer is that both still run, but only AI Toolkit keeps up with new models: FluxGym's README lists FLUX.1 [dev], a dev2pro variant and FLUX.1 [schnell], and nothing newer.
| Trainer | What it is | Latest activity | Notes |
|---|---|---|---|
| AI Toolkit (ostris/ai-toolkit), MIT | Trainer with a YAML command line and a web UI | Commits almost daily, the latest on 2026-09-27 | Trains FLUX.1 [dev] and [schnell], FLUX.1 Kontext, FLUX.2 [dev] and FLUX.2 [klein], plus Qwen-Image and Wan |
| FluxGym (cocktailpeanut/fluxgym), MIT | Gradio web UI over kohya's sd-scripts | Last commit on 2026-07-28 | Its README still installs a CUDA 12.1 PyTorch nightly and the sd-scripts sd3 branch, which last changed on 2026-06-12 |
| kohya-ss/sd-scripts, Apache-2.0 | Command-line scripts that FluxGym runs underneath | v0.12.0 released 2026-09-24 | Trains FLUX.1 from its main branch now |
FluxGym's selling point was its low-VRAM presets for 12, 16 and 20 GB cards. Every QuantaCloud GPU has 48 GB or more, so that advantage does not apply here. Its default install also has no kernels for Blackwell cards such as the RTX PRO 6000, because it uses a CUDA 12.1 PyTorch build. The README's alternative for RTX 50-series cards, a CUDA 12.8 nightly, is the route a Blackwell card needs.
What you need#
You need a QuantaCloud account with credit, a Hugging Face account, and a folder of captioned images of your subject or style. Before anything else, open the FLUX.1 [dev] page on Hugging Face and accept its licence: the repository is gated, and AI Toolkit downloads 33.75 GB of it on the first run (our calculation from the Hugging Face file sizes).
| GPU | Memory | How the example config fits |
|---|---|---|
| RTX A6000, RTX 6000 Ada, L40, L40S | 48 GB | The example is written for 24 GB cards by loading FLUX in 8-bit, so 48 GB leaves headroom |
| A100 80GB, H100 PCIe | 80 GB | The same config, with room for larger batches |
| RTX PRO 6000 Blackwell | 96 GB | Room to try training with quantization off, against the 16-bit weights |
| GPU | Memory | From | Available now |
|---|---|---|---|
| RTX A6000 | 48 GB | $0.48/GPU-hr | Yes |
| RTX 6000 Ada | 48 GB | $0.78/GPU-hr | Yes |
| L40S | 48 GB | $1.09/GPU-hr | Yes |
| A100 SXM4 80GB | 80 GB | $1.49/GPU-hr | Yes |
| RTX PRO 6000 Blackwell | - | Not listed | No |
Disk is not a constraint: in the 2026-09-27 catalog an RTX A6000 1x came with 256 GB and an RTX PRO 6000 1x with 725 GB. LoRA, QLoRA and full fine-tuning VRAM explains why a LoRA needs so much less memory than full training.
Launch a GPU and install AI Toolkit#
The Bare Metal template gives you Ubuntu 22.04 with the NVIDIA driver and Docker, and you connect as ubuntu (SSH docs).
-
Check the driver. AI Toolkit's README installs PyTorch built for CUDA 13.0, and its own Dockerfile notes that those wheels need driver 580 or newer:
nvidia-smiThe header should show CUDA Version 13.0 or higher. Check your GPU, driver and CUDA version explains the header.
-
Install AI Toolkit in its own virtual environment, as its README shows. Ubuntu 22.04's Python 3.10 meets the README's minimum of 3.10:
sudo apt-get update && sudo apt-get install -y git python3-venv git clone https://github.com/ostris/ai-toolkit.git cd ai-toolkit python3 -m venv venv source venv/bin/activate pip3 install --no-cache-dir torch==2.13.0 torchvision==0.28.0 torchaudio==2.11.0 --index-url https://download.pytorch.org/whl/cu130 pip3 install -r requirements.txt -
Log in to Hugging Face so the gated download works, with a read token from Settings, Access Tokens on huggingface.co:
hf auth login
If you would rather click than edit YAML, AI Toolkit also has a web UI on port 8675, which its README's manager installs with python3 -m manager install and starts with python3 -m manager launch. The UI answers on every interface, so set AI_TOOLKIT_AUTH to a long random password before you start it, as the README recommends for servers, and open it through an SSH tunnel: ssh -N -L 8675:127.0.0.1:8675 ubuntu@<instance-ip>.
Prepare the dataset#
The rules come from AI Toolkit's README: a folder of images in jpg, jpeg or png, and next to each one a text file with the same name holding only its caption, so photo01.jpg gets photo01.txt. Webp has issues, so convert those first. You do not need to crop or resize anything, because AI Toolkit sorts images into size buckets and scales large ones down, and it never scales small ones up.
I write captions that describe what changes from image to image and put a trigger word where the subject appears, for example [trigger] standing in a kitchen, morning light. AI Toolkit replaces [trigger] with the trigger_word from the config, and adds the trigger word to any caption that lacks it.
Copy the folder from your own computer to the VM:
rsync -avP ./my_subject/ ubuntu@<instance-ip>:ai-toolkit/datasets/my_subject/
Set up the config#
The quickest reliable config is AI Toolkit's own FLUX example. Copy it, then change five lines:
cp config/examples/train_lora_flux_24gb.yaml config/my_flux_lora.yaml
| Key in the example | Change it to | Why |
|---|---|---|
name | my_flux_lora_v1 | The output folder and file name |
trigger_word (commented out) | Uncomment it and set your word, such as p3r5on | Replaces [trigger] in captions and sample prompts |
folder_path | /home/ubuntu/ai-toolkit/datasets/my_subject | Your dataset |
steps | 2000, or fewer for a style with many images | The example's comment gives 500 to 4,000 as a good range |
prompts under sample | Two or three prompts that use [trigger] | The example renders 10 samples every 250 steps, which adds up |
Leave the rest as the example sets it: LoRA rank 16 and alpha 16, batch size 1, learning rate 1e-4, the adamw8bit optimizer, gradient checkpointing on, bf16, EMA on, and quantize: true, which loads FLUX in 8-bit mixed precision. The example trains at 512, 768 and 1024 pixels, saves a checkpoint every 250 steps and keeps the last four. On the RTX PRO 6000 you can try quantize: false to train against the 16-bit weights.
Train and watch the memory#
Start the run in the background so a dropped SSH connection does not stop it, and follow its log:
nohup python run.py config/my_flux_lora.yaml > train.log 2>&1 &
tail -f train.log
In a second SSH session, watch the GPU:
nvidia-smi --query-gpu=memory.used,utilization.gpu --format=csv -l 5
The first minutes go to downloading FLUX.1 [dev], quantizing it and caching your images' latents, then the progress bar starts. Every 250 steps AI Toolkit saves a checkpoint and renders your sample prompts into output/my_flux_lora_v1/samples, which is how you judge when to stop. If you stop a run with Ctrl+C, wait until no save is in progress, because the README warns that interrupting a save corrupts that checkpoint. Running the same command again resumes from the last checkpoint.
Move the LoRA off the VM before you stop#
Stopping a QuantaCloud instance terminates it and deletes its disk, and there are no volumes or snapshots, so the LoRA has to leave the VM first. The output folder holds the final my_flux_lora_v1.safetensors, the last few step checkpoints such as my_flux_lora_v1_000001750.safetensors, the samples folder and an optimizer.pt you only need to resume training. From your own computer:
rsync -avP --exclude optimizer.pt ubuntu@<instance-ip>:ai-toolkit/output/my_flux_lora_v1/ ./my_flux_lora_v1/
Or push the final file to a private Hugging Face repository from the VM, in the ai-toolkit folder with the venv active:
hf upload your-username/my-flux-lora output/my_flux_lora_v1/my_flux_lora_v1.safetensors --private
Then stop the instance. Billing runs from launch to stop: the first hour is charged at launch, each further hour when the previous one is used up, and the unused seconds of the current hour are refunded when you stop. Moving files to and from a GPU server covers rsync and other tools in more depth.
Use the LoRA in ComfyUI#
The LoRA loads like any other. Put my_flux_lora_v1.safetensors in ComfyUI's models/loras folder, add a Load LoRA node (LoraLoaderModelOnly) after Load Diffusion Model in a FLUX.1 [dev] workflow, and use your trigger word in the prompt. FLUX in ComfyUI has the base workflow, the files and the settings, and loading models on a new ComfyUI instance shows how to copy a file into the container of QuantaCloud's ComfyUI template.
What the FLUX.1 [dev] licence allows#
The licence is the part people skip, and for a FLUX.1 [dev] LoRA it matters. The FLUX.1 [dev] Non-Commercial License v1.1.1 counts a fine-tuned version as a Derivative, and every restriction on the model applies to it, so your LoRA carries the model's non-commercial terms.
"Commercial" is broad: using the model or your LoRA for revenue-generating activity, in anything that interacts with or affects end users, or to train models for commercial use all need a licence from Black Forest Labs. Companies may still test and evaluate in a non-production environment. You may use the images you generate commercially, except to train a model that competes with FLUX.1 [dev]. If you share the LoRA, include a copy of the licence and an attribution notice that says you modified the model. The licence also forbids uses that violate anyone's rights of publicity or digital replica rights, which matters for a LoRA of a real person, and it asks you to filter or review outputs before you distribute them.
For a LoRA with no such limits, train on an Apache-2.0 base. AI Toolkit's train_lora_flux_schnell_24gb.yaml trains FLUX.1 [schnell] with Ostris's schnell training adapter, also Apache-2.0, which the config downloads for you. The FLUX.1 [schnell] repository is gated too, so accept its terms on Hugging Face first. AI Toolkit also supports FLUX.2 [klein] base 4B, which Black Forest Labs publishes under Apache-2.0. Treat this section as a summary, not legal advice.
FLUX LoRA questions#
Is FluxGym still maintained?
It still gets occasional commits, the last on 2026-07-28, but its install instructions have not kept up: they use a CUDA 12.1 PyTorch nightly and a sd-scripts branch that stopped changing in June 2026. AI Toolkit is the one I would install on a new VM.
How much VRAM does FLUX LoRA training need?
AI Toolkit's example config is written for 24 GB, by loading FLUX in 8-bit. Every QuantaCloud GPU has 48 GB or more, so the example runs as written, and the 96 GB RTX PRO 6000 leaves room to try training without quantization.
How many steps should I train?
The example uses 2,000 and calls 500 to 4,000 a good range. Watch the samples every 250 steps and keep the checkpoint where your subject looks right, before the samples start copying your training images.
Can I sell images made with my FLUX.1 [dev] LoRA?
The licence lets you use outputs commercially, but running the model or the LoRA for revenue, or in a product your users touch, needs a licence from Black Forest Labs. Read the licence text before you build on it.
My rule for a first FLUX LoRA: AI Toolkit's example config on a 48 GB RTX A6000, samples every 250 steps, and the LoRA copied off before you stop. Move to the RTX PRO 6000 Blackwell when you want to train without quantization, and train on FLUX.1 [schnell] or FLUX.2 [klein] 4B when the result has to be used commercially without a licence from Black Forest Labs. The RTX A6000 page lists every configuration with its price, and fine-tuning on QuantaCloud covers LLM training too.
Launch an RTX A6000 with Bare Metal