TRL documentation
Training with Jobs
Training with Jobs
Hugging Face Jobs runs your training on Hugging Face GPUs. You pick the hardware for each run and pay only for the seconds it runs. With --push_to_hub, the trained model is uploaded to the Hub at the end.
In this guide, you’ll learn how to:
- Launch a TRL training script with one command
- Follow a run and see its loss curves
- Run your own script, with its hardware settings in the file
- Keep checkpoints and resume an interrupted run
- Train on several GPUs
For how Jobs works in general, including hardware and pricing, see the Jobs documentation.
Requirements
- A Hugging Face account with a positive credit balance. Jobs is pay-as-you-go: you only pay for the seconds you use.
- Logged in to the Hugging Face Hub (
hf auth login)
A first run
This command fine-tunes Qwen2-0.5B-Instruct on the Capybara chat dataset with the TRL SFT script, then pushes the model to your account:
hf jobs uv run \
--flavor a10g-small \
--timeout 30m \
--secrets HF_TOKEN \
-- \
https://raw.githubusercontent.com/huggingface/trl/refs/heads/main/trl/scripts/sft.py \
--model_name_or_path Qwen/Qwen2-0.5B-Instruct \
--dataset_name trl-lib/Capybara \
--max_steps 100 \
--output_dir Qwen2-0.5B-SFT \
--push_to_hubIt takes a few minutes and costs well under $1 on a10g-small. The options before -- are for Jobs: the hardware, a time limit and the token used to push the model. --secrets HF_TOKEN sends the token you logged in with. In Python, pass the token itself. The arguments after the script URL go to the script. hf jobs uv run runs a Python script with uv. sft.py lists its dependencies in a header at the top of the file, so Jobs installs TRL before the script starts. The Job stops when the script exits, and billing stops with it.
--max_steps 100 keeps this first run short. Remove it for the full run: three epochs, the script’s default, take a few hours on a10g-small, so raise --timeout to match, for example --timeout 3h. Jobs stops a run when it reaches its timeout, which is 30 minutes by default. A larger GPU such as a100-large finishes sooner. See Hardware for the flavors and their prices.
Follow the run
hf jobs uv run streams the logs and holds your terminal until the run ends. Ctrl+C stops only the log stream: the Job keeps running. Add --detach to get the Job ID back straight away, then:
hf jobs logs -f <job_id> # stream the logs
hf jobs ps # list your running Jobs
hf jobs inspect <job_id> # status, and the error message if it failed
hf jobs cancel <job_id> # stop itSee Manage Jobs for more.
For loss curves, TRL logs to Trackio. Add --env TRACKIO_SPACE_ID=<your-username>/trackio before -- to name a Space for the dashboard, and --report_to trackio after the script URL. Trackio creates the Space, and a bucket for the metrics, if they do not exist. Both are public by default. Other trackers that Transformers supports work too. For Weights & Biases, add --with wandb --secrets WANDB_API_KEY before -- and --report_to wandb after the script URL. --secrets WANDB_API_KEY reads the value from WANDB_API_KEY in your local environment. See environment variables and secrets for other ways to pass them.
Run your own script
Write your training code in a Python file (for example, train.py). List its dependencies at the top of the file, in a script header, the same way the TRL scripts do. Jobs installs them before the script starts:
# /// script
# dependencies = [
# "trl",
# "peft",
# ]
# ///
from datasets import load_dataset
from peft import LoraConfig
from trl import SFTConfig, SFTTrainer
dataset = load_dataset("trl-lib/Capybara", split="train")
trainer = SFTTrainer(
model="Qwen/Qwen2.5-0.5B",
train_dataset=dataset,
peft_config=LoraConfig(),
args=SFTConfig(output_dir="Qwen2.5-0.5B-SFT", max_steps=100),
)
trainer.train()
trainer.push_to_hub()Then launch it with the hf jobs CLI or the Python API:
hf jobs uv run \
--flavor a10g-small \
--timeout 30m \
--secrets HF_TOKEN \
train.pyFor a script without a header, list the dependencies at launch instead with --with, once per package: --with trl --with peft (dependencies=["trl", "peft"] in Python). The script can also be a URL, such as a GitHub raw link, a Gist or a file in a public Hub repo.
The launch settings can also live in the script. A [tool.hf-jobs] table in the header sets the hardware, timeout and secrets:
# /// script
# dependencies = [
# "trl",
# "peft",
# ]
#
# [tool.hf-jobs]
# flavor = "a10g-small"
# timeout = "30m"
# secrets = ["HF_TOKEN"]
# ///hf jobs uv run train.py then needs no options, and an option you pass still overrides the script. The script carries everything it needs to run, so you can share it, or come back to it later, without remembering which GPU and timeout it needs. hf jobs uv run --dry-run train.py shows the resolved settings and marks the values that come from the script. The table is read by the hf CLI only: run_uv_job() ignores it. See Define the launch config in the script.
Run any TRL script
The same pattern runs every TRL trainer script (SFT, DPO, GRPO, reward modeling and others) and the standalone .py examples in the Examples Index. Each declares its dependencies in a script header. An example that reads other files next to it needs its folder mounted into the Job. See Local directories. Notebook-only examples cannot be submitted this way. The script arguments are the same as when you run the script locally.
Keep checkpoints
A Job’s disk is deleted when the Job ends, including when it reaches its timeout or fails. For a long run, save checkpoints outside the Job so they survive it.
The simplest way is to push each checkpoint to the output model repo as it is saved:
hf jobs uv run \
--flavor a10g-small \
--timeout 3h \
--secrets HF_TOKEN \
-- \
https://raw.githubusercontent.com/huggingface/trl/refs/heads/main/trl/scripts/sft.py \
--model_name_or_path Qwen/Qwen2-0.5B-Instruct \
--dataset_name trl-lib/Capybara \
--output_dir Qwen2-0.5B-SFT \
--push_to_hub \
--save_steps 500 \
--hub_strategy checkpointWith --hub_strategy checkpoint, the most recent checkpoint, including the optimizer state, is kept in a last-checkpoint folder of the repo. To continue an interrupted run, mount the repo into a new Job and resume from that folder: launch the same command with --volume hf://<your-username>/Qwen2-0.5B-SFT:/previous before -- and --resume_from_checkpoint /previous/last-checkpoint after the script URL.
To keep several checkpoints outside the model repo, write them to a Storage Bucket instead. Create one with hf buckets create <your-username>/checkpoints --private, mount it into the Job and point --output_dir at it:
hf jobs uv run \
--flavor a10g-small \
--timeout 3h \
--secrets HF_TOKEN \
--volume hf://buckets/<your-username>/checkpoints:/checkpoints \
-- \
https://raw.githubusercontent.com/huggingface/trl/refs/heads/main/trl/scripts/sft.py \
--model_name_or_path Qwen/Qwen2-0.5B-Instruct \
--dataset_name trl-lib/Capybara \
--output_dir /checkpoints/Qwen2-0.5B-SFT \
--save_steps 500To continue an interrupted run, launch the same command again with --resume_from_checkpoint /checkpoints/Qwen2-0.5B-SFT/checkpoint-<step>. See Write to a bucket as you go for more and Volumes for the mount options.
Each checkpoint holds the model and optimizer state, several times the model size. Add
--save_total_limit 2to keep only the latest ones. If you also pass--push_to_hub, set--hub_model_id, or the repo is named after the last part of the output path.
Multiple GPUs
A flavor with several GPUs shortens a run. For this, use hf jobs run, which runs a command in a Docker image instead of a script, with the huggingface/trl image, which has TRL and Accelerate installed. Then start one process per GPU. This works for the TRL scripts and for your own training code.
Speed up a TRL script
The SFT run from the start of this guide, on four GPUs:
hf jobs run \
--flavor l4x4 \
--timeout 30m \
--secrets HF_TOKEN \
huggingface/trl \
-- \
trl sft --num_processes 4 \
--model_name_or_path Qwen/Qwen2-0.5B-Instruct \
--dataset_name trl-lib/Capybara \
--max_steps 100 \
--output_dir Qwen2-0.5B-SFT \
--push_to_hubThe trl CLI passes --num_processes to accelerate launch. Set it to the number of GPUs in the flavor, here four L4s. Each process takes its own batch, so the effective batch size is four times that of a single-GPU run. This run takes a few minutes. Without --max_steps, the full three epochs take roughly an hour, so raise --timeout, for example to 2h.
Your own script
You are not limited to the TRL scripts. Training code that runs with accelerate launch locally runs the same way on Jobs, as long as the image has the packages it needs. For example, save a GRPO script with your own reward function as train_grpo.py in a folder, and mount the folder into the Job. Jobs uploads the folder to a private bucket and mounts it read-only:
hf jobs run \
--flavor l4x4 \
--timeout 30m \
--secrets HF_TOKEN \
--volume ./my-project:/code \
huggingface/trl \
-- \
accelerate launch --num_processes 4 /code/train_grpo.pyThe script runs with the TRL installed in the image, so it needs no script header. GRPOTrainer generates completions with transformers by default, so this runs in a single Job with no vLLM server. For faster generation, see vLLM Integration. The next section explains how hf jobs run uses the image.
Docker Images
Jobs runs your script with uv, which installs the dependencies declared in its # /// script header into a fresh environment. The TRL your script imports therefore comes from that header, not from the image, and the hf jobs uv run examples above need no --image at all.
A Docker image with TRL preinstalled is available at huggingface/trl. Passing it to hf jobs uv run gives the job the image’s system layer, such as its CUDA toolchain, which matters for dependencies that compile against it:
hf jobs uv run \
--flavor a100-large \
--secrets HF_TOKEN \
--image huggingface/trl \
train.pyTo run the TRL that is installed in the image, use hf jobs run instead. It runs a command in the image directly, with no script header to resolve, so the image’s own TRL is what executes. The -- separates the command from the Jobs options, which is needed whenever the command itself takes options:
hf jobs run \
--flavor a100-large \
--secrets HF_TOKEN \
huggingface/trl \
-- \
trl sft --model_name_or_path Qwen/Qwen2-0.5B-Instruct --dataset_name trl-lib/Capybara --output_dir Qwen2-0.5B-SFTThe image is published under three kinds of tag:
| Tag | Contents |
|---|---|
X.Y.Z | The TRL release of that version, built when that version is released |
latest | The most recent release |
dev | The main branch, rebuilt on every merge |
Use dev to run the development version, which is useful for trying a fix that has landed on main but is not released yet:
hf jobs run \
--flavor a100-large \
--secrets HF_TOKEN \
huggingface/trl:dev \
-- \
trl sft --model_name_or_path Qwen/Qwen2-0.5B-Instruct --dataset_name trl-lib/Capybara --output_dir Qwen2-0.5B-SFTTags only select what is installed in the image, so they matter for hf jobs run. With hf jobs uv run the version comes from the script header instead, and the tag makes no difference to which TRL is imported.
Combining the two, so that uv resolves the script header while some imports still come from the image, needs extra flags. See Reuse the image’s packages and add dependencies with UV for that form and the paths it requires.
Jobs runs on a Docker image from Hugging Face Spaces or Docker Hub, so you can also specify any custom image:
hf jobs uv run \
--flavor a100-large \
--secrets HF_TOKEN \
--image <docker-image> \
train.pyUpdate on GitHubThe TRL Jobs wrapper launches the TRL scripts on Jobs with preset configurations for some models, for example
trl-jobs sft --model_name Qwen/Qwen3-0.6B --dataset_name trl-lib/Capybara.