Skip to content

Getting Started

  • Python 3.10+
  • An NVIDIA GPU with CUDA. 8 GB trains the 1B to 3B presets, 16 GB a 7B, 32 GB up to 32B with QLoRA.
  • PyTorch 2.0+
Terminal window
# Recommended: isolated venv with PATH integration (no conflicts with system Python)
pipx install "backpropagate[standard]"
# Alternative isolated installer (same shape, different tool)
uv tool install "backpropagate[standard]"
# Standard pip (if you manage your own venv)
pip install "backpropagate[standard]"

The [standard] extra is the right default for most operators — Unsloth (2× faster training) + the Reflex web UI. Drop the bracket if you only want the core Python API:

Extra What you get
backpropagate Core API only — minimal footprint
[unsloth] 2x faster training, 50% less VRAM
[ui] Reflex (Radix UI) web interface
[validation] Pydantic config validation
[export] GGUF export for Ollama
[monitoring] WandB + system monitoring
[logging] Structured logging via structlog
[security] JWT auth + token generation
[standard] unsloth + ui (recommended default)
[production] unsloth + ui + validation + logging + security
[full] Everything
from backpropagate import Trainer
trainer = Trainer("Qwen/Qwen2.5-7B-Instruct")
trainer.train("examples/quickstart.jsonl", steps=10)
trainer.export("gguf", quantization="q4_k_m") # Ready for Ollama

The repo ships a small examples/quickstart.jsonl (5 ShareGPT examples) so the snippet above runs end-to-end on a clean install. Qwen/Qwen2.5-7B-Instruct is the canonical default — what Trainer() resolves with no model argument. For your own training, see the dataset format docs.

Terminal window
backprop train --data my_data.jsonl --model Qwen/Qwen2.5-7B-Instruct --steps 100
backprop export ./output/lora --format gguf --quantization q4_k_m --ollama --ollama-name my-model

See CLI reference for every flag.

Terminal window
backprop info

This prints your Python version, PyTorch version, CUDA status, GPU name and VRAM, which optional features are installed, and current configuration defaults. Run it before your first training to confirm everything is working.

backprop ships with argcomplete integration so tab-completion works in bash / zsh / fish. Enable it by adding one line to your shell rc file:

bash (~/.bashrc):

Terminal window
eval "$(register-python-argcomplete backprop)"

zsh (~/.zshrc):

Terminal window
autoload -Uz compinit && compinit
eval "$(register-python-argcomplete backprop)"

fish (~/.config/fish/config.fish):

Terminal window
register-python-argcomplete --shell fish backprop | source

Re-open your shell. backprop tr<TAB> now completes to train; backprop export --format <TAB> completes to lora / merged / gguf. Completion is opt-in (the eval line is what activates it) — without the line, the CLI behaves exactly the same.

If you installed via pipx, argcomplete came along with the package; you only need to add the eval / source line.

Terminal window
backprop ui --port 7862

The UI opens in dark mode. The Web UI tour walks through it with screenshots. From it you can:

  • Train (Single run): pick a model preset or any model, a dataset, the method (SFT, ORPO, SimPO, KTO) and the mode (QLoRA, LoRA, or full fine-tuning). The estimate next to Start training says whether it fits your card, and Measure on this GPU replaces the estimate with a measurement of that model on your own card (a minute or two). Then watch steps, loss (raw and smoothed), time left and the GPU. Stop and save checkpoint finishes the current step and saves. Every setting is a backprop train flag, and the form opens on the same defaults.
  • Multi-run: a SLAO sweep of several short runs merged as it goes, with one progress bar for the whole sweep. Stop ends the sweep after the current run and keeps what was merged.
  • Export a run’s adapter as LoRA, merged weights or GGUF (and register it with Ollama). A run’s page has Export the model, which opens Export with that run filled in. GGUF needs llama.cpp’s converter, as on the CLI.
  • Dataset: upload a .jsonl or .json file to see its layout, its first examples and how long they are. The clean-up settings (remove repeats and empty examples, length limits, order from short to long) show what they would keep as you change them; Save a cleaned copy writes that to a new file and leaves your upload alone. Use in Single run or Use in Multi-run fills in the training form for you.
  • Runs lists past runs with their real outcome (completed, stopped, failed); Models lists your Hugging Face cache (it follows HF_HOME).

Each job runs in its own process, never inside the UI server, and only one runs at a time. Closing the UI stops it. Reloading the page reattaches to a running job, and if a job fails, the page shows its error code and the last lines of its log.

For remote access, use SSH port-forwarding rather than --share:

Terminal window
ssh -L 7860:localhost:7860 you@gpu-host
# Then on your laptop: http://localhost:7860

backprop ui --auth user:pass runs through the v1.2.0 FastAPI auth middleware and gates every HTTP route + the /_event WebSocket upgrade on the supplied credentials. backprop ui --share and backprop ui --host 0.0.0.0 without --auth exit 2 with [RUNTIME_UI_AUTH_NOT_ENFORCED] — a public URL or non-loopback bind without credentials is the v1.1.x foot-gun closed by GHSA-f65r-h4g3-3h9h. See the troubleshooting page for the full contract and the security advisory.

A successful first training run prints something like:

run_started run_id=8f3a2c1d-9e4b-4c5a-...
Trainer initialized: Qwen/Qwen2.5-7B-Instruct
LoRA: r=256, alpha=512
Batch: 2, LR: 0.0002
Degradation knobs: oom_recovery=True, unsloth_fallback=True
Training: [####################] 100% loss=0.42 steps=10
Saved to ./output/lora
run_ended run_id=8f3a2c1d-... duration_seconds=43.2

After the run completes:

output/
├── lora/
│ ├── adapter_config.json
│ ├── adapter_model.safetensors
│ └── tokenizer.json
└── checkpoints/
└── manifest.json

To verify the adapter loaded cleanly, run backprop info — it lists which optional features are detected and which model the resolved config points at. If you see the Saved to ... line and the directory tree above exists, your first run worked.

If something failed during installation or first run, head to the troubleshooting page — it’s a symptoms-first reverse index keyed by what you actually saw in stderr. For CUDA-specific issues (OOM, CUBLAS, driver mismatch), see the dedicated CUDA troubleshooting page. For the full unredacted traceback when filing a bug, re-run with BACKPROPAGATE_DEBUG=1 (or --verbose) — but review the output for any secrets before posting (Bearer tokens, sk-*, hf_*, AWS keys, password= / token= are auto-redacted in non-verbose mode, but the verbose mode disables redaction).

If you are coming from v1.1.0 / v1.1.1, see the migration page for the breaking changes (notably the --share refuse-to-start contract and the removed [observability] extra) and the behavioural fixes that may surface bugs in your callback code.

For SSH port-forwarding, host/origin allowlists, the v1.1.x GHSA advisory, and the four-layer auth defense — see the security page.