Skip to content
BPBackpropagate
Python · PyPI

Fine-tune LLMs in 3 lines.

Headless LLM fine-tuning with smart defaults. Adapter and batch sizes chosen to fit your GPU, multi-run SLAO training to prevent catastrophic forgetting, and one-click GGUF export for Ollama. First-class Windows and CUDA support. Prefer not to write code? Use it in your browser.

Quickstart

pip install backpropagate[standard] # Train in 3 lines from backpropagate import Trainer trainer = Trainer("unsloth/Qwen2.5-7B-Instruct-bnb-4bit") trainer.train("my_data.jsonl", steps=100) trainer.export("gguf", quantization="q4_k_m") # Ready for Ollama

Multi-run SLAO

from backpropagate.multi_run import MultiRunTrainer runner = MultiRunTrainer( model="unsloth/Llama-3.2-3B-Instruct-bnb-4bit", num_runs=5, steps_per_run=100, merge_mode="slao", ) result = runner.run("my_data.jsonl")

Export to Ollama

from backpropagate.export import export_gguf, register_with_ollama result = export_gguf(model, tokenizer, "./output", quantization="q4_k_m") register_with_ollama(result.path, model_name="my-model")

Or use it in your browser

The same training, with no code. Drop in your examples, pick a model, press Start, and watch it learn. Every setting explains itself.

The Backpropagate web UI during a training run: step 40 of 60, the loss falling, GPU memory and temperature, and an estimate that says the run fits.
Take the tour

pip install "backpropagate[standard]", then backprop ui --open-browser. It runs on your computer; your data stays there.

Fine-tuning without the friction

Built for developers who want results, not configuration.

Smart defaults

Learning rate, batch size and LoRA adapter size start from values that suit your model, method and GPU. Change any of them; none is required.

VRAM-aware training

The adapter size and batch size are chosen to fit the memory free on your GPU, and an estimate says whether a run fits before it starts. Works from 8 GB cards (1B to 3B models) to 32 GB (up to 32B with QLoRA).

First-class Windows

Developed on Windows with CUDA. It handles the usual PyTorch and Unsloth pitfalls there for you. One feature, full fine-tuning with offload, needs Linux or WSL2.

Modular installation

Install only the dependencies you need.

ExtraWhat you getKey dependencies
backpropagateCore API only — minimal footprint—
[unsloth]2× faster training, 50% less VRAMunsloth
[ui]Reflex (Radix UI) web interfacereflex
[validation]Pydantic config validationpydantic, pydantic-settings
[export]GGUF export for Ollamallama-cpp-python
[monitoring]WandB + system monitoringwandb, psutil
[logging]Structured logging (2026 best practices)structlog
[security]JWT auth + secure token generationPyJWT, cryptography
[standard]unsloth + ui (recommended)unsloth, reflex
[production]unsloth + ui + validation + logging + securityproduction deployment
[full]Everythingall extras

Get started

Install

# Recommended
pip install backpropagate[standard]

# Minimal core only
pip install backpropagate

# All extras
pip install backpropagate[full]

# Requires: Python 3.10+ · CUDA GPU (8GB+ VRAM)

Basic training

from backpropagate import Trainer

# Smart defaults — no config needed
trainer = Trainer("unsloth/Qwen2.5-7B-Instruct-bnb-4bit")
trainer.train("my_data.jsonl", steps=100)
trainer.save("./my-model")

Multi-run SLAO

from backpropagate.multi_run import MultiRunTrainer

runner = MultiRunTrainer(
    model="unsloth/Llama-3.2-3B-Instruct-bnb-4bit",
    num_runs=5, steps_per_run=100,
    merge_mode="slao",
)
result = runner.run("my_data.jsonl")

Export to Ollama

from backpropagate.export import export_gguf, register_with_ollama

result = export_gguf(model, tokenizer, "./output", quantization="q4_k_m")
register_with_ollama(result.path, model_name="my-model")
# ollama run my-model

Production-ready by design

Built for CI/CD pipelines, automated workflows, and long training runs.

Headless by design

No UI required. Runs in CI/CD pipelines, SSH sessions, and automated workflows. Full Python API with structured logging. Callbacks for progress tracking and early stopping.

Multi-run SLAO

Single LoRA Continual Learning via Asymmetric Merging (arXiv:2512.23017) prevents catastrophic forgetting during extended fine-tuning campaigns via orthogonal init, asymmetric A/B handling, and time-aware scaling. Checkpoint-and-resume keeps long runs recoverable after crashes.

LoRA + QLoRA + full FT + Unsloth

LoRA, QLoRA (4-bit) and full fine-tuning: up to about 6B on a 32 GB card, and a 7B-class model with offload on Linux or WSL2. Unsloth-accelerated when installed. Export to GGUF at q2_k, q4_k_m, q8_0 or f16.

Quality scorecard

Ship Gate audit — 27/37 checked, 10 skipped (each with justification), 100% pass on every applicable item.

CategoryScoreNotes
A. Security5/8SECURITY.md, trust model, no secrets/telemetry, safe_path(), output-directory denylist; the 3 SKIPs cover destructive-action / MCP rows that do not apply.
B. Error Handling3/7Structured exception shape (code/message/hint/cause/retryable) via ERROR_CODES registry; CLI exit codes 0/1/2/3; no raw stack traces without --verbose; run_id correlation; redacted stderr; --share+--auth gating; the 4 SKIPs cover MCP / desktop / VS Code rows that do not apply.
C. Operator Docs4/7README, CHANGELOG, LICENSE, --help; the 3 SKIPs cover formal log-tier / MCP / operational-complexity rows.
D. Shipping Hygiene7/9verify.sh, version=tag, 5 scanners in CI, dependabot, npm publish with Sigstore provenance; the 2 SKIPs cover VS Code extension / desktop app rows.
E. Identity5/6Logo, translations, landing page, metadata; soft gate, does not block ship.