Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions Cargo.lock

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

2 changes: 2 additions & 0 deletions Cargo.toml
Original file line number Diff line number Diff line change
Expand Up @@ -27,6 +27,7 @@ training = [
"dep:rand_chacha",
"dep:serde_yaml",
"dep:sha2",
"dep:cubecl",
]
# Experimental CPU configurations retained for backend measurements; neither changes the default path.
# burn-flex's default features already include simd (macerator/gemm) and rayon; x86-v4 adds the
Expand All @@ -40,6 +41,7 @@ burn-cpu = { version = "=0.21.0-pre.4", optional = true }
burn-flex = "=0.21.0-pre.4"
burn-store = { version = "=0.21.0-pre.4", default-features = false, features = ["std", "pytorch", "burnpack", "memmap"], optional = true }
burn-wgpu = { version = "=0.21.0-pre.4", optional = true }
cubecl = { version = "=0.10.0-pre.4", default-features = false, optional = true }
ciborium = { version = "0.2", optional = true }
clap = { version = "4", features = ["derive"] }
fast_image_resize = { version = "=6.0.0", features = ["image"] }
Expand Down
5 changes: 5 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -79,6 +79,9 @@ montgomery train --architecture yolo26n --data dataset.yaml --epochs 100
# Pretrained initialization
montgomery train --model yolo26n.bpk --data dataset.yaml --epochs 100

# Automatically choose a hardware-specific batch size
montgomery train --model yolo26n.bpk --data dataset.yaml --batch -1 --epochs 100

# Exact continuation (model and dataset come from the training checkpoint)
montgomery train --resume runs/train/checkpoints/last
```
Expand All @@ -96,6 +99,8 @@ Every run contains:
Only the best and latest resumable models are retained.
Use `--save-period` to control recovery checkpoints and `--workers` to override automatic CPU
worker selection.
`--batch -1` runs isolated WGPU optimizer-step probes, finds the largest fitting microbatch up to
the training-set size (capped at 1024), and uses 80% of that verified maximum for runtime headroom.

## Export ONNX

Expand Down
Loading
Loading