Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
31 changes: 17 additions & 14 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,7 +9,7 @@ For a new family, scale, or task, follow [docs/MODEL_BRINGUP.md](docs/MODEL_BRIN

## Repository map

- `src/models/yolox/`: stable YOLOX nano/tiny/s/m/l/x implementation and official `.pth` import.
- `src/models/yolox/`: stable YOLOX nano/tiny/s/m/l/x implementation and tensor-state import.
- `src/models/yolov3_tiny/`, `yolov8/`, `yolov10/`, `yolo11/`, `yolo12/`, `yolo26/`:
experimental Ultralytics-family graphs and native Burnpack loaders.
- YOLOv8, YOLO11, and YOLO26 also provide `-seg` variants; YOLOv8, YOLO11, and YOLO26 provide
Expand All @@ -32,30 +32,37 @@ user-facing text.
Run Python tools from the repository root with the tools project selected:

```console
uv run --project tools tools/export_ultralytics_state.py yolo26n.pt target/yolo26n-state.pt
uv run --project tools tools/export_checkpoint_state.py yolo26n.pt target/yolo26n-state.pt
```

Stable YOLOX and Ultralytics-family models both run from native Burnpacks:

```console
montgomery pack-weights --model yolox-nano --input target/yolox_nano.pth
montgomery predict --model yolox-nano --source docs/dog_bike_man.jpg
uv run --project tools tools/export_checkpoint_state.py target/yolox_nano.pth target/yolox-nano-state.pt
montgomery pack-weights --architecture yolox-nano --state target/yolox-nano-state.pt
montgomery predict --model yolox-nano.bpk --source docs/dog_bike_man.jpg

montgomery pack-weights --model yolo26n --input target/yolo26n-state.pt
montgomery predict --model yolo26n --source docs/dog_bike_man.jpg
montgomery pack-weights --architecture yolo26n --state target/yolo26n-state.pt
montgomery predict --model yolo26n.bpk --source docs/dog_bike_man.jpg
```

Python/PyTorch is conversion- and development-time only; normal inference is Rust/Burn.

## Verification

Run before handing off changes:
CI installs the current stable Rust toolchain on every run. Before handing off changes, update the
local stable toolchain (`rustup update stable`), confirm `rustc --version` matches current stable,
and run the exact CI sequence below. Do not substitute `cargo check` for Clippy, filter the training
tests, or omit `cargo build`:

```console
cargo fmt --check
cargo build
cargo test
cargo clippy --all-targets -- -D warnings
cargo check --no-default-features --lib
cargo test --features training
cargo clippy --features training --all-targets -- -D warnings
cargo clippy --no-default-features --lib -- -D warnings
```

When external checkpoints and fixtures are available:
Expand All @@ -64,12 +71,8 @@ When external checkpoints and fixtures are available:
cargo test -- --ignored
```

Training is opt-in. For training changes, run:

```console
cargo test --features training training
cargo clippy --features training --all-targets -- -D warnings
```
Training is opt-in at runtime, but its full test and Clippy commands above are mandatory before
handing off any change because Linux CI always runs them.

Real training, hardware smoke tests, and latency measurements must use `--release`. Single-image
latency tests must use `--test-threads 1` to avoid CPU contention. When touching a runtime/backend
Expand Down
36 changes: 24 additions & 12 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@
<img alt="Montgomery" src="/docs/logo.svg" width="58%">
</picture>

Native object detection, instance segmentation, and image classification in Rust with [Burn](https://burn.dev).
Native object detection, instance segmentation, and image classification in Rust with [Burn](https://burn.dev)

<h3>

Expand Down Expand Up @@ -38,22 +38,24 @@ install [uv](https://docs.astral.sh/uv/) and run:

```console
uv run --project tools python -c "from ultralytics import YOLO; YOLO('yolo26n.pt')"
uv run --project tools tools/export_ultralytics_state.py yolo26n.pt target/yolo26n-state.pt
montgomery pack-weights --model yolo26n --input target/yolo26n-state.pt
uv run --project tools tools/export_checkpoint_state.py yolo26n.pt target/yolo26n-state.pt
montgomery pack-weights --architecture yolo26n --state target/yolo26n-state.pt
```

This creates `yolo26n.bpk`. Architecture, dataset, upstream version, format version, and licensing
stay inside the artifact metadata. Run inference with the same short model name:
stay inside the artifact metadata. Run inference by naming that artifact explicitly:

```console
montgomery predict --model yolo26n --source image.jpg
montgomery predict --model yolo26n.bpk --source image.jpg
```

Useful options: `--json`, `--confidence 0.30`, `--output result.png`, `--masks`, and `--device gpu`.
Pass `--weights another-name.bpk` only when you deliberately use a nonstandard filename.
There is no implicit filename and no separate architecture flag: `predict --model` only accepts a
self-describing Montgomery `.bpk` Burnpack.

YOLOX accepts its official `.pth` directly through `pack-weights`. Other families use the
tensor-only conversion shown above. Prediction always uses a Montgomery `.bpk` Burnpack.
The conversion workflow is identical for every family, including YOLOX: pass the trusted upstream
`.pt` or `.pth` checkpoint through `export_checkpoint_state.py`, then pass its tensor-only output
through `pack-weights`. Normal prediction and training never accept upstream checkpoint formats.

## Supported models

Expand Down Expand Up @@ -88,9 +90,20 @@ fn main() -> montgomery::Result<()> {
## Train

```console
montgomery train --model yolo26n --data dataset.yaml --epochs 100
# Fresh initialization
montgomery train --architecture yolo26n --data dataset.yaml --epochs 100

# Pretrained initialization
montgomery train --model yolo26n.bpk --data dataset.yaml --epochs 100

# Exact continuation (model and dataset come from the training checkpoint)
montgomery train --resume runs/train/checkpoints/last
```

Exactly one initialization mode is required: `--architecture` means scratch, `--model` requires a
pretrained `.bpk`, and `--resume` requires a full native training checkpoint. A Burnpack initializes
a new run; it is not a resumable optimizer checkpoint.

Every run contains:

- `results.csv`, `results.svg`, and `validation.jsonl`
Expand All @@ -109,11 +122,10 @@ and longer detection/segmentation runs still need work. Full methodology and lim
## Export ONNX

```console
montgomery export-onnx --model yolo26n
montgomery export-onnx --model yolo26n.bpk
```

This reads `yolo26n.bpk` and writes `yolo26n.onnx`; use `--weights` or `--output` only to override
those names.
This reads the explicit Burnpack and writes `yolo26n.onnx`; use `--output` to select another path.

The offline exporter validates the graph with ONNX Runtime. Setup details are in
[tools/onnx/README.md](tools/onnx/README.md).
Expand Down
18 changes: 9 additions & 9 deletions docs/MODEL_BRINGUP.md
Original file line number Diff line number Diff line change
Expand Up @@ -17,9 +17,9 @@ Orientation first:
`src/models/yolo26/` (DFL-free end-to-end), or `src/models/yolov3_tiny/` (classic NMS path).
- Inference modes in the wild: one2one end-to-end heads are top-k selected + confidence filtered
(`end2end_topk_detections` in `src/lib.rs`); plain heads go through class-aware NMS.
- Non-Ultralytics families need the same ground-truth discipline with different tooling. YOLOX
imports its official `.pth` through `pack-weights` and consumes the resulting native `.bpk` at
runtime. Its
- Non-Ultralytics families need the same ground-truth discipline. Every family converts its
trusted upstream checkpoint to a tensor-only state with `tools/export_checkpoint_state.py`, then
packs that state into the native `.bpk` consumed at runtime. YOLOX's
golden fixtures come from the *official YOLOX repository sources* instead of the Ultralytics
package: `tools/export_yolox_fixtures.py` assembles a small import package from a plain YOLOX
checkout under `target/yolox-ref/`, loads the checkpoint with `strict=True`, and dumps per-stage
Expand Down Expand Up @@ -104,11 +104,11 @@ existing templates. Hard rules:
- the model preprocessing profile and appropriate shared runtime dispatch trait;
- `pack_weights` arm;
- the model-name, input-size, and packer-extension tests.
- `src/main.rs`: `--model` help text lists the new name.
- `src/main.rs`: keep `--architecture` for catalog identifiers and `--model` for `.bpk` paths.

## 5. Conversion tooling

- `tools/export_ultralytics_state.py` is model-agnostic; run it against the official `.pt`.
- `tools/export_checkpoint_state.py` is model-agnostic; run it against the official `.pt` or `.pth`.
- Copy `tools/export_yolo26_fixtures.py` to `tools/export_<id>_fixtures.py` and adjust: the hooked
body layer indices, the head index, the `preds["one2one"]` (or plain-tensor) access, and all
`<id>` filenames. It writes the source/preprocessed reference PNGs and the golden JSON under
Expand All @@ -126,8 +126,8 @@ cargo check --no-default-features --lib
Then the parity loop (checkpoint and fixtures live under `target/`):

```console
uv run --project tools tools\export_ultralytics_state.py target/<id>.pt target/<id>-state.pt
montgomery pack-weights --model <id> --input target/<id>-state.pt
uv run --project tools tools\export_checkpoint_state.py target/<id>.pt target/<id>-state.pt
montgomery pack-weights --architecture <id> --state target/<id>-state.pt
uv run --project tools tools\export_<id>_fixtures.py target/<id>.pt docs/dog_bike_man.jpg target
cargo test <id> -- --ignored
```
Expand Down Expand Up @@ -180,7 +180,7 @@ YOLO26-cls bring-up (n/s/m/l/x) is the classification template.
one place — `Yolo11SegHead` wraps `Yolo11Head`) and add the task head modules with checkpoint
key names as field names. The mask tensors join the model output (`SegmentOutput`); the decode,
NMS, and mask assembly stay in the runtime.
5. **Weights path** is unchanged: `tools/export_ultralytics_state.py` is model-agnostic (it dumps
5. **Weights path** is unchanged: `tools/export_checkpoint_state.py` is model-agnostic (it dumps
whatever `state_dict()` holds), so only the Rust key remaps need the new head rules (one per
path-segment pattern), plus `ModelId` arms, packer arms, and verified artifact bytes/SHA-256.
6. **Fixtures and parity**: extend the family's fixture exporter for the task (the seg fixture adds
Expand All @@ -190,7 +190,7 @@ YOLO26-cls bring-up (n/s/m/l/x) is the classification template.
ignored Rust test can compare per-detection mask IoU (target >= 0.95).
7. **Public API**: a new result type (`SegmentationDetection` with `InstanceMask`), a new predictor
method that does not disturb `predict()`, letterbox geometry shared with the boxes, and CLI
wiring (`--model <id>-seg`, `--masks`) that leaves detect-model behavior untouched.
wiring (`--model <id>-seg.bpk`, `--masks`) that leaves detect-model behavior untouched.

Classification (YOLO26-cls template) differs in these ways: the input is 224 px with Ultralytics'
classify transform (anti-aliased shortest-edge resize + centered crop, no letterbox), the class
Expand Down
96 changes: 69 additions & 27 deletions src/data/letterbox.rs
Original file line number Diff line number Diff line change
Expand Up @@ -2,38 +2,80 @@ use fast_image_resize as fir;
use image::{DynamicImage, ImageBuffer, Rgb, RgbImage, imageops};

fn resize_opencv_linear(source: &DynamicImage, width: u32, height: u32) -> RgbImage {
let source = source.to_rgb8();
let source_width = source.width();
let source_height = source.height();
// Borrow the RGB8 buffer when the source already is RGB8 to avoid a full-frame clone;
// otherwise convert once (same pixels `to_rgb8()` produced).
let owned;
let (source_raw, source_width, source_height) = match source.as_rgb8() {
Some(buffer) => (buffer.as_raw().as_slice(), buffer.width(), buffer.height()),
None => {
owned = source.to_rgb8();
(owned.as_raw().as_slice(), owned.width(), owned.height())
}
};
let scale_x = source_width as f64 / width as f64;
let scale_y = source_height as f64 / height as f64;

ImageBuffer::from_fn(width, height, |x, y| {
let max_x = source_width as f64 - 1.0;
let max_y = source_height as f64 - 1.0;

// Hoist the per-column and per-row sampling math out of the inner loop. The formulas are
// unchanged from the per-pixel version; only the evaluation points move.
let width = width as usize;
let height = height as usize;
let source_width = source_width as usize;
let mut x0_table = Vec::with_capacity(width);
let mut x1_table = Vec::with_capacity(width);
let mut wx_table = Vec::with_capacity(width);
for x in 0..width {
let source_x = (x as f64 + 0.5) * scale_x - 0.5;
let source_y = (y as f64 + 0.5) * scale_y - 0.5;
let x0 = source_x.floor().clamp(0.0, source_width as f64 - 1.0) as u32;
let y0 = source_y.floor().clamp(0.0, source_height as f64 - 1.0) as u32;
let x0 = source_x.floor().clamp(0.0, max_x) as usize;
let x1 = (x0 + 1).min(source_width - 1);
let y1 = (y0 + 1).min(source_height - 1);
let weight_x = source_x.clamp(0.0, source_width as f64 - 1.0) - x0 as f64;
let weight_y = source_y.clamp(0.0, source_height as f64 - 1.0) - y0 as f64;
let top_left = source.get_pixel(x0, y0).0;
let top_right = source.get_pixel(x1, y0).0;
let bottom_left = source.get_pixel(x0, y1).0;
let bottom_right = source.get_pixel(x1, y1).0;
let mut output = [0_u8; 3];

for channel in 0..3 {
let top =
top_left[channel] as f64 * (1.0 - weight_x) + top_right[channel] as f64 * weight_x;
let bottom = bottom_left[channel] as f64 * (1.0 - weight_x)
+ bottom_right[channel] as f64 * weight_x;
output[channel] = (top * (1.0 - weight_y) + bottom * weight_y)
.round()
.clamp(0.0, 255.0) as u8;
x0_table.push(x0);
x1_table.push(x1);
wx_table.push(source_x.clamp(0.0, max_x) - x0 as f64);
}
let source_height_usize = source_height as usize;
let mut y0_table = Vec::with_capacity(height);
let mut y1_table = Vec::with_capacity(height);
let mut wy_table = Vec::with_capacity(height);
for y in 0..height {
let source_y = (y as f64 + 0.5) * scale_y - 0.5;
let y0 = source_y.floor().clamp(0.0, max_y) as usize;
let y1 = (y0 + 1).min(source_height_usize - 1);
y0_table.push(y0);
y1_table.push(y1);
wy_table.push(source_y.clamp(0.0, max_y) - y0 as f64);
}

let mut output = vec![0u8; width * height * 3];
for y in 0..height {
let y0 = y0_table[y] * source_width;
let y1 = y1_table[y] * source_width;
let weight_y = wy_table[y];
let inv_weight_y = 1.0 - weight_y;
let row = y * width * 3;
for x in 0..width {
let x0 = x0_table[x];
let x1 = x1_table[x];
let weight_x = wx_table[x];
let inv_weight_x = 1.0 - weight_x;
let tl = y0 + x0;
let tr = y0 + x1;
let bl = y1 + x0;
let br = y1 + x1;
let out = row + x * 3;
for channel in 0..3 {
let top = source_raw[tl * 3 + channel] as f64 * inv_weight_x
+ source_raw[tr * 3 + channel] as f64 * weight_x;
let bottom = source_raw[bl * 3 + channel] as f64 * inv_weight_x
+ source_raw[br * 3 + channel] as f64 * weight_x;
output[out + channel] = (top * inv_weight_y + bottom * weight_y)
.round()
.clamp(0.0, 255.0) as u8;
}
}
Rgb(output)
})
}
ImageBuffer::from_raw(width as u32, height as u32, output)
.expect("pre-sized resize buffer matches dimensions")
}

/// Resize an RGB8 image with the `fast_image_resize` crate (runtime-dispatched SIMD kernels).
Expand Down
Loading
Loading