Skip to content

Optimize current RF chroma shift pass - #202

Merged
JunliangRen merged 1 commit into
mainfrom
codex/perf-phase67-wavefront
Aug 26, 2026
Merged

JunliangRen merged 1 commit into
mainfrom
codex/perf-phase67-wavefront

Conversation

@JunliangRen

Copy link
Copy Markdown
Owner

Summary

  • fuse the first float32 quantization into the current RF chroma roll/DC pass and vectorize the final centering conversion
  • add an independent frozen oracle with AVX-required and AVX-disabled CI gates, including heap-wrap and IEEE special-value coverage
  • refresh the English, Chinese, and Japanese performance tables and detailed evidence from one 60-run matrix

Compatibility and performance

  • all 60 Exact/IPP-fast matrix runs were deterministic within and across default/1/5/10/20 worker settings
  • luma, chroma, raw JSON, stdout, normalized stderr/logs, ordered fileLoc, and field count remained identical
  • the 32,768-sample kernel improved from 28.025 ms to 16.166 ms over 1,000 calls (42.317%)
  • three 1,000-frame Exact current --threads 20 A/B pairs retained all compatibility surfaces; paired median wall/CPU gains were 0.5385%/0.6860%

Validation

  • Release build: 0 warnings, 0 errors
  • xUnit v3: 1,615 discovered, 1,612 passed, 3 expected environment skips
  • AVX-required focused test: 1/1 passed
  • AVX-disabled chroma reduction tests: 4/4 passed
  • localized README tests: 2/2 passed
  • local independent review findings addressed

@JunliangRen
JunliangRen marked this pull request as ready for review August 26, 2026 12:27
@JunliangRen
JunliangRen merged commit d09e5d8 into main Aug 26, 2026
2 checks passed
@JunliangRen
JunliangRen deleted the codex/perf-phase67-wavefront branch August 26, 2026 12:36
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant