Skip to content

vello_cpu: Add fast path for pixel-aligned alpha-mask image fills - #6

Open
nicoburns wants to merge 1 commit into
mainfrom
devin/1790445884-alpha-mask-fast-path
Open

nicoburns wants to merge 1 commit into
mainfrom
devin/1790445884-alpha-mask-fast-path

Conversation

@nicoburns

Copy link
Copy Markdown
Member

Summary

Cached glyphs are drawn as image fills from the glyph atlas with the following settings:

  • Tint { mode: AlphaMask }
  • Pad extend
  • an identity-scale paint transform with an integer offset

Before this change, those fills took the generic painter path:

  1. sample RGBA into paint_buf
  2. run a separate apply_tint pass
  3. composite the buffer

That path was ~18% of single-threaded cached rendering.

This PR adds a fast path to Fine's EncodedPaint::Image arm. It is taken when all of the following hold:

  • the tint mode is AlphaMask
  • both extends are Pad
  • EncodedImageExt::integer_translation() returns Some((dx, dy)), i.e. the transform is exactly [1, 0, 0, 1] with near-integer offsets

On that path:

pad_alpha_mask(pixmap, x + dx, y + dy, width, strip_alphas, &mut mask_buf)
    // reads atlas alpha into strip-alpha layout (SIMD 4×4 transpose),
    // multiplied by the strip coverage; out-of-bounds uses Pad semantics
alpha_composite_solid(dest, tint_color, Some(mask_buf))   // or T::blend(...) for non-default blend / masks

The converted tint colour is cached across commands. Everything else still uses the existing painters.

A unit test compares pad_alpha_mask against a scalar reference, including offsets outside the image.

Part 3 of 3 of the CPU glyph-cache speedups. It conflicts with the Fine image-resolve cache PR in fine/mod.rs; whichever merges second needs a trivial rebase. On its own it speeds up single-threaded rendering. Multithreaded scaling additionally needs the resolve-cache PR, because this path still resolves the image, and so bumps the shared Arc refcount, on every command.

Measurements

Throwaway probe: 1600×1200 target, ~9.3k glyphs, 14px, atlas_cache(true), mean of 60 warm frames on a noisy VM. Render time, main → this branch:

threads main this branch
1 4.50 ms 2.73 ms
4 6.32 ms 5.74 ms (still contended, see above)

All three PRs combined, at 14px:

  • 1 thread: 7.3 → 5.5 ms total (render 3.35 → 2.56 ms)
  • 8 threads: 10.8 → 3.9 ms total (render 7.64 → 1.05 ms)

For comparison, uncached rendering at 8 threads is 8.0 ms total.

Probe output is pixel-identical to main, with the cache both on and off. cargo test -p vello_cpu -p vello_common passes.

Link to Devin session: https://dioxus.staging.devinenterprise.com/sessions/55344a6e161d48b4a3a8688fabe11d4c
Open in Devin Desktop: https://dioxus.staging.devinenterprise.com/desktop/session/55344a6e161d48b4a3a8688fabe11d4c?variant=devin-insiders
Requested by: @nicoburns

@staging-devin-ai-integration

Copy link
Copy Markdown

I'll fix CI failures and address comments from users with write access that start with 'Devin'.

  • Disable automatic comment, CI, and merge conflict monitoring

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant