Skip to content

[aat] Cache resolved morx subtable views - #475

Closed
behdad wants to merge 1 commit into
mainfrom
perf/cache-morx-views
Closed

behdad wants to merge 1 commit into
mainfrom
perf/cache-morx-views

Conversation

@behdad

@behdad behdad commented Sep 6, 2026

Copy link
Copy Markdown
Member

Summary

  • resolve read-fonts morx subtable views once while preparing AAT tables
  • reuse the borrowed views during shaping instead of reconstructing them after every subtable screen
  • retain the existing reconstruction path for malformed or unavailable cached views
  • keep contextual views covariant so the cache does not restrict hb_font_t lifetimes

No font-table arrays are decoded or copied. On x86_64, the cache is 112 bytes per morx subtable: about 4.4 KiB for Devanagari Sangam MN. The stripped library grows by about 11 KiB.

This is substantially smaller than the previously reverted compiled-state-machine cache, which used 88--152 KiB for the tested fonts.

Performance

Measured shaping-time changes against current main, with the same optimized Fontations dependency:

Workload Change
Devanagari Sangam MN, Hindi words about -4.5%
Geeza Pro, Persian words about -2.7%
Lucida Grande, English text about -3.1%
Menlo, English text about -2.3%

For Devanagari Sangam MN, hardware counters also showed about 3.9% fewer instructions, 5.5% fewer branches, and 6.4% fewer branch misses.

Testing

  • cargo fmt --all -- --check
  • cargo build --no-default-features --features=libm
  • cargo clippy --all-features --all-targets -- -D warnings
  • cargo test --workspace

The lifetime-free morx cache still reconstructs each read-fonts subtable
view after screening on every shape. Resolve those borrowed views once
when preparing the AAT tables and reuse them during shaping.

Keep contextual subtables covariant by storing their state table and
lookup-array data separately. This avoids making hb_font_t invariant and
does not copy any font-table arrays.

The cache costs 112 bytes per morx subtable on x86_64 (about 4.4 KiB for
Devanagari Sangam MN). It reduces measured shaping time by roughly 2--4.5%
across Devanagari Sangam MN, Geeza Pro, Lucida Grande, and Menlo.

Assisted-by: OpenAI Codex <noreply@openai.com>
@behdad
behdad requested a review from dfrg September 6, 2026 22:31
@dfrg

dfrg commented Sep 7, 2026

Copy link
Copy Markdown
Collaborator

Would be nice if this was viable but any cache that holds a lifetime won’t persist across shaping calls for any real users.

@dfrg dfrg closed this Sep 7, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants