Skip to content

docs: release v1.21.0 update, SEO optimization, and FAQ section - #1

Merged
sebishogun merged 13 commits into
mainfrom
feature/v1.21.0-release-and-seo
Aug 22, 2026
Merged

sebishogun merged 13 commits into
mainfrom
feature/v1.21.0-release-and-seo

Conversation

@sebishogun

Copy link
Copy Markdown
Owner

Summary

  • Version Bump: Update current shipped release version to v1.21.0 in README.md, CHANGELOG.md, AGENTS.md, CLAUDE.md, and docs/assets/badges/release.svg.
  • Search Optimization (SEO): Optimize H1 title and subtitle in README.md with high-intent search keywords (Go SIMD, Vector Primitives, AVX-512, NEON, SVE2, RVV).
  • Structured FAQ: Add Frequently Asked Questions section addressing key developer queries (WASM, GOEXPERIMENT=simd, minimum slice sizes).
  • Verification: All doc verification tests (go test ./internal/tests/docs) and make verify pass 100% green.

sebishogun and others added 13 commits August 14, 2026 23:35
…embly

complexOps returned the portable reference set directly, and the generator
emitted no complex entries into the per-tier dispatch tables. So DotComplex,
DotComplexConj, ScaleComplex, SumComplex and the complex arithmetic set ran
ordinary Go on every machine, while nine tiers of generated complex assembly
-- sse2, avx2, avx512, neon, sve2, rvv, vsx, vx, lasx, across six
architectures -- sat linked into a test-only aggregator and were never called.

Measured on amd64/avx512, scalar against the tier in one binary: 1.9x to
12.9x depending on operation and length. Results are unchanged, and that is
checkable rather than asserted: the numerical contract is the same fixed
accumulation order, and the reference and every tier agree bit for bit.

The cache had to grow a second shape. opsCache is generic over the ELEMENT
type and kernel.Complex and kernel.ComplexParts are not kernel.Ops, so
groupCache is generic over the struct itself -- same lazy per-tier merge, one
type parameter further out.

The generator now REFUSES to emit a dispatch file for a kernel group it
cannot route. That is the part worth keeping: this defect was invisible
because everything downstream of it worked. The tests passed, the assembly
assembled, the emission check counted the kernels as emitted, and the only
symptom was that the fast path was never taken -- which nothing measured.
A gate that asserts "the group is in the table" would not have caught it
either, since the group was in no table; the assertion is on the generator's
own completeness.

docs/wrong.md carries the finding, the measurement, and the n-ary closure
research: a general CombineInto taking a Go closure is measured, designed and
NOT implemented, because a closure call per element defeats the vectorization
that is the reason to call it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The ROADMAP lists a three-way partition as the fix for Sort's one losing
case. extractEqual in sort.go already does that split, on shipped kernels
(EqualScalarInto, CountTrue, NotMask, CompressInto), so the item names
work that exists.

The 34% loss it cites could not be re-measured on this machine: at load
2.2-2.5 the minimum of three 300-iteration runs put simd at 28,531 ns
against 23,000 ns for slices.Sort, a 1.24x loss rather than 1.34x, and at
load 12.0 the numbers were unusable. No ROADMAP number is changed --
re-measuring needs a quiet machine.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…lane

Probed: `Sqrt: sqrt[T]` deleted from floatOps in internal/ref/ref.go — one
entry, the shape of a wiring somebody forgets when adding a kernel.

  TestDeclaredKernelsAreWired, normal lane   PASS (confirmed it ran, with -v)
  TestDeclaredKernelsAreWired, purego lane   PASS
  whole suite, normal — all 14 packages      PASS
  whole suite, purego                        FAIL, TestInPlaceMatchesInto

The completeness gate that exists for exactly this did not see it, and the
only thing that did was a functional test in a lane nobody runs by default.

The cause is scope, and the older test's own doc comment says it: it walks
backend.Inventory against the dispatch tables and archSets() — the GENERATED
side — and deliberately not ref.Set(), because ref leaves every Fast slot nil
until a backend registers. That reasoning is correct about Fast slots and was
taken to cover the whole reference.

The reference is not optional. It is the live fallback on architectures with
no backend, in purego builds, and below every kernel's element threshold, so a
nil there is a nil call in the small-n path of an operation that works fine on
a big slice.

TestDeclaredKernelsAreWiredInTheReference is the other half: every non-Fast
entry in backend.Inventory must be non-nil in refBase. 853 today. Probed with
the same deletion — red in the normal lane now.

Also settled, and it argues against a planned change: the plan's item 1
proposes generating the internal/ref ops entry and the kernel.Ops struct
field, justified as removing "the two easiest things to forget". Half of that
is wrong — a forgotten kernel.Ops FIELD cannot ship, because the generator
emits code referencing it and the build fails. Only the ref wiring could, and
a gate fixes that without moving the numerical contract out of the
kernel.Ops field comments and into a manifest.

README's wrong.md count moved 81 -> 82; the count gate caught it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A document naming a test claims a guarantee is enforced, and it rots like a
link: the test is renamed, the sentence goes on naming the old one, and a
reader who greps finds nothing and cannot tell whether the guarantee was
dropped or the name moved.

Swept the whole tree first — 46 documents, 617 declared tests — and only two
citations named nothing:

- CHANGELOG.md's `TestReadmeTableHasExamples`, renamed to
  `TestOperationCatalogHasExamples`. The entry is left as written, with the
  current name added in parentheses: a changelog entry is accurate about the
  release it describes, and rewriting it to match today is what docs/wrong.md
  exists to prevent.
- docs/research's `BenchmarkNoopAsm`, which is mmcloughlin's benchmark quoted
  from a blog post, not this repository's.

Both are in categories the gate excludes for those reasons, so nothing in the
gated set was wrong — the sweep is the finding.

The subject is activeDocs plus docs/verification.md and
internal/tests/README.md, the two documents that describe the suite itself and
carry most citations. Kept as its own list rather than by widening activeDocs,
because CLAUDE.md describes that set as what the LINK gate covers and says the
LLDs, plans and records are checked by hand — which they were, on 2026-08-16,
and all resolve.

Five citations today. The gate fails loudly if either side comes back empty,
and was probed by adding a fabricated name to docs/verification.md.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The recipe named two and the comment said "Two fuzz targets". There are three,
and only ONE was being fuzzed:

  FuzzKernelsAgainstReference        fuzzed, correctly
  FuzzDifferential                   named, and run in the ROOT package while
                                     the target lives in internal/tests/arrays
  FuzzJSONMasksMatchesSeparateCalls  never named; only its seed corpus ever ran

The middle one is why this is an entry rather than a tidy-up:

  $ go test -run '^$' -fuzz FuzzDifferential -fuzztime 3s .
  PASS
  ok  github.com/sebishogun/simd  0.002s
  EXIT=0

`go test -fuzz X` in a package with no target called X exits 0 and prints PASS.
No error, no warning. Half of `make fuzz` had been reporting success for a step
that fuzzed nothing, and the only tell was two milliseconds where sixty seconds
were asked for.

Targets are discovered per package now, so a target can only be fuzzed where
`go test -list` found it and the wrong-package shape is unreachable rather than
fixed. The recipe fails outright if it runs nothing at all, because a discovery
loop over an empty list is the same silent green one level up.

Found by sweeping the family for fuzz targets nothing runs. simdlogs already
discovered its targets and its workflow says why; simdhttp hand-listed three of
four and is fixed in its own tree; simdcsv (3 targets) and simdparquet (13)
have no fuzz recipe at all, and simdjson names 11 of 14 -- recorded in
docs/wrong.md, not yet fixed.

docs/verification.md's table said `make fuzz` runs two named targets; it now
says every target in the tree, and names the three there are today. README's
wrong.md count 82 -> 83.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…n it

The plan for internal/fastpath asserts that Go's own simd/archsimd beats the
portable loop where the dispatch guards fall back, from a table of AddInto
timings. Two things have to be true before any of that is worth writing, and
both are measurable today -- Go 1.26.5 ships simd/archsimd behind
GOEXPERIMENT=simd -- so they were measured first.

BIT-IDENTITY, which the plan makes a hard gate: identical at every index over
n = 0, 1, 3, 7, 8, 9, 15, 16, 17, 31, 32, 33, with cancellation-prone values
(i*0.1 + 1e-8 against i*-0.3 + 1e30), covering every tail length the
partial-load helpers handle.

THE WIN, instructions retired and cycles, perf stat, three interleaved rounds,
minimum -- instructions rather than wall clock because the load average was
above 1 and this repo's wall-clock noise floor is 8.3%:

   n    scalar instr    archsimd instr           cycles
   8     227,775,443     183,973,765   -19.2%    -13.0%
  16     387,771,780     284,004,004   -26.8%    -21.1%
  32     708,189,014     483,774,777   -31.7%    -37.5%

Every range disjoint: the slowest archsimd round beats the fastest scalar round
at all three sizes.

What it does NOT measure is written down beside it, because that is the
tempting number to quote: it is a DIRECT call and not the guarded one, so the
end-to-end figure will be smaller; it is ONE elementwise operation, the
friendliest possible case, and the reductions are where bit-identity is
genuinely at risk because kernel.CombineTree's sixteen-accumulator shape would
have to be reproduced exactly rather than merely correctly; archsimd is amd64
and arm64 only, so four architectures gain nothing; and the build tag is viral,
so every lane runs twice or the tagged path becomes the vacuously-green lane
wrong.md entry 41 warned about.

docs/research/08-goexperiment-simd-small-n.md. Nothing shipped; the open item
stays in ROADMAP.md.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…not inline

docs/research/08 measured GOEXPERIMENT=simd intrinsics at -19.2% (n=8), -26.8%
(16) and -31.7% (32) instructions against the portable loop, and the plan put
the implementation in an internal/fastpath package that every generated guard
would call below its threshold.

Built it, gated it on bit-identity, and measured. perf stat -e instructions:u,
2,000,000 iterations, interleaved, disjoint ranges:

    f32 n=8    124.06M ref   155.85M archsimd   +25.6%
    f32 n=9    136.12M       288.08M            +111.6%
    f32 n=16   220.20M       228.13M            +3.6%
    f32 n=17   232.14M       360.57M            +55.3%
    f32 n=32   412.03M       371.91M            -9.7%
    f64 n=8    124.17M       228.07M            +83.7%
    f64 n=16   220.05M       372.34M            +69.2%

It loses across the whole band the guards use and wins only at f32 n>=32, above
where any guard falls back.

The cause is inlining. `-gcflags=-m` says "can inline AddFloat32" for the
one-line forward to ref and says nothing for the archsimd loop -- Go does not
inline a function containing one. So the portable fallback costs the guard
nothing to reach and the intrinsics version is a real call, which at n=8 is most
of the work. The research document named that risk and did not measure it: "the
end-to-end figure will be smaller than 19-32%". Smaller was the wrong word.

Its table also measured only exact multiples of the vector width, which hid the
tail entirely: n=9 and n=17 -- one vector plus a one-element tail, the shape the
band mostly sees -- cost +111.6% and +55.3%. And float64 loses everywhere for a
structural reason the f32-only table could not show: 4 lanes means the call is
amortised over half as much arithmetic at the same length.

Reverted. The next attempt has to emit the intrinsics INSIDE the generated
backend, in the same package as the guard where they can inline into it, not
behind a package boundary.

The bit-identity half stands and was re-verified: elementwise archsimd is
bit-identical to ref at every length 0-40 over IEEE specials -- both zeros, both
infinities, NaN, the subnormal and finite extremes -- for add, sub, mul and div,
on both lanes, compared on Float32bits/Float64bits rather than ==.

docs/wrong.md entry 75; README count 83 -> 84.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…see scalar

Measured before writing any pragma. 336 (kernel, tier) refusals over 132
kernels, not 346; the permute/scatter family is 48% of them and needs
instructions the targets lack; there are ZERO fast_* refusals, so the plan's
"114 float/fast_* -- the plausible wins" bucket does not exist. Only ten loops
in all of csrc are refused for the reason the pragma addresses, and clang's own
remark says that pragma grants REORDERING -- which the numerical contract
forbids and which the simd_fast_* prefix already exists to house.

Those same ten loops are the accumulation inside convolve, correlate and
matmul, and those kernels are refused on zero tiers while shipping scalar:
simd_convolve_f32 has zero packed arithmetic and ten scalar float ops on sse2,
avx2 and avx512 alike, and its only vector instructions on sse2 are a movups
and an xorps. verify.vectorWidth counts any %xmm/%ymm/%zmm operand as vector,
and every scalar float instruction on x86-64 uses those registers, so
RequireVector structurally cannot see a scalar float kernel on amd64. Scanning
every symbol for "does arithmetic, none of it packed": 26 of 913 on sse2, 8 of
913 on avx2.

No code change: fixing the gate drops the eight to internal/ref, and whether
that is faster is unmeasured -- the machine was not quiet.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
vectorWidth answers "does any instruction touch a vector register", and on
amd64 every scalar float instruction does -- scalar SSE arithmetic lives in
%xmm. RequireVector, whose whole purpose is to stop a kernel shipping as an
acceleration while running scalar code, therefore cannot see that case on the
one architecture that runs natively here. arm64 has the same blind spot in a
different spelling: `fmul s0, s1, s2` and `fmul v0.4s, v1.4s, v2.4s` differ in
the operand, not the register file.

arithKind classifies each instruction as packed arithmetic, scalar arithmetic,
or neither. Moves are neither, and so is `xorps %xmm0, %xmm0` -- those four
instructions are the entirety of what made simd_convolve_f32 look vectorized at
sse2. The classifier runs on amd64 and arm64 only, and ArithKnown says so, so a
zero elsewhere reads as "not measured" rather than as a finding.

check-emission now prints SCALAR-ONLY per kernel that does arithmetic with none
of it in lanes: 38 instances over 13 kernels -- convolve, correlate, movavg,
polyeval, minr, rolling_min/max_f64, dtoa_f64 -- across avx512 9, sse2 8, avx2
8, sve2 8, neon 5. Nothing is dropped: falling back to internal/ref instead is
a behaviour change that wants a benchmark, and the machine was not quiet.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… vectorize

LLVM vectorizes only the innermost loop, and in all four the innermost loop was
the reduction -- a serial float dependency it cannot break without
reassociating. So they shipped scalar on every tier: simd_convolve_f32 had zero
packed arithmetic at sse2, avx2 and avx512 alike, while passing the emission
gate on the strength of a movups and an xorps (entry 76).

Blocking over the outputs with a fixed 16-element accumulator makes the
innermost loop the independent one. Every output still accumulates in the same
order, so this is bit-identical rather than close -- which is the difference
between this and the pragma, whose whole effect is to permit the reassociation
the numerical contract forbids. Nothing is allocated: 16 doubles is 128 bytes
against a 512-byte frame budget.

All six symbols now carry packed arithmetic on both amd64 tiers, SCALAR-ONLY
drops from 13 kernels to 5, and loong64 gains two: movavg f32/f64 were refused
there with "LLVM did not vectorize it for lasx" and now emit, so the kernel
total moves 6,931 -> 6,933.

Bit-identity is checked by the conformance differential against internal/ref,
and the check was probed for being vacuous: eleven of its eighteen length pairs
are at or above the guard's threshold of 16, and poisoning the blocked path
alone reddens it at dst=16, the first blocked case. Speed is NOT measured -- the
machine was not quiet.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`go test ./...` was red at e2ab049: entry 77 was added to docs/wrong.md
without bumping README's "records 85 things", and that one line takes
`make verify`, `make test-tiers` (aborting at the first tier) and
`make test-cross` down with it.

Entry 77's numbers were wrong in three places, all from generalising one
tier's output:

  - the avx2 packed-arithmetic column was wrong in all six rows it listed.
    Re-derived twice, once against `git show c1d08c4:csrc/numeric.c` and
    once against the blocked file, so the before column is measured too.
    The table now gives all eight symbols at sse2/avx2/avx512.
  - "the ten scalar instructions that remain are the tail loop" -- the
    scalar count is 10 at sse2 and 18 at avx2/avx512, and it is identical
    before and after the change on every amd64 tier. Blocking added packed
    work beside those instructions; it did not convert them.
  - "eleven of eighteen length pairs are at or above the guard's threshold"
    -- it is eight.

Added what the record was missing: a review sweep of 3,286,542 cases per
tier across all nine tiers (cancellation at 1e16, mid-accumulation overflow,
block-boundary magnitude switches, NaN payloads) finding zero differing
bits, with a poisoned `ref` proving it non-vacuous; and the arm64 cost,
where LLVM vectorizes over the taps instead of the block and six kernels go
from a zero frame to 80-320 bytes with 7-13x the scalar float arithmetic --
bit-identical, inside the frame budget, and measured for speed nowhere.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@sebishogun
sebishogun merged commit ff8c9a9 into main Aug 22, 2026
2 checks passed
@sebishogun
sebishogun deleted the feature/v1.21.0-release-and-seo branch August 22, 2026 00:30
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant