docs: release v1.21.0 update, SEO optimization, and FAQ section - #1
Merged
Merged
Conversation
…embly complexOps returned the portable reference set directly, and the generator emitted no complex entries into the per-tier dispatch tables. So DotComplex, DotComplexConj, ScaleComplex, SumComplex and the complex arithmetic set ran ordinary Go on every machine, while nine tiers of generated complex assembly -- sse2, avx2, avx512, neon, sve2, rvv, vsx, vx, lasx, across six architectures -- sat linked into a test-only aggregator and were never called. Measured on amd64/avx512, scalar against the tier in one binary: 1.9x to 12.9x depending on operation and length. Results are unchanged, and that is checkable rather than asserted: the numerical contract is the same fixed accumulation order, and the reference and every tier agree bit for bit. The cache had to grow a second shape. opsCache is generic over the ELEMENT type and kernel.Complex and kernel.ComplexParts are not kernel.Ops, so groupCache is generic over the struct itself -- same lazy per-tier merge, one type parameter further out. The generator now REFUSES to emit a dispatch file for a kernel group it cannot route. That is the part worth keeping: this defect was invisible because everything downstream of it worked. The tests passed, the assembly assembled, the emission check counted the kernels as emitted, and the only symptom was that the fast path was never taken -- which nothing measured. A gate that asserts "the group is in the table" would not have caught it either, since the group was in no table; the assertion is on the generator's own completeness. docs/wrong.md carries the finding, the measurement, and the n-ary closure research: a general CombineInto taking a Go closure is measured, designed and NOT implemented, because a closure call per element defeats the vectorization that is the reason to call it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The ROADMAP lists a three-way partition as the fix for Sort's one losing case. extractEqual in sort.go already does that split, on shipped kernels (EqualScalarInto, CountTrue, NotMask, CompressInto), so the item names work that exists. The 34% loss it cites could not be re-measured on this machine: at load 2.2-2.5 the minimum of three 300-iteration runs put simd at 28,531 ns against 23,000 ns for slices.Sort, a 1.24x loss rather than 1.34x, and at load 12.0 the numbers were unusable. No ROADMAP number is changed -- re-measuring needs a quiet machine. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…lane Probed: `Sqrt: sqrt[T]` deleted from floatOps in internal/ref/ref.go — one entry, the shape of a wiring somebody forgets when adding a kernel. TestDeclaredKernelsAreWired, normal lane PASS (confirmed it ran, with -v) TestDeclaredKernelsAreWired, purego lane PASS whole suite, normal — all 14 packages PASS whole suite, purego FAIL, TestInPlaceMatchesInto The completeness gate that exists for exactly this did not see it, and the only thing that did was a functional test in a lane nobody runs by default. The cause is scope, and the older test's own doc comment says it: it walks backend.Inventory against the dispatch tables and archSets() — the GENERATED side — and deliberately not ref.Set(), because ref leaves every Fast slot nil until a backend registers. That reasoning is correct about Fast slots and was taken to cover the whole reference. The reference is not optional. It is the live fallback on architectures with no backend, in purego builds, and below every kernel's element threshold, so a nil there is a nil call in the small-n path of an operation that works fine on a big slice. TestDeclaredKernelsAreWiredInTheReference is the other half: every non-Fast entry in backend.Inventory must be non-nil in refBase. 853 today. Probed with the same deletion — red in the normal lane now. Also settled, and it argues against a planned change: the plan's item 1 proposes generating the internal/ref ops entry and the kernel.Ops struct field, justified as removing "the two easiest things to forget". Half of that is wrong — a forgotten kernel.Ops FIELD cannot ship, because the generator emits code referencing it and the build fails. Only the ref wiring could, and a gate fixes that without moving the numerical contract out of the kernel.Ops field comments and into a manifest. README's wrong.md count moved 81 -> 82; the count gate caught it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A document naming a test claims a guarantee is enforced, and it rots like a link: the test is renamed, the sentence goes on naming the old one, and a reader who greps finds nothing and cannot tell whether the guarantee was dropped or the name moved. Swept the whole tree first — 46 documents, 617 declared tests — and only two citations named nothing: - CHANGELOG.md's `TestReadmeTableHasExamples`, renamed to `TestOperationCatalogHasExamples`. The entry is left as written, with the current name added in parentheses: a changelog entry is accurate about the release it describes, and rewriting it to match today is what docs/wrong.md exists to prevent. - docs/research's `BenchmarkNoopAsm`, which is mmcloughlin's benchmark quoted from a blog post, not this repository's. Both are in categories the gate excludes for those reasons, so nothing in the gated set was wrong — the sweep is the finding. The subject is activeDocs plus docs/verification.md and internal/tests/README.md, the two documents that describe the suite itself and carry most citations. Kept as its own list rather than by widening activeDocs, because CLAUDE.md describes that set as what the LINK gate covers and says the LLDs, plans and records are checked by hand — which they were, on 2026-08-16, and all resolve. Five citations today. The gate fails loudly if either side comes back empty, and was probed by adding a fabricated name to docs/verification.md. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The recipe named two and the comment said "Two fuzz targets". There are three,
and only ONE was being fuzzed:
FuzzKernelsAgainstReference fuzzed, correctly
FuzzDifferential named, and run in the ROOT package while
the target lives in internal/tests/arrays
FuzzJSONMasksMatchesSeparateCalls never named; only its seed corpus ever ran
The middle one is why this is an entry rather than a tidy-up:
$ go test -run '^$' -fuzz FuzzDifferential -fuzztime 3s .
PASS
ok github.com/sebishogun/simd 0.002s
EXIT=0
`go test -fuzz X` in a package with no target called X exits 0 and prints PASS.
No error, no warning. Half of `make fuzz` had been reporting success for a step
that fuzzed nothing, and the only tell was two milliseconds where sixty seconds
were asked for.
Targets are discovered per package now, so a target can only be fuzzed where
`go test -list` found it and the wrong-package shape is unreachable rather than
fixed. The recipe fails outright if it runs nothing at all, because a discovery
loop over an empty list is the same silent green one level up.
Found by sweeping the family for fuzz targets nothing runs. simdlogs already
discovered its targets and its workflow says why; simdhttp hand-listed three of
four and is fixed in its own tree; simdcsv (3 targets) and simdparquet (13)
have no fuzz recipe at all, and simdjson names 11 of 14 -- recorded in
docs/wrong.md, not yet fixed.
docs/verification.md's table said `make fuzz` runs two named targets; it now
says every target in the tree, and names the three there are today. README's
wrong.md count 82 -> 83.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…n it The plan for internal/fastpath asserts that Go's own simd/archsimd beats the portable loop where the dispatch guards fall back, from a table of AddInto timings. Two things have to be true before any of that is worth writing, and both are measurable today -- Go 1.26.5 ships simd/archsimd behind GOEXPERIMENT=simd -- so they were measured first. BIT-IDENTITY, which the plan makes a hard gate: identical at every index over n = 0, 1, 3, 7, 8, 9, 15, 16, 17, 31, 32, 33, with cancellation-prone values (i*0.1 + 1e-8 against i*-0.3 + 1e30), covering every tail length the partial-load helpers handle. THE WIN, instructions retired and cycles, perf stat, three interleaved rounds, minimum -- instructions rather than wall clock because the load average was above 1 and this repo's wall-clock noise floor is 8.3%: n scalar instr archsimd instr cycles 8 227,775,443 183,973,765 -19.2% -13.0% 16 387,771,780 284,004,004 -26.8% -21.1% 32 708,189,014 483,774,777 -31.7% -37.5% Every range disjoint: the slowest archsimd round beats the fastest scalar round at all three sizes. What it does NOT measure is written down beside it, because that is the tempting number to quote: it is a DIRECT call and not the guarded one, so the end-to-end figure will be smaller; it is ONE elementwise operation, the friendliest possible case, and the reductions are where bit-identity is genuinely at risk because kernel.CombineTree's sixteen-accumulator shape would have to be reproduced exactly rather than merely correctly; archsimd is amd64 and arm64 only, so four architectures gain nothing; and the build tag is viral, so every lane runs twice or the tagged path becomes the vacuously-green lane wrong.md entry 41 warned about. docs/research/08-goexperiment-simd-small-n.md. Nothing shipped; the open item stays in ROADMAP.md. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…not inline
docs/research/08 measured GOEXPERIMENT=simd intrinsics at -19.2% (n=8), -26.8%
(16) and -31.7% (32) instructions against the portable loop, and the plan put
the implementation in an internal/fastpath package that every generated guard
would call below its threshold.
Built it, gated it on bit-identity, and measured. perf stat -e instructions:u,
2,000,000 iterations, interleaved, disjoint ranges:
f32 n=8 124.06M ref 155.85M archsimd +25.6%
f32 n=9 136.12M 288.08M +111.6%
f32 n=16 220.20M 228.13M +3.6%
f32 n=17 232.14M 360.57M +55.3%
f32 n=32 412.03M 371.91M -9.7%
f64 n=8 124.17M 228.07M +83.7%
f64 n=16 220.05M 372.34M +69.2%
It loses across the whole band the guards use and wins only at f32 n>=32, above
where any guard falls back.
The cause is inlining. `-gcflags=-m` says "can inline AddFloat32" for the
one-line forward to ref and says nothing for the archsimd loop -- Go does not
inline a function containing one. So the portable fallback costs the guard
nothing to reach and the intrinsics version is a real call, which at n=8 is most
of the work. The research document named that risk and did not measure it: "the
end-to-end figure will be smaller than 19-32%". Smaller was the wrong word.
Its table also measured only exact multiples of the vector width, which hid the
tail entirely: n=9 and n=17 -- one vector plus a one-element tail, the shape the
band mostly sees -- cost +111.6% and +55.3%. And float64 loses everywhere for a
structural reason the f32-only table could not show: 4 lanes means the call is
amortised over half as much arithmetic at the same length.
Reverted. The next attempt has to emit the intrinsics INSIDE the generated
backend, in the same package as the guard where they can inline into it, not
behind a package boundary.
The bit-identity half stands and was re-verified: elementwise archsimd is
bit-identical to ref at every length 0-40 over IEEE specials -- both zeros, both
infinities, NaN, the subnormal and finite extremes -- for add, sub, mul and div,
on both lanes, compared on Float32bits/Float64bits rather than ==.
docs/wrong.md entry 75; README count 83 -> 84.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…see scalar Measured before writing any pragma. 336 (kernel, tier) refusals over 132 kernels, not 346; the permute/scatter family is 48% of them and needs instructions the targets lack; there are ZERO fast_* refusals, so the plan's "114 float/fast_* -- the plausible wins" bucket does not exist. Only ten loops in all of csrc are refused for the reason the pragma addresses, and clang's own remark says that pragma grants REORDERING -- which the numerical contract forbids and which the simd_fast_* prefix already exists to house. Those same ten loops are the accumulation inside convolve, correlate and matmul, and those kernels are refused on zero tiers while shipping scalar: simd_convolve_f32 has zero packed arithmetic and ten scalar float ops on sse2, avx2 and avx512 alike, and its only vector instructions on sse2 are a movups and an xorps. verify.vectorWidth counts any %xmm/%ymm/%zmm operand as vector, and every scalar float instruction on x86-64 uses those registers, so RequireVector structurally cannot see a scalar float kernel on amd64. Scanning every symbol for "does arithmetic, none of it packed": 26 of 913 on sse2, 8 of 913 on avx2. No code change: fixing the gate drops the eight to internal/ref, and whether that is faster is unmeasured -- the machine was not quiet. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
vectorWidth answers "does any instruction touch a vector register", and on amd64 every scalar float instruction does -- scalar SSE arithmetic lives in %xmm. RequireVector, whose whole purpose is to stop a kernel shipping as an acceleration while running scalar code, therefore cannot see that case on the one architecture that runs natively here. arm64 has the same blind spot in a different spelling: `fmul s0, s1, s2` and `fmul v0.4s, v1.4s, v2.4s` differ in the operand, not the register file. arithKind classifies each instruction as packed arithmetic, scalar arithmetic, or neither. Moves are neither, and so is `xorps %xmm0, %xmm0` -- those four instructions are the entirety of what made simd_convolve_f32 look vectorized at sse2. The classifier runs on amd64 and arm64 only, and ArithKnown says so, so a zero elsewhere reads as "not measured" rather than as a finding. check-emission now prints SCALAR-ONLY per kernel that does arithmetic with none of it in lanes: 38 instances over 13 kernels -- convolve, correlate, movavg, polyeval, minr, rolling_min/max_f64, dtoa_f64 -- across avx512 9, sse2 8, avx2 8, sve2 8, neon 5. Nothing is dropped: falling back to internal/ref instead is a behaviour change that wants a benchmark, and the machine was not quiet. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… vectorize LLVM vectorizes only the innermost loop, and in all four the innermost loop was the reduction -- a serial float dependency it cannot break without reassociating. So they shipped scalar on every tier: simd_convolve_f32 had zero packed arithmetic at sse2, avx2 and avx512 alike, while passing the emission gate on the strength of a movups and an xorps (entry 76). Blocking over the outputs with a fixed 16-element accumulator makes the innermost loop the independent one. Every output still accumulates in the same order, so this is bit-identical rather than close -- which is the difference between this and the pragma, whose whole effect is to permit the reassociation the numerical contract forbids. Nothing is allocated: 16 doubles is 128 bytes against a 512-byte frame budget. All six symbols now carry packed arithmetic on both amd64 tiers, SCALAR-ONLY drops from 13 kernels to 5, and loong64 gains two: movavg f32/f64 were refused there with "LLVM did not vectorize it for lasx" and now emit, so the kernel total moves 6,931 -> 6,933. Bit-identity is checked by the conformance differential against internal/ref, and the check was probed for being vacuous: eleven of its eighteen length pairs are at or above the guard's threshold of 16, and poisoning the blocked path alone reddens it at dst=16, the first blocked case. Speed is NOT measured -- the machine was not quiet. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`go test ./...` was red at e2ab049: entry 77 was added to docs/wrong.md without bumping README's "records 85 things", and that one line takes `make verify`, `make test-tiers` (aborting at the first tier) and `make test-cross` down with it. Entry 77's numbers were wrong in three places, all from generalising one tier's output: - the avx2 packed-arithmetic column was wrong in all six rows it listed. Re-derived twice, once against `git show c1d08c4:csrc/numeric.c` and once against the blocked file, so the before column is measured too. The table now gives all eight symbols at sse2/avx2/avx512. - "the ten scalar instructions that remain are the tail loop" -- the scalar count is 10 at sse2 and 18 at avx2/avx512, and it is identical before and after the change on every amd64 tier. Blocking added packed work beside those instructions; it did not convert them. - "eleven of eighteen length pairs are at or above the guard's threshold" -- it is eight. Added what the record was missing: a review sweep of 3,286,542 cases per tier across all nine tiers (cancellation at 1e16, mid-accumulation overflow, block-boundary magnitude switches, NaN payloads) finding zero differing bits, with a poisoned `ref` proving it non-vacuous; and the arm64 cost, where LLVM vectorizes over the taps instead of the block and six kernels go from a zero frame to 80-320 bytes with 7-13x the scalar float arithmetic -- bit-identical, inside the frame budget, and measured for speed nowhere. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
v1.21.0inREADME.md,CHANGELOG.md,AGENTS.md,CLAUDE.md, anddocs/assets/badges/release.svg.README.mdwith high-intent search keywords (Go SIMD,Vector Primitives,AVX-512,NEON,SVE2,RVV).GOEXPERIMENT=simd, minimum slice sizes).go test ./internal/tests/docs) andmake verifypass 100% green.