Symptom
TestSearchModelsController#test_search_data (in ontologies_api) intermittently fails on the Virtuoso CI backend: Solr's totalCount for the :ontology_data collection comes up short of the SPARQL triple-store id count, by a varying amount.
The two failures ran concurrently (push + pull_request events of the same commit at 21:42 UTC on 2026-07-27); the varying deficit (147 vs 143 missing documents) is the key clue.
Root cause
fetch_ids in lib/ontologies_linked_data/services/submission_process/operations/submission_all_data_indexer.rb (line 81 on develop, introduced Jan 2026 in 836b2dc) paginates the ids to index with LIMIT/OFFSET and no ORDER BY:
query = Goo.sparql_query_client.select(:id)
.distinct
.from(RDF::URI.new(@submission.id))
.where(%i[id p v])
.limit(size)
.offset((page - 1) * size)
SPARQL guarantees no result ordering without ORDER BY, so OFFSET windows over an unordered result set are undefined. Across the page queries, an id can appear in two windows (harmless: Solr dedupes by document id) or in no window at all (document silently never indexed). The observed shortfall is the second case.
Aggravating factors, and why only Virtuoso fails in practice:
index_all_data (line 21) uses page size 100 for Virtuoso vs 1000 for other backends, so the same dataset produces ~15 offset windows on vo vs 2 elsewhere: many more window boundaries to misalign.
- Pages are fetched by 10 parallel threads (
Parallel.map(..., in_threads: 10), line 40), so the paginated queries interleave under load.
- Virtuoso's row ordering without ORDER BY is the least stable of the four CI backends, particularly under concurrent load.
Impact
- CI: intermittent vo failures in ontologies_api (and any suite exercising
index_all_data).
- Production: any deployment that populates the
:ontology_data collection can silently under-index documents. (The NCBO deployment currently does not create this collection, so the practical impact today is CI noise plus a correctness trap for future use.)
Suggested fix
Make the pagination deterministic. Options, in rough order of preference:
- Add
ORDER BY ?id to the fetch_ids query: one-line change; deterministic windows regardless of thread interleaving. Check Virtuoso performance for large graphs (ORDER BY + OFFSET forces a sort per page; the per-submission graph scope bounds it).
- Keyset pagination (
FILTER(?id > <last>) + ORDER BY ?id): avoids deep-OFFSET cost but serializes page fetching, so it conflicts with the Parallel.map design unless ids are pre-fetched.
- Fetch the full id list once (single unpaginated DISTINCT query), then split in Ruby and keep parallel triple fetching per slice: removes the offset race entirely; memory-bounded by id count.
Acceptance criteria
- After indexing, Solr document count equals the SPARQL DISTINCT id count for the submission graph (this is exactly what
test_search_data asserts at test/controllers/test_search_models_controller.rb:393 in ontologies_api).
- Repeated runs on the Virtuoso backend are stable. To reproduce the flake pre-fix, run the ontologies_api test with
GOO_BACKEND_NAME=VO (or however the vo matrix leg is configured in ci.yml) in a loop; the failure is probabilistic, not deterministic.
Context
Diagnosed during the ontologies_api /search performance work (ncbo/ontologies_api#244, PR ncbo/ontologies_api#245); the failures appeared on that PR's CI runs but the defect predates it and is unrelated (the compared counts are produced entirely during OLD submission processing, before any API code runs).
Symptom
TestSearchModelsController#test_search_data(in ontologies_api) intermittently fails on the Virtuoso CI backend: Solr'stotalCountfor the:ontology_datacollection comes up short of the SPARQL triple-store id count, by a varying amount.The two failures ran concurrently (push + pull_request events of the same commit at 21:42 UTC on 2026-07-27); the varying deficit (147 vs 143 missing documents) is the key clue.
Root cause
fetch_idsinlib/ontologies_linked_data/services/submission_process/operations/submission_all_data_indexer.rb(line 81 on develop, introduced Jan 2026 in 836b2dc) paginates the ids to index with LIMIT/OFFSET and no ORDER BY:SPARQL guarantees no result ordering without ORDER BY, so OFFSET windows over an unordered result set are undefined. Across the page queries, an id can appear in two windows (harmless: Solr dedupes by document id) or in no window at all (document silently never indexed). The observed shortfall is the second case.
Aggravating factors, and why only Virtuoso fails in practice:
index_all_data(line 21) uses page size 100 for Virtuoso vs 1000 for other backends, so the same dataset produces ~15 offset windows on vo vs 2 elsewhere: many more window boundaries to misalign.Parallel.map(..., in_threads: 10), line 40), so the paginated queries interleave under load.Impact
index_all_data).:ontology_datacollection can silently under-index documents. (The NCBO deployment currently does not create this collection, so the practical impact today is CI noise plus a correctness trap for future use.)Suggested fix
Make the pagination deterministic. Options, in rough order of preference:
ORDER BY ?idto thefetch_idsquery: one-line change; deterministic windows regardless of thread interleaving. Check Virtuoso performance for large graphs (ORDER BY + OFFSET forces a sort per page; the per-submission graph scope bounds it).FILTER(?id > <last>)+ORDER BY ?id): avoids deep-OFFSET cost but serializes page fetching, so it conflicts with theParallel.mapdesign unless ids are pre-fetched.Acceptance criteria
test_search_dataasserts at test/controllers/test_search_models_controller.rb:393 in ontologies_api).GOO_BACKEND_NAME=VO(or however the vo matrix leg is configured in ci.yml) in a loop; the failure is probabilistic, not deterministic.Context
Diagnosed during the ontologies_api /search performance work (ncbo/ontologies_api#244, PR ncbo/ontologies_api#245); the failures appeared on that PR's CI runs but the defect predates it and is unrelated (the compared counts are produced entirely during OLD submission processing, before any API code runs).