Summary
Select, download, generate, and standardize official benchmark datasets across varying sizes and topologies (synthetic and real-world) for systematic evaluation.
Objectives and Technical Scope
1. Dataset Selection and Sizing
• Synthetic benchmarks:
• Berlin SPARQL Benchmark (BSBM): e-commerce domain, mix of simple and complex queries.
• Lehigh University Benchmark (LUBM): university domain, ontology with reasoning potential.
• BowlognaBench: student/university relationships.
• Real-world benchmarks:
• Curated DBpedia samples across standard distributions.
• Establish standardized scale tiers: Small (~50k triples), Medium (~1M triples), Large (~10M+ triples).
2. Storage and Preprocessing
• Provide automated download/generation scripts and verify checksums.
• Normalize serialization formats (N-Triples, Turtle) to ensure parsing consistency across different evaluated stores.
Acceptance Criteria
[ ] Dataset generation and download scripts are automated and committed to the repository.
[ ] Datasets across small, medium, and large scales are formatted, validated, and documented.
Summary
Select, download, generate, and standardize official benchmark datasets across varying sizes and topologies (synthetic and real-world) for systematic evaluation.
Objectives and Technical Scope
1. Dataset Selection and Sizing
• Synthetic benchmarks:
• Berlin SPARQL Benchmark (BSBM): e-commerce domain, mix of simple and complex queries.
• Lehigh University Benchmark (LUBM): university domain, ontology with reasoning potential.
• BowlognaBench: student/university relationships.
• Real-world benchmarks:
• Curated DBpedia samples across standard distributions.
• Establish standardized scale tiers: Small (~50k triples), Medium (~1M triples), Large (~10M+ triples).
2. Storage and Preprocessing
• Provide automated download/generation scripts and verify checksums.
• Normalize serialization formats (N-Triples, Turtle) to ensure parsing consistency across different evaluated stores.
Acceptance Criteria
[ ] Dataset generation and download scripts are automated and committed to the repository.
[ ] Datasets across small, medium, and large scales are formatted, validated, and documented.