Medical question and answer dataset gathered from the web.
-
Updated
Dec 9, 2020
Medical question and answer dataset gathered from the web.
Tiny QA Benchmark++ a micro-benchmark suite (52-item gold + on-demand multilingual synthetic packs), generator CLI, and CI-ready eval harness for ultra-fast LLM smoke-testing & regression-catching.
Additional Videos Data and QA pairs for Balancing Original MUSIC-AVQA Dataset
Flipper Zero Sub-GHz RF dataset (280-1100 MHz, 9 countries) with 1500 Q&A pairs for LLM fine-tuning, fact-checked allocations, and a GPU-accelerated validation pipeline (Ollama Qwen 32B + DeBERTa NLI).
Data for HindiRC
Welfare-QA : Korean WelfareDomain Dataset (QA from docs)
Command-line tool to split documents into chunks and automatically generate question–answer datasets, designed for preparing data to fine-tune large language models (LLMs).
QATorch is a Python tool for in-depth analysis of Q&A datasets, designed to prepare data for Retrieval Augmented Generation (RAG) systems. It performs data quality checks, deduplication, metadata analysis, and generates detailed, customizable reports to streamline AI/NLP dataset preparation.
🤖 Transform unstructured policies into structured conversation and evaluation data for LLMs with Synkro's versatile framework.
Open Q&A dataset for the Swedish construction industry. 300 Q&As in v1.0, target 1000+. Multi-format (JSON, JSONL, Alpaca, ShareGPT, CSV). CC BY 4.0. Maintained by Zaragoza AB, Helsingborg.
F1-Dataset 1950–2025 — 1,149 races, 25,784 results & 10656 QA pairs from Jolpica + Wikipedia for LLM finetuning.
Curated bilingual long-context QA dataset with 20 questions, 160 rollouts, evidence, scores, and validation.
Automated QA data generation pipeline - three-stage logistic regression filtering, zero-dependency Node backend, single-file frontend
To associate your repository with the qa-dataset topic, visit your repo's landing page and select "manage topics."