diff --git a/backend/app/evaluation/README.md b/backend/app/evaluation/README.md index 639bae5..ba33b81 100644 --- a/backend/app/evaluation/README.md +++ b/backend/app/evaluation/README.md @@ -425,7 +425,18 @@ All `metrics.json` files include a `performance` block: ## Dataset Structure -Download all datasets from [Google Drive](https://drive.google.com/drive/u/0/folders/1Rd1LH-oEwCkU0pBDRrYYedExorwmXA89). The Drive contains one folder per validator. Download the CSV files and place them in `backend/app/evaluation/datasets/`. +Datasets are hosted on Google Drive across two folders. Download the CSV files and place them in `backend/app/evaluation/datasets/`. + +[Current datasets](https://drive.google.com/drive/folders/1bM8GxH2lVlcT9Q79oSZhxYaWIP9Xhpya) — subfolders here are named by product area rather than by validator: + +| Drive folder | Validator | +| -------------------- | ---------------------- | +| `Privacy protection` | PII Remover | +| `Content Safety` | Toxicity | +| `Gender neutrality` | Gender Assumption Bias | +| `Scope Control` | Topic Relevance | + +[Earlier datasets](https://drive.google.com/drive/u/0/folders/1Rd1LH-oEwCkU0pBDRrYYedExorwmXA89) — ban list and multiple validators, with one subfolder per validator. Each evaluation script expects a specific filename — files must be named exactly as listed below: