Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
61 changes: 61 additions & 0 deletions docs/architecture/assessment/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,61 @@
# AI Assessments — Getting Started

**An assessment uses an LLM to grade your items against a rubric and gives you back a structured result** (scores, reasoning, feedback) for every item — not free text, but a fixed JSON shape you choose.

You give Kaapi two things:

1. A **config** — your rubric (the grading instructions), the model to use, and the exact result shape you want back.
2. Your **items** — the rows you want graded (text and/or image/PDF URLs).
Comment thread
coderabbitai[bot] marked this conversation as resolved.

Kaapi grades every item and delivers the results to your **webhook**.

---

## The whole flow in three steps

| Step | You do | Kaapi does |
|---|---|---|
| **1. Create a config** | Save an `ASSESSMENT` config once (`POST /configs`) | Stores it, versioned |
| **2. Submit items** | `POST /assessments` with your rows + a `callback_url` | Returns an `assessment_id`, starts grading in the background |
| **3. Get results** | Wait for the webhook | POSTs the finished results to your `callback_url` |

You never poll or wait on the request — submitting returns immediately, and the results arrive later at your webhook.

![BATCH assessment flow](assets/batch-flow.png)

**BATCH is fully batched.** Both stages run as provider **batch jobs** — the
pre-filters run as a batch, and the assessment runs as a batch. Results are
delivered to your **webhook** when everything completes (no polling).

---

## Two methods (Kaapi picks for you)

You never set a "mode". Kaapi looks at your input and decides:

| Method | When | Input shape | Status |
|---|---|---|---|
| **BATCH** | Many items at once | `data` is a list of rows | ✅ Available |
| **RESPONSE** | A single item, fast | a single item's `attachments` (no `data`) | 🚧 WIP (returns `501` today) |

This guide covers **BATCH**, the method that is live.

---

## Supported models

Pick the provider per config (and per pre-filter):

| Provider | Value in config | Status |
|---|---|---|
| OpenAI | `openai` | ✅ |
| Google (AI Studio / Gemini) | `google` | ✅ |
| Anthropic (Claude) | `anthropic` | ✅ |
| Google Cloud / Vertex | — | 🚧 WIP |

---

## Where to go next

1. **[Configuration and versioning](configuration-and-versioning.md)** — build your rubric, choose the model, define the result shape, and manage versions.
2. **[API contract](api-contract.md)** — request/response fields, types, status values, and error codes.
205 changes: 205 additions & 0 deletions docs/architecture/assessment/api-contract.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,205 @@
# API Contract — `POST /assessments`

Precise request and response shapes for the BATCH assessment API. For a
walkthrough with context, see the [overview](README.md).

Everything is delivered by **webhook** — there is no status or result poll
endpoint. RESPONSE-shaped input returns `501` (WIP).

**Sample input / output JSON files:**
https://drive.google.com/drive/folders/1BCaauUuXr9DaZTWI-_-x101SDT4ktwp5?usp=share_link

---

## Request

`POST /assessments`

| Field | Type | Required | Notes |
|---|---|---|---|
| `config` | object | ✅ | which saved config version to run |
| `config.id` | UUID | ✅ | config id (must be tagged `ASSESSMENT`) |
| `config.version` | int ≥ 1 | ✅ | config version to pin |
| `input` | object | ✅ | a `data` list ⇒ BATCH; `attachments` only (no `data`) ⇒ RESPONSE (501) |
| `input.data` | array (≥ 1) | ✅ | rows; each row is a flat `{ column: string }` object |
| `callback_url` | URL (**HTTPS**) | ✅ | webhook the result is POSTed to |
| `request_metadata` | object | optional | echoed back unchanged in the result |

Rules:

- **Strict input** — no extra keys are allowed on `input`; a body carrying both
`data` and `attachments` is rejected.
- **Rows match the config's `input_schema`** — every declared column present, no
undeclared columns, `image`/`pdf` values must be URLs. Otherwise `422`.
- **`callback_url`** must be HTTPS and public (private/loopback hosts are rejected).

```json
{
"config": { "id": "a9015dbf-…", "version": 1 },
"input": {
"data": [
{ "submission_id": "s1", "answer_sheet": "https://cdn.example.com/s1.jpg" }
]
},
"callback_url": "https://your-app.example.com/webhooks/assessment",
"request_metadata": { "batch": "class7-term1" }
}
```

### Building the batch input

The `input` object carries only your rows; the prompt template lives in the config
(`config_blob.assessment.params.submission`), not in the request. Each row's keys
match the config's top-level `input_schema`:

1. **One object per item** goes in `input.data`. Each object's keys are the column
names declared in the config's `input_schema`, and the values are strings.
2. **Attachment columns** (`image` / `pdf`) take a URL string; text columns take
plain text.
3. **The prompt is the config's `submission` template.** Any `{column}` placeholder
in it is replaced with that row's value at grading time, so one template applies
to every row. The request no longer carries a `query`.
4. **Match the schema exactly** — every declared column present, no extra columns.

Example: for `input_schema = { submission_id: text, answer_sheet: image(url) }`,
each row is `{ "submission_id": "...", "answer_sheet": "https://..." }` and the
config's `submission` template can reference `{submission_id}` and `{answer_sheet}`.

---

## Response — submit acknowledgement (`200`)

Returned immediately; contains no results. Wrapped in the standard envelope
`{ success, data, error, metadata }`.

| Field (`data`) | Type | Notes |
|---|---|---|
| `assessment_id` | UUID | correlate with the webhook |
| `status` | enum | `PROCESSING` on accept |
| `message` | string | human-readable |
| `inserted_at` / `updated_at` | timestamp | ISO-8601 |

```json
{
"success": true,
"data": {
"assessment_id": "8a2a7bc1-…",
"status": "PROCESSING",
"message": "Your assessment is being processed",
"inserted_at": "2026-08-12T10:15:30Z",
"updated_at": "2026-08-12T10:15:30Z"
},
"error": null,
"metadata": null
}
```

---

## Webhook — the result (POST to `callback_url`)

Delivered once, on completion. Wrapped in the same standard envelope as the ack —
`{ success, data, error, metadata }` — with the `AssessmentCallback` payload
nested under `data`. `status` lives only inside that nested `data`, never at the
envelope's top level (unlike the ack, where `status` sits directly under `data`).

| Field (`data`) | Type | Notes |
|---|---|---|
| `assessment_id` | UUID | matches the ack |
| `status` | enum | terminal (see below) |
| `data` | object | the `AssessmentBatchResult` (BATCH) |
| `request_metadata` | object \| null | echoed from the request |

Comment thread
vprashrex marked this conversation as resolved.
`data.data` (`AssessmentBatchResult`):

| Field | Type | Notes |
|---|---|---|
| `total_items` | int | number of input rows |
| `counts.assessed` | int | rows graded |
| `counts.filtered` | int | rows gated out by a pre-filter |
| `counts.errors` | int | rows with an error |
| `items` | array | one `AssessmentResult` per input row, in order |

`items[]` (`AssessmentResult`):

| Field | Type | Notes |
|---|---|---|
| `output.assessment` | object \| string \| null | your `json_output_schema` filled in; string for free-text; `null` if gated out / failed |
| `output.pre_filter.topic_relevance` | `{verdict: bool, reasoning: string}` \| null | null if not configured |
| `error` | string \| null | per-row error |

```json
{
"success": true,
"data": {
"assessment_id": "8a2a7bc1-…",
"status": "COMPLETED",
"data": {
"total_items": 2,
"counts": { "assessed": 1, "filtered": 1, "errors": 0 },
"items": [
{
"output": {
"assessment": { "score": 20, "feedback": "…" },
"pre_filter": { "topic_relevance": { "verdict": true, "reasoning": "…" } }
},
"error": null
},
{
"output": {
"assessment": null,
"pre_filter": { "topic_relevance": { "verdict": false, "reasoning": "off-topic" } }
},
"error": null
}
]
},
"request_metadata": { "batch": "class7-term1" }
},
"error": null,
"metadata": null
}
Comment thread
vprashrex marked this conversation as resolved.
```

### Verifying the webhook signature

When a webhook signing secret is configured for the project (organization +
project credential of provider `webhook_secret`), each delivery carries two
headers:

| Header | Value |
|---|---|
| `X-Webhook-Signature` | hex HMAC-SHA256 digest |
| `X-Webhook-Timestamp` | Unix timestamp in milliseconds, at send time |

To verify: rebuild the signing string as `"<timestamp_ms>.".encode() + raw_body`,
where `raw_body` is the exact compact-JSON bytes received (no re-serialization),
compute HMAC-SHA256 over it with the shared secret, and compare the hex digest
to `X-Webhook-Signature` using a constant-time comparison. Reject deliveries
whose `X-Webhook-Timestamp` is too far from the current time (replay
protection). No secret configured ⇒ no signature headers are sent.
```

---

## Status values

| Status | Meaning |
|---|---|
| `PENDING` | accepted, not started |
| `PROCESSING` | grading in progress (the ack status) |
| `COMPLETED` | all rows graded, no errors |
| `COMPLETED_WITH_ERRORS` | finished, some rows errored |
| `FAILED` | the run failed |

`status` lives inside the outer envelope's `data` — for both the ack and the
webhook — and is never duplicated inside the nested `AssessmentBatchResult`.

## Error codes (at submit)

| Code | When |
|---|---|
| `422` | invalid body, or a row doesn't match `input_schema`, or a non-HTTPS/private `callback_url` |
| `404` | config id not found |
| `501` | RESPONSE-shaped input (`attachments` only, no `data`) — WIP |
| `503` | failed to dispatch for processing (retry) |
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading
Loading