-
Notifications
You must be signed in to change notification settings - Fork 10
feat(docs): Add overview for AI Assessments #1017
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Open
vprashrex
wants to merge
13
commits into
main
Choose a base branch
from
feat/doc-assessment-architecture
base: main
Could not load branches
Branch not found: {{ refName }}
Loading
Could not load tags
Nothing to show
Loading
Are you sure you want to change the base?
Some commits from the old base branch may be removed from the timeline,
and old review comments may become outdated.
Open
Changes from all commits
Commits
Show all changes
13 commits
Select commit
Hold shift + click to select a range
2dd81a8
feat(architecture-doc): Add comprehensive overview for AI Assessments…
vprashrex 5fb8cdf
Merge branch 'main' into feat/doc-assessment-architecture
vprashrex e35ced1
Refactor code structure for improved readability and maintainability
vprashrex 87823fd
Add documentation for configuration and versioning of assessments
vprashrex 22ab98c
docs: Add sample input/output JSON files and config field reference l…
vprashrex 425e061
docs: Update README to clarify BATCH processing and enhance navigatio…
vprashrex ae701f2
docs: Update batch and response flow diagrams for clarity
vprashrex 58eea27
chore(docs): Update API contract and architecture documentation to cl…
vprashrex 9820ab4
Merge branch 'main' into feat/doc-assessment-architecture
vprashrex d81cae2
Merge branch 'main' into feat/doc-assessment-architecture
Ayush8923 6b02032
Merge branch 'main' into feat/doc-assessment-architecture
vprashrex e7dbb09
docs: Enhance API contract and architecture documentation for webhook…
vprashrex 133a528
docs: Add section on GCS attachment resolution and include diagram
vprashrex File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,61 @@ | ||
| # AI Assessments — Getting Started | ||
|
|
||
| **An assessment uses an LLM to grade your items against a rubric and gives you back a structured result** (scores, reasoning, feedback) for every item — not free text, but a fixed JSON shape you choose. | ||
|
|
||
| You give Kaapi two things: | ||
|
|
||
| 1. A **config** — your rubric (the grading instructions), the model to use, and the exact result shape you want back. | ||
| 2. Your **items** — the rows you want graded (text and/or image/PDF URLs). | ||
|
|
||
| Kaapi grades every item and delivers the results to your **webhook**. | ||
|
|
||
| --- | ||
|
|
||
| ## The whole flow in three steps | ||
|
|
||
| | Step | You do | Kaapi does | | ||
| |---|---|---| | ||
| | **1. Create a config** | Save an `ASSESSMENT` config once (`POST /configs`) | Stores it, versioned | | ||
| | **2. Submit items** | `POST /assessments` with your rows + a `callback_url` | Returns an `assessment_id`, starts grading in the background | | ||
| | **3. Get results** | Wait for the webhook | POSTs the finished results to your `callback_url` | | ||
|
|
||
| You never poll or wait on the request — submitting returns immediately, and the results arrive later at your webhook. | ||
|
|
||
|  | ||
|
|
||
| **BATCH is fully batched.** Both stages run as provider **batch jobs** — the | ||
| pre-filters run as a batch, and the assessment runs as a batch. Results are | ||
| delivered to your **webhook** when everything completes (no polling). | ||
|
|
||
| --- | ||
|
|
||
| ## Two methods (Kaapi picks for you) | ||
|
|
||
| You never set a "mode". Kaapi looks at your input and decides: | ||
|
|
||
| | Method | When | Input shape | Status | | ||
| |---|---|---|---| | ||
| | **BATCH** | Many items at once | `data` is a list of rows | ✅ Available | | ||
| | **RESPONSE** | A single item, fast | a single item's `attachments` (no `data`) | 🚧 WIP (returns `501` today) | | ||
|
|
||
| This guide covers **BATCH**, the method that is live. | ||
|
|
||
| --- | ||
|
|
||
| ## Supported models | ||
|
|
||
| Pick the provider per config (and per pre-filter): | ||
|
|
||
| | Provider | Value in config | Status | | ||
| |---|---|---| | ||
| | OpenAI | `openai` | ✅ | | ||
| | Google (AI Studio / Gemini) | `google` | ✅ | | ||
| | Anthropic (Claude) | `anthropic` | ✅ | | ||
| | Google Cloud / Vertex | — | 🚧 WIP | | ||
|
|
||
| --- | ||
|
|
||
| ## Where to go next | ||
|
|
||
| 1. **[Configuration and versioning](configuration-and-versioning.md)** — build your rubric, choose the model, define the result shape, and manage versions. | ||
| 2. **[API contract](api-contract.md)** — request/response fields, types, status values, and error codes. | ||
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,205 @@ | ||
| # API Contract — `POST /assessments` | ||
|
|
||
| Precise request and response shapes for the BATCH assessment API. For a | ||
| walkthrough with context, see the [overview](README.md). | ||
|
|
||
| Everything is delivered by **webhook** — there is no status or result poll | ||
| endpoint. RESPONSE-shaped input returns `501` (WIP). | ||
|
|
||
| **Sample input / output JSON files:** | ||
| https://drive.google.com/drive/folders/1BCaauUuXr9DaZTWI-_-x101SDT4ktwp5?usp=share_link | ||
|
|
||
| --- | ||
|
|
||
| ## Request | ||
|
|
||
| `POST /assessments` | ||
|
|
||
| | Field | Type | Required | Notes | | ||
| |---|---|---|---| | ||
| | `config` | object | ✅ | which saved config version to run | | ||
| | `config.id` | UUID | ✅ | config id (must be tagged `ASSESSMENT`) | | ||
| | `config.version` | int ≥ 1 | ✅ | config version to pin | | ||
| | `input` | object | ✅ | a `data` list ⇒ BATCH; `attachments` only (no `data`) ⇒ RESPONSE (501) | | ||
| | `input.data` | array (≥ 1) | ✅ | rows; each row is a flat `{ column: string }` object | | ||
| | `callback_url` | URL (**HTTPS**) | ✅ | webhook the result is POSTed to | | ||
| | `request_metadata` | object | optional | echoed back unchanged in the result | | ||
|
|
||
| Rules: | ||
|
|
||
| - **Strict input** — no extra keys are allowed on `input`; a body carrying both | ||
| `data` and `attachments` is rejected. | ||
| - **Rows match the config's `input_schema`** — every declared column present, no | ||
| undeclared columns, `image`/`pdf` values must be URLs. Otherwise `422`. | ||
| - **`callback_url`** must be HTTPS and public (private/loopback hosts are rejected). | ||
|
|
||
| ```json | ||
| { | ||
| "config": { "id": "a9015dbf-…", "version": 1 }, | ||
| "input": { | ||
| "data": [ | ||
| { "submission_id": "s1", "answer_sheet": "https://cdn.example.com/s1.jpg" } | ||
| ] | ||
| }, | ||
| "callback_url": "https://your-app.example.com/webhooks/assessment", | ||
| "request_metadata": { "batch": "class7-term1" } | ||
| } | ||
| ``` | ||
|
|
||
| ### Building the batch input | ||
|
|
||
| The `input` object carries only your rows; the prompt template lives in the config | ||
| (`config_blob.assessment.params.submission`), not in the request. Each row's keys | ||
| match the config's top-level `input_schema`: | ||
|
|
||
| 1. **One object per item** goes in `input.data`. Each object's keys are the column | ||
| names declared in the config's `input_schema`, and the values are strings. | ||
| 2. **Attachment columns** (`image` / `pdf`) take a URL string; text columns take | ||
| plain text. | ||
| 3. **The prompt is the config's `submission` template.** Any `{column}` placeholder | ||
| in it is replaced with that row's value at grading time, so one template applies | ||
| to every row. The request no longer carries a `query`. | ||
| 4. **Match the schema exactly** — every declared column present, no extra columns. | ||
|
|
||
| Example: for `input_schema = { submission_id: text, answer_sheet: image(url) }`, | ||
| each row is `{ "submission_id": "...", "answer_sheet": "https://..." }` and the | ||
| config's `submission` template can reference `{submission_id}` and `{answer_sheet}`. | ||
|
|
||
| --- | ||
|
|
||
| ## Response — submit acknowledgement (`200`) | ||
|
|
||
| Returned immediately; contains no results. Wrapped in the standard envelope | ||
| `{ success, data, error, metadata }`. | ||
|
|
||
| | Field (`data`) | Type | Notes | | ||
| |---|---|---| | ||
| | `assessment_id` | UUID | correlate with the webhook | | ||
| | `status` | enum | `PROCESSING` on accept | | ||
| | `message` | string | human-readable | | ||
| | `inserted_at` / `updated_at` | timestamp | ISO-8601 | | ||
|
|
||
| ```json | ||
| { | ||
| "success": true, | ||
| "data": { | ||
| "assessment_id": "8a2a7bc1-…", | ||
| "status": "PROCESSING", | ||
| "message": "Your assessment is being processed", | ||
| "inserted_at": "2026-08-12T10:15:30Z", | ||
| "updated_at": "2026-08-12T10:15:30Z" | ||
| }, | ||
| "error": null, | ||
| "metadata": null | ||
| } | ||
| ``` | ||
|
|
||
| --- | ||
|
|
||
| ## Webhook — the result (POST to `callback_url`) | ||
|
|
||
| Delivered once, on completion. Wrapped in the same standard envelope as the ack — | ||
| `{ success, data, error, metadata }` — with the `AssessmentCallback` payload | ||
| nested under `data`. `status` lives only inside that nested `data`, never at the | ||
| envelope's top level (unlike the ack, where `status` sits directly under `data`). | ||
|
|
||
| | Field (`data`) | Type | Notes | | ||
| |---|---|---| | ||
| | `assessment_id` | UUID | matches the ack | | ||
| | `status` | enum | terminal (see below) | | ||
| | `data` | object | the `AssessmentBatchResult` (BATCH) | | ||
| | `request_metadata` | object \| null | echoed from the request | | ||
|
|
||
|
vprashrex marked this conversation as resolved.
|
||
| `data.data` (`AssessmentBatchResult`): | ||
|
|
||
| | Field | Type | Notes | | ||
| |---|---|---| | ||
| | `total_items` | int | number of input rows | | ||
| | `counts.assessed` | int | rows graded | | ||
| | `counts.filtered` | int | rows gated out by a pre-filter | | ||
| | `counts.errors` | int | rows with an error | | ||
| | `items` | array | one `AssessmentResult` per input row, in order | | ||
|
|
||
| `items[]` (`AssessmentResult`): | ||
|
|
||
| | Field | Type | Notes | | ||
| |---|---|---| | ||
| | `output.assessment` | object \| string \| null | your `json_output_schema` filled in; string for free-text; `null` if gated out / failed | | ||
| | `output.pre_filter.topic_relevance` | `{verdict: bool, reasoning: string}` \| null | null if not configured | | ||
| | `error` | string \| null | per-row error | | ||
|
|
||
| ```json | ||
| { | ||
| "success": true, | ||
| "data": { | ||
| "assessment_id": "8a2a7bc1-…", | ||
| "status": "COMPLETED", | ||
| "data": { | ||
| "total_items": 2, | ||
| "counts": { "assessed": 1, "filtered": 1, "errors": 0 }, | ||
| "items": [ | ||
| { | ||
| "output": { | ||
| "assessment": { "score": 20, "feedback": "…" }, | ||
| "pre_filter": { "topic_relevance": { "verdict": true, "reasoning": "…" } } | ||
| }, | ||
| "error": null | ||
| }, | ||
| { | ||
| "output": { | ||
| "assessment": null, | ||
| "pre_filter": { "topic_relevance": { "verdict": false, "reasoning": "off-topic" } } | ||
| }, | ||
| "error": null | ||
| } | ||
| ] | ||
| }, | ||
| "request_metadata": { "batch": "class7-term1" } | ||
| }, | ||
| "error": null, | ||
| "metadata": null | ||
| } | ||
|
vprashrex marked this conversation as resolved.
|
||
| ``` | ||
|
|
||
| ### Verifying the webhook signature | ||
|
|
||
| When a webhook signing secret is configured for the project (organization + | ||
| project credential of provider `webhook_secret`), each delivery carries two | ||
| headers: | ||
|
|
||
| | Header | Value | | ||
| |---|---| | ||
| | `X-Webhook-Signature` | hex HMAC-SHA256 digest | | ||
| | `X-Webhook-Timestamp` | Unix timestamp in milliseconds, at send time | | ||
|
|
||
| To verify: rebuild the signing string as `"<timestamp_ms>.".encode() + raw_body`, | ||
| where `raw_body` is the exact compact-JSON bytes received (no re-serialization), | ||
| compute HMAC-SHA256 over it with the shared secret, and compare the hex digest | ||
| to `X-Webhook-Signature` using a constant-time comparison. Reject deliveries | ||
| whose `X-Webhook-Timestamp` is too far from the current time (replay | ||
| protection). No secret configured ⇒ no signature headers are sent. | ||
| ``` | ||
|
|
||
| --- | ||
|
|
||
| ## Status values | ||
|
|
||
| | Status | Meaning | | ||
| |---|---| | ||
| | `PENDING` | accepted, not started | | ||
| | `PROCESSING` | grading in progress (the ack status) | | ||
| | `COMPLETED` | all rows graded, no errors | | ||
| | `COMPLETED_WITH_ERRORS` | finished, some rows errored | | ||
| | `FAILED` | the run failed | | ||
|
|
||
| `status` lives inside the outer envelope's `data` — for both the ack and the | ||
| webhook — and is never duplicated inside the nested `AssessmentBatchResult`. | ||
|
|
||
| ## Error codes (at submit) | ||
|
|
||
| | Code | When | | ||
| |---|---| | ||
| | `422` | invalid body, or a row doesn't match `input_schema`, or a non-HTTPS/private `callback_url` | | ||
| | `404` | config id not found | | ||
| | `501` | RESPONSE-shaped input (`attachments` only, no `data`) — WIP | | ||
| | `503` | failed to dispatch for processing (retry) | | ||
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Oops, something went wrong.
Oops, something went wrong.
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
Uh oh!
There was an error while loading. Please reload this page.