Skip to content

Agent: Natural-language queries interface #1210

Description

@Ayush8923

Describe the current behavior

Kaapi currently have no way to ask natural-language questions about their own data. We want a /agent endpoint that lets a user ask things like:

  • How many eval runs have I done?
  • Summarize my last three evals.
  • Which documents were in the knowledge base for this eval run?

...and have the agent figure out which Kaapi endpoints to call (possibly several, in sequence) and answer in plain language.

Reference: Open Chat Studio (OCS)
We looked at OCS as a reference, not a blueprint:

  • OCS built-in agent tools are static/generic (web search, etc.).
  • OCS does support "custom actions" (base URL + API schema), which is conceptually close to what we want, but we are not adopting OCS architecture wholesale. we will skip parts of it and build things it doesn't have.

Describe the enhancement you'd like

  • Right now, we just need to implement a read-only agent, scoped per-NGO, that answers questions about eval runs, datasets, collections, and documents.
  • Agent decides autonomously which tool(s) to call and in what order.

Non-goals (v1)

  • No write access. POST endpoints (e.g. triggering eval runs) are explicitly excluded from the toolset. A write-capable agent is a later step.
  • Not need any frontend changes for now. first we build the backend flow then we can integrate this in the Kaapi console.
  • No MCP server.

Architecture

Endpoint

  • New POST /agent on the Kaapi API.
  • all tool calls are scoped to the NGO whose API key the requesting user is authenticated with. The tool-executor node injects this key into the request headers before calling the real Kaapi endpoint, the LLM never sees or handles the key. This closes off prompt-injection-driven key leakage and enforces tenant isolation at the harness level, not the model level.

Agent loop (LangGraph)
for this, thinking to use the LangGraph to implement this as a stateful loop rather than a bespoke while loop, mainly for: conditional branching (iteration cap, tool-vs-final-answer routing), built-in state history for debugging, and a clean upgrade path if we later need human-in-the-loop interrupts or multiple agents.

High-level shape:

START → agent_node → (tool call?) → tool_executor_node → agent_node → ... → END
                    ↘ (final answer) ────────────────────────────────→ END

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions