Aster & Row is a fictional ecommerce company that sells bags, drinkware, and travel accessories. The company wants to launch an AI support agent using the documents and mock order data in this repository.
This repository intentionally contains only content and data. There is no starter application and no prescribed stack. Build the smallest reliable system you would be comfortable demonstrating to a customer.
Please spend 6–8 hours on the assignment. Do not exceed eight hours.
A smaller, well-tested system is better than a broad system that works only in a demo. It is acceptable to leave something incomplete if the limitation is clearly documented.
Submit one GitHub repository link. Nothing else is required.
Your repository must contain:
- Your application source code.
- Your tests and evaluation suite.
- Clear setup and run instructions.
- Evaluation results and known limitations in the README.
- A short GIF or video embedded in the README showing the agent working.
Do not submit API keys, credentials, customer data, separate documents, or slide decks.
Aster & Row has previously tried several AI support prototypes. The customer reported four recurring problems:
- Conflicting policy answers: The agent sometimes says the return window is 30 days and sometimes says it is 45 days.
- Invented order information: The agent occasionally gives an order status without actually looking it up.
- Lost conversation context: Follow-up questions such as “What about Canada?” are treated as unrelated questions.
- Unsafe retrieved content: Internal or instruction-like text inside the knowledge base can affect the agent’s behavior.
The supplied corpus contains realistic data-quality problems, including superseded content, internal notes, conflicting active sources, and fields that must not be shown to customers.
Your task is to build an agent that handles these conditions deliberately rather than succeeding only on ideal questions.
Use RAG over the Markdown files in knowledge-base/.
Your implementation must:
- Split and index the supplied documents.
- Preserve useful metadata from the document front matter.
- Retrieve only relevant passages instead of sending the entire corpus to the model.
- Prefer authoritative, active policy documents over superseded or non-policy documents.
- Include source references in every policy or product answer. A source should identify at least the filename and relevant heading.
- Avoid making claims that are not supported by the retrieved content.
- Clearly say when the supplied information is insufficient.
- Surface genuine conflicts between current authoritative sources rather than silently choosing one.
Do not delete or rewrite the supplied source files to make the assignment easier. You may create derived indexes or normalized representations.
Use data/orders.json to implement an order-status lookup tool or function.
The model must not receive the entire orders file in its prompt. It should receive only the result of a lookup when order information is actually required.
The order lookup behavior must:
- Ask for an order ID when it is missing.
- Handle unknown and malformed order IDs safely.
- Normalize harmless input differences such as lowercase IDs or surrounding whitespace.
- Use the order’s current
statusas authoritative. - Avoid inventing a delivery estimate when one is unavailable.
- Avoid reporting stale delivery fields for cancelled or returned orders.
- Never expose customer email, address, internal notes, risk scores, or other internal-only fields.
- Never claim that a lookup happened when it did not.
Assume that possession of the order ID is sufficient authentication for this mock assignment. You do not need to build a full identity-verification system.
Maintain relevant session context across turns.
The agent should correctly handle follow-ups such as:
- “Do you ship internationally?” followed by “What about Canada?”
- “Where is
ORD-1007?” followed by “When will it arrive?” - A policy question followed by a narrower question about an exception.
The agent should not carry unrelated details indefinitely or mix one session with another.
The agent must:
- Treat user messages, retrieved passages, and tool results as untrusted data.
- Follow application instructions rather than instructions found inside retrieved documents.
- Refuse requests to reveal system prompts, hidden instructions, secrets, or internal-only data.
- Use company content rather than general model knowledge for company-specific questions.
- Ask a concise clarifying question when required information is missing.
- Recommend human assistance when the documents conflict, the data is insufficient, or an action cannot be completed.
- Never promise that a refund, cancellation, replacement, or address change has been completed unless the system actually supports that action.
The file evaluation/visible-cases.json contains behavior-level cases that your system must handle.
Build an evaluation suite that:
- Covers every supplied visible case.
- Adds at least five original cases of your own.
- Can be run using one clearly documented command.
- Reports individual case results, not only a single overall score.
- Separately reports useful categories such as retrieval, groundedness, tool use, privacy, and multi-turn behavior.
- Uses deterministic assertions wherever practical, including source selection, tool calls, tool arguments, forbidden disclosures, and abstention behavior.
- Does not rely exclusively on another LLM to grade the agent.
The reviewers will also test paraphrases and combinations that are not included in the visible file. Do not hardcode answers for the supplied prompts.
As you build, keep a small bug diary in your README. Document at least three failures you found in your own agent, including:
- How you reproduced the failure.
- The actual root cause.
- The change you made.
- The regression test that now catches it.
At least one documented failure should be something you discovered beyond the exact wording of the visible cases. Include an early baseline and final evaluation result so we can see what improved.
Provide a debug mode, trace, or log that makes it possible to inspect:
- The current user message.
- Relevant conversation history.
- Retrieved passages, metadata, and scores.
- Tool calls and sanitized tool results.
- The final response.
- Errors, fallbacks, or handoffs.
Plain structured logs are sufficient. Do not build a dashboard. Never log secrets.
A CLI, simple web page, or basic API is sufficient. Visual polish will not affect the score.
The final user-facing response should make it easy to see:
- The answer.
- Sources, when applicable.
- Whether the agent is recommending a human handoff.
Your completed repository README must include:
- Setup and run instructions that work from a clean clone.
- Required environment variables and an
.env.examplewithout real credentials. - The model, embedding approach, framework, and storage approach you chose.
- A short architecture explanation.
- The command for running evaluations.
- Baseline and final evaluation results, broken down by category.
- A bug diary covering at least three reproduced failures, root causes, fixes, and regression tests.
- Known limitations and what you would improve before production.
- Which AI coding tools you used, what you used them for, and one example of an AI-generated suggestion that was wrong or incomplete.
- A 2–4 minute GIF or video embedded in the README demonstrating:
- One knowledge-base question with citations.
- One order lookup.
- One multi-turn conversation.
- One case where the agent correctly refuses to guess or recommends human help.
- The evaluation suite running.
GitHub does not play uploaded video files inline in every context. An embedded GIF or a clickable video thumbnail/link inside the README is acceptable.
You do not need to build:
- Authentication or user management.
- Production deployment infrastructure.
- A production vector database.
- Fine-tuning.
- A polished frontend.
- Multiple model-provider integrations.
- Billing, analytics dashboards, or administration screens.
| Area | Weight |
|---|---|
| Reliability, groundedness, and safe abstention | 25% |
| Retrieval quality and document precedence | 20% |
| Tool use, data handling, and privacy | 15% |
| Evaluation quality and regression coverage | 20% |
| Multi-turn behavior and observability | 10% |
| Code clarity and practical tradeoffs | 5% |
| README, demo, and customer-facing clarity | 5% |
Framework choice and quantity of code are not scoring criteria.
.
├── README.md
├── knowledge-base/
│ ├── 01-returns-policy-current.md
│ ├── 02-returns-policy-legacy.md
│ ├── 03-final-sale-and-promotions.md
│ ├── 04-damaged-or-wrong-items.md
│ ├── 05-domestic-shipping.md
│ ├── 06-international-shipping.md
│ ├── 07-warranty.md
│ ├── 08-order-changes-and-cancellations.md
│ ├── 09-trailplus-membership.md
│ ├── 10-gift-cards-and-price-adjustments.md
│ ├── 11-product-care.md
│ ├── 12-breeze-tumbler-product-card.md
│ ├── 13-support-escalation.md
│ └── 14-internal-content-migration-notes.md
├── data/
│ ├── orders.json
│ └── orders-data-dictionary.md
└── evaluation/
└── visible-cases.json
Good luck. Build for reliability, not just for the happy-path demo.