The Applied AI Field Guide: fieldwork, value, engineering, and operations for AI that works beyond the demo.
-
Updated
Sep 15, 2026 - JavaScript
The Applied AI Field Guide: fieldwork, value, engineering, and operations for AI that works beyond the demo.
Core engine behind Calibrate, a framework for evaluating AI agents: speech-to-text, text-to-speech, LLM evaluation, end-to-end simulations
An end-to-end AI agent project that transcribes audio files, embeds user queries, and searches in Qdrant and web browser via the Brave API. A Streamlit interface powered by OpenAI GPT models delivers actionable health insights from both the archive and the latest research.
Evaluation of Multi-Agent Systems on Cloud Run with the GenAI Client in Vertex AI SDK
Frontend for Calibrate, a framework for evaluating AI agents: speech-to-text, text-to-speech, LLM evaluation, end-to-end simulations
Record, replay, evaluate, and regression-test AI agents with deterministic traces, calibrated LLM-as-a-judge evals, OpenTelemetry observability, and CI quality gates.
Evaluation harness for LLM browser agents — real Chromium tasks, DOM-based verifiers, step-level tracing, multi-provider benchmarks and CI regression checks.
Test, record, and replay AI agents locally and in CI with a private pytest-style framework
Compare OpenClaw setups against the same scenario suite. Run prompts across multiple configurations, capture answers, latency, token usage, tool calls, and file reads, then generate a single comparison report.
Open-source AI agent governance and evaluation: consent, human approval, LangGraph loops, graph-assisted retrieval, and a read-only GitHub pilot.
Run AI agent evaluations in your CI/CD pipeline using Calibrate and catch regressions
Agentic Workflow Evaluation: Text Summarization Agent. This project includes an AI agent evaluation workflow using a text summarization model with OpenAI API and Transformers library. It follows an iterative approach: generate summaries, analyze metrics, adjust parameters, and retest to refine AI agents for accuracy, readability, and performance.
Open-source framework for evaluating AI agent performance: task completion rate, accuracy, latency, cost efficiency, and reliability metrics for enterprise workflows.
Specification and documentation for the Evaluation Context Protocol
Reproducible benchmark framework for testing hypotheses about AI coding agents
To associate your repository with the ai-agent-evaluation topic, visit your repo's landing page and select "manage topics."