Skip to content
View caleb-love's full-sized avatar
😄
You're great.
😄
You're great.

Block or report caleb-love

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
caleb-love/README.md

Caleb Love

Building infrastructure for autonomous AI agents
Solutions Engineer

caleb-love.com · LinkedIn


I work on the systems that make AI agents reliable in production — tracing, evaluation, cost controls, regression detection. Not the demo, the infrastructure underneath it.

By day I'm a Solutions Engineer at Pendo, running technical evaluations and building proof-of-concepts for enterprise teams. By night I'm deep in agentic AI, building the tooling that turns autonomous agents from impressive demos into things you can actually trust.


Currently

  • 🦊 Shipping eval + drift detection for Foxhound
  • 🧪 Pressure-testing multi-agent orchestration patterns across LangGraph, CrewAI, and Claude Agent SDK
  • 📝 Writing up field notes on what actually breaks when agents go to production
  • 🤝 Open to chats about agent observability, evals, and SE work in AI

Foxhound

Compliance-grade observability for AI agent fleets.

Trace every decision. Evaluate every response. Budget every dollar.

GitHub Stars Last commit MIT License Docs

Foxhound is an open-source observability platform I built for teams running AI agents in production. Generic APM tools treat agent workflows as black boxes — Foxhound traces the full decision tree so you can see exactly what happened, why, and where it broke.

What it does:

  • Tracing — Structured span trees for every agent run, session replay to reconstruct state at any point, run diffs to compare executions side-by-side
  • Evaluation — LLM-as-judge evaluators, dataset curation from production traces, experiment runner with A/B comparison
  • Cost & SLA — Per-agent budgets with SDK callbacks, SLA monitoring, automated alerting across Slack, PagerDuty, Linear, and webhooks
  • Regression Detection — Behavioral baselines per agent version, automated drift detection across deployments
  • Prompt Management — Versioned templates with label-based promotion from staging to production
  • Quality Gate — GitHub Actions integration that blocks PRs failing eval thresholds

Architecture:

11-package TypeScript monorepo · Fastify API (~80 endpoints) · PostgreSQL + Redis
7 background job queues · Python & TypeScript SDKs · 37 MCP debugging tools

Auto-instrumentation for: LangGraph · CrewAI · Pydantic AI · OpenAI Agents · Claude Agent SDK · Google ADK · AWS Bedrock AgentCore · Mastra · OpenTelemetry


Other work

  • Dex — Open-source personal knowledge system built on Claude. Active contributor, daily user, testbed for agentic patterns.
  • OpenClaw — Exploring how different agent frameworks handle orchestration, tool use, and long-running autonomous tasks.
  • MCP servers & integrations — Connecting AI agents to the tools people actually use: Slack, calendars, CRMs, project management. Making agents do more than generate text.

Lessons from the field

Clarity > cleverness Agents don't read between the lines. Vague specs produce vague output.
Fluent ≠ correct AI fails confidently. The skill is catching what looks right but isn't.
Orchestration is design Multi-agent systems only work with right-sized tasks and clean handoffs.
Know the failure modes Context degradation, spec drift, silent failures. You learn fast once burned.
Context is the unlock Right information, right agent, right time. That's what separates demos from systems.
Not everything needs an agent Sometimes a script is the right call. Half the job is knowing the economics.

Stack

Languages TypeScript Python SQL
Backend Node.js Fastify PostgreSQL Redis Drizzle
Frontend Next.js React Tailwind
Infra Fly.io Cloudflare Turborepo GitHub Actions
AI Claude OpenAI LangGraph MCP

Also — I help companies with AI

Outside of Pendo, I consult and build under my own name — Caleb Love. If you need someone to make AI actually work for your business (strategy, building agents, integrations, or just an outside technical read on what's real vs. hype), that's what I do.

caleb-love.com


Let's talk

Always up for a conversation about agent infrastructure, evals, or what enterprise AI actually looks like once the demo is over.

LinkedIn Website

Popular repositories Loading

  1. servo_app servo_app Public

    First project for team InitToWinIt - Caleb, Mohammed, Nikki & Simon

    JavaScript 3

  2. react-flights react-flights Public

    Forked from BigBBazz/react-flights

    JavaScript

  3. caleb-love caleb-love Public

  4. MSSQLEXPRESS-M1-Install MSSQLEXPRESS-M1-Install Public

    Forked from jimm98y/MSSQLEXPRESS-M1-Install

    Installers and installation scripts for Microsoft SQL Server Express on ARM64.

    Batchfile

  5. pendo-snippets pendo-snippets Public

    Forked from pendo-io/snippets

    AS IS WITH NO WARRANTY OR SLA. Generic Tools, Libraries and Snippets created for use by Techincal Customer Success and Sales Engineering.

    JavaScript

  6. vibe-code-camp vibe-code-camp Public

    Forked from EveryInc/vibe-code-camp

    The full transcript of our marathon livestream with the world's best vibe coders.