Every Monday at 12:00 UTC, this project asks the largest AI models, eleven of them at last count, the same question, records what they said and how long they took, and publishes the results as a completely straight-faced industry analyst report.
One word answer.
Live: https://hotdogbenchmark.lol
The models are also asked about a hamburger and a taco, because a benchmark with one question is a demo and a benchmark with three is a research program.
Every question is also asked under two more framings: once with a system prompt that says "A hot dog is a sandwich." and once with one that says it is not. The report records how far each model's answer moved when it was told the answer. That property, suggestibility under instruction, generalizes to every real evaluation; the sandwich question does not.
Planned and built with n-dx, En Dash's AI-powered development toolkit. An En Dash Consulting research program.
The question is deliberately silly. Nothing else about it is.
Building a real cross-provider LLM benchmark means solving a specific set of unglamorous problems: seven vendors with seven wire formats, token counts that do not mean the same thing across providers, latency that depends on where your runner happens to be, models that answer differently on Tuesday than they did on Monday, and one provider having an outage in the middle of your data collection.
This repository solves those in the smallest honest way and documents how. Throw away the hot
dog, change questions.json, and you have a benchmark.
One finding, as a sample of what is in here. A live call asking one model the hot dog question returned 647 prompt tokens, 1 completion token, and a billed total of 1,295 — the difference being 647 reasoning tokens counted outside the completion count. Deriving the total as input plus output, the obvious implementation, would have understated that call by half. The schema stores the vendor's own total because of it. That kind of thing is what the tutorial is about.
No API keys required. Mock mode replays recorded provider responses, so the whole pipeline runs offline.
git clone https://github.com/en-dash-consulting/hotdogbenchmark.git
cd hotdogbenchmark
nvm use && npm install
npm run bench -- run --mock --out tmp/mock-run.json # ask every model, from recorded fixtures
npm run dev # serve the report siteOpen http://localhost:4321. That is the whole loop, and it takes about two minutes on a fresh
clone. The site renders the committed edition in data/runs/, which is real data; the mock run
proves the pipeline works on your machine without touching it. Mock mode refuses to overwrite a
real edition — on a fork with no data yet, drop --out and the mock run becomes the site's
first edition, clearly labeled as sample data.
Everything downstream of the network call is real: answer classification, aggregation, cost
estimation, schema validation, and the site build. Set BENCH_SEED=1 to make mock timings
deterministic.
That includes the report itself. No model writes any of the prose. The executive summary, the
key findings, the vendor profiles and the scores are pure functions over the committed edition,
evaluated at build time, so the same data always renders the same words and building the site
needs no API key at all. Ask a different question and the report follows, because it is derived
from questions.json rather than authored.
npm run bench -- run --dry-run # print the plan without calling anythingcp .env.example .env # fill in whatever keys you have
npm run bench -- providers # which keys are configured (never prints a key)
npm run bench:smoke -- --provider anthropic # one live call: text, usage, timing
npm run bench -- run # the real thingMissing keys are skipped with a warning, not recorded as failures, so a partial key set still
produces a usable report. Per-provider setup, free tiers and real costs are in
docs/providers.md — the whole benchmark runs for cents a month.
You do not need to. docs/self-hosting.md takes you from fork to live
site in about fifteen minutes: enable Pages, add secrets, done. A scheduled GitHub Action runs the
benchmark, commits the data, and redeploys the site.
docs/fork-this.md is the longer walkthrough: fork, replace
questions.json, rewrite the framings, record fixtures, run a real edition, and deploy it. The
running example is "Is a burrito a sandwich?" because the point is that the hot dog is
replaceable.
questions.json ─┐
├─► runner ──► data/runs/<iso-week>.json ──► Astro build ──► GitHub Pages
models.json ────┘ │ ▲
▼ │
provider adapters data/index.json
(one file per vendor)
questions.jsonholds the questions. Adding one is a data change. The schema enforces that every question ends withOne word answer., which is what makes the compliance metric mean anything.models.jsonholds the models. Every entry records the docs page its ID was verified against and the date its pricing was read.- Provider adapters (
src/providers/) turn each vendor's API into one small shape. They receive credentials andfetchby injection and are forbidden by lint from importing Node builtins, so the same code can run in a browser. - The runner (
src/runner/) asks every model every question three times, with bounded concurrency and never more than one in-flight call per provider — otherwise the benchmark ends up measuring its own rate limiting. data/runs/stores one versioned JSON file per ISO week. Re-running a week corrects it rather than duplicating it.- The site reads
data/at build time and emits static HTML. The only client JavaScript is the answer-board replay and the framing explorer, a few kilobytes against a 30 KB budget; every page works with scripts off.
Longer version: docs/tutorial/ — eight pages, each mapping a concept to the
file that implements it.
| Tutorial | Build a benchmark like this, in eight steps |
| Self-hosting | Fork to live site in fifteen minutes |
| DNS and hosting | Records, HTTPS, verification, and the traffic plan |
| Fork this | Fork to live site asking your own question |
| Providers | Keys, free tiers, rate limits, real costs |
| Data schema | Every field, its units, and why it may be null |
| Usage normalization | Why token counts are not comparable across vendors |
| Accessibility | What was verified, how, and what still needs a human |
| Contributing | Adding a question, a model, or a provider |
| Security | How API keys are handled |
| Path | What lives there |
|---|---|
site.json |
The site's name, publisher, repository and contact route: a fork changes this file |
questions.json |
The questions, their framings, and the copy the site derives from them |
models.json |
The models: provider, id, pricing, and whether each is enabled |
src/providers/ |
One adapter per vendor, no Node imports, credentials by injection |
src/runner/ |
Asks every model every question under every framing; classifies and aggregates |
src/schema/ |
Zod schemas for runs, questions, models, conditions; the data contract |
src/data/ |
Reading, migrating and indexing data/runs/ |
src/cli/ |
The bench command: run, smoke, record, providers, init |
src/site/ |
The Astro site: pages, components, styles, and the SEO and prose helpers |
data/runs/ |
One JSON file per edition; superseded editions kept under superseded/ |
scripts/ |
Build-time renderers (OG cards, PDFs, screenshots) and the audit scripts |
tests/ |
Vitest suites, including the ones that build the site and drive it in a browser |
docs/ |
Tutorial, provider setup, self-hosting, DNS, data schema, accessibility |
proxy/ |
A Cloudflare Worker for the deferred "run your own" page; not deployed by default |
.rex/, .hench/ |
Product requirements and work records kept by n-dx |
One file. Adapters stay under about 150 lines because they are tutorial examples before they are
infrastructure. Start by reading
src/providers/anthropic.ts, then see
CONTRIBUTING.
nvm use # Node version is pinned in .nvmrc
npm install| Script | What it does |
|---|---|
npm run dev |
Serve the site locally with live reload |
npm run build |
Build the site, OG images and PDF editions into dist/ |
npm run bench |
Run the benchmark (-- --help for usage) |
npm run bench:smoke |
One live call to one provider |
npm run bench:record |
Capture fresh mock fixtures from a provider |
npm run data:validate |
Check every file under data/ against the schema |
npm run data:index |
Regenerate data/index.json |
npm test |
Vitest unit and integration suite |
npm run test:a11y |
axe-core over every built page, both themes |
npm run test:audit |
Keyboard, focus, 320px reflow, zoom, forced colors |
npm run test:responsive |
Overflow, pointer targets, text spacing at seven widths |
npm run test:budget |
Client JavaScript size budget |
npm run lint |
ESLint |
npm run typecheck |
tsc --noEmit |
npm run validate |
lint + typecheck + test, in one command |
- typescript — strict types. Also why
tsconfig.jsonsetserasableSyntaxOnly: Node runs these.tsfiles directly by stripping types rather than compiling them. - eslint, @eslint/js, typescript-eslint — lint, plus the load-bearing rule keeping
src/providersandsrc/runnerfree of Node builtins andprocess.env. - prettier, prettier-plugin-astro — one formatting answer, no debate.
- vitest — fast unit tests, no configuration.
- astro — static site, zero client JS by default. The sitemap, robots.txt and llms.txt are endpoints of our own, built from one page list.
- playwright, @axe-core/playwright — accessibility checks, the PDF edition, OG cards and README screenshots, all from one browser rather than four tools.
- @lhci/cli — Lighthouse budgets in CI.
- yaml — parsing workflows and issue forms in tests, so a malformed one fails locally.
- zod — the schema that everything else trusts.
- @types/node — types for the Node APIs in the CLI and build scripts.
Beyond the usual: that the pull-request workflow references no secrets, that no file under
src/providers reads process.env, that every committed fixture is free of key-shaped strings,
that regenerating data/index.json produces no diff, that every color pair meets its WCAG
ratio, that no page ships an emoji, and that the client JavaScript budget is not exceeded.
The launch test also holds a content floor: with scripts and styles stripped, the front page
and every report must still carry the model names, their answers and a few hundred words. The
manual equivalent, against the live site, is curl -s https://hotdogbenchmark.lol/ | wc -w,
which is what a crawler sees before any JavaScript runs.
Read CONTRIBUTING.md. Bug reports, new providers, and screenshots of models being weird are all welcome.
MIT — see LICENSE. The data under data/ is published under the same terms; if you
cite the Hotdog Benchmark in your own work, that is between you and your
conscience.

