From cd00305eb270d9f204abe1e616b090598ab59de8 Mon Sep 17 00:00:00 2001 From: Roger Chappel Date: Wed, 15 Jul 2026 13:23:13 +1000 Subject: [PATCH 1/2] docs: add tool expansion review demo --- README.md | 10 ++++ demo/run-tool-expansion-review.sh | 38 +++++++++++++ docs/tutorials/tool-expansion-quality-gate.md | 54 +++++++++++++++++++ 3 files changed, 102 insertions(+) create mode 100755 demo/run-tool-expansion-review.sh create mode 100644 docs/tutorials/tool-expansion-quality-gate.md diff --git a/README.md b/README.md index ab387b1..46bdccd 100644 --- a/README.md +++ b/README.md @@ -107,6 +107,16 @@ npm run smoke For a reviewer-facing walkthrough, see [`docs/tutorials/review-agent-tool-expansion.md`](docs/tutorials/review-agent-tool-expansion.md). It demonstrates a prompt revision that expands browser and shell tool language, removes an explicit secret-handling guardrail, and changes the output contract. +For a one-command local demo that writes Markdown and JSON review artifacts: + +```bash +bash demo/run-tool-expansion-review.sh +``` + +See [`docs/tutorials/tool-expansion-quality-gate.md`](docs/tutorials/tool-expansion-quality-gate.md) +for the quality-gate flow and [`docs/promo/social-hooks.md`](docs/promo/social-hooks.md) +for grounded launch copy. + ## Development ```bash diff --git a/demo/run-tool-expansion-review.sh b/demo/run-tool-expansion-review.sh new file mode 100755 index 0000000..97fd21a --- /dev/null +++ b/demo/run-tool-expansion-review.sh @@ -0,0 +1,38 @@ +#!/usr/bin/env bash +set -euo pipefail + +repo_root="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" +cd "$repo_root" + +npm run build >/dev/null + +rm -rf .tmp/promptdiff-demo +mkdir -p .tmp/promptdiff-demo + +node dist/cli.js compare \ + examples/prompts/tool-expansion-old.md \ + examples/prompts/tool-expansion-new.md \ + --out .tmp/promptdiff-demo/tool-expansion.md + +node dist/cli.js compare \ + examples/prompts/tool-expansion-old.md \ + examples/prompts/tool-expansion-new.md \ + --format json \ + --out .tmp/promptdiff-demo/tool-expansion.json + +set +e +node dist/cli.js check examples/prompts/*.md --rules examples/rules.json --fail-on high \ + --out .tmp/promptdiff-demo/rules-check.md +gate_code=$? +set -e + +if [ "$gate_code" -ne 2 ]; then + printf 'expected rules quality gate to exit 2, got %s\n' "$gate_code" >&2 + exit 1 +fi + +grep -q "Tool" .tmp/promptdiff-demo/tool-expansion.md +grep -q "highestSeverity" .tmp/promptdiff-demo/tool-expansion.json +grep -q "PromptDiff" .tmp/promptdiff-demo/rules-check.md + +printf 'Wrote PromptDiff demo artifacts under .tmp/promptdiff-demo\n' diff --git a/docs/tutorials/tool-expansion-quality-gate.md b/docs/tutorials/tool-expansion-quality-gate.md new file mode 100644 index 0000000..040de64 --- /dev/null +++ b/docs/tutorials/tool-expansion-quality-gate.md @@ -0,0 +1,54 @@ +# Tool Expansion Quality Gate Demo + +This walkthrough creates reviewer-facing artifacts for a prompt change that +expands tool language and then runs the checked-in rules gate. + +## Build the CLI + +```sh +npm run build +``` + +## Compare the prompt revision + +```sh +node dist/cli.js compare \ + examples/prompts/tool-expansion-old.md \ + examples/prompts/tool-expansion-new.md \ + --out .tmp/promptdiff-demo/tool-expansion.md +``` + +Write JSON for automation: + +```sh +node dist/cli.js compare \ + examples/prompts/tool-expansion-old.md \ + examples/prompts/tool-expansion-new.md \ + --format json \ + --out .tmp/promptdiff-demo/tool-expansion.json +``` + +## Run the rules gate + +```sh +node dist/cli.js check examples/prompts/*.md --rules examples/rules.json --fail-on high \ + --out .tmp/promptdiff-demo/rules-check.md +``` + +Exit code `2` means the configured prompt quality gate failed. In this demo, +that is expected evidence for the risky fixture set, not a runtime error. + +## One-command demo + +```sh +bash demo/run-tool-expansion-review.sh +``` + +The script writes Markdown and JSON artifacts, verifies expected report text, +and confirms the rules gate exits with code `2`. + +## Boundaries + +- PromptDiff is deterministic and heuristic; it is not an LLM judge. +- Reports are review artifacts, not final safety decisions. +- Do not claim hosted scanning, telemetry, or network behavior. From 8dfdf1bbb930cb66c4ad21f7ef07d202d94eb427 Mon Sep 17 00:00:00 2001 From: Roger Chappel Date: Wed, 15 Jul 2026 13:23:13 +1000 Subject: [PATCH 2/2] docs: add promptdiff promotion hooks --- docs/promo/social-hooks.md | 40 +++++++++++++++++++------------------- 1 file changed, 20 insertions(+), 20 deletions(-) diff --git a/docs/promo/social-hooks.md b/docs/promo/social-hooks.md index d61dc21..28c759c 100644 --- a/docs/promo/social-hooks.md +++ b/docs/promo/social-hooks.md @@ -1,26 +1,26 @@ -# Social Hooks +# PromptDiff Promotion Hooks -These drafts are grounded in the current README, examples, and CLI behavior. +## Grounded facts -## Prompt Tool Review +- PromptDiff compares prompt revisions from local files. +- It emits Markdown or JSON reports. +- It redacts common secret-like values by default. +- It includes a `check` command backed by a JSON rules file. +- It is deterministic and heuristic, not an LLM judge. -Prompt edits can quietly change tool access, safety language, and output contracts. +## Short posts -PromptDiff gives those changes names in a local Markdown or JSON report, so reviewers can discuss the actual risk instead of eyeballing a wall of text. +1. Prompt changes can expand tool access without looking dramatic. PromptDiff + turns that revision into a local review artifact. +2. Treat prompts like code: compare revisions, name risky categories, and keep + a Markdown report with the PR. +3. Demo angle: old support prompt, new tool-expanded prompt, one report showing + what changed and one rules gate that exits non-zero. -Demo: compare `examples/prompts/tool-expansion-old.md` with `examples/prompts/tool-expansion-new.md`. +## Video outline -## CI Angle - -PromptDiff has two useful modes: - -- `compare` for prompt revision reports -- `check` for required phrases, forbidden phrases, and section rules - -It is deterministic, local-first, and built for review evidence rather than scoring prompts with another model. - -## Limitation-Aware Post - -PromptDiff is not an LLM judge and does not claim to understand every semantic change. - -That is the point: it catches concrete review signals such as risky instruction language, removed guardrails, tool references, output-contract shifts, and secret-like values. +1. Open `examples/prompts/tool-expansion-old.md`. +2. Open `examples/prompts/tool-expansion-new.md`. +3. Run `bash demo/run-tool-expansion-review.sh`. +4. Show `.tmp/promptdiff-demo/tool-expansion.md`. +5. Show the expected quality-gate exit behavior from the script.