Skip to content

feat(libsy): route capability decisions by relative advantage - #901

Merged
nachiketb-nvidia merged 3 commits into
mainfrom
nachiketb/switch-1687-decision-judge
Oct 2, 2026
Merged

nachiketb-nvidia merged 3 commits into
mainfrom
nachiketb/switch-1687-decision-judge

Conversation

@nachiketb-nvidia

@nachiketb-nvidia nachiketb-nvidia commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

What

Add a provider-neutral decision judge to the existing capability router. It asks whether the capable candidate will succeed where the efficient candidate will fail, then routes capable only when that event's Choice score is strictly above a configured cutoff. Equality routes efficient.

Why

This follows the relative-advantage policy used in the Jev task-routing experiment. The score describes capable-only success, not either model's independent solve probability. It therefore has its own policy instead of using the LLM judge's p_solve and capability-boundary thresholds.

Implements SWITCH-1687.

How

  • Select CapabilityJudgeConfig::Decision(DecisionJudgeConfig { ... }) inside the existing capability mode. Reuse conversation selection, affinity, classification frequency, and fallback.
  • Send one typed decision call with structured setting/evidence/comparison/boundary/policy instructions and two Choice options: advantage and its complement, no_advantage.
  • Context contains the selected conversation, candidate labels, the compared pair, and supplied evidence. A candidate-label-to-target map keeps transport IDs out of the generated candidate list. Collections can hold N candidates; this policy compares only the first runtime capable and efficient targets.
  • Read the advantage entry from the Choice distribution. Missing, wrongly typed, nonfinite, or out-of-range scores abstain. Provider errors respect fail_open; configuration mapping errors remain errors.
  • Preserve unknown reference outcomes and supplied evidence. Emit bounded score/cutoff evidence and reuse redacted failure reporting. Provider confidence and the provider's selected option do not replace the application's cutoff.

Rust configuration, inside TaskClassifierConfig:

judge: CapabilityJudgeConfig::Decision(DecisionJudgeConfig {
    cutoff: evaluated_cutoff,
    instructions: None, // packaged relative-advantage instructions
    candidates: BTreeMap::from([
        ("a".into(), capable_target),
        ("b".into(), efficient_target),
    ]),
    evidence: supplied_evidence,
}),

evidence can contain candidate descriptions, independent reference outcomes, capability/cost summaries, and selection notes keyed by the candidate labels. It is passed unchanged to the judge and interpreted using the instructions; the router does not parse its fields. Callers supply this data; the router does not derive training evidence or calibrate the cutoff.

Validation

  • One new table-driven test covers request shape, three-candidate evidence, instruction overrides, strict cutoff/equality, absent or invalid scores, provider confidence separation, error recovery/propagation, dropped replies, missing target mappings, and retained routing.
  • All 27 focused capability tests passed, including the existing LLM prompt, policy, config, and affinity coverage.
  • Four existing composite tests and three stage-judge tests passed. Formatting, affected-crate Clippy with warnings denied, and strict libsy Rustdoc passed.
  • No live Jev calls or full local test suite. The production client still needs the separate Jev transport work.

Notes for reviewers

Based on main, following the merged refactor in #900. Start with llm_class/decision.rs, then the small capability-construction change and its single test.

No protocol changes or new public algorithm. Runner target wiring, Python decision support, N-model selection, and transport metrics remain separate work.

Diff: five files, +479 / −20 lines. Added Rust lines include 199 functionality lines and 238 test lines; the remainder is comments, spacing, and the prompt asset.

Summary by CodeRabbit

  • New Features
    • Capability routing can now use a decision-based judge to compare capable and efficient candidates using task evidence and a configurable cutoff.
    • When the judge returns a valid advantage score, routing selects the corresponding candidate. Missing or invalid scores can produce an ambiguous result, while judge errors follow the configured fail-open behavior.
    • Existing LLM-based capability judging remains available alongside the new decision-based option.

@github-actions

github-actions Bot commented Oct 1, 2026 •

Copy link
Copy Markdown
PR Preview Action v1.8.1

🚀 View preview at
https://NVIDIA-NeMo.github.io/Switchyard/pr-preview/pr-901/

Built to branch gh-pages at 2026-10-02 17:45 UTC.
Preview will be ready when the GitHub Pages deployment is complete.

Base automatically changed from nachiketb/switch-1687-judge-config to main October 1, 2026 23:08
@coderabbitai

coderabbitai Bot commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository: NVIDIA-NeMo/Switchyard/.coderabbit.yaml

Review profile: CHILL

Plan: Enterprise

Run ID: 8b6ab86f-a501-46bc-8ed2-7da3a4b8670e

📥 Commits

Reviewing files that changed from the base of the PR and between 16cbe59 and e645303.

📒 Files selected for processing (10)
  • crates/libsy-llm-client/tests/observability.rs
  • crates/libsy/src/algorithms/composite.rs
  • crates/libsy/src/algorithms/llm_class.rs
  • crates/libsy/src/algorithms/llm_class/decision.rs
  • crates/libsy/src/algorithms/stage.rs
  • crates/libsy/src/algorithms/util/llm_judge.rs
  • crates/libsy/src/lib.rs
  • crates/libsy/src/prompts/capability-classifier/relative_advantage.json
  • crates/switchyard-py/src/libsy_bindings.rs
  • crates/switchyard-runner/src/algorithm.rs

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 9 remain after this review.


Walkthrough

The capability classifier now supports LLM and Decision judges through a judge-selection configuration. The Decision judge compares capable and efficient candidates and applies a probability cutoff. Runner and Python bindings construct the updated LLM configuration.

Changes

Capability classifier judge selection

Layer / File(s) Summary
Judge configuration and classifier construction
crates/libsy/src/algorithms/llm_class.rs, crates/libsy/src/lib.rs, crates/libsy/src/algorithms/llm_class/decision.rs
TaskClassifierConfig now selects an LLM or Decision judge. LLM settings use LlmCapabilityConfig. Validation and classifier construction support both judge types.
Decision judge scoring and fail-open behavior
crates/libsy/src/algorithms/llm_class/decision.rs, crates/libsy/src/algorithms/util/llm_judge.rs, crates/libsy/src/prompts/capability-classifier/relative_advantage.json, crates/libsy/src/algorithms/llm_class.rs
DecisionClassifier compares candidate models, applies the cutoff, and records classification evidence. Tests cover scores, invalid responses, failures, and routing behavior.
LLM configuration integrations and regression tests
crates/switchyard-runner/src/algorithm.rs, crates/switchyard-py/src/libsy_bindings.rs, crates/libsy/src/algorithms/stage.rs, crates/libsy/src/algorithms/composite.rs, crates/libsy/src/algorithms/llm_class.rs, crates/libsy-llm-client/tests/observability.rs
Runner and Python binding code now creates nested LLM judge settings. Existing configuration and routing tests use the updated shape.

Priority: ⬇️ Low

Estimated code review effort: 3 (Moderate) | ~25 minutes

Merge Risk: ⚪ Minimal · up to e6453

No actionable merge-blocking risk is identified. Production use of the Decision judge still requires the separately planned transport integration; existing integrations continue selecting the LLM judge.

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 35.71% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 42 functions across 9 files. (1 skipped: … Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and concisely describes the main change: adding relative-advantage routing for capability decisions in libsy.
Full details: Docstring Coverage

Explanation

Docstring coverage is 35.71% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 42 functions across 9 files. (1 skipped: 1 unsupported.)

  • Fix all pre-merge checks with AI
✨ Finishing Touches 💡 1
⚔️ Resolve merge conflicts 💡
  • Resolve merge conflict in branch nachiketb/switch-1687-decision-judge
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Autopilot is currently an internal CodeRabbit preview.


A rabbit checks the scores in view
And nests the judge settings anew
The capable path may win the race
If its score clears the cutoff’s place
The efficient route may guide the way
Then hops the rabbit off to play

Comment @coderabbitai help to get the list of available commands.

Signed-off-by: nachiketb <nachiketb@nvidia.com>
@nachiketb-nvidia
nachiketb-nvidia force-pushed the nachiketb/switch-1687-decision-judge branch from e645303 to b394b6e Compare October 1, 2026 23:15
@grahamking

Copy link
Copy Markdown
Contributor

The implementation is focused, reuses existing routing behavior, and covers the important failure cases.

Astra upvote.

Comment thread crates/libsy/src/algorithms/llm_class/decision.rs Outdated
Comment thread crates/libsy/src/algorithms/llm_class/decision.rs
Comment thread crates/libsy/src/algorithms/llm_class/decision.rs
Comment thread crates/libsy/src/algorithms/llm_class/decision.rs
Comment thread crates/libsy/src/algorithms/llm_class/decision.rs
Signed-off-by: nachiketb <nachiketb@nvidia.com>
Signed-off-by: nachiketb <nachiketb@nvidia.com>
@nachiketb-nvidia
nachiketb-nvidia merged commit 4177133 into main Oct 2, 2026
18 checks passed
@nachiketb-nvidia
nachiketb-nvidia deleted the nachiketb/switch-1687-decision-judge branch October 2, 2026 18:38
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants