feat(libsy): route capability decisions by relative advantage - #901
Conversation
|
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Repository: NVIDIA-NeMo/Switchyard/.coderabbit.yaml Review profile: CHILL Plan: Enterprise Run ID: 📒 Files selected for processing (10)
Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 9 remain after this review. WalkthroughThe capability classifier now supports LLM and Decision judges through a judge-selection configuration. The Decision judge compares capable and efficient candidates and applies a probability cutoff. Runner and Python bindings construct the updated LLM configuration. ChangesCapability classifier judge selection
Priority: ⬇️ Low Estimated code review effort: 3 (Moderate) | ~25 minutes Merge Risk: ⚪ Minimal · up to No actionable merge-blocking risk is identified. Production use of the Decision judge still requires the separately planned transport integration; existing integrations continue selecting the LLM judge. 🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
Full details: Docstring CoverageExplanation Docstring coverage is 35.71% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 42 functions across 9 files. (1 skipped: 1 unsupported.)
✨ Finishing Touches 💡 1⚔️ Resolve merge conflicts 💡
A rabbit checks the scores in view Comment |
Signed-off-by: nachiketb <nachiketb@nvidia.com>
e645303 to
b394b6e
Compare
Astra upvote. |
Signed-off-by: nachiketb <nachiketb@nvidia.com>
Signed-off-by: nachiketb <nachiketb@nvidia.com>
What
Add a provider-neutral decision judge to the existing capability router. It asks whether the capable candidate will succeed where the efficient candidate will fail, then routes capable only when that event's Choice score is strictly above a configured cutoff. Equality routes efficient.
Why
This follows the relative-advantage policy used in the Jev task-routing experiment. The score describes capable-only success, not either model's independent solve probability. It therefore has its own policy instead of using the LLM judge's
p_solveand capability-boundary thresholds.Implements SWITCH-1687.
How
CapabilityJudgeConfig::Decision(DecisionJudgeConfig { ... })inside the existing capability mode. Reuse conversation selection, affinity, classification frequency, and fallback.advantageand its complement,no_advantage.advantageentry from the Choice distribution. Missing, wrongly typed, nonfinite, or out-of-range scores abstain. Provider errors respectfail_open; configuration mapping errors remain errors.Rust configuration, inside
TaskClassifierConfig:evidencecan contain candidate descriptions, independent reference outcomes, capability/cost summaries, and selection notes keyed by the candidate labels. It is passed unchanged to the judge and interpreted using the instructions; the router does not parse its fields. Callers supply this data; the router does not derive training evidence or calibrate the cutoff.Validation
Notes for reviewers
Based on
main, following the merged refactor in #900. Start withllm_class/decision.rs, then the small capability-construction change and its single test.No protocol changes or new public algorithm. Runner target wiring, Python decision support, N-model selection, and transport metrics remain separate work.
Diff: five files, +479 / −20 lines. Added Rust lines include 199 functionality lines and 238 test lines; the remainder is comments, spacing, and the prompt asset.
Summary by CodeRabbit