RTX Recon improvements - #8
Conversation
… spec, doctrine + README refresh - R1: evals/ golden task set (20 tasks), pass@k scorer, LLM-as-judge skill, runbook - R2: skills/ — ecosystem templates repackaged as Agent Skills (SKILL.md) - R3: agent-harness/ — Ralph-loop initializer/coder/evaluator harness - R4: Vision + comparison rewritten around agentic/loop/context engineering - R5: Commander/Soldier multi-agent doctrine + fan-out cost warning - R6: Romanized Tax tokenizer nuance + Hinglish hallucination gates - R7: spec-kit-bridge/ conversion layer + OpenSpec delta template - R8: memory-layer-spec.md (L1-L4) - R9: Evidence-Before-Completion rule in Precision Protocol - R10: Model Matrix reset to 2026 models, all scores UNEVALUATED (no invented numbers) - .gitignore: whitelist new paths
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. Note Currently processing new changes in this PR. This may take a few minutes, please wait... ⚙️ Run configuration
📒 Files selected for processing (19)
✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
PR Summary by QodoAdd RTX evaluation kit, agent harness, skills, and protocol guidance
AI Description
Diagram
High-Level Assessment
Files changed (19)
|




⚡ RTX Recon — all 10 recommendations implemented
This PR implements every actionable recommendation from the Operation RTX Recon report, built by 4 parallel workers and assembled + reviewed by the commander.
What shipped (per recommendation)
Tier 1
evals/): 20 golden tasks (golden-tasks.yaml), workingpass@kscorer (score.py, self-test passes), LLM-as-judge Agent Skill (evals/llm-judge/SKILL.md, rubrics R1–R7), andRUNBOOK.mdwith the official matrix (re-)scoring methodology. All score fields deliberately empty — unevaluated, run the kit.skills/): the 4 ecosystem templates repackaged asSKILL.md(progressive disclosure, trigger-oriented descriptions). Legacytemplates/untouched.agent-harness/): workinginit.sh(verified live),PROMPT.md,PROGRESS-template.md,EVALUATOR.md(third-role, anti-self-grading, hard caps: 10 iters/run, 3 retries/feature), worked example.Tier 2
spec-kit-bridge/): Precision Protocol ↔ Spec-Kit conversion layer + OpenSpecADDED/MODIFIED/REMOVEDdelta template.memory-layer-spec.md): L1 session → L2 bounded profile → L3 episodic → L4 workflow memory, anti-pattern: raw transcript replay.Tier 3
UNEVALUATED — run evals/ kit. No invented numbers.Also:
.gitignorewhitelists the new paths (repo uses whitelist-style ignores).Deliberately deferred
👀 Needs your eye
rtx-cursoretc.) — rename to capability names if you prefer.Verification
python3 evals/score.py --self-test→ ALL PASSEDagent-harness/init.sh→ live scaffold + edge cases verified