Skip to content

RTX Recon improvements - #8

Merged
PsProsen-Dev merged 1 commit into
masterfrom
rtx-recon-improvements
Oct 3, 2026
Merged

PsProsen-Dev merged 1 commit into
masterfrom
rtx-recon-improvements

Conversation

@PsProsen-Dev

Copy link
Copy Markdown
Owner

⚡ RTX Recon — all 10 recommendations implemented

This PR implements every actionable recommendation from the Operation RTX Recon report, built by 4 parallel workers and assembled + reviewed by the commander.

What shipped (per recommendation)

Tier 1

  • R1 — Eval Kit (evals/): 20 golden tasks (golden-tasks.yaml), working pass@k scorer (score.py, self-test passes), LLM-as-judge Agent Skill (evals/llm-judge/SKILL.md, rubrics R1–R7), and RUNBOOK.md with the official matrix (re-)scoring methodology. All score fields deliberately empty — unevaluated, run the kit.
  • R2 — Agent Skills (skills/): the 4 ecosystem templates repackaged as SKILL.md (progressive disclosure, trigger-oriented descriptions). Legacy templates/ untouched.
  • R3 — Ralph-loop harness (agent-harness/): working init.sh (verified live), PROMPT.md, PROGRESS-template.md, EVALUATOR.md (third-role, anti-self-grading, hard caps: 10 iters/run, 3 retries/feature), worked example.
  • R4 — Industry vocabulary: Vision + RTX-vs-Alternatives rewritten around agentic engineering / loop engineering / context engineering.
  • R5 — Multi-agent doctrine: Commander/Soldier rule in Architecture Flow (fan-out reads, single-threaded writes) + the 48%-of-bill cost warning.

Tier 2

  • R6 — Romanized Tax nuance: 3–5× labeled tokenizer-specific (GPT-4o o200k: Devanagari now cheaper than romanized) + Hinglish hallucination gates (stronger verification for native sessions) + open research gap note.
  • R7 — Spec-Kit bridge (spec-kit-bridge/): Precision Protocol ↔ Spec-Kit conversion layer + OpenSpec ADDED/MODIFIED/REMOVED delta template.
  • R8 — Memory layer spec (memory-layer-spec.md): L1 session → L2 bounded profile → L3 episodic → L4 workflow memory, anti-pattern: raw transcript replay.

Tier 3

  • R9 — Evidence-Before-Completion: review checklists now require artifacts (terminal output, test results, screenshots) — agent self-report is never evidence.
  • R10 — Model Matrix reset: 2026 models (Gemini 3.1 Pro, GPT-5.x, Claude 4.x/5.x, DeepSeek V4-class), every score UNEVALUATED — run evals/ kit. No invented numbers.

Also: .gitignore whitelists the new paths (repo uses whitelist-style ignores).

Deliberately deferred

  • Actual model benchmarking — needs API access to proprietary models; methodology shipped, scores stay UNEVALUATED until someone runs the kit.
  • README.hi.md — Hindi translation needs a separate localization pass (English README only in this PR).

👀 Needs your eye

  • Creator's Story mother-tongue: corrected Bengali → Hindi per your confirmed record (twice-confirmed). Revert if you want the old wording.
  • Skill naming uses tool names (rtx-cursor etc.) — rename to capability names if you prefer.

Verification

  • python3 evals/score.py --self-test → ALL PASSED
  • agent-harness/init.sh → live scaffold + edge cases verified
  • All 5 SKILL.md frontmatter YAML-valid; all README/new-file relative links resolve; no secrets; code fences balanced.

… spec, doctrine + README refresh

- R1: evals/ golden task set (20 tasks), pass@k scorer, LLM-as-judge skill, runbook
- R2: skills/ — ecosystem templates repackaged as Agent Skills (SKILL.md)
- R3: agent-harness/ — Ralph-loop initializer/coder/evaluator harness
- R4: Vision + comparison rewritten around agentic/loop/context engineering
- R5: Commander/Soldier multi-agent doctrine + fan-out cost warning
- R6: Romanized Tax tokenizer nuance + Hinglish hallucination gates
- R7: spec-kit-bridge/ conversion layer + OpenSpec delta template
- R8: memory-layer-spec.md (L1-L4)
- R9: Evidence-Before-Completion rule in Precision Protocol
- R10: Model Matrix reset to 2026 models, all scores UNEVALUATED (no invented numbers)
- .gitignore: whitelist new paths
Copilot AI balanced review requested due to automatic review settings October 3, 2026 10:45

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@qodo-code-review

qodo-code-review Bot commented Oct 3, 2026 •

Copy link
Copy Markdown

Code Review by Qodo

Grey Divider

Sorry, something went wrong

We weren't able to complete the code review on our side. Please try again manually by commenting /agentic_review on this PR.

Grey Divider

Qodo Logo

@PsProsen-Dev
PsProsen-Dev merged commit 655eede into master Oct 3, 2026
0 of 2 checks passed
@coderabbitai

coderabbitai Bot commented Oct 3, 2026

Copy link
Copy Markdown

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

Note

Currently processing new changes in this PR. This may take a few minutes, please wait...

⚙️ Run configuration
  • Configuration used: defaults
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: 98b175a1-f54d-4f1e-9e22-5276823f7256
📥 Commits

Reviewing files that changed from the base of the PR and between c86041c and cb589b4.

📒 Files selected for processing (19)
  • .gitignore
  • README.md
  • agent-harness/EVALUATOR.md
  • agent-harness/PROGRESS-template.md
  • agent-harness/PROMPT.md
  • agent-harness/README.md
  • agent-harness/init.sh
  • evals/RUNBOOK.md
  • evals/golden-tasks.yaml
  • evals/llm-judge/SKILL.md
  • evals/score.py
  • memory-layer-spec.md
  • skills/README.md
  • skills/rtx-claude/SKILL.md
  • skills/rtx-cline/SKILL.md
  • skills/rtx-copilot/SKILL.md
  • skills/rtx-cursor/SKILL.md
  • spec-kit-bridge/CONVERSION.md
  • spec-kit-bridge/openspec-delta-template.md
 ________________________________________________________________________
< Making the Death Star fully operational, with zero exhaust port flaws. >
 ------------------------------------------------------------------------
  \
   \   \
        \ /\
        ( )
      .( o ).
✨ Finishing Touches
📝 Generate docstrings
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@sonarqubecloud

sonarqubecloud Bot commented Oct 3, 2026

Copy link
Copy Markdown

Quality Gate Failed Quality Gate failed

Failed conditions
C Security Rating on New Code (required ≥ A)
C Reliability Rating on New Code (required ≥ A)

See analysis details on SonarQube Cloud

Catch issues before they fail your Quality Gate with our IDE extension SonarQube for IDE

@qodo-code-review

Copy link
Copy Markdown

PR Summary by Qodo

Add RTX evaluation kit, agent harness, skills, and protocol guidance

✨ Enhancement 📝 Documentation ⚙️ Configuration changes 🕐 40+ Minutes

Grey Divider

AI Description

• Add golden tasks, transcript grading, and pass@k scoring so model compatibility can be measured
 rather than guessed.
• Introduce a bounded coder–evaluator harness and tool-specific Agent Skills for repeatable,
 context-conscious workflows.
• Document evidence gates, memory layers, spec conversion, and updated guidance for multilingual and
 multi-agent work.
Diagram

graph TD
  Core["RTX framework"] --> Tasks["Golden tasks"] --> Judge["Transcript judge"] --> Scorer["Pass@k scorer"] --> Matrix["Model matrix"]
  Core --> Harness["Agent harness"]
  Core --> Skills["Agent Skills"]
Loading
High-Level Assessment

The following are alternative approaches to this PR:

1. Fully automated agent runner and grader
  • ➕ Could enforce iteration caps and produce trial records without manual handoffs.
  • ➕ Could make repeated model evaluations easier to reproduce.
  • ➖ Requires model integrations, execution isolation, and substantially more infrastructure.
  • ➖ Would add cost and complexity before the evaluation criteria have been validated.

Recommendation: Keep the lightweight, human-operated harness and hybrid deterministic/judge grading for this first kit. Before publishing scores or treating the harness caps as enforced, validate full runs: the caps are currently instructions and state fields, not an automated controller. Also align the runbook's Python prerequisite with the scorer's float | None syntax and reconcile the older numeric scorecard in evals/EVALUATION-SUITE.md with the new no-unmeasured-scores policy.

Files changed (19) +1839 / -34

Enhancement (6) +375 / -0
PROGRESS-template.mdTemplate the shared run progress record +39/-0

Template the shared run progress record

• Defines the mission, phase, feature status, decisions, issues, and append-only session log read by subsequent coder sessions.

agent-harness/PROGRESS-template.md

init.shScaffold a bounded agent run +99/-0

Scaffold a bounded agent run

• Validates a run name and creates a new directory containing rendered progress, a feature checklist, JSON state, and an evaluation-report folder. Refuses to overwrite an existing run.

agent-harness/init.sh

SKILL.mdPackage Claude Code RTX directives +53/-0

Package Claude Code RTX directives

• Adds trigger-oriented instructions for build/test verification, vertical layout, language blending, and resumable output tracking.

skills/rtx-claude/SKILL.md

SKILL.mdPackage Cline RTX directives +68/-0

Package Cline RTX directives

• Adds the identity and output protocol, Hinglish blend, English-only technical assets, and guarded autonomous execution guidance.

skills/rtx-cline/SKILL.md

SKILL.mdPackage Copilot RTX directives +57/-0

Package Copilot RTX directives

• Adds Copilot-oriented identity, output formatting, language-blend, and technical-content exemption instructions.

skills/rtx-copilot/SKILL.md

SKILL.mdPackage Cursor RTX directives +59/-0

Package Cursor RTX directives

• Adds baseline-reference, vertical-spacing, guarded execution, and language-precision guidance for Cursor sessions.

skills/rtx-cursor/SKILL.md

Tests (1) +293 / -0
golden-tasks.yamlAdd twenty RTX compliance scenarios +293/-0

Add twenty RTX compliance scenarios

• Defines prompts and pass/fail checks covering spec-first work, native-language assertions, review evidence, formatting, drift, boot behavior, and protocol integrity. Contains no measured scores.

evals/golden-tasks.yaml

Documentation (9) +854 / -34
README.mdPresent the new workflows and reset model scores +189/-34

Present the new workflows and reset model scores

• Introduces the eval kit, evidence rule, skills, spec bridge, memory layers, and agent harness. Reframes industry terminology and multilingual guidance, corrects the creator's stated mother tongue, and replaces legacy matrix scores with unevaluated entries.

README.md

EVALUATOR.mdDefine independent live-app evaluation +90/-0

Define independent live-app evaluation

• Specifies a separate evaluator session, evidence-backed verdict reports, stop conditions, and iteration limits intended to prevent coder self-grading.

agent-harness/EVALUATOR.md

PROMPT.mdFix the coder's per-session instructions +58/-0

Fix the coder's per-session instructions

• Directs each coder session to build one feature, verify it with artifacts, update shared state, and hand off to the evaluator.

agent-harness/PROMPT.md

README.mdExplain the initializer–coder–evaluator loop +115/-0

Explain the initializer–coder–evaluator loop

• Provides setup instructions, the on-disk state contract, a worked todo-app example, and the harness's operating rules.

agent-harness/README.md

RUNBOOK.mdDefine reproducible model evaluation +133/-0

Define reproducible model evaluation

• Describes fresh-session trials, transcript retention, grading, results JSON, pass@k interpretation, and the requirements for publishing a matrix score.

evals/RUNBOOK.md

memory-layer-spec.mdSpecify bounded, retrievable agent memory +124/-0

Specify bounded, retrievable agent memory

• Defines ephemeral session context, a capped always-on profile, retrieved episodic facts, and trigger-based workflow lessons. Prescribes shared run state and distilled summaries instead of raw transcript replay.

memory-layer-spec.md

README.mdIndex and explain Agent Skills +43/-0

Index and explain Agent Skills

• Documents installation, intended tool targets, progressive disclosure, and the continued availability of legacy flat templates.

skills/README.md

CONVERSION.mdMap RTX specifications to Spec-Kit +53/-0

Map RTX specifications to Spec-Kit

• Documents conversion in both directions while retaining bilingual acceptance criteria, architecture decisions, task evidence gates, and brownfield deltas.

spec-kit-bridge/CONVERSION.md

openspec-delta-template.mdTemplate brownfield change proposals +49/-0

Template brownfield change proposals

• Requires ADDED, MODIFIED, and REMOVED buckets alongside impact, rollback, test evidence, and reviewer sign-off.

spec-kit-bridge/openspec-delta-template.md

Other (3) +317 / -0
.gitignoreAllow new RTX resources into the whitelist +14/-0

Allow new RTX resources into the whitelist

• Unignores the skills, spec bridge, evaluation assets, agent harness, and memory specification in this whitelist-style repository.

.gitignore

SKILL.mdAdd evidence-quoting transcript rubrics +105/-0

Add evidence-quoting transcript rubrics

• Packages seven grading rubrics as an Agent Skill. Requires applicable pass/fail verdicts with quoted transcript evidence and leaves numeric aggregation to the scorer.

evals/llm-judge/SKILL.md

score.pyCalculate pass@k from recorded trials +198/-0

Calculate pass@k from recorded trials

• Loads the task YAML and results JSON, computes per-task and mean pass@k, and renders missing evaluations as TBD. Includes a built-in scorer and task-set sanity check.

evals/score.py

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants