Conversation
merlerm
added this pull request to stack #232
September 19, 2026 14:53
merlerm
marked this pull request as ready for review
September 19, 2026 14:54
merlerm
force-pushed
the
pr/apptainer-network-audit
branch
from
September 19, 2026 14:57
1ab0ace to
3949017
Compare
merlerm
force-pushed
the
pr/apptainer-full-red-team
branch
from
September 19, 2026 14:57
ffb0114 to
e149ead
Compare
merlerm
force-pushed
the
pr/apptainer-network-audit
branch
from
September 19, 2026 15:22
3949017 to
fa791fb
Compare
merlerm
force-pushed
the
pr/apptainer-full-red-team
branch
from
September 19, 2026 15:22
e149ead to
86fade0
Compare
merlerm
removed this pull request from stack #232
September 19, 2026 15:30
merlerm
force-pushed
the
pr/apptainer-network-audit
branch
from
September 19, 2026 15:31
fa791fb to
c1ebb2a
Compare
Collaborator
Author
|
Consolidated into #230. Retained evidence and full-suite execution now live in the original red_team_sandbox.py; the separate coordinator has been removed. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The existing agent probes did not provide a complete retained Apptainer run across both backends, and some verdicts confused echoed markers or broken requests with evidence.
Add a coordinator that runs every catalog case, keeps per-case raw streams/scripts/broker logs, continues after failures, and distinguishes failed, inconclusive, and passed results. Add strict interpreter/path/source inventories and a rendered-policy import attack. Randomize undisclosed canaries, recognize Codex tool events, require genuine result markers, and verify raw strict commands with fresh working connections and explicit denials.
Validation: 7 coordinator/checker tests passed at this layer. Both original full suites completed (44 cases plus network audit): Codex 38 passed / 6 inconclusive / 1 failed; Claude 35 / 8 / 2. Review found real package/bytecode defects fixed by earlier stack layers, checker false positives, and unresolved refusals/budget exhaustion. Targeted reruns and the final deterministic/network audits are documented; the original runs are not represented as all green. The new rendered-policy case expands the catalog to 45 plus network audit.
Paid probes are opt-in. Raw local audit logs, images, credentials, and unrelated campaign files are not committed.
Stack integration validation: 258 tests passed, 4 image/runtime tests skipped in the isolated worktree; pylint passed and mypy passed for 11 changed modules. Paid cluster audit results are summarized above where relevant.
Stack created with GitHub Stacks CLI