Roadmap: harden the existing execution model before expanding scope
背景
sshx 当前已经具备较完整的 Agent 远程执行能力,包括:
- 单主机 / 多主机远程执行
- 稳定的 JSON / JSONL 执行契约
- dry-run
- 结构化 audit
- host / credential 管理
- SFTP / transfer
- guarded
apply
- guarded
sql
- host inspection / plugin
- stdio MCP
- human-only interactive login
- source address binding
- release signing / SBOM / provenance
当前阶段不应继续横向增加大量新功能。
接下来的重点应该是:
提高执行确定性、安全性、可验证性、稳定性和长期维护能力。
目标是让 sshx 从“功能完整的 Agent SSH CLI”进一步收敛为:
A predictable, auditable and safe remote execution primitive for agents.
核心原则
新增能力应优先满足以下条件之一:
- 减少 Agent 判断次数
- 减少执行结果的不确定性
- 提高执行前后的可验证性
- 提高错误分类和恢复能力
- 提高审计证据质量
- 提高跨平台一致性
- 降低长期维护成本
避免仅仅因为某个远端工具存在,就增加新的 sshx verb。
P0 — Execution Plan Integrity
目前 --dry-run 可以展示执行计划,但计划与实际执行之间还缺少明确的完整性绑定。
目标:
resolve
↓
plan
↓
review / agent decision
↓
execute exactly that plan
↓
audit
增加稳定的 execution plan fingerprint。
例如:
{
"schema": "sshx.plan.v1",
"target": "db-01",
"action": "apply",
"risk": "mutation",
"plan_hash": "sha256:..."
}
执行时允许:
sshx apply ... --expect-plan=sha256:...
如果执行环境或影响执行语义的输入发生变化,应拒绝执行。
Tasks
Acceptance
相同输入产生相同 plan hash。
任何会改变实际执行目标、权限或副作用的关键输入发生变化后,旧 plan hash 不得继续执行。
P0 — Unified Risk Model
目前不同执行路径已经存在安全限制,但风险语义还可以进一步统一。
建议引入统一风险等级:
read
mutation
privileged
destructive
风险不是单纯根据命令字符串判断,而应该来自:
action
+ target
+ privilege
+ operation type
+ known side effects
例如:
| Operation |
Risk |
uname -a |
read |
| download file |
read |
| upload new file |
mutation |
apply |
mutation |
| sudo read |
privileged |
| sudo mutation |
privileged |
| destructive SQL |
destructive |
| disk / filesystem destructive command |
destructive |
Tasks
Acceptance
Agent 不需要根据不同 verb 单独推断风险。
所有产生副作用的执行入口使用相同风险语义。
P0 — Preconditions & Postconditions
对于 mutation,sshx 应尽可能提供执行前后的事实,而不是只报告命令成功。
例如 apply:
before_sha256
expected_sha256
after_sha256
backup_path
changed
对于其他操作也可以提供适当的 precondition / postcondition。
目标:
Execution success != desired effect confirmed.
Tasks
P1 — Execution Fingerprint
除了 plan hash,实际执行结果也应产生 fingerprint。
例如:
{
"plan_hash": "sha256:...",
"execution_id": "...",
"execution_fingerprint": "sha256:...",
"started_at": "...",
"finished_at": "...",
"verified": true
}
用于关联:
plan
execution
result
audit
Tasks
P1 — Cancellation & Deadline Semantics
多主机 fan-out 已经存在,下一步重点应放在失败和取消语义。
Tasks
P1 — Fan-out Failure Policy
增加少量、明确的 fan-out 控制能力:
--fail-fast
--max-failures=N
不要演化成 workflow engine。
Tasks
P1 — Error Taxonomy Review
系统能力已经增长较多,需要重新审视 error_kind。
目标:
Agent 可以根据 error kind 明确决定:
retry
fix config
request credential
request approval
abort
inspect target
Tasks
P1 — Audit Evidence Quality
当前 audit 已支持 query / export。
下一阶段重点不是增加日志数量,而是提升证据质量。
建议保证能够回答:
Who/what initiated it?
Which host?
Which resolved address?
Which credential role?
Which plan?
Which risk?
Was a bypass used?
What was executed?
Was state changed?
Was the result verified?
What failed?
Tasks
P1 — Test Coverage & Reliability
目前继续扩功能的收益已经低于补测试的收益。
优先覆盖:
execution contract
safety
credentials
audit
cross-platform
failure paths
Tasks
不追求单纯的 coverage 数字。
优先保证关键执行路径和 failure path 被覆盖。
P2 — Internal Architecture Cleanup
随着功能增长,应继续降低 package 间耦合。
重点关注:
resolve
plan
policy
execute
verify
audit
这些阶段是否存在隐式交叉。
Tasks
原则:
不为“架构漂亮”重构,只重构已经产生重复语义或行为漂移的部分。
P2 — Cross-platform Parity
继续明确 Linux / macOS / Windows 的能力矩阵。
Tasks
禁止 silent degradation。
Non-goals
本阶段明确不做:
- daemon
- resident remote agent
- HTTP/SSE MCP server
- Web UI
- scheduler
- cron
- workflow / playbook
- desired-state reconciliation
- connection pool
- persistent task queue
- fleet heartbeat
- SOCKS proxy
- generic tunnel manager
- Kubernetes management platform
- Docker management platform
- Redis / Kafka / MongoDB 等“一工具一个 verb”的横向扩张
如果某项能力不能明显减少 Agent 判断成本或提高执行可信度,默认不进入 sshx core。
Suggested release sequence
v0.14
Execution integrity:
- plan hash
--expect-plan
- unified risk model
- preconditions / postconditions
v0.15
Execution lifecycle:
- execution fingerprint
- cancellation
- timeout semantics
- fan-out failure policy
v0.16
Audit & contract hardening:
- error taxonomy
- audit evidence
- frozen schema tests
- MCP parity
v0.17
Reliability:
- failure-path E2E
- cross-platform parity
- internal lifecycle cleanup
- security review
之后再评估进入 v1.0。
Definition of Done
这一阶段完成后,sshx 应满足:
最终目标不是让 sshx 支持更多事情。
最终目标是让 Agent 更放心地执行已经支持的事情。
Roadmap: harden the existing execution model before expanding scope
背景
sshx 当前已经具备较完整的 Agent 远程执行能力,包括:
applysql当前阶段不应继续横向增加大量新功能。
接下来的重点应该是:
目标是让 sshx 从“功能完整的 Agent SSH CLI”进一步收敛为:
核心原则
新增能力应优先满足以下条件之一:
避免仅仅因为某个远端工具存在,就增加新的 sshx verb。
P0 — Execution Plan Integrity
目前
--dry-run可以展示执行计划,但计划与实际执行之间还缺少明确的完整性绑定。目标:
增加稳定的 execution plan fingerprint。
例如:
{ "schema": "sshx.plan.v1", "target": "db-01", "action": "apply", "risk": "mutation", "plan_hash": "sha256:..." }执行时允许:
如果执行环境或影响执行语义的输入发生变化,应拒绝执行。
Tasks
sshx.plan.v1--dry-run --json输出plan_hash--expect-plan=<hash>error_kindplan_hashAcceptance
相同输入产生相同 plan hash。
任何会改变实际执行目标、权限或副作用的关键输入发生变化后,旧 plan hash 不得继续执行。
P0 — Unified Risk Model
目前不同执行路径已经存在安全限制,但风险语义还可以进一步统一。
建议引入统一风险等级:
风险不是单纯根据命令字符串判断,而应该来自:
例如:
uname -aapplyTasks
run使用统一风险分类apply使用统一风险分类sql使用统一风险分类sftp/ transfer 使用统一风险分类--force/ bypass 与 risk 模型统一Acceptance
Agent 不需要根据不同 verb 单独推断风险。
所有产生副作用的执行入口使用相同风险语义。
P0 — Preconditions & Postconditions
对于 mutation,sshx 应尽可能提供执行前后的事实,而不是只报告命令成功。
例如
apply:对于其他操作也可以提供适当的 precondition / postcondition。
目标:
Tasks
apply完善 before / after fingerprintchangedexecuted与verifiederror_kindP1 — Execution Fingerprint
除了 plan hash,实际执行结果也应产生 fingerprint。
例如:
{ "plan_hash": "sha256:...", "execution_id": "...", "execution_fingerprint": "sha256:...", "started_at": "...", "finished_at": "...", "verified": true }用于关联:
Tasks
P1 — Cancellation & Deadline Semantics
多主机 fan-out 已经存在,下一步重点应放在失败和取消语义。
Tasks
error_kindP1 — Fan-out Failure Policy
增加少量、明确的 fan-out 控制能力:
不要演化成 workflow engine。
Tasks
--fail-fast--max-failuresP1 — Error Taxonomy Review
系统能力已经增长较多,需要重新审视
error_kind。目标:
Agent 可以根据 error kind 明确决定:
Tasks
retryable字段P1 — Audit Evidence Quality
当前 audit 已支持 query / export。
下一阶段重点不是增加日志数量,而是提升证据质量。
建议保证能够回答:
Tasks
P1 — Test Coverage & Reliability
目前继续扩功能的收益已经低于补测试的收益。
优先覆盖:
Tasks
不追求单纯的 coverage 数字。
优先保证关键执行路径和 failure path 被覆盖。
P2 — Internal Architecture Cleanup
随着功能增长,应继续降低 package 间耦合。
重点关注:
这些阶段是否存在隐式交叉。
Tasks
internal/appresponsibilities原则:
P2 — Cross-platform Parity
继续明确 Linux / macOS / Windows 的能力矩阵。
Tasks
禁止 silent degradation。
Non-goals
本阶段明确不做:
如果某项能力不能明显减少 Agent 判断成本或提高执行可信度,默认不进入 sshx core。
Suggested release sequence
v0.14
Execution integrity:
--expect-planv0.15
Execution lifecycle:
v0.16
Audit & contract hardening:
v0.17
Reliability:
之后再评估进入
v1.0。Definition of Done
这一阶段完成后,sshx 应满足:
最终目标不是让 sshx 支持更多事情。
最终目标是让 Agent 更放心地执行已经支持的事情。