Skip to content

Latest commit

 

History

History
272 lines (220 loc) · 11.1 KB

File metadata and controls

272 lines (220 loc) · 11.1 KB

Python and Gymnasium API

The package under python/ is a typed synchronous client for protocol 0.3, an isolated one-child supervisor, and Gymnasium adapters for map play and native character generation. It has no LLM dependency.

The supported 1.x import, signature, model, and Gymnasium boundaries are defined in the compatibility and deprecation policy. Supported names should be imported from qud_agent rather than its submodules.

Install for development

python3 -m venv .venv
.venv/bin/pip install -e './python[dev]'
./scripts/check.sh

The repository ignores the virtual environment and all Python/.NET build intermediates.

Typed client

from qud_agent import QudClient

with QudClient.from_user_dir("/absolute/Qud/user/directory") as client:
    hello = client.get_capabilities()
    observation = client.observe()
    result = client.step({"type": "move", "direction": "north"})

Observations and step results are Pydantic models. The client validates schema 0.3, decision-boundary invariants, monotonic IDs, exact advertised actions, and pre/post step identity. Read-only operations may reconnect once. A transmitted step is never replayed: loss of its response raises QudStepUncertainError, requiring reconnect plus observation before recovery.

Restricted environment

from qud_agent import QudEnv, QudEnvConfig, heuristic_action

config = QudEnvConfig.from_user_dir(
    "/absolute/Qud/user/directory",
    max_episode_steps=500,
)
env = QudEnv(config)
observation, info = env.reset()

while True:
    action = heuristic_action(observation)
    observation, reward, terminated, truncated, info = env.step(action)
    if terminated or truncated:
        break
env.close()

Discrete actions are stable: north through northwest are indices 0..7, and wait is 8. action_mask marks actions advertised by the current observation; it is a protocol-validity mask, not a collision oracle. The compact observation also includes neighbor occupancy (0 unseen, 1 visibly empty, 2 visibly occupied), HP, position, mode, turn, IDs, and message/cell counts. The complete typed observation and semantic action result remain available in info.

Reward is always 0.0. Player death sets terminated. Bridge truncation, leaving the restricted map mode, and an optional max_episode_steps limit set truncated with a reason in info.

Without a supervisor, reset() attaches to an already loaded actionable game. With a supervisor, reset accepts exactly one of baseline, character_spec, or episode_handle; omitting all three selects the configured Ksekra baseline. See Getting started for isolated setup and QudChargenEnv handoff.

Bootstrap recipes

Versioned bootstrap recipes drive native character generation through semantic targets and verify the module landmark before and after every action. The built-in input-free Joppa recipe returns an auditable result containing the seeded character specification, duration, chosen lifecycle actions, and reached landmarks:

from qud_agent import (
    CLASSIC_PRAETORIAN_JOPPA_V1,
    QudEnv,
    QudEnvConfig,
    QudSupervisor,
    QudSupervisorConfig,
    BootstrapPreparationConfig,
    prepare_bootstrap_recipe,
)

supervisor = QudSupervisor(QudSupervisorConfig(...))
result = prepare_bootstrap_recipe(supervisor, CLASSIC_PRAETORIAN_JOPPA_V1)
env = QudEnv(QudEnvConfig(), supervisor=supervisor)
observation, info = env.reset(
    options={"character_spec": result.character_spec}
)

BootstrapRecipe and BootstrapStep define additional immutable recipes. resolve_bootstrap_action maps a semantic step to QudChargenEnv's fixed action space. A failed target, transition, required selection, or graceful shutdown raises QudBootstrapError; its recipe_id, zero-based step_index, and final lifecycle fields identify the failed landmark when available. get_bootstrap_recipe resolves a stable built-in recipe ID and rejects unknown IDs with QudConfigurationError before any game process is launched.

prepare_bootstrap_recipe is the audited preparation entry point. By default it executes chargen and writes a mode-0600 JSON record below agent_root/bootstrap-audits. Success records contain the recipe identity, duration, chosen semantic actions, reached landmarks, normalized invariants, and generated-spec digest. Failure records add a sanitized final lifecycle observation; native build data is deliberately excluded. QudBootstrapError also exposes the audit path when recording succeeds.

Locally generated character specifications can be reused as an explicit acceleration/debugging option:

result = prepare_bootstrap_recipe(
    supervisor,
    CLASSIC_PRAETORIAN_JOPPA_V1,
    config=BootstrapPreparationConfig(use_generated_cache=True),
)

The cache lives below agent_root/generated-snapshots, remains outside the source tree, and is partitioned by expected game build and the canonical digest of every recipe input. Its content-addressed generated-spec digest is validated before each hit. A build or recipe change therefore misses the cache. This cache skips native chargen only; every environment reset still starts a new isolated episode from the cached character specification.

Use character_spec_digest to log a canonical SHA-256 identity without printing the native build code. observation_fingerprint supports the default decision-boundary policy and a structural Joppa policy that retains visible geometry and object counts while ignoring mobile-object positions.

Interactive scenario recipes

get_interactive_scenario_recipe resolves the built-in fresh-Joppa setup flow for each supported interactive mode and the separate disconnect-recovery case. Unknown IDs raise QudConfigurationError before launch. Each immutable InteractiveScenarioRecipe names its bootstrap recipe, semantic setup actions, per-step mode and position landmarks, probe action, and expected outcome. InteractiveRecipeAction and VisibleTargetSelector expose the typed pieces used by manifest runners without embedding opaque target IDs.

The IDs are fresh-joppa-<mode>-v1, with hyphens in world-travel, plus fresh-joppa-popup-disconnect-v1. These are live validation recipes for the known-working build, not environment-reset options; prepare their referenced bootstrap recipe and pass the resulting character_spec through the existing reset API.

Interactive environment

QudInteractiveEnv retains a fixed spaces.Dict action containing operation, target index, direction, and value. Each observation supplies a padded table and mask of exact valid tuples plus a padded visible-target catalog. Complete text and typed descriptors remain in info.

from qud_agent import QudInteractiveConfig, QudInteractiveEnv

env = QudInteractiveEnv(QudInteractiveConfig.from_user_dir("/absolute/Qud/user/directory"))
observation, info = env.reset()
row = observation["available_action_table"][0]
action = dict(zip(("operation", "target_index", "direction", "value"), map(int, row)))
observation, reward, terminated, truncated, info = env.step(action)

Defaults allow 512 visible targets and 2,048 exact actions. Exceeding either limit fails explicitly. Reward remains zero, while supported modal and world-travel transitions continue the episode.

Trajectory recording

Isolated supervisor episodes record automatically beneath <agent_root>/recordings. Add a scenario label and scalar tags at reset:

from qud_agent import TrajectoryMetadata

observation, info = env.reset(
    options={
        "baseline": "ksekra",
        "trajectory_metadata": TrajectoryMetadata(
            label="survive-50-turns",
            tags={"policy": "cautious", "seed": 7},
        ),
    }
)

EpisodeHandle.trajectory_path points to trajectory.ndjson while the run is active, and trajectory_summary_path points to the atomically created summary.json after shutdown. The summary reports initial/final turns and HP, action and outcome counts, turns consumed, Gym reward total, termination state, valid byte count, and a SHA-256 digest.

Direct clients and manually attached environments must opt in with an absolute owned root:

from pathlib import Path
from qud_agent import QudClient, TrajectoryConfig, TrajectoryMetadata

client = QudClient.from_user_dir(
    "/absolute/Qud/user/directory",
    trajectory=TrajectoryConfig(root=Path("/absolute/trajectory-root")),
    trajectory_metadata=TrajectoryMetadata(label="manual-evaluation"),
)

Use iter_trajectory(path) for validated schema-0.1 events and load_trajectory_summary(path) for the typed summary. A directory or its specific data file is accepted. The reader ignores only an unterminated final frame by default, allowing complete events from a crashed process to be read; pass allow_incomplete=False for strict rejection.

Every gameplay or lifecycle mutation writes an intent before dispatch and a result afterward. QudRecordingError.command_state is not_sent, completed, uncertain, or not_applicable; callers must not retry a command reported as completed or uncertain. Supervised environments stop their exact child on a recording failure. Manually attached clients disconnect without quitting the user-launched game.

Recording is structured NDJSON, not video or executable replay. Public fields are the default; include_privileged=True retains explicitly enabled debug fields but still excludes authentication and secret material.

Trajectory inspection, simulation, and replay

load_trajectory_document strictly pairs gameplay intents and results and checks the summary digest, pre/post snapshots, outcomes, and completeness. inspect_trajectory returns provenance and replayability diagnostics, while diff_trajectories compares aligned steps under decision_boundary, structural_joppa, or strict_public normalization:

from qud_agent import (
    TrajectoryComparisonProfile,
    diff_trajectories,
    inspect_trajectory,
)

inspection = inspect_trajectory("/absolute/recording")
comparison = diff_trajectories(
    "/absolute/left-recording",
    "/absolute/right-recording",
    comparison_profile=TrajectoryComparisonProfile.DECISION_BOUNDARY,
)

TrajectorySimulationClient.from_trajectory runs the recorded branch through an in-process raw-NDJSON protocol peer. It opens no socket and can be passed to QudEnv or QudInteractiveEnv. Only the exact recorded action branch is available; a different advertised action reports simulation_action_mismatch.

replay_trajectory is dry by default. Live mutation requires both a client and ReplayConfig(dispatch_actions=True). It compares the live state before every action, dispatches the action once using the live observation ID, compares the semantic outcome and post-state, and stops on the first mismatch. An uncertain transport result is never retried. Incomplete, crashed, recording-error, or uncertain source trajectories cannot be replayed.

The original schema-0.1 iter_trajectory and load_trajectory_summary readers remain non-dispatching. Lifecycle intents in a recording are inspected but never replayed; the live CLI reconstructs an isolated reset separately.