Skip to content

Repository files navigation

OmniBioAI Control Center

README last reviewed: 2026-08-31

Operational health dashboard, ecosystem report server, and observability hub for the OmniBioAI stack.

The Control Center is a FastAPI service that aggregates health status across all OmniBioAI components, serves an interactive ecosystem report, exposes Prometheus metrics, and auto-generates reports on a configurable schedule.


What It Does

  • Health monitoring — TCP, HTTP, and disk checks across all ecosystem services
  • Enterprise Admin Console — the operational and administrative interface, served as a separate frontend build at admin.omnibioai.org: Organizations, Users, Teams, Roles & Permissions, Security (Security Overview, Security Posture, MFA Policy, IAM/SSO Management, SAML, Audit Logs, Audit Explorer, Compliance Report, Sessions, Interactions, API Keys/Service Accounts), HIPAA Compliance, Billing, Usage Analytics, Workflows, Tool Execution, AI Models, Agentic AI, RAG/PubMed, Integrations, Settings, plus an Operations → Infrastructure group (Health, Regression Health, Deployment Health, Integration Health, Docker, Ecosystem Report, Config, LLMs, Cloud, Actions, Scheduled Jobs, Known Issues) — see Admin Console below
  • Regression Health — reads a reviewed, promoted end-to-end certification artifact (never inferred from a pytest run) and exposes it read-only — see Regression Health
  • Deployment Health (V1, certified) — read-only, dependency-aware deployment and runtime health for the ecosystem, combining Compose metadata, Docker state, and application probes — see Deployment Health
  • Ecosystem report — interactive HTML report (architecture · projects · languages · coverage · health) served at /; /dashboard redirects here (its live per-service cards and generate button were folded into the report's header status chip and Admin tab)
  • JSON API — machine-readable health summary at /summary for CI/CD and external monitoring
  • Scheduled report generation — auto-regenerates the ecosystem report every N hours (configurable via REPORT_SCHEDULE_HOURS)
  • Admin controls — an Admin tab in the report for triggering report/coverage regeneration, pausing/rescheduling the 7 host cron jobs, and tracking known issues; every write action is JWT-role-gated (admin role required), enforced both by nginx and independently by the app itself
  • Prometheus metrics/metrics endpoint scraped by Prometheus for Grafana dashboards
  • Docker inventory — platform containers, tool SIF images, and plugin Docker images via /docker/* endpoints
  • Structured JSON logging — all key events logged as JSON to stdout for log aggregation
  • LLM monitoring — local Ollama models and API key status via /llms
  • Reference genome registry — 12 organisms, indexes, variants via /reference
  • AI Knowledge Base — 28M+ PubMed abstracts, FAISS indexes via /knowledge-base
  • Storage monitoring — disk usage, per-organism reference indexes via /storage
  • Cloud backends — execution backend status via /cloud

Control Center preview

The Control build is the internal operations surface for platform health, Docker, ecosystem status, configuration, LLMs, and cloud infrastructure. The current public health view is shown below; authenticated operational views require the backend services and an authorized session.

Current Control Center login screen

Architecture

Architecture

Authentication

Control Center delegates all authentication to omnibioai-auth; it never verifies a password or issues a token itself. It plays two distinct roles:

  1. Same-origin proxy for the browserroutes_auth_proxy.py relays /auth/login, /auth/refresh, /auth/logout, and /auth/validate straight through to auth-service, existing purely because control.omnibioai.org's ingress reaches this service directly, bypassing the nginx router that fronts every other domain's /auth/* path.
  2. Local verifier for its own admin-gated endpointsrequire_admin (core/auth.py) independently checks every write endpoint's bearer token, rather than trusting the proxy hop above.

Login

The Admin tab's login form (or the omnibioai-studio SPA sharing the same origin) POSTs credentials through this service's /auth/login proxy. auth-service authenticates and returns an access_token in the JSON body and sets the omnibioai_session cookie on the response — the proxy relays that Set-Cookie back to the browser rather than dropping it.

Browser session

The frontend (frontend/cc-ui/src/auth.ts) keeps only the short-lived access token client-side, in localStorage (omnibioai_access_token) — attached as Authorization: Bearer <token> on every gated request. It never reads, stores, or forwards the refresh token: that lives solely in the HttpOnly omnibioai_session cookie the browser attaches automatically to same-origin requests, invisible to page JavaScript.

Refresh flow

The frontend decodes the access token's own exp claim (no signature check needed — it's opaque either way until the server re-validates it) and schedules a silent refresh ~60 seconds before it expires. That refresh call sends an empty body to /auth/refresh — the browser's omnibioai_session cookie is what actually authorizes the rotation; the proxy relays the incoming Cookie header upstream and the rotated Set-Cookie back downstream. A failed refresh (missing/expired/revoked session) forces a return to the login screen; a network-level failure fails open and simply lets the current access token run out naturally.

Logout

POST /auth/logout (via the proxy) sends only the access token in the body — the refresh token is never available to send, since it's cookie-only. The proxy fills in refresh_token server-side from the omnibioai_session cookie before forwarding to auth-service, which revokes it, blacklists the access token's jti, and clears the cookie; the frontend clears its own localStorage entry regardless of whether the network call itself succeeded.

Admin authorization

require_admin (core/auth.py) gates every write endpoint (see the "Admin-gated" markers in API Endpoints):

verify token
      │
      ▼
check roles
      │
      ▼
allow / deny

It parses the Authorization header, delegates signature/expiry/type/ claim/revocation verification to core/jwt_verify.py, and then makes its own decision: 401 if the token itself is invalid, 403 if it's valid but lacks the admin role. This runs independently of nginx's own auth_request check in front of /_svc/control — so a misconfigured nginx rule alone can never expose a write endpoint, since this service checks the specific role itself either way.

JWT verification

core/jwt_verify.py is this service's own local copy of the shared verification logic (structurally identical to omnibioai-security-audit's) — it does not call back into auth-service on every request. It checks signature, expiry, token type (rejects a presented refresh token), the required sub claim, and Redis jti-blacklist revocation, dispatched by each token's own alg header (see RS256 readiness below).

Redis revocation

The same jti-blacklist auth-service writes to on logout (blacklist:jti:{jti}) is checked here directly against the same Redis instance — fail-open on a Redis error, matching auth-service's own documented tradeoff: a Redis blip must not 401 every admin request in this service either.

RS256 readiness

core/jwt_verify.py verifies both HS256 (today's production default) and RS256 (auth-service's /.well-known/jwks.json, once JWT_ALGORITHM=RS256 is enabled there), dispatched by each token's own alg header — never by local configuration. An RS256 token's kid is resolved against a cached, auto-refreshing JWKS client; any signature or JWKS-fetch failure fails closed. No production deployment has switched issuance to RS256 yet — see omnibioai-auth's README and the ecosystem root README's Deployment Notes.

Authentication sequence

Browser
   │
   ▼
Control Center            (routes_auth_proxy.py — same-origin relay)
   │
   ▼
Auth Service               (omnibioai-auth — authenticates, issues tokens)
   │
   ▼
JWT                        access_token in JSON body (Authorization: Bearer, 15 min)
   │
   ▼
Refresh Cookie              omnibioai_session (HttpOnly, Secure, SameSite=Lax) —
                             relayed browser ↔ Control Center ↔ Auth Service,
                             drives the silent-refresh flow above

Repository Structure

omnibioai-control-center/
│
├── scripts/
│   └── generate_report.py          # Ecosystem report generator (CLI)
│
├── backend/
│   ├── pyproject.toml              # Package definition and dependencies
│   ├── src/control_center/
│   │   ├── main.py                 # FastAPI app — registers all routers
│   │   ├── api/
│   │   │   ├── routes_health.py    # GET /health
│   │   │   ├── routes_services.py  # GET /services
│   │   │   ├── routes_summary.py   # GET /summary
│   │   │   ├── routes_report.py    # GET /report, /report/generate, /report/status, /report/data
│   │   │   ├── routes_config.py    # GET /config, POST /config/service
│   │   │   ├── routes_cron.py      # /cron/jobs + pause/resume/schedule
│   │   │   ├── routes_docker.py    # /docker/*
│   │   │   ├── routes_regression_health.py # GET /regression-health (SPA: /regression-health/data via nginx)
│   │   │   ├── routes_deployment_health.py # GET /deployment-health (SPA: /deployment-health/data via nginx)
│   │   │   ├── routes_known_issues.py  # /known-issues CRUD
│   │   │   ├── routes_llm.py       # GET /llms, /knowledge-base
│   │   │   ├── routes_cloud.py     # GET /cloud
│   │   │   ├── routes_reference.py # GET /reference
│   │   │   ├── routes_storage.py   # GET /storage
│   │   │   ├── routes_infra.py     # GET /gpu, /celery, /database, /license, /usage,
│   │   │   │                       # /gateway-traffic, /audit-trail, /activity,
│   │   │   │                       # /image-freshness, /integrity
│   │   │   ├── routes_dashboard.py # GET /dashboard/summary (Overview page stat cards)
│   │   │   ├── routes_auth_proxy.py    # /auth/login, /auth/refresh, /auth/logout, /auth/validate
│   │   │   │                           # — same-origin relay to auth-service
│   │   │   └── routes_{org,user,role,team,service_accounts,org_sso,org_mfa,audit,
│   │   │       billing,tes,model_registry,workflow_bundles,rag,platform_config}_proxy.py
│   │   │                           # Admin Console enterprise proxy layer (15 files) —
│   │   │                           # see "Admin Console" below
│   │   ├── regression_health.py    # Reads/validates the promoted RH-1 certification artifact
│   │   ├── deployment_health.py    # DH-1: static Compose service/dependency model (no I/O)
│   │   ├── deployment_health_runtime.py # DH-2: Docker/probe/dependency merge -> effective health
│   │   ├── checks/
│   │   │   ├── http.py / tcp.py / disk.py          # Core checks — HTTP, TCP (MySQL/Redis), disk
│   │   │   ├── cron_jobs.py        # Host-crontab status + pause/resume/reschedule logic
│   │   │   ├── known_issues.py     # Known-issue store (known_issues.json)
│   │   │   └── gpu.py / celery_status.py / database_status.py / license_status.py /
│   │   │       usage_status.py / audit_trail.py / gateway_traffic.py /
│   │   │       image_freshness.py / integrity.py / activity.py
│   │   │                           # Infra/observability checks backing routes_infra.py
│   │   ├── core/
│   │   │   ├── runner.py           # Dispatches checks per service type
│   │   │   ├── settings.py         # Loads control_center.yaml
│   │   │   ├── auth.py             # require_admin — JWT role gate for write endpoints
│   │   │   └── jwt_verify.py       # Local JWT verification (HS256 + RS256/JWKS)
│   │   ├── notifications/
│   │   │   └── discord.py          # Discord webhook alerts (known issues, GPU temp)
│   │   └── utils/
│   │       └── summary_client.py   # Fetches /summary for report generation
│   └── tests/                      # 1,554 collected tests at last review — see "Running Tests" below
│
├── frontend/cc-ui/src/
│   ├── apps/
│   │   ├── ControlApp.tsx          # Ops console root — built when VITE_APP_MODE=control
│   │   ├── AdminApp.tsx            # Admin Console root — built when VITE_APP_MODE=admin
│   │   └── AuthGate.tsx            # Shared login/session state machine (both apps)
│   ├── navigation.ts               # Single source of truth for Admin Console's sectioned nav
│   └── pages/, components/         # Page and component implementations
│
├── compose/
│   └── docker-compose.control-center.yml
├── config/
│   ├── control_center.yaml         # Active configuration
│   └── control_center.example.yaml # Reference configuration
└── docker/
    ├── Dockerfile
    └── nginx/
        ├── control-center.conf     # Host-based split: control.omnibioai.org / admin.omnibioai.org
        └── api-proxy.conf          # Shared API proxy rules, included by both server blocks

API Endpoints

Endpoint Method Description
/ GET Ecosystem report (auto-refreshes)
/dashboard GET Redirects to / (retired — see Admin tab)
/health GET Control Center self-check
/services GET Per-service health status (JSON)
/summary GET Full ecosystem summary — services + disk (JSON)
/report GET Redirects to / (serves report with live header)
/report/generate POST Admin-gated. Trigger background report generation
/report/status GET Poll report job state (running/done/error/idle)
/report/data GET Structured report data as JSON
/coverage/generate POST Admin-gated. Trigger coverage collection, scoped to control-center itself (see Design Principles)
/coverage/status GET Poll coverage job state
/config GET Raw control_center.yaml contents (plain text)
/config/service POST Append a new monitored service to the config
/cron/jobs GET Status of the 7 known host-crontab jobs
/cron/jobs/{id}/pause POST Admin-gated. Pause one of the 7 known jobs (whitelist-only)
/cron/jobs/{id}/resume POST Admin-gated. Resume a paused job
/cron/jobs/{id}/schedule PUT Admin-gated. Reschedule a job in the real crontab
/known-issues GET List tracked known issues
/known-issues POST Admin-gated. Create a known issue
/known-issues/{id} PUT Admin-gated. Update a known issue
/known-issues/{id} DELETE Admin-gated. Delete a known issue
/docker/containers GET Platform container list with status
/docker/sif-images GET Tool SIF image inventory and sizes
/docker/plugin-images GET Plugin Docker image inventory
/regression-health GET platform.manage_infra-gated. Reviewed end-to-end certification status, read from the promoted RH-1 artifact — see Regression Health
/deployment-health GET platform.manage_infra-gated. Read-only, dependency-aware deployment/runtime health — see Deployment Health
/metrics GET Prometheus metrics endpoint
/llms GET Local LLM models + API key status
/cloud GET Cloud/HPC execution backend status
/reference GET Reference genome registry (12 organisms)
/knowledge-base GET AI knowledge base stats (PubMed + FAISS)
/storage GET Disk usage + per-organism index sizes
/dashboard/summary GET Overview page stat cards (both Admin Console and ops console)
/gpu GET GPU temperature/utilization via nvidia-smi
/celery GET Celery worker + recent task-queue status
/database GET MySQL/Redis/Neo4j data-layer connectivity
/license GET Seat usage and license expiry status
/usage GET Product usage — active users, run counts
/gateway-traffic GET API Gateway request/route traffic (7-day window)
/audit-trail GET Auth/policy/HPC audit log stream (7-day window)
/activity GET Live CPU/memory/network via Prometheus + cAdvisor
/image-freshness GET Deployed image digests vs. latest on GHCR
/integrity GET Configured symlink/mount integrity checks

"Admin-gated" endpoints require a valid JWT carrying the admin role, checked twice independently: once by nginx's auth_request (any valid JWT) and once inside the app itself via require_admin (the specific role) — so an nginx misconfiguration alone can't expose a write endpoint. All other endpoints above are fully open, no auth required.

/regression-health and /deployment-health above are the backend routes. Each also has a human-facing Admin Console SPA page at the same path (/regression-health, /deployment-health) that a browser navigates to directly — nginx keeps these distinct by rewriting a separate /regression-health/data / /deployment-health/data request path to the backend route (docker/nginx/api-proxy.conf), rather than letting the SPA route and the API route collide (the exact mistake, internally tracked as REG-010, that this split was introduced to fix). The frontend always calls the /data path; only a browser's own document navigation should ever hit the bare path.

This table covers the original ops-console surface. A separate, larger set of /orgs/*, /platform/*, /billing/*, /tes/*, /model-registry/*, /workflow-bundles/*, /rag/*, and /auth/config proxy routes backs the Admin Console — see Admin Console → Enterprise proxy routes below rather than duplicating all ~35 of them here; each is gated by whatever permission its owning service already requires (omnibioai-auth's manage_all_orgs/org-membership checks, omnibioai-billing's org-scoped IAM, etc.) — Control Center's own require_admin/nginx layer above doesn't apply to them.

/health

{ "status": "ok" }

/summary

{
  "overall_status": "UP",
  "generated_at": "2026-03-20T02:44:00+00:00",
  "services": [
    {
      "name": "omnibioai",
      "type": "http",
      "target": "http://omnibioai:8000/",
      "status": "UP",
      "latency_ms": 12,
      "message": "HTTP 200"
    },
    {
      "name": "mysql",
      "type": "mysql",
      "target": "mysql:3306",
      "status": "UP",
      "latency_ms": 3,
      "message": "TCP connect ok"
    }
  ],
  "system": {
    "disk": [
      {
        "name": "disk:/workspace/out",
        "type": "disk",
        "target": "/workspace/out",
        "status": "UP",
        "latency_ms": null,
        "message": "45.2% free"
      }
    ]
  }
}

Status values: UP | DOWN | WARN


Configuration

All monitored services and disk paths are defined in config/control_center.yaml.

services:
  mysql:
    type: mysql
    host: mysql
    port: 3306

  redis:
    type: redis
    host: redis
    port: 6379

  toolserver:
    type: http
    url: http://toolserver:9090/health
    timeout_s: 2

  tes:
    type: http
    url: http://tes:8081/health
    timeout_s: 2

  omnibioai:
    type: http
    url: http://omnibioai:8000/
    timeout_s: 2

  lims-x:
    type: http
    url: http://lims-x:7000/
    timeout_s: 2

  model-registry:
    type: http
    url: http://model-registry:8095/health
    timeout_s: 2

system:
  disk_checks:
    - path: /workspace/out
      warn_pct_free_below: 15
    - path: /workspace/tmpdata
      warn_pct_free_below: 10
    - path: /workspace/local_registry
      warn_pct_free_below: 10

Supported check types

Type Required fields Description
http url, timeout_s HTTP GET — UP if 2xx, WARN if 3xx/4xx/5xx
mysql host, port TCP connect to MySQL port
redis host, port TCP connect to Redis port

Adding a new service

Add a block to config/control_center.yaml and restart the container:

services:
  my-new-service:
    type: http
    url: http://my-service:8080/health
    timeout_s: 2

No code changes required.


Running

Via Docker (repository-local)

The repository root Dockerfile builds both frontend bundles and the FastAPI service. Build and run it directly from this checkout:

docker build -t omnibioai-control-center -f Dockerfile .
docker run --rm -p 7070:7070 \
  -e CONTROL_CENTER_CONFIG=/workspace/config/control_center.yaml \
  -e WORKSPACE_ROOT=/workspace \
  -v "$PWD:/workspace" \
  omnibioai-control-center

The checked-in compose/docker-compose.control-center.yml is an ecosystem deployment template and currently references an external deploy/control-center build context; do not use it as a standalone command from this checkout until that deployment layout is present.

Access the direct development service at http://127.0.0.1:7070. In the full ecosystem deployment, nginx exposes the service at /_svc/control; read-only endpoints are public there, while other endpoints require JWT authentication and write endpoints additionally require the admin role.

Standalone (development)

cd backend
pip install -e ".[dev]"

CONTROL_CENTER_CONFIG=../config/control_center.yaml \
WORKSPACE_ROOT=~/Desktop/machine \
uvicorn control_center.main:app --host 0.0.0.0 --port 7070 --reload

Environment variables

For the owner-only Open LIMS launch, configure the public client values at frontend build time. The client secret is never placed in Control Center or browser configuration; it remains server-side in Auth and LIMS.

VITE_LIMS_SSO_CLIENT_ID=<registered-client-id>
VITE_LIMS_SSO_REDIRECT_URI=https://lims.omnibioai.org/sso/callback

The redirect URI must exactly match Auth's LIMS_SSO_REDIRECT_URI value.

Variable Default Description
CONTROL_CENTER_CONFIG /config/control_center.yaml Path to YAML config
WORKSPACE_ROOT /workspace Ecosystem root as seen inside the container or local process
CONTROL_CENTER_PORT 7070 Service port
REPORT_SCHEDULE_HOURS 6 Auto-regenerate report every N hours
WORK_DIR /workspace/omnibioai-work Work/output directory; use a path valid for the selected local or container deployment
JWT_SECRET change-me Shared HS256 secret for validating admin JWTs locally (require_admin) — same value as AUTH_SECRET_KEY used by omnibioai-auth/workbench/api-gateway/model-registry
JWKS_URL https://auth.omnibioai.org/.well-known/jwks.json RS256 verification (not yet enabled in production) — see Authentication
JWKS_TIMEOUT_SECONDS / JWKS_CACHE_TTL_SECONDS 5 / 300 JWKS fetch timeout and key-set cache lifetime
CRONTAB_SPOOL_PATH /var/spool/cron/crontabs/manish Host crontab spool file, bind-mounted in so /cron/jobs/{id}/pause|resume|schedule can read/write it directly
DISCORD_ALERT_WEBHOOK_URL (empty) Discord webhook for new high-severity known-issue alerts — empty disables alerting gracefully, same pattern as SENTRY_DSN
REGRESSION_HEALTH_ARTIFACT_PATH $WORKSPACE_ROOT/omnibioai-ecosystem-regression/status/regression-health.json Read-only promoted certification artifact — see Regression Health
REGRESSION_HEALTH_STALE_AFTER_HOURS 168 Freshness threshold; a stale artifact keeps its certification fields, only freshness changes
DEPLOYMENT_HEALTH_COMPOSE_PATH (unset — required) Path to the authoritative Compose baseline — see Deployment Health
DEPLOYMENT_HEALTH_BASELINE_SOURCE unknown development / release / unknown
DEPLOYMENT_HEALTH_DEPLOYMENT_REPOSITORY (unset) Optional logical repo name for ownership attribution on a relative build context

Ecosystem Report

The report is a single interactive HTML file with a left sidebar-nav layout (not flat tabs) — top-level groups, some of which expand into sub-tabs:

Group Sub-tab Contents
Architecture SVG lane diagram of all services
Projects Code Summary Code line distribution across repositories
Languages Language breakdown across the ecosystem
Code Coverage Per-repo pytest coverage with progress bars
Ecosystem Status Per-repo git working-tree status (branch, clean/dirty, modified/untracked/unpushed) across every repo under the ecosystem root — same scan as bash omnibioai-utils/ecosystem_status.sh. Also surfaced as its own tab on the Admin Console's Ecosystem Report page (EcosystemPage.tsx), both reading the same gitStatus array from /report/data
Health Status Overview KPI summary + status donut + per-service latency bars
Services Live per-service health cards
Disk & Mounts Disk usage checks + symlink/mount integrity
GPU nvidia-smi temperature/utilization panel
Activity Live CPU/memory/network via Prometheus + cAdvisor
Audit Trail Auth/policy/HPC audit log stream
Errors Aggregated Sentry error counts
Usage Product Usage Active users, run counts (30-day window)
API Gateway Gateway request/route traffic (7-day window)
LLMs & Cloud LLMs Local Ollama models + API key configuration
Cloud Cloud/HPC execution backend status
Cost Tracking Placeholder — not yet implemented
Reference Data 12 organism genomes, indexes, variant databases
AI Knowledge Base 28M+ PubMed abstracts, FAISS index stats
Model Registry Registered model versions, real vs. synthetic data classification
Docker Images Platform Containers Running/stopped platform container inventory
Tool SIF Images Singularity image build status and sizes
Plugin Docker Images Plugin image inventory vs. local Docker images
Miscellaneous Active Runs Currently running/queued workflow runs
Storage Disk usage bar, data categories, organism indexes
Catalog Plugin/tool/workflow-bundle counts and breakdowns
Data Layer MySQL/Redis/Neo4j connectivity status
Task Queue Celery worker + recent task status
License Seat usage and license expiry
Secrets Audit Compose file scan for exposed secrets
Image Freshness Deployed image digests vs. latest on GHCR
Exposed Ports Compose file scan for host-exposed ports
CI/CD Health GitHub Actions status + dependency vulnerability scan per repo (npm audit requires Node in the runtime image — see Requirements)
CVE Trend Vulnerability count history over time, charted
Backup Status Backup job recency and status
Admin Actions Regenerate report / refresh coverage — admin-gated, hidden behind a login form for everyone else
Scheduled Jobs Status of the 7 host cron jobs; pause/resume/reschedule for admins, read-only otherwise
Known Issues Tracked open/acknowledged/resolved issues — live CRUD for admins (moved here from Miscellaneous), read-only otherwise. Creating a high-severity entry fires a Discord alert (see below); medium/low never do

Admin access

See Authentication above for the full login/refresh/logout flow. In short: when embedded as an iframe under omnibioai-studio's web build (same origin, /_svc/control), the Admin tab reads the same omnibioai_access_token localStorage entry Studio already wrote, so an existing Studio login is recognized automatically with no separate sign-in. If no token is present, the tab shows a login form that posts directly to /auth/login; if a token is present but lacks the admin role, the tab renders everything read-only with an explicit "admin access required" message instead of showing controls that would just fail. The scheduled silent refresh (see Authentication) keeps the session alive across the 15-minute access-token TTL; only an unrecoverable 401 (refresh itself failed) clears the stored token and re-prompts for login.

Known-issue Discord alerts

Every genuinely new high-severity known issue (create only — never on updates, and never for medium/low, to keep this high-signal) posts a Discord embed via DISCORD_ALERT_WEBHOOK_URL (a separate webhook from DISCORD_WEBHOOK_URL, which is used for GPU temperature alerts). This fires on the same code path regardless of whether the entry was created by a human via the Admin tab or automatically by one of the host cron self-check scripts (check_cron_health.py, check_disk_space.py, check_domain_health.py) — closing the loop between "the system detected a problem" and "a human finds out." Unset/unreachable webhook, or any error in sending the alert, never blocks the actual known-issue from being recorded — the alert is fire-and-forget.

Generate

# From the ecosystem root — with live health data
python omnibioai-control-center/scripts/generate_report.py \
    --root ~/Desktop/machine

# Skip health check (faster, offline)
python omnibioai-control-center/scripts/generate_report.py \
    --root ~/Desktop/machine \
    --skip-health

# Skip coverage collection (code stats only, very fast)
python omnibioai-control-center/scripts/generate_report.py \
    --root ~/Desktop/machine \
    --skip-coverage

# All options
python omnibioai-control-center/scripts/generate_report.py \
    --root ~/Desktop/machine \
    --control-center-url http://127.0.0.1:7070 \
    --out out/reports/omnibioai_ecosystem_report.html \
    --title "OmniBioAI Ecosystem Report"

Requirements

# cloc for code counting
sudo apt-get install cloc        # Ubuntu/Debian
conda install -c conda-forge cloc  # Conda

# Python dependencies
pip install pandas

# For coverage collection (best-effort)
pip install pytest pytest-cov

# For npm-audit vulnerability scanning of JS-manifest repos
# (already installed in the backend Docker image; only needed if running
# the report generator standalone outside the container)
sudo apt-get install nodejs npm  # Ubuntu/Debian — pins whatever major
                                  # version your distro repo ships; the
                                  # Docker image pins Node 20 explicitly

View

  • File: ~/Desktop/machine/out/reports/omnibioai_ecosystem_report.html
  • Browser: Open directly — no server needed
  • Live: http://localhost/_svc/control when Control Center is running

The report generates gracefully even if the Control Center is offline or coverage collection fails — those tabs show a clear unavailable state rather than breaking the whole report.


Admin Console

The detailed, maintained Admin Console guide is available here. It covers the feature catalog, authentication and authorization model, service ownership, local development, deployment, testing, troubleshooting, and current boundaries. The section below remains the repository-level architecture summary.

One repository, two frontend builds, two domains. control.omnibioai.org serves the ops console described above (Health, Docker, Ecosystem Report, Config, LLMs, Cloud — no enterprise administration UI at all, not hidden, genuinely absent from that build's JavaScript). admin.omnibioai.org serves a separate enterprise administration SPA covering Organizations, Users, Teams, Roles & Permissions, Security (MFA Policy, IAM/SSO Management, Audit Logs, API Keys/Service Accounts), Billing, Workflows, Tool Execution, AI Models, RAG/PubMed, and platform Settings.

Same repository, same FastAPI backend, same auth system, same permission checks either way — the domain only decides which pre-built frontend bundle nginx serves. The build split is a deployment/UX optimization, never an authorization boundary: every require_permission/ require_admin check and every org-membership check in omnibioai-auth is unchanged regardless of which domain a request came from, and nothing about the serving hostname is itself authenticated. Full design record: docs/admin-console-build.md.

Build

src/main.tsx picks a root component at build time from VITE_APP_MODE:

VITE_APP_MODE=admin   npm run build:admin    → src/apps/AdminApp.tsx   → dist-admin/
VITE_APP_MODE=control npm run build:control  → src/apps/ControlApp.tsx → dist-control/

AdminApp.tsx is the original, full-featured app; ControlApp.tsx imports only the six ops pages and has no reference anywhere to OrganizationsPage, UsersPage, or any components/organizations/ components/roles/components/teams code. Vite/Rollup constant-folds and tree-shakes the unused branch at build time (verified against real output — dist-control's bundle is ~45KB smaller, and the losing app's strings don't appear in it at all), so this is a genuinely smaller bundle, not hidden-but-shipped UI. Both apps share AuthGate.tsx (session/login state machine) and Header.tsx.

Navigation and feature catalog

Single source of truth: frontend/cc-ui/src/navigation.ts. Every entry below is functional: true — this app's <ComingSoon/> placeholder convention still exists in the framework (an unwired entry would render it, never hidden), but no current entry uses it.

Section Page(s) Notes
Overview Stat cards via /dashboard/summary
Administration Organizations, Users, Teams, Roles & Permissions
Operations → Infrastructure Health, Regression Health, Deployment Health, Integration Health, Docker, Ecosystem Report, Config, LLMs, Cloud, Actions, Scheduled Jobs, Known Issues One expandable parent; Regression Health, Deployment Health, and Integration Health require platform.manage_infra; the remaining pages require general admin access
Operations Workflows, Tool Execution, AI Models, Agentic AI Proxy omnibioai-workflow-bundles, omnibioai-tes, omnibioai-model-registry, and omnibioai-workbench's agent-orchestrator service directly — authorization is entirely each upstream service's own, per-request
Security Security Overview, Security Posture, MFA Policy, IAM/SSO Management, SAML Settings, Audit Logs, Audit Explorer, Compliance Report, Sessions, Interactions, API Keys/Service Accounts Audit Logs is Auth's identity-audit ledger; Audit Explorer is the read-only Security Audit event query surface. Compliance Report is the org-scoped HIPAA usage/access-log export, distinct from the platform-engineering HIPAA Compliance section below
Compliance HIPAA Compliance The platform's own HIPAA remediation history (which PRs closed which control gaps) — not org data
Business Billing, Usage Analytics Billing proxies omnibioai-billing's read APIs; Usage Analytics is scoped server-side to the caller (platform_admin/org_admin/team_admin)
Knowledge RAG, PubMed Both point at one page — RAG's only indexed corpus today is PubMed abstracts
Platform Integrations, Settings Integrations reports env-derived third-party status (Sentry, Discord); Settings is a read-only view of omnibioai-auth's GET /auth/config — the corresponding PUT (platform-wide LLM/cloud credentials) is deliberately not proxied yet

Every page above is wired to a real backend; see docs/admin-console/README.md for the full navigation/feature catalog with per-page authorization reasoning.

Supported routes and health surfaces

The Admin Console uses the browser History API for its supported direct routes. /workflows is the Workflows page route and supports an authenticated direct deep link, hard refresh, sidebar navigation, and browser Back/Forward history. The separate /workflow-operations path is not a supported committed Admin Console route; Workflow Operations functionality is represented by the Workflows page and the workflow-bundles proxy.

The implemented health/status surfaces are:

  • Health — generic infrastructure and service reachability.
  • Deployment Health — Compose, Docker runtime, dependency, and application probe evidence.
  • Regression Health — the reviewed external certification artifact.
  • Integration Health — configured integration inventory and provider readiness; inventory is derived from the configured Workbench plugin registry, not a hard-coded count.
  • Security Posture — evidence-backed security implementation, test, live, certification, and freshness status.

Audit Logs and Audit Explorer

Audit Logs is the Auth-backed identity/audit ledger exposed through /platform/audit-events. Audit Explorer is a different, read-only surface: it queries Security Audit's durable audit_events SQL store through the safe event contract. The browser never calls Security Audit directly:

Browser → Control Center → Security Audit GET /audit/events/safe
                                      ↓
                              durable audit_events SQL store

Control Center forwards authenticated safe queries, while Security Audit remains authoritative for tenant scope, authorization, filtering, and the safe projection. Organization callers cannot widen tenant scope. Where required by that upstream contract, GLOBAL and UNKNOWN events are excluded. Metadata is allowlisted, the interface is read-only, and freshness/retention remain UNKNOWN when authoritative evidence is absent. See the Audit Explorer design and evidence.

Admin Console architecture

flowchart TD
    Browser[User Browser] --> Admin[Admin Console]
    Admin --> Control[Control Center API]
    Control --> Auth[Auth]
    Control --> Audit[Security Audit]
    Control --> Bundles[Workflow Bundles]
    Control --> TES[TES]
    Control --> RAG[RAG]
    Control --> Models[Model Registry]
    Control --> Other[Other supported services]

    Producers[Gateway · TES · RAG · Workflow Bundles · LIMS] --> Signed[Signed ingestion]
    Signed --> Audit
    Audit --> SQL[(SQL audit_events)]
    SQL --> Safe[GET /audit/events/safe]
    Safe --> Control
    Control --> Explorer[Audit Explorer]
Loading

Security Audit SAT status

The Security Audit implementation is reconciled against the dedicated audit documentation; these statuses describe implementation and evidence without inventing freshness, retention, or ingestion-lag claims:

SAT Current state
SAT-1 Tenant contract (organization_id, tenant_scope), signing/integrity, and durable SQL persistence are implemented; legacy behavior remains UNKNOWN where evidence is absent.
SAT-2 Producer propagation is merged for confirmed producers: Gateway, TES, RAG, Workflow Bundles, and LIMS. Gateway/RAG/LIMS have live evidence; TES/Workflow Bundles remain fixture-limited. Model Registry is not a confirmed producer; Security SDK adoption remains future work.
SAT-3 Safe tenant-aware audit query authorization is implemented server-side.
SAT-4 Evidence and source semantics are explicit; freshness, retention, and ingestion lag are not inferred.

Production status matrix

Implementation status and live evidence are separate dimensions. The Admin Console as a whole is not yet production-certified.

Surface Implementation Live/evidence status
Admin authentication IMPLEMENTED LIVE-CERTIFIED
Deployment Health IMPLEMENTED LIVE-CERTIFIED
Regression Health IMPLEMENTED LIVE-CERTIFIED
Integration Health IMPLEMENTED LIVE-CERTIFIED
Security Posture IMPLEMENTED LIVE-CERTIFIED
Audit Explorer MERGED PARTIALLY LIVE-CERTIFIED — platform-admin flow certified; populated organization browser evidence remains limited
Workflows MERGED LIVE-CERTIFIED — direct route, hard refresh, sidebar, and history navigation
Tool Execution IMPLEMENTED EVIDENCE LIMITED
AI Models IMPLEMENTED EVIDENCE LIMITED
Agentic AI IMPLEMENTED EVIDENCE LIMITED
Organization administration IMPLEMENTED EVIDENCE LIMITED — whole-console certification remains incomplete

Architecture and certification limitations

  • Full-application TestClient coverage retains pre-existing lifecycle and scheduler teardown debt; isolated async proxy tests pass.
  • Remaining unrelated Playwright assertion/regression debt is outside this README reconciliation and does not invalidate the /workflows routing certification.
  • TES and Workflow Bundles producer propagation has fixture-limited live evidence. Workflow Bundles run history also remains an upstream limitation: its current API is in-memory and does not enforce organization-level filtering.
  • Audit Explorer has no populated organization browser case in the available evidence, and its browser source-failure state lacks a safe temporary live fixture; both are documented as evidence limitations rather than failures.
  • Whole-console production certification, beyond the explicitly tested surfaces above, remains outstanding.

For the implementation details and per-surface authorization boundaries, use the maintained Admin Console guide, Audit Explorer record, and Workflows record.

Enterprise proxy routes

The Admin Console's enterprise service-backed pages use thin routes_*_proxy.py layers — Control Center holds no organization/user/role/ billing data of its own; it forwards those requests to the service that owns them. The local health and report surfaces are implemented by Control Center itself:

Router Example paths Proxies to (env var)
routes_org_proxy.py /orgs, /orgs/{id}, /orgs/{id}/members, /platform/orgs IAM_URL → auth-service
routes_user_proxy.py /platform/users, /platform/users/{id}, /platform/users/{id}/mfa/reset IAM_URL → auth-service
routes_role_proxy.py /platform/roles, /orgs/{id}/roles, /organizations/{id}/permissions IAM_URL → auth-service
routes_team_proxy.py /orgs/{id}/teams, /orgs/{id}/teams/{id}/members IAM_URL → auth-service
routes_service_accounts_proxy.py /orgs/{id}/api-keys, /orgs/{id}/oauth-clients, /platform/permissions IAM_URL → auth-service
routes_org_sso_proxy.py /orgs/{id}/sso, /orgs/{id}/sso/override IAM_URL → auth-service
routes_org_mfa_proxy.py /orgs/{id}/mfa-policy, /orgs/{id}/mfa-policy/override IAM_URL → auth-service
routes_audit_proxy.py /platform/audit-events IAM_URL → auth-service
routes_platform_config_proxy.py /auth/config (read-only) IAM_URL → auth-service
routes_billing_proxy.py /billing/organizations/{id}/usage, /summary, /invoices, /cost-breakdown, /subscription BILLING_URL → billing-service
routes_tes_proxy.py /tes/tools, /tes/tools/capabilities, /tes/runs TES_URL → tes
routes_model_registry_proxy.py /model-registry/models, /health, /auth-status MODEL_REGISTRY_URL → model-registry
routes_workflow_bundles_proxy.py /workflow-bundles/workflows, /categories, /runs WORKFLOW_BUNDLES_URL → workflow-bundles
routes_rag_proxy.py /rag/studies, /rag/cache-stats, /rag/health RAG_URL → rag

Auth forwarding, a uniform 10s timeout, and error mapping (httpx.RequestError → 503, non-JSON upstream body → same status + {"error": ...}) are consistent across all 15 files — audited line-by-line in docs/pr-e-admin-console-production-hardening.md. No secret (RAGBIO_API_KEY, etc.) is ever echoed into a response or logged; each is injected only into the outbound Authorization header.

Deployment status

PR14.7B (merged): dist-admin/dist-control are wired into an actual nginx serving image (docker/nginx/{control-center.conf,api-proxy.conf} + a dedicated control-center-web container), host-based on the Host header, verified on localhost:5174. control-center's FastAPI process itself is unchanged — it never serves static files itself.

Current deployment status: the Admin Console hostname is reachable for the authenticated live-certified flows recorded in the release evidence, including /workflows and Audit Explorer. This confirms the deployed serving path, not production certification of every Admin Console surface. See docs/admin-console-build.md for the build split and deployment architecture.


Regression Health

Reviewed, end-to-end certification status for the OmniBioAI ecosystem — distinct from both /health's live reachability check and Deployment Health's deployment/runtime health below. Certification is a human/process judgment about whether a capability has actually been validated end-to-end; it is never inferred directly from a pytest exit code inside this service.

Reviewed regression run (omnibioai-ecosystem-regression)
        │  promotes a single JSON artifact once reviewed
        ▼
Read-only artifact mount (REGRESSION_HEALTH_ARTIFACT_PATH)
        │
        ▼
Control Center backend (regression_health.py — reads and validates only,
        │                never regenerates or infers the artifact)
        ▼
GET /regression-health   (platform.manage_infra)
        │
        ▼
Admin Console → Operations → Infrastructure → Regression Health

The artifact reports, per P0/P1/P2 phase: status and certification status; a list of capabilities, each with implementation/test/live/ certification status; REG-* findings (fixed/open/closed, with a validation status); and technical debt/paused items. A freshness field — CURRENT / STALE / UNKNOWN, derived from generated_at against REGRESSION_HEALTH_STALE_AFTER_HOURS (default 168h) — is added by Control Center at read time; a stale artifact keeps its certification fields exactly as promoted, never silently downgraded.

Any failure to read or validate the artifact (missing file, malformed JSON, schema mismatch, or a value that looks like it could be sensitive) returns a safe 503 {"status": "STATUS_UNAVAILABLE", ...} — never the configured path, the raw exception, or a stack trace. GET /regression-health requires platform.manage_infra; there is no write path anywhere in this feature, in either direction — the artifact is produced entirely outside Control Center.


Deployment Health

V1, certified. Read-only, dependency-aware deployment and runtime health for every Compose-defined service in the ecosystem — what's declared (Compose), what's actually running (Docker), and what's actually responding (application probes), merged into one inventory with an explicit evidence trail for every fact.

Compose deployment baseline (an explicitly-configured file — see below)
        │
        ▼
Static service/dependency model (deployment_health.py — ownership,
        │                        category, HARD/SOFT dependency edges;
        │                        no Docker, no network, no live service)
        ▼
Runtime merge (deployment_health_runtime.py) ──┬── Docker (existing
        │                                      │    routes_docker.py
        │                                      │    inspection, reused
        │                                      │    verbatim, via the
        │                                      │    Docker Socket Proxy —
        │                                      │    never a raw socket)
        │                                      ├── Application probes
        │                                      │    (existing
        │                                      │    core.runner, the same
        │                                      │    source /services and
        │                                      │    /summary expose)
        │                                      ├── Regression Health
        │                                      │    (compact, global
        │                                      │    context only)
        │                                      └── Prometheus (bare
        │                                           availability flag —
        │                                           see below)
        ▼
GET /deployment-health   (platform.manage_infra)
        │
        ▼
Admin Console → Operations → Infrastructure → Deployment Health

Intrinsic vs. effective health

Every service carries two independent health values — HEALTHY / DEGRADED / UNHEALTHY / UNKNOWN — never collapsed into one:

  • Intrinsic health is the service's own observed health, in strict precedence order: an application probe result, then a Docker healthcheck, then bare "container is running," then UNKNOWN. A running container with no probe and no healthcheck is UNKNOWN, not HEALTHY — a container running is not proof a service is healthy.
  • Effective health adjusts intrinsic health for HARD dependency evidence only (from Compose depends_on conditions; service_healthy/service_completed_successfully are HARD, service_started or no condition is SOFT). A failing SOFT dependency never changes effective health. UNHEALTHY is reserved exclusively for intrinsic failure — a dependency problem can push a service to DEGRADED at most, never rewrite its own intrinsic state. Propagation looks only at each dependency's intrinsic value, one hop, which is what makes it safe against dependency cycles by construction rather than something detected at runtime.

UNKNOWN does not mean unhealthy. It means there wasn't enough evidence to say either way — a service with no configured application probe and no Docker healthcheck, a one-shot init container that has already exited, or a source (Docker, Regression Health) that's temporarily unavailable. None of these are ever silently reported as HEALTHY.

Ownership and third-party services

Each service's owning repository is derived from real evidence — a Compose build-context path segment, a published image reference, or a small curated fallback table — never guessed from a name pattern. Genuine third-party infrastructure (MySQL, Redis, Neo4j, Prometheus, Grafana, OPA, and similar) legitimately has no OmniBioAI repository and reports repository: null — this is a normal, expected state, entirely independent of and never implying anything about that service's health.

Evidence and data-source availability

Every health determination is backed by an evidence entry naming its source and a safe (never-sensitive) detail. The response's data_sources block reports each source's own current state:

Source States Notes
Compose available The only source whose failure is fatal to the whole endpoint (503) — every other one below degrades independently
Docker available / unavailable Via the existing Docker Socket Proxy; unavailable → every dependent service's runtime falls to UNKNOWN, never a fabricated HEALTHY
Application Probe available / unavailable The existing config/control_center.yaml-driven checks
Regression Health available / unavailable Compact context only — see Regression Health; never mapped to individual services, since certification is capability-oriented, not per-service
Prometheus not_configured (current deployments) / available / unavailable No scrape configuration exists anywhere in this workspace as of V1 — reported honestly as deferred, not presented as integrated. No per-service Prometheus correlation exists yet

Configuration

Variable Description
DEPLOYMENT_HEALTH_COMPOSE_PATH Path to the authoritative Compose file. Unset → the endpoint returns the same safe unavailable response as a missing artifact; there is no built-in default path
DEPLOYMENT_HEALTH_BASELINE_SOURCE development / release / unknown (default) — which Compose baseline the supplied path represents; never guessed
DEPLOYMENT_HEALTH_DEPLOYMENT_REPOSITORY Optional logical repo name (e.g. omnibioai-studio) used only to attribute ownership for a service whose build context is a genuinely relative path inside that same repository

Deployment Health does not require its own dedicated mount — in every currently certified deployment, the Compose file it reads was already reachable through a mount that existed for unrelated, pre-existing reasons; these three variables only point the parser at it. No .env value is ever read into the model, and no environment value, secret, absolute host path, container ID, or backend handle is ever part of a response — every serialized field is drawn from an explicit allowlist, never a recursive dump of a Docker or Compose structure.

V1 certification

Deployment Health V1 has been certified against a real, live deployment — not only automated tests: a live authorization matrix (401 anonymous, 403 authenticated-without-permission, 200 authorized), the real service inventory, Docker/runtime merge, intrinsic-vs-effective health (including a real HARD-dependency-triggered degradation), Docker-unavailable and Compose-missing/invalid failure semantics, Regression Health source failure, image-reference comparison (MATCH/MISMATCH/UNKNOWN, latest never treated as a verified release), Admin Console rendering, and read-only enforcement. At certification time the live baseline had 41 services and 87 dependency edges — counts like these are a snapshot of that run, not a fixed architecture fact, and will change as the underlying Compose baseline does. Full certification record: docs/DEPLOYMENT_HEALTH_DH1.md through DH4.md.

Deployment Health V1 is strictly read-only. It does not provide restart, stop, kill, deploy, recreate, scale, image-pull, arbitrary Docker control, arbitrary Slurm control, or any health-override action — GET is the only method the endpoint accepts (verified live: every other method returns 405).

Not yet built, deliberately: source/commit/image-digest drift beyond the narrow tag comparison above, multi-file Compose overlay resolution, and per-service Prometheus correlation.


Running Tests

cd backend
pip install -e ".[dev]"
pytest tests/ -v

Test coverage

File What it tests
test_checks.py TCP, HTTP, and disk check modules; /health and /report routes
test_discord.py Discord webhook notification helper
test_gpu.py GPU checks — nvidia-smi temperature polling and full /gpu status (memory, utilization, processes, Ollama models)
test_main.py FastAPI app lifecycle — job state machine, dashboard/report rendering, background report job runner, scheduler loop, startup hook
test_routes_cloud.py /cloud — execution backend status
test_routes_config.py Config-loading routes backing the dashboard/report UI
test_routes_docker.py /docker/* — container inventory and tool SIF image status
test_routes_llm.py /llms and /knowledge-base — Ollama/API key status, PubMed abstract and index-size scanning
test_routes_reference.py /reference — reference genome registry
test_routes_storage.py /storage — disk usage and per-organism index sizes
test_runner.py Check-runner service-type dispatch, settings loading, /services and /summary routes
test_summary_client.py /summary fetch/parse helpers, report-generator health-parsing helpers
test_check_activity.py /activity — Prometheus-backed container/host resource metrics
test_check_audit_trail.py /audit-trail — Redis audit-stream aggregation (event/decision/reason breakdowns)
test_check_celery_status.py /celery — worker online/offline detection, recent-task parsing from the Redis result backend
test_check_database_status.py /database — MySQL, Redis, and Neo4j live status
test_check_gateway_traffic.py /gateway-traffic — API gateway request/latency/status-code aggregation from the audit stream
test_check_image_freshness.py /image-freshness — local vs. registry :latest image digest comparison
test_check_integrity.py /integrity — configured symlink/mount health checks
test_check_license_status.py /license — license-seat/expiry status derivation
test_check_usage_status.py /usage — user activity, session counts, plugin-run success-rate stats
test_routes_infra.py Wiring for all /gpu, /celery, /database, /image-freshness, /license, /usage, /gateway-traffic, /audit-trail, /activity, /integrity routes
test_core_auth.py require_admin — JWT decode, expiry, missing/invalid token, role check
test_check_cron_jobs.py Cron-job status derivation and the pause/resume/reschedule crontab-mutation logic
test_routes_cron.py /cron/jobs and its admin-gated mutation routes
test_check_known_issues.py Known-issue load/create/update/delete logic, including UUID backfill
test_routes_known_issues.py /known-issues CRUD routes, read-open/write-admin-gated
test_jwt_verify.py core/jwt_verify.py — HS256 + RS256/JWKS verification, revocation, alg-header dispatch
test_routes_dashboard.py /dashboard/summary — Overview page stat cards
test_routes_auth_proxy.py /auth/login, /auth/refresh, /auth/logout, /auth/validate proxy relay + cookie forwarding
test_routes_org_proxy.py /orgs, /platform/orgs proxy routes
test_routes_user_proxy.py /platform/users proxy routes, including MFA reset
test_routes_role_proxy.py /platform/roles, /orgs/{id}/roles, /organizations/{id}/roles proxy routes
test_routes_team_proxy.py /orgs/{id}/teams proxy routes
test_routes_service_accounts_proxy.py /orgs/{id}/api-keys, /orgs/{id}/oauth-clients proxy routes
test_routes_org_sso_proxy.py /orgs/{id}/sso proxy routes
test_routes_org_mfa_proxy.py /orgs/{id}/mfa-policy proxy routes
test_routes_audit_proxy.py /platform/audit-events proxy route
test_routes_platform_config_proxy.py /auth/config read-only proxy route
test_routes_billing_proxy.py /billing/* proxy routes (usage, summary, invoices, cost-breakdown, subscription)
test_routes_tes_proxy.py /tes/* proxy routes
test_routes_model_registry_proxy.py /model-registry/* proxy routes
test_routes_workflow_bundles_proxy.py /workflow-bundles/* proxy routes
test_routes_rag_proxy.py /rag/* proxy routes
test_routes_regression_health.py regression_health.py artifact reading/validation + GET /regression-health (freshness, safe unavailable, authorization)
test_deployment_health.py deployment_health.py — DH-1 static Compose service/dependency parsing
test_deployment_health_runtime.py deployment_health_runtime.py — DH-2 Docker/probe merge, intrinsic/effective health, dependency propagation
test_routes_deployment_health.py GET /deployment-health — authorization, safe failure semantics, data-source degradation
test_nginx_config.py Static assertions on docker/nginx/api-proxy.conf — the resolved-upstream $control_center_upstream pattern, and the Regression/Deployment Health SPA-vs-API route separation

Most tests are self-contained (in-process HTTP servers, real temp-dir filesystem fixtures, no real external services). The checks/*.py and routes_docker.py/routes_llm.py suites additionally mock subprocess (docker/nvidia-smi CLI calls) and network clients (httpx, redis, pymysql, neo4j, celery) at the call site — no real database, broker, GPU, or Docker daemon is required at test time.


Design Principles

  • Stateless-ish — no database; the only persistent writes are the YAML config (/config/service), known_issues.json, and — for admins only — the host crontab itself
  • Config-driven — add services via YAML, no code changes
  • Graceful degradation — unreachable services show DOWN, never crash the dashboard
  • Zero mandatory cloud — runs fully offline and air-gapped
  • Focused dependencies — FastAPI, uvicorn, PyYAML, pydantic, JWT, Prometheus, Celery/Redis, database clients, and report-generation libraries
  • stdlib HTTP in reporturllib used for health fetching in report generator, no extra deps
  • Design-token driven — CSS uses @omnibioai/design-tokens vocabulary; zero hardcoded hex values in the report or dashboard
  • Structured logging — all key events (startup, report triggered/finished/failed, scheduler) emitted as JSON to stdout
  • Defense-in-depth on writes — every admin-gated endpoint checks the JWT's role independently inside the app (require_admin), rather than trusting nginx's auth_request alone
  • Honest scope over convenience/coverage/generate only runs on control-center itself rather than faking full-ecosystem coverage from inside a container that can't actually run the other repos' test suites (see /coverage/generate in API Endpoints)
  • No raw Docker socket — every Docker-touching endpoint, including Deployment Health, goes through the Docker Socket Proxy's own restricted allowlist; this service never bind-mounts /var/run/docker.sock directly
  • Read-only where read-only is claimed — Regression Health and Deployment Health each expose only GET; neither has a write path in this service (Regression Health's artifact is produced entirely outside it), and Deployment Health V1 has no restart/stop/deploy/scale/config-edit/health-override action of any kind

Planned Enhancements (Post-Beta)

  • Historical uptime tracking
  • Alert hooks (Slack, email)
  • /auth/login audit trail — it currently bypasses api-gateway's AuditMiddleware entirely (nginx proxies /auth/* straight to auth-service), so there's no IP-level or attempt-level record of login attempts anywhere in the ecosystem. Needs deliberate design (what to log, where, without creating a new PII/security concern in the audit stream itself), not a quick patch — tracked as a known issue in the Admin tab

Current Status — repository snapshot (2026-08-31)

Feature Status
HTTP health checks ✓ Stable
TCP checks (MySQL, Redis) ✓ Stable
Disk usage checks ✓ Stable
JSON summary API ✓ Stable
Ecosystem report — Architecture ✓ Stable
Ecosystem report — Projects ✓ Stable
Ecosystem report — Languages ✓ Stable
Ecosystem report — Coverage ✓ Stable
Ecosystem report — Health tab ✓ Stable
Unit tests ✓ Frontend and scoped backend suites pass; full-app TestClient lifecycle/scheduler teardown debt remains
Docker deployment ✓ Root Dockerfile is current; Compose file is an ecosystem template with an external build context
Prometheus metrics (/metrics) ✓ Stable
Scheduled report generation ✓ Stable
JWT authentication (via nginx) ✓ Stable
Browser session cookie + silent refresh ✓ Stable
RS256/JWKS verification (local jwt_verify.py) ✓ Ready — HS256 still the production default
Admin-role gating (app-level, defense-in-depth) ✓ Stable
Admin tab — Actions (report/coverage regen) ✓ Stable
Admin tab — Scheduled Jobs (7 cron jobs, pause/resume/reschedule) ✓ Stable
Admin tab — Known Issues (live CRUD) ✓ Stable
Coverage collection (/coverage/generate, control-center-scoped) ✓ Stable
npm-audit vulnerability scanning (CI/CD Health tab) ✓ Stable
Background report job API ✓ Stable
Docker inventory endpoints ✓ Stable
Structured JSON logging ✓ Stable
Design token CSS alignment ✓ Stable
LLM monitoring (/llms) ✓ Stable
Cloud backend status (/cloud) ✓ Stable
Reference genome registry ✓ Implemented; organism/index availability is deployment-dependent
AI Knowledge Base (/knowledge-base) ✓ Implemented; corpus size is deployment-dependent
Storage monitoring (/storage) ✓ Stable
Report — LLMs tab ✓ Stable
Report — Cloud tab ✓ Stable
Report — Reference Data tab ✓ Stable
Report — AI Knowledge Base tab ✓ Stable
Report — Storage tab ✓ Stable
Sidebar navigation (report UI) ✓ Stable
GPU health detection (/gpu) ✓ Stable
Audit Trail (/audit-trail) ✓ Stable
CVE Trend ✓ Stable
Admin Console — Organizations, Users, Teams, Roles & Permissions ✓ Stable
Admin Console — Security (Security Overview, Security Posture, MFA Policy, IAM/SSO, SAML, Audit Logs, Audit Explorer, Compliance Report, Sessions, Interactions, API Keys/Service Accounts) ✓ Implemented; selected surfaces live-certified
Admin Console — HIPAA Compliance ✓ Stable
Admin Console — Billing, Usage Analytics, Workflows, Tool Execution, AI Models, Agentic AI, RAG/PubMed, Integrations, Settings ✓ Implemented; evidence varies by surface
Regression Health V1 Implemented / Certified — see Regression Health
Deployment Health V1 Implemented / Certified — see Deployment Health. Certification-time live baseline: 41 services, 87 dependency edges — a snapshot of that run, not a fixed count
Admin Console dual-build (dist-admin/dist-control, nginx host-based split) ✓ Implemented
Admin Console tested production flows (admin.omnibioai.org) LIVE-CERTIFIED — selected authenticated routes; whole-console certification remains incomplete
Historical tracking Planned
Alert hooks (Slack, email) Planned
/auth/login audit trail Planned — needs deliberate design, see Planned Enhancements
Deployment Health — source/commit/image drift beyond image-tag comparison Future
Deployment Health — multi-file Compose overlay resolution Future
Deployment Health — Prometheus per-service correlation Future — no scrape config exists anywhere in this workspace yet
Deployment Health — historical health/trend views Future

License

Apache License 2.0

About

FastAPI-based control plane for OmniBioAI — provides HTTP/TCP/disk health monitoring across all platform services, live operational dashboards, and ecosystem-wide report generation (architecture diagrams, codebase stats, test coverage, system health summaries). Single pane of glass for platform operators.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages