A python library built on top of UKAISI Inspect to support Control research, by Redwood
-
Updated
Aug 20, 2026 - Python
A python library built on top of UKAISI Inspect to support Control research, by Redwood
Research developing AI control protocols using task decomposition
🚦🗺️ UrbanFlow AI is a web app for generating 3D traffic simulations from real OpenStreetMap areas. It builds SUMO scenarios, runs microscopic vehicles, pedestrians, buses and trams, edits road events, controls real traffic lights with TraCI, trains JSON AI policies, saves models, and shows live metrics, charts, and notebooks.
An intelligent traffic management system that dynamically adjusts highway lane configurations using AI-powered congestion detection and a movable median barrier.
In-depth exploration of Large Language Models (LLMs), their potential biases, limitations, and the challenges in controlling their outputs. It also includes a Flask application that uses an LLM to perform research on a company and generate a report on its potential for partnership opportunities.
Judge-first framework where LLM outputs must converge under explicit, adversarial oracles.
Python client for Aegis — stabilize AI systems instantly with a simple API call.
AI Governance — human-controlled code injection with oversight
Independent AI governance and control standard
White-box detection of collusion in an untrusted monitor: model organisms, linear probes, and a control evaluation that prices what they buy.
Kho lưu trữ này chứa tài liệu, bài tập, và mã nguồn liên quan đến môn Trí tuệ nhân tạo trong điều khiển. Môn học tập trung vào ứng dụng AI trong các hệ thống điều khiển tự động, bao gồm lý thuyết và thực hành.
An attacker × monitor factorial in ControlArena: which model you pick as your trusted monitor matters more than its capability tier.
Lichtarbeit
v0.8 preview · 2,288 multi-step agent trajectories · 513 hand-authored gold + 1,775 provenance-flagged augmented. A per-step benchmark that scores whether a verifier catches drift inside an agent's trajectory — not whether a prompt is harmful.
GG Tank Watch - frozen public-information archive of a resolved May 2026 chemical emergency. Conduit-only design; responsible-AI safety patterns enforced in code and tests.
AutoRed: Measuring the Elicitation Gap via Automated Red-Blue Optimization — AI Control Hackathon 2026
The project evaluates whether a lightweight hardening layer can recover robustness without redesigning the surrounding protocol. The method combines transcript sanitization, explicit instruction-data separation, and an optional prompt-diversity ensemble.
A compact AI-control benchmark for trusted monitoring of untrusted coding agents.
Control protocols robust to a monitor-aware adversary — empirical AI-control eval on the APPS backdoor testbed (open-weight models).
Trace-based agent safety monitor trained on malicious PyPI packages
Add a description, image, and links to the ai-control topic page so that developers can more easily learn about it.
To associate your repository with the ai-control topic, visit your repo's landing page and select "manage topics."