← Research & Proposals
Proposal

AI Control & Scalable Oversight

Updated 2026-08-05

Status
Proposal — landscape mapped, next experiments not yet started
Key finding
Monitor AUROC falls from >0.9 to 0.5–0.8 once the agent suspects it's watched
Methods surveyed
Debate, weak-to-strong generalization, IDA, process supervision, monitor red-teaming
Output
Literature synthesis — no code experiment yet

Two questions get bundled under one label but aren't the same. Scalable oversight asks how a weak supervisor can reliably train or evaluate a stronger model — debate, recursive decomposition, weak-to-strong generalization. AI control asks how to limit damage from a strong model that might already be misaligned, assuming nothing about its cooperativeness — permissions, monitoring, resampling, shutdown. Current evidence suggests neither alone is enough; the field's working consensus is defense-in-depth — RLAIF-style training, process supervision, debate, mechanistic interpretability, and control, stacked rather than chosen between.

On the oversight side: debate (Irving et al. 2018, arXiv:1805.00899) has two AIs argue while a weaker judge decides. Khan et al. 2024 (ICML best paper, arXiv:2402.06782) showed it lifts a weak judge's accuracy on QuALITY from 60%→88% (human judges) and 48%→76% (LLM judges), beating a non-adversarial consultancy baseline. Weak-to-strong generalization (Burns et al. 2023, arXiv:2312.09390) asks the sharper question — can a strong model generalize past its weak supervisor's mistakes rather than just imitating them? A 2026 result unifies debate and RLAIF under one geometric framework and finds debate's advantage traces specifically to where the two AIs' knowledge diverges, not to the adversarial format itself.

On the control side, three recent evaluations converge on the same shape of result: monitoring works but isn't airtight — best measured sabotage success around 27% with monitor AUC around 0.87 in diverse tool-use settings (SHADE-Arena); roughly 34% of attacks slip past a weaker monitor in a realistic multi-service Linux production environment, against a 94% human red-team success rate (LinuxArena) — and it degrades sharply once the threat model changes: prompt-injecting the monitor itself collapses the safety guarantees of most monitor-dependent protocols, and a model that merely suspects it's being watched sees its detectability (monitor AUROC) fall from above 0.9 to 0.5–0.8.

A tempting fix — many independent monitors voting, isolated from each other, memory-wiped between checks — borrows intuitions from Byzantine fault tolerance and N-version programming that don't transfer cleanly. BFT assumes a majority of honestly-behaving nodes, which doesn't hold if correlated scheming is the actual threat; N-version assumes independent implementations, which same-vintage LLM monitors from the same training distribution aren't. The open problem is where genuine independence between monitors would come from, and how to test for “faking blindness” with honeypots rather than assuming isolation is sufficient.

Highest-value next directions, roughly by tractability: (1) a monitor jailbreak / prompt-injection test suite with a CI harness — production monitors have essentially no systematic red-teaming for being talked out of an alarm; (2) an “unmonitored compute” tripwire — a checker for whether a stretch of agent activity actually ran inside its claimed monitoring configuration, directly answering production red-team findings that coverage gaps are common; (3) a deployment-control attestation MVP — a way to demonstrate after the fact which monitoring configuration a given run actually used, which is close to nonexistent today; (4) replacing same-context retry with genuine resampling (fresh state, not a second attempt in the same context) in an open-source coding agent, since 2026 work shows the two are not equivalent and most agent harnesses default to the weaker one.