Theory of Mind / Empathy Dissociation
Does an open-weight LLM represent cognitive theory-of-mind and affective empathy as separate, dissociable directions — the way TPJ/mPFC and insula/ACC are separate in the human brain?
Read more →Shortcut Forensics
Does a single interpretable direction in a coding agent's residual stream drive shortcut-taking under pressure, or is it several independent mechanisms wearing the same disguise?
Read more →AI Control & Scalable Oversight
Weak models supervising strong ones has real, documented failure modes — and real proposed fixes. Mapping the landscape toward the highest-value next experiments.
Read more →Open-Weight Model Governance
Open-weight models are mostly a public good — until they aren't. Mapping where the real policy disagreement lies, and what's actually being proposed instead of a blanket ban.
Read more →Tamper-Induced Capability Collapse
Refusal-direction steering makes an open-weight model jailbreak easily — a persona swap alone gets ~92% harmful compliance. A training-time defense that welds safety into capability flips that: tampering with one wrecks the other.
Read more →U.S.–China & International AI Governance
Government-to-government AI dialogue between the US and China has happened exactly once. Almost everything else people cite as “coordination” is track-two, scientist-led, or a side conversation at a multilateral summit.
Read more →Compute-Pause Verification
Can cross-node communication size and shape separate benign LLM pretraining from inference? A hardware-based mechanism for verifying AI compute-governance agreements under low trust.
Read more →Claims-Bench
Characterizing language-model agents' implicit moral and stakeholder commitments — what value profile does a model reveal when stakes are unclear and reasonable people disagree?
Read more →AI Futures Simulator
A Monte Carlo simulator of AI-related futures (2026–2050) — capability growth, plot events, and society feedback compounding into emergent outcome regions.
Read more →Governing the Machine Economy
AI agents are becoming autonomous economic actors. The infrastructure for that economy is being built now — the governance layer that ensures it benefits human society is not.
Read more →