AI Alignment
Systems That Pursue the Right Goals
The most capable AI system in the world is dangerous if it is pursuing the wrong objective. AI alignment is the field dedicated to ensuring that AI systems genuinely pursue what we intend, not a flawed approximation of it.
This is harder than it sounds. Human values are complex, context-dependent, and often difficult to articulate, let alone formalize into code. An AI system instructed to ‘maximize housing placements’ might optimize the speed and volume of assignments while overlooking the stability, safety, and neighborhood fit that determine whether a placement works for a family. At scale, these misalignments can compound in ways that are difficult to detect and costly to reverse.
The Lab’s alignment research focuses on how AI systems learn, represent, and act on human values, developing tools for interpretability (understanding what a system has learned to optimize), methods for robust value specification, and frameworks for evaluating whether deployed systems are behaving as intended. The stakes are high: as AI systems take on greater autonomy and decision-making authority, getting alignment right is not a technical nicety. It is a precondition for trust.
AI Policy
Translating Research into Governance That Holds
Technical solutions alone are insufficient. A field can produce world-class research on alignment, safety, and security and still fail to protect the public if that research does not translate into enforceable standards, effective regulation, and accountable institutions. AI policy is the bridge between what we know and what society actually requires.
This is not a peripheral concern. Across the United States and globally, governments are actively shaping the legal and regulatory environment for AI, often without adequate technical input. The decisions being made now about liability frameworks, disclosure requirements, procurement standards, and international coordination will shape the trajectory of AI development for decades. The window to influence those decisions thoughtfully, with evidence and rigor, is narrow.
The Lab engages at the policy level through applied research, expert testimony, practitioner convenings, and direct engagement with legislators and regulators. We translate technical insights from our alignment, safety, and security work into recommendations that policymakers can act on, and we help institutions build the internal capacity to govern AI responsibly, not just comply with minimum requirements.
AI Safety
Ensuring Reliable Behavior in a Complex World
Even well-intentioned systems fail. They encounter situations their designers didn’t anticipate. They are deployed in contexts far removed from their training environments. They interact with other systems in ways that produce unpredictable outcomes. AI safety research addresses how AI behaves under these conditions and how to build systems that remain reliable, correctable, and constrained in their impact when things go wrong.
This is the engineering discipline of responsible AI. Where alignment asks what a system is trying to do, safety asks whether it will do so reliably and without causing harm, especially in high-stakes environments like housing, employment, criminal justice, medical diagnosis, infrastructure management, and autonomous vehicles, where errors carry real consequences for real people.
The Lab’s safety work spans robustness testing (how systems perform under distribution shift and edge cases), corrigibility research (ensuring AI systems can be monitored, corrected, and shut down), and the development of evaluation benchmarks that practitioners and regulators can apply across sectors. Safety is the discipline that transforms alignment from a design intention into a deployable guarantee.
AI Security
Protecting Systems From Those Who Would Exploit Them
AI systems are now becoming targets. Bad actors actively probe weaknesses, seeking to manipulate, deceive, and exploit AI systems for harmful ends. A model that behaves safely under normal conditions may be vulnerable to adversarial attacks that cause it to fail precisely the moments that matter most.
AI security is the field that addresses this threat landscape. It covers prompt injection, i.e. tricking a system into ignoring its instructions; data poisoning, i.e. corrupting training data to embed hidden behaviors; adversarial inputs, i.e. crafted to fool classifiers and perception systems; and model extraction, i.e. reverse-engineering proprietary systems through strategic queries. These are not hypothetical concerns. They are documented, recurring attack vectors that have been demonstrated across commercial, government, and research deployments.
The Lab’s security research assumes the presence of an intelligent adversary. This is a fundamentally different threat model from safety concerns with accidents and errors. Our work produces both technical defenses and operational guidance, equipping organizations to identify vulnerabilities before they are exploited and to respond effectively when they are.