Research on safe, trustworthy, and verifiable AI systems.
Selected reading on narrow finetuning and broad behavioral change.