Leo Zovic
Events
- Could Frontier Labs’ Internal Agents Already Go Rogue? 2026-06-04
- Testing LLM Cooperation in Multi-Agent Simulation 2026-04-16
- Adversarial Defenses for LLMs 2026-03-26
- The Anthropic-Pentagon Stand-Off 2026-03-24
- Open-weight LLMs – Their Strategic Import to Key AI Jurisdictions and The Growing Need for Safety Research, Dialogues, and Coordination 2026-03-10
- Network Topologies for AI and the Implications for Governance 2026-03-05
- AI Safety Thursday: Risks Emerging from Agent Swarms 2026-02-19
- AI Safety Thursday: Claude's New Constitution 2026-02-12
- AI Safety Thursday: Beyond Adversarial Robustness - Rethinking Sociopolitical Safety in AI Systems 2026-01-22
- AI Safety Thursdays: Why Attackers Are Winning and What We Can Do About It 2026-01-08
- AI Safety Thursday: Agentic Bug Detection - Progress and Deployment 2025-12-18
- AI Safety Thursday: Sandbagging - How Models Use Reward-Hacking to Downplay Their True Capabilities 2025-11-27
- AI Safety Thursday: Agentic property-based testing - finding bugs across the Python ecosystem 2025-11-13
- AI Safety Thursday: Monitoring LLMs for deceptive behaviour using probes 2025-11-06
- IABIED Reading Group: "Part 1: Nonhuman Minds" 2025-10-27
- AI Safety Thursday: The Limitations of Reinforcement Learning for LLMs in Achieving AI for Science 2025-10-23
- "If Anyone Builds It..." Reading Group 2025-10-20
- AI Safety Thursday: Modeling and Detecting Deceptive Alignment 2025-10-16
- AI Policy Tuesday: The Case for Regulating AI Companies, Not AI Models 2025-09-30
- AI Policy Tuesdays: Frontier AI Deployments in US National Security and Defence 2025-09-02