Tobias Thaller
Events
- AI Policy Tuesday: Verification Mechanisms for Global AI Governance 2025-12-09
- AI Safety Thursday: Sandbagging - How Models Use Reward-Hacking to Downplay Their True Capabilities 2025-11-27
- Online Animal ACTION PARTY - EU consultation on animal welfare 2025-11-26
- AI Safety Thursday: Introduction to Corrigibility 2025-11-20
- AI Safety Thursday: Agentic property-based testing - finding bugs across the Python ecosystem 2025-11-13
- AI Safety Thursday: Monitoring LLMs for deceptive behaviour using probes 2025-11-06
- AI Policy Tuesday: Debunking the US-Chinese AGI Race 2025-10-28
- AI Safety Thursday: The Limitations of Reinforcement Learning for LLMs in Achieving AI for Science 2025-10-23
- AI Policy Tuesday: Redlines for AI 2025-10-14
- AI Safety Thursday: Building an economic model of AI automation 2025-10-09
- AI Safety Thursday: Attempts and Successes of LLMs Persuading on Harmful Topics 2025-10-02
- AI Safety Thursday: Technical AI Governance - Motivations, Challenges, and Advice 2025-09-18