Shekhar Tiruwa
Speaking appearances
- AI Safety Thursdays: Reasoning Models Don't Always Say What They Think 2025-06-12 — Toronto
- AI Safety Thursdays: Tracing the Thoughts of a Large Language Model 2025-06-05 — Toronto
- AI Safety Thursdays: Understanding The Self-Other Overlap Approach 2025-05-22 — Toronto
- AI Safety Thursdays: When Good Rewards Go Bad - Reward Overoptimization in RLHF 2025-05-15 — Toronto