Eval Engineering for AI app developers - Lesson 1: Hello Evals!
- Date
- 2025-12-09
- Host
- Galileo Events
About this event
Building an AI app is one thing. Knowing whether it actually works, keeps working, and improves over time is another. Eval Engineering for AI app developers - Lesson 1: Hello Evals! is a practical in-person session for developers who want to move beyond vibes and start measuring the quality of AI behavior with more confidence. If you have shipped prompts, agents, copilots, or internal AI tools and found yourself asking, "How do I know this is better?" this event is for you. The focus here is not abstract theory. It is about getting a clear mental model for evals and understanding how they fit into real AI application development. What Is This? This event is the first lesson in a series focused on eval engineering for AI app developers. As the title suggests, this session starts at the beginning: what evals are, why they matter, and how to think about them as part of your development workflow rather than as an afterthought. The format is designed for people who build. Expect a developer-friendly introduction that treats evals as an engineering discipline: something you can structure, reason about, and iterate on. Whether you are working on autonomous systems, chat experiences, or other AI-powered product features, the goal is to help you leave with a more usable framework for testing quality. Because this is an in-person gathering, the event also creates space for conversation with other builders working through similar questions. That matters. Evals are one of those topics that becomes much clearer when you can compare approaches, tradeoffs, and pain points with people doing the work in real settings. What to Expect This session centers on the foundations. You should expect a focused introduction to the role evals play in modern AI development, especially when traditional software testing falls short. AI systems can be probabilistic, context-sensitive, and highly dependent on prompt design, retrieval quality, and model behavior. That means the way you assess them needs to be different too. Topics will likely include: What an eval actually is in the context of AI apps Why evals matter early, not just after launch How developers think about quality when outputs are open-ended Where evals fit in iteration cycles for prompts, agents, and product features Common failure modes that evals are meant to catch You can also expect the event to frame evals as part of a broader engineering process, not a disconnected research concept. For many developers, the biggest unlock is simply learning how to define success more clearly. Once you can describe what "good" looks like, you are in a much better position to improve systems systematically. Because this is labeled Lesson 1, the emphasis is on shared foundations rather than advanced specialization. That makes it a strong entry point if you are new to evals, but it should also be useful if you have been doing ad hoc testing and want a cleaner vocabulary and structure for what you are already attempting. Why Attend If you are building AI features today, evals are quickly becoming core infrastructure. You can write prompts, tune workflows, and swap models endlessly, but without a way to assess outcomes, improvement becomes guesswork. This event helps replace guesswork with a more disciplined way of thinking. Attending can help you: Understand the basic language of eval engineering so you can reason about quality more precisely Spot gaps in your current workflow where testing is too informal or inconsistent Learn how to evaluate AI behavior in a way that matches product reality, not just benchmark abstractions Connect with other developers facing the same reliability and iteration challenges Build a stronger foundation for future lessons or deeper work in AI evaluation There is also a practical career angle. Teams increasingly need developers who can do more than integrate models. They need people who can make AI systems dependable. A solid grasp of evals helps you contribute at that level, whether you are working on internal tools, customer-facing products, or autonomous workflows. Just as importantly, learning evals early can save time. Instead of endlessly debating whether one version feels better than another, you can begin to create clearer criteria, compare changes more intentionally, and make development decisions with more signal. Practical Details This is an in-person event taking place on Tuesday, December 9 at 9:00 AM PST. If you prefer learning in a room with other technically curious people, asking questions live, and having real conversations before or after the session, this format is a strong fit. A few useful things to keep in mind: Arrive a little early so you can get settled before the session starts Bring your current questions about testing prompts, agents, or AI workflows Expect a developer-oriented discussion rather than a broad non-technical overview Come ready to think practically about how evals apply to your own projects This event is especially well suited to builders who want to sharpen their foundations before going deeper. If you have heard people talk about evals and know they matter, but have not yet developed a clear framework for using them, this is the right place to start.
Who should attend
This is for builders who want a more rigorous way to improve AI products without losing sight of real-world development constraints. - You are an **AI app developer** shipping features with LLMs and want a clearer way to measure whether changes actually improve the product. - You are building **agents, copilots, chat interfaces, or autonomous workflows** and keep running into the question: how do we test this reliably? - You already do some informal testing, prompt comparisons, or manual reviews, but you want a **stronger framework and vocabulary** for evals. - You work on a product or internal tool where **output quality, consistency, and failure handling** matter, and you need something more structured than gut feel. - You are technically curious about eval engineering and want a **good entry point** that starts with fundamentals instead of assuming deep prior knowledge. - You value **in-person learning and peer conversation** with other developers navigating the same AI reliability challenges. If you have been moving fast with AI development and want to add more discipline to how you assess quality, this session should feel immediately relevant.