Evals with OpenAI

Date
2026-04-22
Host
Gradient Descending

About this event

If you care about building AI systems that actually work in the real world, evaluation is not optional. Evals with OpenAI is an in-person meetup for people who want a clearer, more practical handle on how to test, compare, and improve model behavior with confidence. This is a community-focused session built around one of the most important topics in modern AI development: knowing whether your system is getting better, not just sounding better. Expect a thoughtful, social, and grounded conversation with others who are actively thinking about quality, reliability, and what good evaluation looks like in practice. About the Event At its core, this event is about making evaluation feel concrete. "Evals" can mean many things depending on what you build, from benchmarking prompts and workflows to reviewing outputs for accuracy, usefulness, consistency, or safety. This meetup creates space to explore those questions with other people who are working through similar challenges. Because this is an in-person community event, the format is likely to feel more interactive than a formal conference session. Rather than sitting through a long series of polished presentations, attendees can expect a meetup atmosphere where ideas are easier to exchange, questions are easier to ask, and conversations can move from high-level concepts into practical detail. The OpenAI focus signals that the conversation will be especially relevant for people working with modern language models and evaluation workflows around them. Whether your interest is technical, product-focused, or strategic, the event is designed to center on how evals help teams make better decisions. What to Expect You should expect a meetup that blends learning, discussion, and networking. The morning timing makes this well suited to people who want to start the day with a substantive session and leave with ideas they can take back into their work right away. Likely themes of conversation may include: What evals are actually for and how different teams define success How to think about quality when model outputs are nuanced, subjective, or context-dependent Ways to compare prompts, workflows, or model behaviors in a structured way Tradeoffs between speed and rigor when building evaluation processes Common pitfalls in judging outputs by intuition alone How teams use feedback loops to improve systems over time Because this is tagged as a community, networking, meetup, and social event, the interpersonal side matters too. Expect opportunities to meet other builders, practitioners, and curious attendees who want to talk seriously about AI quality without turning the conversation into jargon for its own sake. The value of an in-person setting is that you can go beyond abstract ideas. You can ask follow-up questions, pressure-test your thinking, hear how others frame similar problems, and leave with a more useful mental model of what evaluation means in practice. Why Attend If you have ever launched a feature and wondered whether it was truly better, this event speaks directly to that uncertainty. Evaluation is how teams move from vibes to evidence, and from one-off demos to systems they can trust. Spending time on this topic can sharpen how you build, measure, and iterate. You do not need to arrive with a perfect evaluation framework already in place. In fact, this kind of meetup is especially useful if you are still figuring out what "good" looks like for your use case. Hearing how others think about tests, benchmarks, review criteria, and iteration can save time and prevent avoidable mistakes. Attending can help you: Clarify your own evaluation challenges by hearing them discussed in shared language Learn practical ways to think about model quality beyond surface-level impressions Meet people facing similar problems across product, engineering, research, or experimentation work Build better questions to take back to your team, project, or workflow Strengthen your judgment around what to measure and why it matters Even if you are early in your AI journey, evals are worth understanding now. The sooner you start thinking clearly about measurement and reliability, the better your decisions will be as your projects grow in complexity. Practical Details This is an in-person event taking place on Wednesday, April 22 at 8:30 AM GMT+1. The morning start suggests a focused, energetic session that fits well before a full workday, especially for attendees who want a high-signal conversation without committing to a full-day program. Since the event is in person, plan to come ready to engage. Meetups like this tend to reward participation: bring your questions, examples, and opinions about what makes an evaluation useful, credible, or actionable. If you are currently building with AI tools, it may help to think ahead about the exact places where your team struggles to judge output quality. A few ways to prepare: Arrive with one real evaluation question you would like to discuss Think about your current workflow for testing prompts, outputs, or model changes Be ready to network with people across technical and non-technical roles Plan for conversation, not just observation If you want a sharper understanding of evals and a better network of people thinking seriously about them, this meetup is a strong place to start.

Who should attend

This is for people who want a more practical, grounded understanding of how to evaluate AI systems and improve them over time. - **You’re building with language models** and need better ways to judge whether outputs are actually improving, not just changing. - **You work in product, engineering, research, or applied AI** and want clearer thinking around quality, testing, and iteration. - **You’ve run into the limits of intuition** and want more structured ways to compare prompts, workflows, or model behavior. - **You care about reliability and decision-making** and want evaluation approaches that can support real product or workflow choices. - **You learn best through conversation** and want to meet other people wrestling with similar questions in a community setting. - **You’re AI-curious but serious** and want an accessible entry point into one of the most important topics in modern AI practice. If you want to leave with stronger questions, better mental models, and useful new connections, you’ll likely feel at home here.

Topics

Registration

Register / Get tickets