Evals Code Sprint: Add evaluation to your own app in one day (limited capacity)
- Date
- 2025-01-11
- Location
- San Francisco, CA, USA
- Host
- Weights & Biases
About this event
If you have an app and keep meaning to “add evals later,” this is the day to stop postponing it. Evals Code Sprint: Add evaluation to your own app in one day is a hands-on, limited-capacity working session designed to help you make real progress, not just talk about best practices. Bring your own app, your own use case, and the questions that are blocking you. By the end of the day, you should leave with a clearer evaluation setup, a better sense of what to measure, and momentum you can carry straight back into your product work. What Is This? This is an in-person code sprint in San Francisco for people who want to build or improve evaluation workflows for their own applications. Rather than a passive meetup or a broad conference-style event, the format is focused on practical implementation: show up ready to work, test ideas, and get feedback in the room. The core idea is simple: evaluation is one of the most important parts of building reliable AI-powered products, but it often gets pushed aside because it feels abstract, time-consuming, or hard to scope. This event is built to make that process more concrete by turning it into a structured day of coding, discussion, and peer learning. Because capacity is limited, the atmosphere should feel more focused and collaborative than a large public event. You can expect a room of people who are actively building, experimenting, and trying to solve similar problems from different angles. Whether you are starting from scratch or already have some evaluation approach in place, the goal is the same: leave with something more useful than notes. The emphasis is on getting evaluation into your app in a way that is grounded in your actual product and workflow. What to Expect The day starts on Saturday, January 11 at 10:00 AM PST with an in-person gathering in San Francisco, USA. From there, expect a working-session format built around making progress. This is less about polished presentations and more about time spent defining what success looks like for your app, translating that into evaluation logic, and actually building. You can expect a mix of: Focused build time where you work directly on your own app or project Structured discussion around evaluation design, tradeoffs, and implementation choices Peer conversation with other builders facing similar questions Informal troubleshooting and idea exchange in a collaborative setting A strong code sprint works because it combines individual concentration with timely input from others. You may spend part of the event narrowing the specific behaviors or outputs you want to evaluate, then shift into implementation, testing, iteration, and comparison. The point is to turn vague goals like “we need better quality control” into concrete systems, checks, or workflows. Since this is a community-oriented event with limited capacity, expect a room where people can actually talk to each other, share context, and get beyond surface-level networking. The social side matters, but it supports the work rather than distracting from it. Why Attend If you are building an app that depends on model behavior, output quality, or workflow reliability, evaluation is not a nice-to-have. It is how you move from intuition to evidence. This sprint gives you dedicated time and a relevant peer environment to work on that part of your stack with intent. A lot of people understand, in theory, that evals matter. Far fewer have set aside a full day to define what they want to measure, identify the right test cases, and put even a first version into practice. That is the gap this event is designed to close. You should come if you want to: Make real implementation progress instead of collecting more general advice Pressure-test your evaluation approach with other technical and product-minded attendees Clarify what quality means for your specific application Leave with a stronger foundation for iteration, debugging, and product decisions There is also real value in doing this work around other serious builders. Seeing how other teams or individuals think about evaluation can sharpen your own approach, reveal blind spots, and help you avoid overengineering too early. Even brief conversations can save hours of trial and error later. Practical Details This is an in-person event in San Francisco, USA, taking place on Saturday, January 11 at 10:00 AM PST. It is best suited for attendees who are ready to engage actively, ideally with a project, prototype, or app context they can bring into the room. The event is marked as limited capacity, which usually means the experience is intentionally kept smaller to support deeper conversation and hands-on work. If this is relevant to what you are building, it is worth treating it as a work session rather than a casual drop-in. To get the most out of the day, come prepared with: A clear sense of the app or workflow you want to evaluate Questions or pain points you have already run into Code, examples, or product context you can reference while working A willingness to share and learn in a collaborative environment This event sits at the intersection of meetup, community session, code sprint, and builder networking. If you want a Saturday that is practical, social, and genuinely useful to your product development, this is a strong fit.
Who should attend
This is for people who want to spend a day actually building evaluation into a real app, not just hearing why it matters. - You are **building an app, prototype, or AI-powered workflow** and want a more reliable way to measure output quality or system behavior. - You have been **meaning to set up evals but keep pushing it off** because the work feels hard to scope, too open-ended, or easy to deprioritize. - You learn best by **working in the room with other builders**, asking concrete questions, and making progress in real time. - You already have some evaluation ideas in place and want to **improve, refine, or pressure-test your current approach** against real product needs. - You value **small, focused, in-person events** where networking happens naturally through shared work instead of quick introductions. - You are ready to show up with a project context, get hands-on, and leave with **something more concrete than a list of notes**.