YC W25 RAG + Evals night
- Date
- 2025-02-21
- Location
- San Francisco, CA, USA
- Host
- Personal
About this event
RAG and evals are where a lot of AI products either become reliable or quietly fall apart. This night is for people who want to talk seriously about what it takes to make retrieval-augmented systems useful in the real world, and how thoughtful evaluation changes the way you build. If you care about moving beyond demos and into systems that actually hold up, this is the room you want to be in. In a city full of AI events, this one is tuned for a more specific conversation: builders, operators, and curious technical minds gathering in person to compare notes on retrieval, quality, testing, and what people are learning right now. Expect a practical, high-signal evening centered on real conversations with others who are working through the same questions. What Is This? YC W25 RAG + Evals night is an in-person meetup in San Francisco focused on two topics that sit at the center of modern AI product development: retrieval-augmented generation and evaluation. These are the areas that shape whether an AI system is trustworthy, debuggable, and genuinely helpful once it leaves the prototype stage. This event is built as a community gathering rather than a formal conference. The emphasis is on meeting people who are actively building, experimenting, or thinking deeply about these problems and creating space for sharper conversations than you usually get in a broad AI meetup. You should expect a room full of people comparing approaches, tradeoffs, and lessons learned. Some attendees may be early-stage founders, some may be engineers, some may be researchers, and some may simply be trying to get much better at understanding how RAG pipelines and eval frameworks work in practice. Because the focus is narrow, the conversations can be more concrete. Instead of talking about AI in general, you can spend your time discussing things like retrieval quality, failure modes, hallucination handling, benchmark design, human review loops, product reliability, and how teams decide what “good” actually means. What to Expect The evening starts at 6:00 PM PST on Thursday, February 20 and is designed as an in-person gathering in San Francisco. You can expect a social, networking-friendly format with plenty of room for direct conversation rather than a packed schedule that keeps everyone sitting silently in rows. While the exact flow may vary, the event is best understood as a focused meetup for people who want to exchange ideas, ask sharper questions, and meet others working on similar challenges. The value comes from the density of relevant people in the room and the specificity of the topic. Likely rhythms of the evening include: Arrivals and informal introductions so people can quickly find others working on adjacent problems Topic-driven conversations around RAG architectures, evaluation design, and operational lessons Peer networking with founders, engineers, researchers, and technically curious attendees Open discussion about what is and is not working in current AI workflows This is the kind of event where a short conversation can turn into a useful new idea, a better framework for evaluating your product, or a new relationship with someone who has already solved a problem you are currently stuck on. Come ready to talk through your thinking, compare implementation decisions, and ask questions that go beyond surface-level hype. Why Attend If you are building with LLMs, RAG and evals are no longer optional side topics. They are often the difference between a product that feels impressive for five minutes and one that users can trust repeatedly. Spending time with other people who care about this layer of the stack can save you weeks of isolated trial and error. This event gives you a chance to pressure-test your assumptions with people who understand the technical and product tradeoffs involved. You may leave with a clearer view of how others think about retrieval quality, how teams structure evaluation loops, or what practical standards people are using to judge performance. There is also a strong community value here. Focused meetups tend to produce better conversations because everyone arrives with a shared vocabulary and a similar level of urgency around the topic. That means less time explaining why this matters and more time discussing how to do it well. A few reasons this night may be worth your time: You want sharper conversations than a general AI mixer usually offers You are working on RAG systems and want to compare notes with peers You care about evals as a real discipline, not just a buzzword You want practical insight from people building and iterating right now You value in-person connection with a community that shares your technical interests Practical Details This is an in-person event in San Francisco, USA, so plan for an evening centered on face-to-face conversation and community. If you do your best thinking in live discussion, this format will suit you well. The event begins at 6:00 PM PST on Thursday, February 20. Since it is an evening meetup, it is a good fit for attendees coming after work or wrapping up the day before heading into focused conversations with other builders and operators. A few useful ways to prepare: Come with a point of view on RAG, evals, or both Be ready to introduce what you are working on in a few clear sentences Bring specific questions if you want useful feedback from the room Expect networking to be a core part of the experience If you have been looking for a San Francisco AI event that is narrower, more technical, and more grounded in real implementation questions, this is likely a strong fit. The best attendees for this night are people who want substance, not spectacle, and who are excited to meet others taking RAG and evals seriously.
Who should attend
This is for people who want smarter, more specific AI conversations and who care about how LLM systems actually perform once they are in use. - You’re **building with LLMs** and want to improve how your product retrieves information, responds, and handles edge cases in the real world. - You’re a **founder, engineer, researcher, or technical operator** who thinks about reliability, quality, and product behavior, not just model output in a demo. - You’re actively working on **RAG pipelines, knowledge systems, search layers, or evaluation workflows** and want to compare approaches with others doing similar work. - You care about **evals as a practical tool** for shipping better systems and want to learn how other people define, measure, and improve quality. - You want an **in-person San Francisco meetup** where the topic is focused enough to create useful conversations, strong connections, and less generic networking. - You may still be early in your learning, but you’re curious, engaged, and ready to ask thoughtful questions about retrieval, testing, failure modes, and what makes AI products dependable. If you want high-signal conversations with people who take RAG and evals seriously, you’ll likely feel at home here.