Dynamo & Dine: High-performance LLM Inference with Baseten and NVIDIA Dynamo

Date
2025-10-23
Location
San Francisco, CA, USA
Host
Baseten
Register

About this event

Running LLMs in production is no longer just about model quality, it is about throughput, latency, reliability, and cost under real demand. Dynamo & Dine brings together operators, infra-minded ML teams, and practitioners in San Francisco for a focused in-person conversation on high-performance LLM inference with Baseten and NVIDIA Dynamo. If you care about what happens after the demo works, this event is built for you. Expect practical discussion, technical context, and the kind of peer-to-peer exchange that helps you make better decisions about serving, scaling, and operating modern AI systems. About the Event This is an in-person community gathering centered on one of the most important questions in applied AI right now: how to run LLM inference well at production scale. The event focuses specifically on high-performance inference through the lens of Baseten and NVIDIA Dynamo, making it especially relevant for teams working on deployment architecture, serving performance, and operational efficiency. Rather than a broad AI meetup that tries to cover everything, this event is narrow in the right way. It is designed for people who want to understand the practical tradeoffs behind fast, reliable inference and how modern infrastructure choices shape user experience, engineering velocity, and unit economics. The format combines technical substance with the ease of an in-person midday gathering. You will have space to learn, ask questions, compare approaches with peers, and continue conversations over a meal in a setting that is meant to be useful, not performative. What to Expect You can expect a program that is grounded in real operational concerns, not abstract AI optimism. The core theme is high-performance LLM inference, with discussion likely centered on how teams think about serving architectures, hardware utilization, latency targets, scaling behavior, and production readiness. The event will also create room for direct interaction. In a room full of operators, builders, and infra-curious practitioners, some of the most valuable moments often come from hearing how other teams are approaching the same constraints from different angles. Likely areas of focus include: Inference performance and what actually moves the needle in production Operational tradeoffs across speed, reliability, and cost Infrastructure strategy for teams deploying and scaling LLM-powered products Tooling and platform considerations when moving from experimentation to dependable serving Peer discussion and networking with others working on similar technical problems Because this is an in-person event, expect a more candid and practical atmosphere than you would get from a webinar. You can ask sharper questions, pressure-test your assumptions, and leave with a clearer sense of what matters most when inference becomes a business-critical system rather than a prototype. Why Attend If you are responsible for making AI applications work under real-world conditions, this event offers a concentrated way to sharpen your thinking. High-performance inference sits at the intersection of model behavior, systems design, hardware constraints, and product expectations. Getting it right can dramatically improve responsiveness, stability, and cost efficiency. This is also a chance to learn in context. Instead of piecing together scattered opinions from blog posts and social feeds, you will be in a room with people actively thinking about the same production challenges. That makes it easier to compare architectures, hear where others have hit bottlenecks, and understand how different teams evaluate tradeoffs. You should attend if you want to: Better understand the infrastructure side of LLM product delivery Get more concrete about the challenges behind low-latency, high-throughput inference Meet practitioners who care about operating AI systems, not just building demos Gather ideas you can bring back to your own roadmap, stack, or deployment strategy For founders, platform engineers, ML engineers, and technical operators, the value is straightforward: you will leave with stronger mental models for inference performance and more grounded conversations about what production excellence actually requires. Practical Details When: Thursday, October 23 at 11:30 AM PDT Where: In person in San Francisco, USA This is a midday, in-person event, making it a good fit if you want a focused technical session without blocking an entire day. The format also naturally supports side conversations, quick introductions, and follow-up discussion over lunch. Because the gathering is in person, it is best suited for attendees who want direct access to the community around LLM infrastructure and operations. If your best learning happens through technical conversation, whiteboard-style thinking, and hearing how peers are solving production problems, you will get more from this format than from a remote session. Plan to come ready to talk specifics. The strongest experience will come from bringing your own questions about inference stacks, deployment bottlenecks, scaling concerns, or performance goals. This event is built for people who want substance, practical insight, and a room full of others working on the hard parts of AI systems.

Who should attend

This is for people who care about what it takes to serve LLMs reliably, efficiently, and at scale. - You run or influence **production AI infrastructure** and want clearer thinking on latency, throughput, and system performance. - You are an **ML engineer, platform engineer, or infra engineer** working on model serving, deployment, or the path from prototype to dependable production use. - You are an **operator or technical leader** evaluating how infrastructure choices affect user experience, reliability, and cost. - You are building **LLM-powered products** and need a better understanding of the tradeoffs behind high-performance inference. - You learn best by talking with peers and want to meet others in San Francisco who are solving real deployment and scaling problems, not just discussing trends. - You are curious about the role of **Baseten and NVIDIA Dynamo** in modern inference workflows and want practical context from a focused community event. If you are looking for broad beginner AI content, this may be too specific. If you want practical conversations about serving and operating LLMs well, you will likely feel right at home.

Topics