Serving and Programming LLMs: Agents, Prefill, and Distributed Systems

Date
2026-05-22
Location
Seattle, WA, USA
Host
Seattle Systems
Register

About this event

Large language models are no longer just demos in notebooks. They are live systems with real users, real latency constraints, real infrastructure tradeoffs, and a fast-growing set of architectural patterns around serving, agents, prefill, and distributed execution. If you want to understand what it actually takes to build and run these systems well, this Seattle gathering is designed to get practical quickly. About the Event This is an in-person community event for people working on, experimenting with, or trying to better understand the engineering side of modern LLM systems. The focus is on how these systems behave in production and near-production settings: how requests are served, where bottlenecks show up, what agent-style workflows change about system design, and why prefill and distributed systems have become central topics for teams building serious LLM products. Expect a technically grounded conversation rather than a high-level trend talk. The theme connects several of the most active areas in applied AI infrastructure: programming models for LLM-powered applications, serving stacks and performance considerations, and the systems thinking required to make these tools reliable at scale. The format is built for people who like learning through discussion, examples, and exchange with peers. Whether you spend your time writing application code, tuning inference paths, designing infrastructure, or just trying to form a sharper mental model of the stack, the event is meant to bring those perspectives into the same room. What to Expect You can expect an evening centered on technical ideas, practical tradeoffs, and conversation with others who care about how LLM systems actually work under the hood. The subject matter spans multiple layers of the stack, so the event should feel relevant whether you think first about APIs, orchestration, model behavior, latency, or cluster-level concerns. Topics likely to shape the discussion include: Serving LLMs in practice: request handling, throughput, latency, concurrency, and where real-world complexity enters the picture Programming with LLMs: how application design changes when the model becomes part of your control flow, interface, and decision-making logic Agents and multi-step workflows: what becomes harder when you move from single prompts to tool use, orchestration, memory, and iterative execution Prefill and performance: why prefill matters, where it affects responsiveness and cost, and how teams think about optimization Distributed systems considerations: coordination, scaling, reliability, scheduling, and the broader infrastructure implications of LLM workloads Because this is a community-oriented event, expect a mix of structured content and room for interaction. You may spend part of the evening listening to technical framing and part of it comparing notes with others building adjacent systems. That combination is especially valuable in a field where patterns are still emerging and everyone is learning in public, even when they are deep in production work. Bring your current questions. The most useful events in this space are rarely passive; they are places where people pressure-test assumptions, swap implementation lessons, and leave with a better map of the problem space. Why Attend If you have been trying to connect the dots between LLM applications and the infrastructure beneath them, this event offers a focused way to do that. Many conversations about AI stay abstract. This one is aimed at the engineering realities: what changes when an LLM is part of a real product, where performance work matters, and how system design choices affect developer experience and end-user outcomes. You should leave with a clearer vocabulary for discussing the current LLM stack, especially around terms that often get mentioned without enough depth. Agents, prefill, and distributed execution are not just buzzwords; they point to concrete design decisions that shape reliability, speed, complexity, and cost. Seeing those topics discussed together can help you understand how they interact rather than treating them as separate specialties. There is also strong value in the room itself. Seattle has a deep technical community, and in-person conversations often surface the kind of implementation detail that never makes it into polished blog posts. If you are evaluating architecture choices, building internal tools, scaling an AI feature, or simply trying to understand what matters beyond the demo layer, talking with others facing similar constraints can save time and sharpen judgment. In short, this event is for people who want more than surface-level AI excitement. It is for builders, researchers, and engineers who want to think clearly about how LLM systems are programmed, served, and scaled. Practical Details The event takes place in person in Seattle, USA on Thursday, May 21 at 5:30 PM PDT. Being in the room matters for this kind of topic: nuanced technical discussion tends to be better live, and it is easier to ask follow-up questions, meet peers, and continue conversations after the formal program. This is an evening event, which makes it a good fit for people coming from work, side projects, or study. If you are local to Seattle or nearby, plan for an in-person meetup atmosphere with a technically engaged audience and a topic set that rewards curiosity and preparation. A few useful ways to get the most out of the event: Come with one or two concrete questions from your own work or learning Be ready to discuss tradeoffs, not just tools If you are building with LLMs today, think about where your biggest pain points are: latency, orchestration, evaluation, scaling, or reliability If you are newer to the systems side, bring the concepts you want to understand more clearly Whether you are already deep in LLM infrastructure or trying to move from application experimentation toward systems thinking, this event offers a strong reason to spend an evening with people working through the same shift.

Who should attend

This event is for people who want to understand the engineering reality behind modern LLM systems, not just the surface-level product story. - You are a **software engineer or platform engineer** building AI features and want a better handle on serving, latency, orchestration, and system design tradeoffs. - You are an **ML engineer or applied AI practitioner** working with LLMs and want to connect model behavior to infrastructure concerns like prefill, throughput, and distributed execution. - You are exploring **agent-style applications** and need a sharper mental model for what changes when workflows become multi-step, stateful, or tool-driven. - You care about **programming models for LLMs** and want to learn how others are structuring applications where the model is part of the core control flow. - You are a **technical founder, researcher, or advanced builder** who wants to compare notes with a serious local community on what is actually working in practice. - You are comfortable with technical discussion and want to spend time with people asking deeper questions about how LLM systems are built, served, and scaled.

Speakers

Topics