Beyond Data Ingestion: What Data Engineering Design Patterns for Open Lakehouse

Date
2026-02-18
Host
Open Lakehouse + AI
Register

About this event

Modern data platforms do not fail because teams cannot ingest data. They fail when ingestion is the only pattern anyone has, and every downstream problem gets pushed into brittle pipelines, duplicated logic, and expensive rework. This event is for people who want to think beyond moving data from point A to point B and design open lakehouse systems that are easier to scale, govern, and actually use. About the Event This in-person session explores data engineering design patterns for the open lakehouse, with a practical focus on how modern teams structure data platforms for analytics, AI, and operational use cases. Rather than treating ingestion as the end goal, the conversation moves up a level: how to design systems, workflows, and interfaces that make data reliable, accessible, and adaptable over time. You can expect a community-driven setting built for technical practitioners, builders, and decision-makers who care about open source and the future of data architecture. The topic sits at the intersection of open lakehouse, AI, and real-world data engineering, making it relevant whether you are modernizing an existing stack or evaluating where your platform should go next. The format is designed to be useful, not abstract. Expect discussion grounded in patterns, tradeoffs, and implementation thinking rather than broad trend talk. If you have ever looked at your data platform and thought, "we can ingest everything, but why is this still so hard to operate?" this event is speaking directly to that problem. What to Expect The core of the event will focus on design patterns that help teams build open lakehouse environments with more intention. That includes how data is organized, how pipelines and transformations are structured, and how engineering choices affect performance, governance, collaboration, and downstream AI readiness. Likely themes include: Moving from ingestion-centric thinking to platform-centric thinking Architectural patterns that support open lakehouse workflows Tradeoffs between flexibility, reliability, and operational simplicity How open-source approaches shape interoperability and long-term maintainability Considerations for supporting analytics and AI on the same data foundation Because this is an in-person gathering, there is also real value in the room itself. You will have the chance to compare approaches with other practitioners, pressure-test your assumptions, and hear how peers are navigating similar platform decisions. The networking component matters here because data architecture is rarely solved in isolation; many of the best ideas come from hearing how other teams handle the same constraints differently. Expect a session that is technical enough to be useful but broad enough to connect architecture, engineering process, and business reality. Whether the discussion gets into table formats, orchestration patterns, data product thinking, or governance boundaries, the throughline is clear: better system design leads to better outcomes than simply adding more pipelines. Why Attend If you work in data, you already know that the hard part is not getting data into storage. The hard part is designing a platform that can support changing requirements without becoming fragile, opaque, or costly to maintain. This event helps you step back from day-to-day pipeline work and think in terms of repeatable patterns that improve how your lakehouse actually functions. You should attend if you want sharper language for architectural decisions, a clearer mental model of open lakehouse design, and a more practical sense of what good patterns look like in modern data engineering. This is especially valuable if your team is balancing open-source tools, growing AI demands, and pressure to deliver trusted data faster. You may leave with: Better ways to evaluate lakehouse design choices beyond simple ingestion throughput Ideas for reducing complexity in data pipelines and downstream consumption A stronger understanding of how open architectures support evolution over time Useful peer context on what others in the community are building and prioritizing Questions worth taking back to your own team, roadmap, or platform review The value is not only in hearing about patterns. It is in seeing where those patterns fit, when they break down, and how to think more deliberately about the next stage of your data platform. For anyone building toward a more open, scalable, AI-ready foundation, that perspective is worth making time for. Practical Details This is an in-person event taking place on Wednesday, February 18 at 11:00 AM EST. If you prefer conversations with real practitioners over passive webinar attendance, the in-person format is a strong reason to be there. The event is well suited to attendees who want both technical substance and community connection. You should expect a setting where listening, discussion, and networking all play a role, especially for people working through active architecture or platform questions. A few practical notes: Format: in person Date: Wednesday, February 18 Time: 11:00 AM EST Focus: open lakehouse, open-source data engineering, AI-adjacent platform design, and community networking If this topic is close to your current work, come ready with your own examples, constraints, and opinions. The most useful conversations at events like this usually start with a specific problem, not a generic interest in the space.

Who should attend

This is for people who are actively shaping modern data platforms and want stronger patterns than "just ingest more data." - You are a **data engineer** or **analytics engineer** building pipelines, transformation layers, or platform workflows and want a better architectural framework for open lakehouse environments. - You are a **data architect** or **platform engineer** thinking about interoperability, governance, performance, and maintainability across an open-source-oriented stack. - You lead or influence a **data team** and need to make platform decisions that support both today’s analytics needs and tomorrow’s AI use cases. - You work close to the boundary between **data infrastructure and applied AI** and want to understand how design choices upstream affect model readiness, data quality, and reuse. - You are evaluating how to move from fragmented tools and one-off ingestion jobs toward a more coherent, scalable operating model. - You value **technical community conversations** and want to meet others working through similar lakehouse and open-source challenges in real environments. If you are looking for practical thinking, architecture-level perspective, and peer discussion grounded in current data engineering realities, you will likely feel at home here.

Speakers

Topics