OSO Reading Group: OmniSQL: Synthesizing High-quality Text-to-SQL Data at Scale

Date
2025-12-02
Host
Public Goods
Register

About this event

Text-to-SQL systems live or die on the quality of their training data. This reading group takes a close look at OmniSQL, a paper focused on synthesizing high-quality text-to-SQL data at scale, and creates space to unpack what that actually means for people building, evaluating, or researching natural language interfaces to databases. If you care about LLM reliability, data generation, or the practical challenges behind getting models to produce useful SQL, this session is designed to be worth your morning. About the Event This is an in-person OSO Reading Group session centered on a single technical paper: OmniSQL: Synthesizing High-quality Text-to-SQL Data at Scale. The goal is not just to summarize the work, but to read it critically as a community and surface the ideas that matter most in practice. Reading groups work best when they combine focused discussion with a range of perspectives. Whether you come from machine learning, data engineering, product, research, or simply have a strong interest in language-model systems, this event gives you a structured way to understand the paper beyond the abstract. Expect a format that is conversational, technical, and grounded. Rather than treating the paper as settled truth, the group will look at its assumptions, methods, strengths, limitations, and implications for real-world text-to-SQL workflows. What to Expect The session will revolve around a shared discussion of the paper, with attention to both the core contribution and the broader context. That likely includes questions such as: what counts as “high-quality” synthetic text-to-SQL data, how scale changes the problem, what signals matter for evaluation, and how synthetic datasets affect downstream model behavior. You can expect a mix of close reading and practical interpretation. Discussion may touch on: The motivation behind generating text-to-SQL data synthetically How the paper approaches quality control at scale What the method suggests about dataset bottlenecks in text-to-SQL research How this work might translate to production settings or internal tooling Where the paper seems strong, and where it deserves skepticism Because this is a reading group rather than a lecture, participation matters. You do not need to arrive with a polished opinion, but it helps to come ready to ask questions, compare interpretations, and connect the paper to systems you’ve seen or built. There is also a community element built into the format. In-person discussion creates room for the side conversations that often make technical events useful: clarifying a concept with someone nearby, hearing how another team thinks about structured data interfaces, or meeting people who care about similar problems. Why Attend If you work with LLM applications, analytics tools, or data-access interfaces, text-to-SQL is one of the clearest tests of whether natural language systems can be trusted on structured tasks. This paper offers a concrete lens into a key issue in the field: not just model architecture, but the data pipeline that shapes model performance. Attending gives you more than a paper summary. You will leave with a sharper understanding of how synthetic data can be used to improve text-to-SQL systems, what tradeoffs come with scaling that process, and what kinds of evaluation questions should be asked before trusting the results. This session is also useful if you want to strengthen your research taste. Reading groups help you practice distinguishing between an impressive headline and a durable contribution. You get to pressure-test methods, inspect claims, and learn how others read technical work critically. For many attendees, the value will also be in the people in the room. The event brings together a community interested in technical depth, shared learning, and thoughtful discussion. If you want a space that feels more substantial than passive content consumption, this is exactly that. Practical Details This event is in person, which makes it a good fit for attendees who want direct conversation and a more engaged discussion environment. If you tend to get more out of technical sessions when you can ask follow-up questions in real time and meet people face to face, the format should work in your favor. The reading group takes place on Tuesday, December 2 at 8:30 AM PST. Since it starts in the morning, plan to arrive a few minutes early so you can settle in and be ready for the discussion from the beginning. A few ways to prepare: Read or skim the paper in advance if you can Bring questions about the methodology, evaluation, or applicability Come ready to discuss both technical details and broader implications Expect an interactive session rather than a passive talk If the title of the paper already sparked your curiosity, you are likely the right person for this room. This session is for people who want to understand not only what OmniSQL proposes, but why that proposal matters for the future of text-to-SQL systems and LLM-powered data interfaces.

Who should attend

If you want a more thoughtful, technical conversation about text-to-SQL than you usually get from a talk or social post, this reading group is likely for you. - You work with **LLMs, NLP, or ML systems** and want to better understand how synthetic training data affects downstream performance on structured tasks. - You are a **data engineer, analytics engineer, or backend builder** interested in how natural language interfaces connect to real databases and SQL workflows. - You read research papers but want a **community setting to unpack methods, assumptions, and tradeoffs** instead of doing it alone. - You are exploring **text-to-SQL products, internal tools, or agent workflows** and want a clearer view of what makes generated SQL data actually useful. - You enjoy asking questions like **“Would this hold up in production?”** or **“How should this be evaluated?”** and want to discuss those questions with others who care about technical rigor. - You want to meet people in the **tech and research community** who are interested in language models, data systems, and practical AI applications. You do not need to be a domain expert to participate well. Curiosity, a willingness to engage, and interest in the paper’s core problem are the main things to bring.

Topics