Disseminate: The Computer Science Research Podcast

Mateusz Gienieczko | AnyBlox: A Framework for Self-Decoding Datasets | #69

Tue, 17 Mar 2026 07:00:00 GMT

In this episode of Disseminate: The Computer Science Research Podcast, host Dr. Jack Waudby is joined by Mateusz Gienieczko, PhD researcher at TU Munich and co-author of the VLDB Best Paper Award winning paper AnyBlox.

They dive deep into a fundamental problem in modern data systems: why cutting-edge data encodings and file formats rarely make it from research into real-world systems — and how AnyBlox proposes a radical solution.

Mateusz explains the core idea of self-decoding data, where datasets ship with their own portable, sandboxed decoders, allowing any database system to read any encoding safely and efficiently. Built on WebAssembly, AnyBlox bridges the long-standing gap between database research and practice without sacrificing performance, portability, or security.

This episode is essential listening for database researchers, data engineers, system builders, and industry practitioners interested in the future of data formats, analytics performance, and making research matter in practice

Links:

Paper: https://www.vldb.org/pvldb/vol18/p4017-gienieczko.pdf
GitHub: https://github.com/AnyBlox
Mat's Homepage: https://v0ldek.com/

Hosted on Acast. See acast.com/privacy for more information.

Xiangyao Yu | Disaggregation: A New Architecture for Cloud Databases | #68

Thu, 27 Nov 2025 05:00:00 GMT

In this episode of Disseminate: The Computer Science Research Podcast, host Jack Waudby sits down with Xiangyao Yu (UW–Madison), one of the leading voices shaping the next generation of cloud-native databases.

We dive deep into disaggregation — the architectural shift transforming how modern data systems are built. Xiangyao breaks down:

Why traditional shared-nothing databases struggle in cloud environments
How separating compute and storage unlocks elasticity, scalability, and cost efficiency
The evolution of disaggregated systems, from Aurora and Snowflake through to advanced pushdown processing and new modular services
His team's research on reinventing core protocols like 2-phase commit for cloud-native environments
Real-time analytics, HTAP challenges, and the Hermes architecture
Where disaggregation goes next — indexing, query optimizers, materialized views, multi-cloud architectures, and more

Whether you're a database engineer, researcher, or a practitioner building scalable cloud systems, this episode gives a clear, accessible look into the architecture that’s rapidly becoming the default for modern data platforms.

Links:

Hosted on Acast. See acast.com/privacy for more information.

Navid Eslami | Diva: Dynamic Range Filter for Var-Length Keys and Queries | #67

Thu, 13 Nov 2025 05:00:00 GMT

In this episode of Disseminate: The Computer Science Research Podcast, Jack sits down with Navid Eslami, PhD researcher at the University of Toronto, to discuss his award-winning paper “DIVA: Dynamic Range Filter for Variable Length Keys and Queries”, which earned Best Research Paper at VLDB.

Navid breaks down how range filters extend the power of traditional filters for modern databases and storage systems, enabling faster queries, better scalability, and theoretical guarantees. We dive into:

How DIVA overcomes the limitations of existing range filters
What makes it the “holy grail” of filtering for dynamic data
Real-world integration in WiredTiger (the MongoDB storage engine)
Future challenges in data distribution smoothing and hybrid filtering

Whether you're a database engineer, systems researcher, or student exploring data structures, this episode reveals how cutting-edge research can transform how we query, filter, and scale modern data systems.

Links:

Hosted on Acast. See acast.com/privacy for more information.

Adaptive Factorization in DuckDB with Paul Groß

Thu, 06 Nov 2025 05:00:00 GMT

In this episode of the DuckDB in Research series, host Jack Waudby sits down with Paul Groß, PhD student at CWI Amsterdam, to explore his work on adaptive factorization and worst-case optimal joins - techniques that push the boundaries of analytical query performance.

Paul shares insights from his CIDR'25 paper “Adaptive Factorization Using Linear Chained Hash Tables”, revealing how decades of database theory meet modern, practical system design in DuckDB. From hash table internals to adaptive query planning, this episode uncovers how research innovations are becoming part of real-world systems.

Whether you’re a database researcher, engineer, or curious student, you’ll come away with a deeper understanding of query optimization and the realities of systems engineering.

Links:

Adaptive Factorization Using Linear-Chained Hash Tables

Hosted on Acast. See acast.com/privacy for more information.

Parachute: Rethinking Query Execution and Bidirectional Information Flow in DuckDB - with Mihail Stoian

Thu, 30 Oct 2025 05:00:00 GMT

In this episode of the DuckDB in Research series, host Jack Waudby sits down with Mihail Stoian, PhD student at the Data Systems Lab, University of Technology Nuremberg, to unpack the cutting-edge ideas behind Parachute, a new approach to robust query processing and bidirectional information passing in modern analytical databases.

We explore how Parachute bridges theory and practice, combining concepts from instance-optimal algorithms and semi-join filtering to boost performance in DuckDB, the in-process analytical SQL engine that’s reshaping how research meets real-world data systems.

Mihail discusses:

How Parachute extends semi-join filtering for two-way information flow
The challenges of implementing research ideas inside DuckDB
Practical performance gains on TPC-H and CEB workloads
The future of adaptive query processing and research-driven system design

Whether you're a database researcher, systems engineer, or curious practitioner, this deep-dive reveals how academic innovation continues to shape modern data infrastructure.

Links:

Hosted on Acast. See acast.com/privacy for more information.

Anarchy in the Database: Abigale Kim on DuckDB and DBMS Extensibility

Thu, 23 Oct 2025 05:00:00 GMT

In this episode of the DuckDB in Research series, host Jack Waudby talks with Abigale Kim, PhD student at the University of Wisconsin–Madison and author of VLDB 2025 paper: “Anarchy in the Database: A Survey and Evaluation of DBMS Extensibility”. They explore how database extensibility is reshaping modern data systems — and why DuckDB is emerging as the gold standard for safe, flexible, and high-performance extensions. Abigale shares the inside story of her research, the surprises uncovered when testing Postgres and DuckDB extensions, and what’s next for extensibility and composable database design.

This episode is perfect for researchers, practitioners, and students interested in databases, systems design, and the interplay between academia and industry innovation.

Highlights:

What “extensibility” really means in a DBMS
How DuckDB compares to Postgres, MySQL, and Redis
The rise of GPU-accelerated DuckDB extensions
Why bridging research and engineering matters for the future of databases

Links:

You can find Abigale at:

Hosted on Acast. See acast.com/privacy for more information.

Recursive CTEs, Trampolines, and Teaching Databases with DuckDB - with Prof. Torsten Grust

Thu, 16 Oct 2025 17:30:00 GMT

In this episode of the DuckDB in Research series, host Dr Jack Waudby talks with Professor Torsten Grust from the University of Tübingen. Torsten is one of the pioneers behind DuckDB’s implementation of recursive CTEs.

In the episode they unpack:

The power of recursive CTEs and how they turn SQL into a full-fledged programming language.
The story behind adding recursion to DuckDB, including the using key feature and the trampoline and TTL extensions emerging from Torsten’s lab.
How these ideas are transforming research, teaching, and even DuckDB’s internal architecture.
Why DuckDB makes databases exciting again — from classroom to cutting-edge systems research.

If you’re into data systems, query processing, or bridging research and practice, this episode is for you.

Links:

Hosted on Acast. See acast.com/privacy for more information.

DuckDB in Research S2 Coming Soon!

Thu, 16 Oct 2025 12:35:10 GMT

Hey folks! The DuckDB in Research series is back for S2!

In this season we chat with:

Torsten Grust: Recursive CTEs
Abigale Kim: Anarchy in the Database
Mihail Stoian: Parachute: Single-Pass Bi-Directional Information Passing
Paul Gross: Adaptive Factorization Using Linear-Chained Hash Tables

Whether you're a researcher, engineer, or just curious about the intersection of databases and innovation we are sure you will love this series.

Hosted on Acast. See acast.com/privacy for more information.

Rohan Padhye & Ao Li | Fray: An Efficient General-Purpose Concurrency JVM Testing Platform | #66

Mon, 06 Oct 2025 07:07:31 GMT

In this episode of Disseminate: The Computer Science Research Podcast, guest host Bogdan Stoica sits down with Ao Li and Rohan Padhye (Carnegie Mellon University) to discuss their OOPSLA 2025 paper: "Fray: An Efficient General-Purpose Concurrency Testing Platform for the JVM".

We dive into:

Why concurrency bugs remain so hard to catch -- even in "well-tested" Java projects.
The design of Fray, a new concurrency testing platform that outperforms prior tools like JPF and rr.
Real-world bugs discovered in Apache Kafka, Lucene, and Google Guava.
The gap between academic research and industrial practice, and how Fray bridges it.
What’s next for concurrency testing: debugging tools, distributed systems, and beyond.

If you’re a Java developer, systems researcher, or just curious about how to make software more reliable, this conversation is packed with insights on the future of software testing.

Links & Resources:

- The Fray paper (OOPSLA 2025):

- Fray on GitHub

- Ao Li’s research

- Rohan Padhye’s research

Don’t forget to like, subscribe, and hit the 🔔 to stay updated on the latest episodes about cutting-edge computer science research.

#Java #Concurrency #SoftwareTesting #Fray #OOPSLA2025 #Programming #Debugging #JVM #ComputerScience #ResearchPodcast

Hosted on Acast. See acast.com/privacy for more information.

Shrey Tiwari | It's About Time: A Study of Date and Time Bugs in Python Software | #65

Tue, 23 Sep 2025 07:55:51 GMT

In this episode, Bogdan Stoica, Postdoctoral Research Associate in the SysNet group at the University of Illinois Urbana-Champaign (UIUC) steps in to guest host. Bogdan sits down with Shrey Tiwari, a PhD student in the Software and Societal Systems Department at Carnegie Mellon University and member of the PASTA Lab, advised by Prof. Rohan Padhye. Together, they dive into Shrey’s award-winning research on date and time bugs in open-source Python software, exploring why these issues are so deceptively tricky and how they continue to affect systems we rely on every day.

The conversation traces Shrey’s journey from industry to research, including formative experiences at Citrix and Microsoft Research, and how those shaped his passion for software reliability. Shrey and Bogdan discuss the surprising complexity of date and time handling, the methodology behind Shrey’s empirical study, and the practical lessons developers can take away to build more robust systems. Along the way, they highlight broader questions about testing, bug detection, and the future role of AI in ensuring software correctness. This episode is a must-listen for anyone interested in debugging, reliability, and the hidden challenges that underpin modern software.

Links:

It’s About Time: An Empirical Study of Date and Time Bugs in Open-Source Python Software 🏆 ACM SIGSOFT Distinguished Paper Award
Shrey's homepage

Hosted on Acast. See acast.com/privacy for more information.

Lessons Learned from Five Years of Artifact Evaluations at EuroSys | #64

Wed, 30 Jul 2025 11:00:00 GMT

In this episode we are joined by Thaleia Doudali, Miguel Matos, and Anjo Vahldiek-Oberwagner to delve into five years of experience managing artifact evaluation at the EuroSys conference. They explain the goals and mechanics of artifact evaluation, a voluntary process that encourages reproducibility and reusability in computer systems research by assessing the supporting code, data, and documentation of accepted papers. The conversation outlines the three-tiered badge system, the multi-phase review process, and the importance of open-source practices. The guests present data showing increasing participation, sustained artifact availability, and varying levels of community engagement, underscoring the growing relevance of artifacts in validating and extending research.

The discussion also highlights recurring challenges such as tight timelines between paper acceptance and camera-ready deadlines, disparities in expectations between main program and artifact committees, difficulties with specialized hardware requirements, and lack of institutional continuity among evaluators. To address these, the guests propose early artifact preparation, stronger integration across committees, formalization of evaluation guidelines, and possibly making artifact submission mandatory. They advocate for broader standardization across CS subfields and suggest introducing a “Test of Time” award for artifacts. Looking to the future, they envision a more scalable, consistent, and impactful artifact evaluation process—but caution that continued growth in paper volume will demand innovation to maintain quality and reviewer sustainability.

Links:

Hosted on Acast. See acast.com/privacy for more information.

Dominik Winterer | Validating SMT Solvers for Correctness and Performance via Grammar-based Enumeration | #63

Fri, 25 Jul 2025 13:57:09 GMT

In this episode of the Disseminate podcast, Dominik Winterer discusses his research on SMT (Satisfiability Modulo Theories) solvers and his recent OOPSLA paper titled "Validating SMT Solvers for Correction and Performance via Grammar Based Enumeration". Dominik shares his academic journey from the University of Freiburg to ETH Zurich, and now to a lectureship at the University of Manchester. He introduces ET, a tool he developed for exhaustive grammar-based testing of SMT solvers. Unlike traditional fuzzers that use random input generation, ET systematically enumerates small, syntactically valid inputs using context-free grammars to expose bugs more effectively. This approach simplifies bug triage and has revealed over 100 bugs—many of them soundness and performance-related—with a striking number having already been fixed. Dominik emphasizes the tool’s surprising ability to identify deep bugs using minimal input and track solver evolution over time, highlighting ET's potential for continuous integration into CI pipelines.

The conversation then expands into broader reflections on formal methods and the future of software reliability. Dominik advocates for a new discipline—Formal Methods Engineering—to bridge the gap between software engineering and formal verification tools. He stresses the importance of building trustworthy verification tools since the reliability of software increasingly depends on them. Dominik also discusses adapting ET to other domains, such as JavaScript engines, and suggests that grammar-based enumeration can be applied widely to any system with a context-free grammar. Addressing the rise of AI, he envisions validation portfolios that integrate formal methods into LLM-based tooling, offering certified assessments of model outputs. He closes with a call for the community to embrace pragmatic, systematic, and scalable approaches to formal methods to ensure these tools can live up to their promises in real-world development settings.

Links:

Hosted on Acast. See acast.com/privacy for more information.

Haralampos Gavriilidis | Fast and Scalable Data Transfer across Data Systems | #62

Mon, 16 Jun 2025 07:00:00 GMT

In this episode of Disseminate, we welcome Harry Gavrilidis back to the podcast to explore his latest research on fast and scalable data transfer across systems, soon to be presented at SIGMOD 2025. Building on his work with XDB, Harry introduces XDBC, a novel data transfer framework designed to balance performance and generalizability. They dive into the challenges of moving data across heterogeneous environments—ranging from cloud systems to IoT devices—and critique the limitations of current generic methods like JDBC and specialized point-to-point connectors.

Harry walks us through the architecture of XDBC, which modularizes the data transfer pipeline into configurable stages like reading, serialization, compression, and networking. The episode highlights how this architecture adapts to varying performance constraints and introduces a cost-based optimizer to automate tuning for different environments. We also touch on future directions, including dynamic reconfiguration, fault tolerance, and learning-based optimizations. If you're interested in systems, performance engineering, or database interoperability, this episode is a must-listen.

Hosted on Acast. See acast.com/privacy for more information.

Haralampos Gavriilidis | SheetReader: Efficient spreadsheet parsing

Thu, 17 Apr 2025 08:00:00 GMT

In this episode of the DuckDB in Research series, Harry Gavriilidis (PhD student at TU Berlin) joins us to discuss Sheet Reader — a high-performance spreadsheet parser that dramatically outpaces traditional tools in both speed and memory efficiency. By taking advantage of the standardized structure of spreadsheet files and bypassing generic XML parsers, Sheet Reader delivers fast and lightweight parsing, even on large files. Now available as a DuckDB extension, it enables users to query spreadsheets directly with SQL and integrate them seamlessly into broader analytical workflows.

Harry shares insights into the development process, performance benchmarks, and the surprisingly complex world of spreadsheet parsing. He also discusses community feedback, feature requests (like detecting multiple tables or parsing colored rows), and future plans — including tighter integration with DuckDB and support for Arrow. The conversation wraps up with a look at Harry’s broader research on composable database systems and data interoperability, highlighting how tools like DuckDB are reshaping modern data analysis.

Hosted on Acast. See acast.com/privacy for more information.

Arjen P. de Vries | faiss: An extension for vector data & search

Thu, 10 Apr 2025 11:00:35 GMT

In this episode of the DuckDB in Research series, we’re joined by Arjen de Vries, Professor of Data Science at Radboud University. Arjen dives into his team’s development of a DuckDB extension for FAISS, a library originally developed at Facebook for efficient similarity search and vector operations.

We explore the growing importance of embeddings and dense retrieval in modern information retrieval systems, and how DuckDB’s zero-copy architecture and tight integration with the Python ecosystem make it a compelling choice for managing large-scale vector data. Arjen shares insights into the technical challenges and architectural decisions behind the extension, comparisons with DuckDB’s native VSS (vector search) solution, and the broader vision of integrating vector search more deeply into relational databases.

Along the way, we also touch on DuckDB's extension ecosystem, its potential for future research, and why tools like this are reshaping how we build and query modern AI-enabled systems.

Hosted on Acast. See acast.com/privacy for more information.

David Justen | POLAR: Adaptive and non-invasive join order selection via plans of least resistance

Thu, 03 Apr 2025 09:00:00 GMT

In this episode, we sit down with David Justen to discuss his work on POLAR: Adaptive and Non-invasive Join Order Selection via Plans of Least Resistance which was implemented in DuckDB. David shares his journey in the database space, insights into performance optimization, and the challenges of working with modern analytical workloads. We dive into the intricacies of query compilation, vectorized execution, and how DuckDB is shaping the future of in-memory databases. Tune in for a deep dive into database internals, industry trends, and what’s next for high-performance data processing!

Links:

Hosted on Acast. See acast.com/privacy for more information.

Daniël ten Wolde | DuckPGQ: A graph extension supporting SQL/PGQ

Thu, 20 Mar 2025 11:40:00 GMT

In this episode, we sit down with Daniël ten Wolde, a PhD researcher at CWI’s Database Architectures Group, to explore DuckPGQ—an extension to DuckDB that brings powerful graph querying capabilities to relational databases. Daniel shares his journey into database research, the motivations behind DuckPGQ, and how it simplifies working with graph data. We also dive into the technical challenges of implementing SQL Property Graph Queries (SQL PGQ) in DuckDB, discuss performance benchmarks, and explore the future of DuckPGQ in graph analytics and machine learning. Tune in to learn how this cutting-edge extension is bridging the gap between research and industry!

Links:

Hosted on Acast. See acast.com/privacy for more information.

Till Döhmen | DuckDQ: A Python library for data quality checks in ML pipelines

Thu, 13 Mar 2025 10:30:00 GMT

In this episode we kick off our DuckDB in Research series with Till Döhmen, a software engineer at MotherDuck, where he leads AI efforts. Till shares insights into DuckDQ, a Python library designed for efficient data quality validation in machine learning pipelines, leveraging DuckDB’s high-performance querying capabilities.

We discuss the challenges of ensuring data integrity in ML workflows, the inefficiencies of existing solutions, and how DuckDQ provides a lightweight, drop-in replacement that seamlessly integrates with scikit-learn. Till also reflects on his research journey, the impact of DuckDB’s optimizations, and the future potential of data quality tooling. Plus, we explore how AI tools like ChatGPT are reshaping research and productivity. Tune in for a deep dive into the intersection of databases, machine learning, and data validation!

Resources:

GitHub
Paper
Slides
Till's Homepage
datasketches extension (released by a DuckDB community member 2 weeks after we recorded!)

Hosted on Acast. See acast.com/privacy for more information.

Disseminate x DuckDB Coming Soon...

Thu, 06 Mar 2025 11:00:00 GMT

Hey folks!

We have been collaborating with everyone's favourite in-process SQL OLAP database management system DuckDB to bring you a new podcast series - the DuckDB in Research series!

At Disseminate our mission is to bridge the gap between research and industry by exploring research that has a real-world impact. DuckDB embodies this synergy—decades of research underpin its design, and now it’s making waves in the research community as a platform for others to build on and this is what the series will focus on!

Join us as we kick off the series with:

📌 Daniel ten Wolde – DuckPGQ, a graph workload extension for DuckDB supporting SQL/PGQ

📌 David Justen – POLAR: Adaptive, non-invasive join order selection

📌 Till Döhmen – DuckDQ: A Python library for data quality checks in ML pipelines

📌 Arjen de Vries – FAISS extension for vector similarity search in DuckDB

📌 Harry Gavriilidis – SheetReader: Efficient spreadsheet parsing

Whether you're a researcher, engineer, or just curious about the intersection of databases and innovation we are sure you will love this series.

Subscribe now and stay tuned for our first episode! 🚀

Hosted on Acast. See acast.com/privacy for more information.

High Impact in Databases with... Anastasia Ailamaki

Mon, 03 Mar 2025 08:53:02 GMT

In this High Impact in Databases episode we talk to Anastasia Ailamaki.

Anastasia is a Professor of Computer and Communication Sciences at the École Polytechnique Fédérale de Lausanne (EPFL). Tune in to hear Anastasia's story!

The podcast is proudly sponsored by Pometry the developers behind Raphtory, the open source temporal graph analytics engine for Python and Rust.

You can find Anastasia on:

Hosted on Acast. See acast.com/privacy for more information.

Anastasiia Kozar | Fault Tolerance Placement in the Internet of Things | #61

Mon, 16 Dec 2024 08:11:58 GMT

In this episode, we chat with Anastasiia Kozar about her research on fault tolerance in resource-constrained environments. As IoT applications leverage sensors, edge devices, and cloud infrastructure, ensuring system reliability at the edge poses unique challenges. Unlike the cloud, edge devices operate without persistent backups or high availability standards, leading to increased vulnerability to failures. Anastasiia explains how traditional methods fall short, as they fail to align resource allocation with fault tolerance needs, often resulting in system underperformance.

To address this, Anastasiia introduces a novel resource-aware approach that combines operator placement and fault tolerance into a unified process. By optimizing where and how data is backed up, her solution significantly improves system reliability, especially for low-end edge devices with limited resources. The result? Up to a tenfold increase in throughput compared to existing methods. Tune to learn more!

Links:

Hosted on Acast. See acast.com/privacy for more information.

Liana Patel | ACORN: Performant and Predicate-Agnostic Hybrid Search | #60

Mon, 11 Nov 2024 08:27:05 GMT

In this episode, we chat with with Liana Patel to discuss ACORN, a groundbreaking method for hybrid search in applications using mixed-modality data. As more systems require simultaneous access to embedded images, text, video, and structured data, traditional search methods struggle to maintain efficiency and flexibility. Liana explains how ACORN, leveraging Hierarchical Navigable Small Worlds (HNSW), enables efficient, predicate-agnostic searches by introducing innovative predicate subgraph traversal. This allows ACORN to outperform existing methods significantly, supporting complex query semantics and achieving 2–1,000 times higher throughput on diverse datasets. Tune in to learn more!

Links:

Hosted on Acast. See acast.com/privacy for more information.

High Impact in Databases with... David Maier

Mon, 04 Nov 2024 08:02:20 GMT

In this High Impact episode we talk to David Maier.

David is the Maseeh Professor Emeritus of Emerging Technologies at Portland State University. Tune in to hear David's story and learn about some of his most impactful work.

The podcast is proudly sponsored by Pometry the developers behind Raphtory, the open source temporal graph analytics engine for Python and Rust.

You can find David on:

Hosted on Acast. See acast.com/privacy for more information.

Raunak Shah | R2D2: Reducing Redundancy and Duplication in Data Lakes | #59

Mon, 28 Oct 2024 08:20:11 GMT

In this episode, Raunak Shah joins us to discuss the critical issue of data redundancy in enterprise data lakes, which can lead to soaring storage and maintenance costs. Raunak highlights how large-scale data environments, ranging from terabytes to petabytes, often contain duplicate and redundant datasets that are difficult to manage. He introduces the concept of "dataset containment" and explains its significance in identifying and reducing redundancy at the table level in these massive data lakes—an area where there has been little prior work.

Raunak then dives into the details of R2D2, a novel three-step hierarchical pipeline designed to efficiently tackle dataset containment. By utilizing schema containment graphs, statistical min-max pruning, and content-level pruning, R2D2 progressively reduces the search space to pinpoint redundant data. Raunak also discusses how the system, implemented on platforms like Azure Databricks and AWS, offers significant improvements over existing methods, processing TB-scale data lakes in just a few hours with high accuracy. He concludes with a discussion on how R2D2 optimally balances storage savings and performance by identifying datasets that can be deleted and reconstructed on demand, providing valuable insights for enterprises aiming to streamline their data management strategies.

Materials:

Hosted on Acast. See acast.com/privacy for more information.

High Impact in Databases with... Aditya Parameswaran

Mon, 21 Oct 2024 07:02:49 GMT

In this High Impact episode we talk to Aditya Parameswaran about his some of his most impactful work.

Aditya is an Associate Professor at the University of California, Berkeley. Tune in to hear Aditya's story!

The podcast is proudly sponsored by Pometry the developers behind Raphtory, the open source temporal graph analytics engine for Python and Rust.

Links:

EPIC Data Lab
Answering Queries using Humans, Algorithms and Databases (CIDR'11)
Potter’s Wheel: An Interactive Data Cleaning System (VLDB'01)
Online Aggregation (SIGMOD'97)
Polaris: A System for Query, Analysis and Visualization of Multi-dimensional Relational Databases (INFOVIS'00)
Coping with Rejection
Ponder

You can find Aditya on:

Hosted on Acast. See acast.com/privacy for more information.

Marco Costa | Taming Adversarial Queries with Optimal Range Filters | #58

Mon, 14 Oct 2024 07:16:32 GMT

In this episode, we sit down with Marco Costa to discuss the fascinating world of range filters, focusing on how they help optimize queries in databases by determining whether a range intersects with a given set of keys. Marco explains how traditional range filters, like Bloom filters, often result in high false positives and slow query times, especially when dealing with adversarial inputs where queries are correlated with the keys. He walks us through the limitations of existing heuristic-based solutions and the common challenges they face in maintaining accuracy and speed under such conditions.

The highlight of our conversation is Grafite, a novel range filter introduced by Marco and his team. Unlike previous approaches, Grafite comes with clear theoretical guarantees and offers robust performance across various datasets, query sizes, and workloads. Marco dives into the technicalities, explaining how Grafite delivers faster query times and maintains predictable false positive rates, making it the most reliable range filter in scenarios where queries are correlated with keys. Additionally, he introduces a simple heuristic filter that excels in uncorrelated queries, pushing the boundaries of current solutions in the field.

SIGMOD' 24 Paper - Grafite: Taming Adversarial Queries with Optimal Range Filters

Hosted on Acast. See acast.com/privacy for more information.

High Impact in Databases with... Ali Dasdan

Tue, 08 Oct 2024 07:05:45 GMT

In this High Impact episode we talk to Ali Dasdan, CTO at Zoominfo. Tune in to hear Ali's story and learn about some of his most impactful work such as his work on "Map-Reduce-Merge".

The podcast is proudly sponsored by Pometry the developers behind Raphtory, the open source temporal graph analytics engine for Python and Rust.

Materials mentioned on this episode:

Map-Reduce-Merge: Simplified Relational Data Processing on Large Clusters (SIGMOD'07)
The Art of Doing Science and Engineering: Learning to Learn, Richard Hamming
How to Solve It, George Polya
Systems Architecting: Creating & Building Complex Systems, Eberhardt Rechtin

You can find Ali on:

Hosted on Acast. See acast.com/privacy for more information.

Matt Perron | Analytical Workload Cost and Performance Stability With Elastic Pools | #57

Mon, 22 Jul 2024 06:24:20 GMT

In this episode, we dive deep into the complexities of managing analytical query workloads with our guest, Matt Perron. Matt explains how the rapid and unpredictable fluctuations in resource demands present a significant challenge for provisioning. Traditional methods often lead to either over-provisioning, resulting in excessive costs, or under-provisioning, which causes poor query latency during demand spikes. However, there's a promising solution on the horizon. Matt shares insights from recent research that showcases the viability of using cloud functions to dynamically match compute supply with workload demand without the need for prior resource provisioning. While effective for low query volumes, this approach becomes cost-prohibitive as query volumes increase, highlighting the need for a more balanced strategy.

Matt introduces us to a novel strategy that combines the best of both worlds: the rapid scalability of cloud functions and the cost-effectiveness of virtual machines. This innovative approach leverages the fast but expensive cloud functions alongside slow-starting yet inexpensive virtual machines to provide elasticity without sacrificing cost efficiency. He elaborates on how their implementation, called Cackle, achieves consistent performance and cost savings across a wide range of workloads and conditions. Tune in to learn how Cackle avoids the pitfalls of traditional approaches, delivering stable query performance and minimizing costs even as demand fluctuates wildly.

Links:

Hosted on Acast. See acast.com/privacy for more information.

High Impact in Databases with... Andreas Kipf

Mon, 15 Jul 2024 07:30:07 GMT

In this High Impact episode we talk to Andreas Kipf about his work on "Learned Cardinalities".

Andreas is the Professor of Data Systems at Technische Universität Nürnberg (UTN). Tune in to hear Andreas's story and learn about some of his most impactful work.

The podcast is proudly sponsored by Pometry the developers behind Raphtory, the open source temporal graph analytics engine for Python and Rust.

Papers mentioned on this episode:

You can find Andreas on:

Hosted on Acast. See acast.com/privacy for more information.

Marvin Wyrich & Justus Bogner | How Software Engineering Research Is Discussed on LinkedIn | #56

Mon, 08 Jul 2024 07:13:17 GMT

In this episode, we delve into the intersection of software engineering (SE) research and professional practice with experts Marvin Wyrich and Justus Bogner. As LinkedIn stands as the largest professional network globally, it serves as a critical platform for bridging the gap between SE researchers and practitioners. Marvin and Justus explore the dynamics of how research findings are shared and discussed on LinkedIn, providing both quantitative and qualitative insights into the effectiveness of these interactions. They reveal that a significant portion of SE research posts on LinkedIn are authored by individuals outside the original research team and that a majority of comments on these posts come from industry professionals, highlighting a vibrant but underutilized avenue for science communication.

Our guests shed light on the current state of this metaphorical bridge, emphasizing the potential for LinkedIn to enhance collaboration and knowledge exchange between academia and industry. Despite the promising engagement from practitioners, the discussion reveals that only half of the SE research posts receive any comments, indicating room for improvement in fostering more interactive dialogues. Marvin and Justus offer practical advice for researchers to better engage with practitioners on LinkedIn and suggest strategies for making research dissemination more impactful. This episode provides valuable insights for anyone interested in leveraging social media for advancing software engineering knowledge and practice.

Links:

Hosted on Acast. See acast.com/privacy for more information.

High Impact in Databases with... Joe Hellerstein

Mon, 01 Jul 2024 06:25:22 GMT

In this High Impact episode we talk to Joe Hellerstein.

Joe is the Jim Gray Professor of Computer Science at UC Berkeley. Tune in to hear Joe's story and learn about some of his most impactful work.

The podcast is proudly sponsored by Pometry the developers behind Raphtory, the open source temporal graph analytics engine for Python and Rust.

Hosted on Acast. See acast.com/privacy for more information.

Harry Goldstein | Property-Based Testing | #55

Tue, 25 Jun 2024 16:37:35 GMT

In this episode, we chat with Harry Goldstein about Property-Based Testing (PBT). Harry shares insights from interviews with PBT users at Jane Street, highlighting PBT's strengths in testing complex code and boosting developer confidence. Harry also discusses the challenges of writing properties and generating random data, and the difficulties in assessing test effectiveness. He identifies key areas for future improvement, such as performance enhancements and better random input generation. This episode is essential for those interested in the latest developments in software testing and PBT's future.

Links:

Hosted on Acast. See acast.com/privacy for more information.

High Impact in Databases with... Raghu Ramakrishnan

Mon, 17 Jun 2024 07:15:18 GMT

In this High Impact episode we talk to Raghu Ramakrishnan.

Raghu is CTO for Data and a Technical Fellow at Microsoft. Tune in to hear Raghu's story and learn about some of his most impactful work.

The podcast is proudly sponsored by Pometry the developers behind Raphtory, the open source temporal graph analytics engine for Python and Rust.

Hosted on Acast. See acast.com/privacy for more information.

Gina Yuan | In-Network Assistance With Sidekick Protocols | #54

Mon, 10 Jun 2024 06:30:00 GMT

Join us as we chat with Gina Yuan about her pioneering work on sidekick protocols, designed to enhance the performance of encrypted transport protocols like QUIC and WebRTC. These protocols ensure privacy but limit in-network innovations. Gina explains how sidekick protocols allow intermediaries to assist endpoints without compromising encryption.

Discover how Gina tackles the challenge of referencing opaque packets with her innovative quACK tool and learn about the real-world benefits, including improved Wi-Fi retransmissions, energy-saving proxy acknowledgments, and the PACUBIC congestion-control mechanism. This episode offers a glimpse into the future of network performance and security.

Links:

Hosted on Acast. See acast.com/privacy for more information.

High Impact in Databases with... Moshe Vardi

Mon, 03 Jun 2024 06:30:03 GMT

Welcome to another episode of the High Impact series - today we talk with Moshe Vardi!

Moshe is the Karen George Distinguished Service Professor in Computational Engineering at Rice University where his research focuses on automated reasoning. Tune in to hear Moshe's story and learn about some of his most impactful work.

The podcast is proudly sponsored by Pometry the developers behind Raphtory, the open source temporal graph analytics engine for Python and Rust.

You can find Moshe on X, LinkedIn, and Mastadon @vardi. Links to all his work can be found on his website here.

Hosted on Acast. See acast.com/privacy for more information.

Tammy Sukprasert | Move Your Workloads To Sweden! | #53

Mon, 27 May 2024 06:20:14 GMT

In this episode, we dip our toes into the world of sustainable computing and interview Tammy Sukprasert about her research on reducing carbon emissions in cloud computing through workload scheduling. Tammy explores the concept of shifting cloud workloads across different times and locations to coincide with low-carbon energy availability. Unlike previous studies that focused on specific regions or workloads, her comprehensive analysis uses carbon intensity data from 123 regions to assess both batch and interactive workloads. She considers various factors such as job duration, deadlines, and service level objectives (SLOs). Tammy's findings reveal that while spatiotemporal workload shifting can reduce carbon emissions, the practical upper bounds of these reductions are limited and far from ideal. Simple scheduling policies often achieve most of the potential reductions, with more complex techniques offering minimal additional benefits.

Additionally, Tammy's research highlights that as the energy grid becomes greener, the benefits of carbon-aware scheduling over carbon-agnostic approaches decrease. This discussion offers crucial insights for the future of cloud computing and sustainable technology. Whether you're a tech enthusiast, environmental advocate, or cloud industry professional, Tammy's work provides valuable perspectives on the intersection of technology and sustainability. Join us to learn more about how innovative scheduling strategies can contribute to a greener cloud computing landscape.

Links:

Hosted on Acast. See acast.com/privacy for more information.

High Impact in Databases with... Ryan Marcus

Mon, 20 May 2024 06:30:08 GMT

Welcome to the first episode of the High Impact series!

The High Impact series is inspired by a blog post “Most Influential Database Papers" by Ryan Marcus and today we talk to Ryan! Tune in to hear about Ryan's story so far. We chat about his current work before moving on to discuss his most impactful work. We also dig into what motivates him and how he handles setbacks, as well as getting his take on the current trends.

The podcast is proudly sponsored by Pometry the developers behind Raphtory, the open source temporal graph analytics engine for Python and Rust.

Links:

Hosted on Acast. See acast.com/privacy for more information.

Yazhuo Zhang | SIEVE is Simpler than LRU | #52

Mon, 13 May 2024 06:25:04 GMT

In this episode, we explore the world of caching with Yazhuo Zhang, who introduces the game-changing SIEVE algorithm. Traditional eviction algorithms have long struggled with a trade-off between efficiency, throughput, and simplicity. However, SIEVE disrupts this balance by offering a simpler alternative to LRU while outperforming state-of-the-art algorithms in both efficiency and scalability for web cache workloads. Implemented in five production cache libraries with minimal code changes, SIEVE's superiority shines through in a comprehensive evaluation across 1559 cache traces. With up to a remarkable 63.2% lower miss ratio than ARC and surpassing nine other algorithms in over 45% of cases, SIEVE's simplicity doesn't compromise on scalability, doubling throughput compared to optimized LRU implementations. Join us as Yazhuo reveals how SIEVE is set to redefine caching efficiency, promising faster and more streamlined data serving in production systems.

Links:

Hosted on Acast. See acast.com/privacy for more information.

Introducing the High Impact Series...

Mon, 06 May 2024 07:06:26 GMT

Introducing the High Impact Series!

Hey folks, we have a new series coming soon inspired by a blog post “Most Influential Database Papers" by Ryan Marcus. The series will feature interviews with the authors of some of the most impactful work in the field of databases. We will talk about the story behind some of their most impactful work, getting them to reflect on the impact it has had over years, as well as getting their take on the current trends in the field.

Proudly sponsored by Pometry

Hosted on Acast. See acast.com/privacy for more information.

Eleni Zapridou | Oligolithic Cross-task Optimizations across Isolated Workloads | #51

Mon, 29 Apr 2024 07:06:50 GMT

In this episode, we talk to Eleni Zapridou and delve into the challenges of data processing within enterprises, where multiple applications operate concurrently on shared resources. Traditional resource boundaries between applications often lead to increased costs and resource consumption. However, as Eleni explains the principle of functional isolation offers a solution by combining cross-task optimizations with performance isolation. We explore GroupShare, an innovative strategy that reduces CPU consumption and query latency, transforming data processing efficiency. Join us as we discuss the implications of functional isolation with Eleni and its potential to revolutionize enterprise data processing.

Links:

Hosted on Acast. See acast.com/privacy for more information.

Pat Helland | Scalable OLTP in the Cloud: What’s the BIG DEAL? | #50

Mon, 15 Apr 2024 07:55:44 GMT

In this thought-provoking podcast episode, we dive into the world of scalable OLTP (OnLine Transaction Processing) systems with the insightful Pat Helland. As a seasoned expert in the field, Pat shares his insights on the critical role of isolation semantics in the scalability of OLTP systems, emphasizing its significance as the "BIG DEAL." By examining the interface between OLTP databases and applications, particularly through the lens of RCSI (READ COMMITTED SNAPSHOT ISOLATION) SQL databases, Pat talks about the limitations imposed by current database architectures and application patterns on scalability.

Through a compelling thought experiment, Pat explores the asymptotic limits to scale for OLTP systems, challenging the status quo and envisioning a reimagined approach to building both databases and applications that empowers scalability while adhering to established to RCSI. By shedding light on how today's popular databases and common app patterns may unnecessarily hinder scalability, Pat sparks discussions within the database community, paving the way for new opportunities and advancements in OLTP systems. Join us as we delve into this conversation with Pat Helland, where every insight shared could potentially catalyze significant transformations in the realm of OLTP scalability.

Papers mentioned during the episode:

You can find Pat on:

Hosted on Acast. See acast.com/privacy for more information.

Rui Liu | Towards Resource-adaptive Query Execution in Cloud Native Databases | #49

Mon, 01 Apr 2024 07:30:18 GMT

In this episode, we talk to Rui Liu and explore the transformative potential of Ratchet, a groundbreaking resource-adaptive query execution framework. We delve into the challenges posed by ephemeral resources in modern cloud environments and the innovative solutions offered by Ratchet. Rui guides us through the intricacies of Ratchet's design, highlighting its ability to enable adaptive query suspension and resumption, sophisticated resource arbitration for diverse workloads, and a fine-grained pricing model to navigate fluctuating resource availability. Join us as we uncover the future of cloud-native databases and workloads, and discover how Ratchet is poised to revolutionize the way we harness the power of dynamic cloud resources.

Links:

You can find links to all Rui's work from his Google Scholar profile.

Hosted on Acast. See acast.com/privacy for more information.

Yifei Yang | Predicate Transfer: Efficient Pre-Filtering on Multi-Join Queries | #48

Mon, 18 Mar 2024 06:31:13 GMT

In this episode, Yifei Yang introduces predicate transfer, a revolutionary method for optimizing join performance in databases. Predicate transfer builds on Bloom joins, extending its benefits to multi-table joins. Inspired by Yannakakis's theoretical insights, predicate transfer leverages Bloom filters to achieve significant speed improvements. Yang's evaluation shows an average 3.3× performance boost over Bloom join on the TPC-H benchmark, highlighting the potential of predicate transfer to revolutionize database query optimization. Join us as we explore the transformative impact of predicate transfer on database operations.

Links:

Hosted on Acast. See acast.com/privacy for more information.

Vikramank Singh | Panda: Performance Debugging for Databases using LLM Agents | #47

Mon, 04 Mar 2024 08:08:54 GMT

In this episode, Vikramank Singh introduces the Panda framework, aimed at refining Large Language Models' (LLMs) capability to address database performance issues. Vikramank elaborates on Panda's four components—Grounding, Verification, Affordance, and Feedback—illustrating how they collaborate to contextualize LLM responses and deliver actionable recommendations. By bridging the divide between technical knowledge and practical troubleshooting needs, Panda has the potential to revolutionize database debugging practices, offering a promising avenue for more effective and efficient resolution of performance challenges in database systems. Tune in to learn more!

Links:

Hosted on Acast. See acast.com/privacy for more information.

Tamer Eldeeb | Chablis: Fast and General Transactions in Geo-Distributed Systems | #46

Mon, 12 Feb 2024 21:09:12 GMT

In this episode Jinkun Geng talks to us about Nezha, a high-performance consensus protocol. Nezha can be deployed by cloud tenants without support from cloud providers. Nezha bridges the gap between protocols such as MultiPaxos and Raft, which can be readily deployed, and protocols such as NOPaxos and Speculative Paxos, that provide better performance, but require access to technologies such as programmable switches and in-network prioritization, which cloud tenants do not have. Tune in to learn more!

Links:

Hosted on Acast. See acast.com/privacy for more information.

Dimitris Koutsoukos | NVM: Is it Not Very Meaningful for Databases? | #41

Mon, 09 Oct 2023 06:04:25 GMT

Summary:

In this episode, Dimitris Koutsoukos talks to us about Persistent or Non Volatile Memory (PMEM) and we answer the question: Is it Not Very Meaningful for Databases?

PMEM offers expanded memory capacity and faster access to persistent storage. However, (before Dimitris's work) there was no comprehensive empirical analysis of existing database engines under diferent PMEM modes, to understand how databases can benefit from the various hardware configurations. Dimitris and his colleagues have then analyzes multiple diferent engines under common benchmarks with PMEM in AppDirect mode and Memory mode - tune in to hear the findings!

Links:

Hosted on Acast. See acast.com/privacy for more information.

Mohamed Alzayat | Groundhog: Efficient Request Isolation in FaaS | #40

Mon, 11 Sep 2023 07:04:39 GMT

Summary:

Listen to hear the vision for Kùzu and to learn more about Kùzu's factorized query processor!

Lukas Vogel | Data Pipes: Declarative Control over Data Movement | #28

Tue, 28 Mar 2023 07:02:37 GMT

Summary:

Today’s storage landscape offers a deep and heterogeneous stack of technologies that promises to meet even the most demanding data intensive workload needs. The diversity of technologies, however, presents a challenge. Parts of it are not controlled directly by the application, e.g., the cache layers, and the parts that are controlled, often require the programmer to deal with very different transfer mechanisms, such as disk and network APIs. Combining these different abstractions properly requires great skill, and even so, expert-written programs can lead to sub-optimal utilization of the storage stack and present performance unpredictability. In this episode, Lukas Vogel tells us how we can combat these issues with a new programming abstraction called Data Pipes. Tune in to learn more!

Links:

Hosted on Acast. See acast.com/privacy for more information.

Haralampos Gavriilidis | In-Situ Cross-Database Query Processing | #27

Mon, 20 Mar 2023 21:33:11 GMT

Summary:

Today’s organizations utilize a plethora of heterogeneous and autonomous DBMSes, many of those being spread across different geo-locations. It is therefore crucial to have effective and efficient cross-database query processing capabilities. In this episode, Haralampos Gavriilidis tell us about XDB, an efficient middleware system that runs cross database analytics over existing DBMSes. Tune in to learn more!

Links:

Support the podcast here!

Hosted on Acast. See acast.com/privacy for more information.

Paras Jain & Sarah Wooders | Skyplane: Fast Data Transfers Between Any Cloud | #26

Mon, 13 Mar 2023 08:04:40 GMT

Summary:

This week Paras Jain and Sarah Wooders tell us about how you can quickly data transfers between any cloud with Skyplane. Tune in to learn more!

Links:

Audrey Cheng | TAOBench: An End-to-End Benchmark for Social Network Workloads | #15

Mon, 12 Dec 2022 08:00:04 GMT

Summary: This episode features Audrey Cheng talking about TAOBench, a new benchmark that captures the social graph workload at Meta. Audrey tells us about the features of workload, how it compares with other benchmarks, and how it fills a gap in the existing space of benchmark. Also, we hear all about the fantastic real-world impact the benchmark has already had across a range of companies.

Links:

In this episode Matthias Jasny from TU Darmstadt talks about P4DB, a database that uses a programmable switch to accelerate OLTP workloads. The main idea of P4DB is that it implements a transaction processing engine on top of a P4-programmable switch. The switch can thus act as an accelerator in the network, especially when it is used to store and process hot (contended) tuples on the switch. P4DB provides significant benefits compared to traditional DBMS architectures and can achieve a speedup of up to 8x.

Questions:

0:55: Can you set the scene for your research and describe the motivation behind P4DB?

1:42: Can you describe to listeners who may not be familiar with them, what exactly is a programmable switch?

3:55: What are the characteristics of OLTP workloads that make them a good fit for programmable switches?

5:33: Can you elaborate on the key idea of P4DB?

6:46: How do you go about mapping the execution of transactions to the architecture of a programmable switch?

10:13: Can you walk us through the lifecycle of a switch transaction?

11:04: How does P4DB determine the optimal tuple placement on the switch?

12:16: Is this allocation static or is it dynamic, can the tuple order be changed at runtime?

12:55: What happens if a transaction needs to access tuples in a different order then that laid out on the switch?

14:11: Obviously you can’t fit all data on the switch, only the hot data, how does P4DB execute transactions that access some hot and some cold data that’s not on the switch?

16:04: How did you evaluate P4DB? What are the results?

18:28: What was the magnitude of the speed up in the scenarios in which P4DB showed performance gains?

19:29: Are there any situations in which P4DB performs non-optimally and what are the workload characteristics of these situations?

20:36: How many tuples can you get on a switch?

21:23: Where do you see your results being useful? Who will find them the most relevant?

21:57: Across your time working on P4DB, what are the most interesting, perhaps unexpected, lessons that you learned?

22:39: That leads me into my next question, what were the things you tried while working on P4DB that failed? Can you give any words of advice to people who might work with programmable switches in the future?

23:24: What do you have planned for future research?

24:24: Is P4DB publically available?

24:53: What attracted you to this research area?

25:42: What’s the one key thing you want listeners to take away from your research and your work on P4DB?

Links:

Hosted on Acast. See acast.com/privacy for more information.

Tobias Ziegler | ScaleStore: A Fast and Cost-Efficient Storage Engine using DRAM, NVMe, and RDMA | #9

Mon, 01 Aug 2022 08:00:43 GMT

Summary:

In this episode Tobias talks about his work on ScaleStore, a distributed storage engine that exploits DRAM caching, NVMe storage, and RDMA networking to achieve high performance, cost-efficiency, and scalability.

Using low latency RDMA messages, ScaleStore implements a transparent memory abstraction that provides access to the aggregated DRAM memory and NVMe storage of all nodes. In contrast to existing distributed RDMA designs such as NAM-DB or FaRM, ScaleStore stores cold data on NVMe SSDs (flash), lowering the overall hardware cost significantly.

At the heart of ScaleStore is a distributed caching strategy that dynamically decides which data to keep in memory (and which on SSDs) based on the workload. Tobias also talks about how the caching protocol provides strong consistency in the presence of concurrent data modifications.

Questions:

0:56: What is ScaleStore?

2:43: Can you elaborate on how ScaleStore solves the problems you just mentioned? And talk more about its caching protocol?

3:59: How does ScaleStore handle these concurrent updates, where two people want to update the same page?

5:16: Cool, so how does anticipatory chaining work and did you consider any other ways of dealing with concurrent updates to hot pages?

7:13: So over time pages get cached, the workload may change, and the DRAM buffers fill up. How does ScaleStore handle cache eviction?

8:57: As a user, how do I interact with ScaleStore?

10:19: How did you evaluate ScaleStore? What did you compare it against? What were the key results?

12:31: You said that ScaleStore is pretty unique in that there is no other system quite like it, but are there any situations in which it performs poorly or is maybe the wrong choice?

14:09: Where do you see this research having the biggest impact? Who will find ScaleStore useful, who are the results most relevant for?

15:23: What are the most interesting or maybe unexpected lessons that you have learned while building ScaleStore?

16:55: Progress in research is sort of non-linear, so from the conception of the idea to the end, where there things you tried that failed? What were the dead ends you ran into that others could benefit from knowing about so they don’t make the same mistakes?

18:19: What do you have planned for future research?

20:01: What attracted you to this research area? What do you think is the biggest challenge in this area now?

20:21: If the network is no longer the bottleneck, what is the new bottleneck?

22:15: The last word now: what’s the one key thing you want listeners to take away from your research?

Links:

Hosted on Acast. See acast.com/privacy for more information.

Chuzhe Tang | Ad Hoc Transactions in Web Applications: The Good, the Bad, and the Ugly | #8

Mon, 25 Jul 2022 08:00:00 GMT

Summary:

Many transactions in web applications are constructed ad-hoc in the application code. For example, developers might explicitly use locking primitives or validation procedures to coordinate critical code fragments. In this episode, Chuzhe tells us these ad-hoc transactions, database operations coordinated by application code.

Until Chuzhe’s work, little was known about them. In this episode he chats about the first comprehensive study on ad hoc transactions. By studying 91 ad hoc transactions among 8 popular open-source web applications, he and his co-authors found that (i) every studied application uses ad hoc transactions (up to 16 per application), 71 of which play critical roles; (ii) compared with database transactions, concurrency control of ad hoc transactions is much more flexible; (iii) ad hoc transactions are error-prone-53 of them have correctness issues, and 33 of them were confirmed by developers; and (iv) ad hoc transactions have the potential to improve performance in contentious workloads by utilizing application semantics such as access patterns.

During the interview he discusses the implications of ad hoc transactions to the database research community.

Questions:

0.58: What is concurrency control and why is it important for web applications?

3:00: How do applications today use concurrency control? Do they use classical database transactions? Or do they use other approaches?

4:09: How are these ad-hoc transactions used in practice? What was the primary focus of this paper?

5:13: You mentioned you studied various open-source applications to investigate ad-hoc transactions, which applications did you look at?

6:16: So what did you find when studying these different web applications? What do these ad-hoc transactions look like in the wild? Can you elaborate on how they differ

8:59: When you compared ad-hoc transactions vs classic transactions? Are comparing potentially incorrect ad-hoc transactions vs correct transactions, if so are performance gains just not accepting it might be potentially incorrect at some point?

10:25: We’ve spoken about how ad-hoc transactions were incorrect. Can we talk about the root cause of this, what were the common mistakes people were making with ad-hoc transactions?

12:16: What was the performance gain of ad-hoc transactions?

15:47: Are there other studies of transactions in the wild? If so, how do their findings compare to yours?

18:38: What does all this mean in practice? Why don’t people just use database transactions? What puts people off using them and thinking I’ll just roll my own?

21:10: Where do you see your findings having the biggest impact?

24:42: What do you have planned for future research?

26:46: What was the most interesting or perhaps unexpected lesson you learnt whilst working on ad-hoc transactions?

29:13: What attracted you to database concurrency control research?

30:53: What is the one key thing the listener should take away from your research?

Links:

Presentation

Paper

Chuzhe's Website

Feral Concurrency Control

What are we doing with our lives? Nobody cares about our concurrency control research

Hosted on Acast. See acast.com/privacy for more information.

Michael Abebe | Proteus: Autonomous Adaptive Storage for Mixed Workloads | #7

Mon, 18 Jul 2022 08:00:58 GMT

Summary:

Enterprises use distributed database systems to meet the demands of mixed or hybrid transaction/analytical processing (HTAP) workloads that contain both transactional (OLTP) and analytical (OLAP) requests. Distributed HTAP systems typically maintain a complete copy of data in row-oriented storage format that is well-suited for OLTP workloads and a second complete copy in column-oriented storage format optimised for OLAP workloads. Maintaining these data copies consumes significant storage space and system resources. Conversely, if a system stores data in a single format, OLTP or OLAP workload performance suffers.

In this interview, Michael talks about Proteus, a distributed HTAP database system that adaptively and autonomously selects and changes its storage layout to optimize for mixed workloads. Proteus generates physical execution plans that utilize storage-aware operators for efficient transaction execution. For HTAP workloads, Proteus delivers superior performance while providing OLTP and OLAP performance on par with designs specialized for either type of workload.

Questions:

0:56: Can you start off by explaining what a mixed workload is?

1:58: What is the challenge database systems face in trying to support these mixed workloads?

3:23: How have previous database systems tried to support mixed workloads?

5:19: What are the design goals of Proteus?

7:23: Can you elaborate more on the architecture of Proteus and how it makes decisions?

8:46: Can you dig into how you predict the transaction latency, what is the mechanism behind this?

10:35: It feels to me that you are accumulating a lot of metadata, this must have some overhead, how does this impact performance?

12:08: It sounds like the Adaptive Storage Advisor is a centralized coordinator, what are the limitations of this decision choice?

13:35: Are we in the context of a data-center here or can Proteus handle a geo-distributed deployment?

14:34: Changing the storage layout has some implicit cost, how does Proteus decide whether a storage layout change is good or bad?

16:57: How does Proteus predict what the transaction is going to be?

18:46: How did you evaluate Proteus?

20:20: If you had to summarize your work, what is the one key insight the listener can take away?

21:07: Is Proteus publicly available?

21:39: What are the next steps?

22:57: What is the most unexpected lesson you have learned whilst working on distributed database systems?

24:21: Do you think a single system catering for both workload types is better than two specialized engines?

26:10: What attracted you to work on this topic?

Contact:

Website: https://cs.uwaterloo.ca/~mtabebe/
Email: mtabebe@uwaterloo.ca
GitHub: @mtabebe

Hosted on Acast. See acast.com/privacy for more information.

Hani Al-Sayeh | Juggler: Autonomous Cost Optimization and Performance Prediction of Big Data Applications | #6

Mon, 11 Jul 2022 08:00:46 GMT

Summary:

Distributed in-memory processing frameworks accelerate iterative workloads by caching suitable datasets in memory rather than recomputing them in each iteration. Selecting appropriate datasets to cache as well as allocating a suitable cluster configuration for caching these datasets play a crucial role in achieving optimal performance. In practice, both are tedious, time-consuming tasks and are often neglected by end users, who are typically not aware of workload semantics, sizes of intermediate data, and cluster specification. To address these problems, Hani and his colleagues developed Juggler, an end-to-end framework, which autonomously selects appropriate datasets for caching and recommends a correspondingly suitable cluster configuration to end users, with the aim of achieving optimal execution time and cost.

Questions:

1:02 - Can you introduce your work and describe the current workflow for developing big data applications in the cloud?

2:49 - What is the challenge (maybe hidden challenge) facing application developers in this workflow? What harms performance?

5:36 - How does Juggler solve this problem?

11:55 - As an end user, how do I interact with Juggler?

14:07 - Can you talk us through your evaluation of Juggler? What were the key insights?

16:30 - What other tools are similar to Juggler? How do they compare?

18:17 - What are the limitations of Juggler?

21:57 - Who will find Juggler the most useful? Who is it for?

24:05 - Is Juggler publicly available?

24:23 - What is the most interesting (maybe unexpected) lesson you learned while working on this topic?

27:50 - What is next for Juggler? What do you have planned for future research?

28:49 - What attracted you to this research area?

29:45 - What do you think is the biggest challenge now in this area?

Contact:

Email: hani-bassam.al-sayeh@tu-ilmenau.de
LinkedIn
TU Ilmenau Database and Information Systems Group

Hosted on Acast. See acast.com/privacy for more information.

Thomas Hütter | JEDI: These aren’t the JSON documents you’re looking for | #4

Fri, 08 Jul 2022 08:00:53 GMT

Summary:

The JavaScript Object Notation (JSON) is a popular data format used in document stores to natively support semi-structured data.

In this interview, Thomas talks about how he addressed the problem of JSON similarity lookup queries: given a query document and a distance threshold, retrieve all documents that are within the threshold from the query document, i.e., get me all similar documents!. Different from other hierarchical formats such as XML, JSON supports both ordered and unordered sibling collections within a single document which poses a new challenge to the tree model and distance computation. Thomas talks about his proposal JSON tree, a lossless tree representation of JSON documents, and define the JSON Edit Distance (JEDI), the first edit-based distance measure for JSON. He talks about the development of QuickJEDI, an algorithm that computes JEDI by leveraging a new technique to prune expensive sibling matchings. It outperforms a baseline algorithm by an order of magnitude in runtime. Our experimental evaluation shows that our solution scales to databases with millions of documents and JSON trees with tens of thousands of nodes.

Questions:

0:47: Can you explain to the listeners what is JSON?

1:14: What is the problem you're trying to solve in your research?

1:48: What was the reason JSON was under researched?

2:13: What is the motivation for this research? Why do we need it?

2:52: What was the solution you developed to solve this problem?

4:35: How does tree edit distance work?

5:18: How do we go from tree edit distance to JEDI?

6:29: How did you evaluate JEDI?

8:31: Do other database systems provide similar functionality?

9:33: Can you tell the listeners more about AsterixDB?

10:20: What was the most challenge aspect of working on this topic?

10:59: What are the future plans for this research?

11:56: What attracted you to working on similarity queries?

Links:

Hosted on Acast. See acast.com/privacy for more information.

Sainyam Galhotra | Causal Feature Selection for Algorithmic Fairness | #5

Fri, 08 Jul 2022 08:00:41 GMT

Summary:

The use of machine learning (ML) in high-stakes societal decisions has encouraged the consideration of fairness throughout the ML lifecycle. Although data integration is one of the primary steps to generate high-quality training data, most of the fairness literature ignores this stage. In this interview Sainyam discusses why he focuses on fairness in the integration component of data management, aiming to identify features that improve prediction without adding any bias to the dataset. Sainyam works under the causal fairness paradigm and without requiring the underlying structural causal model a priori, we has developed an approach to identify a sub-collection of features that ensure fairness of the dataset by performing conditional independence tests between different subsets of features.

Questions:

0:35: Can you introduce your work and describe the problem you're aiming to solve?

2:39: Can you elaborate on what fairness mean?

3:51: Lets dig into your solution, how does the causal approach work?

4:41: How does your approach compare to other approach into your evaluations?

6:17: How can data scientists apply your findings to the real world?

7:54: What was the most unexpected challenge you faced while working on algorithmic fairness?

8:29: What is next for your research?

9:17: Tell us about your other publications at SIGMOD?

10:57: How can the research get involved in algorithmic fairness?

Links:

Hosted on Acast. See acast.com/privacy for more information.

Draco Xu | TSUBASA: Climate Network Construction on Historical and Real-Time Data | #3

Mon, 04 Jul 2022 08:00:01 GMT

Summary:

A climate network represents the global climate system by the interactions of a set of anomaly time-series. Network science has been applied on climate data to study the dynamics of a climate network. The core task and first step to enable interactive network science on climate data is the efficient construction and update of a climate network on user-defined time-windows. In this interview Draco talks about TSUBASA, an algorithm for the efficient construction of climate networks based on the exact calculation of Pearson’s correlation of large time-series. By pre-computing simple and low-overhead statistics, TSUBASA can efficiently compute the exact pairwise correlation of time-series on arbitrary time windows at query time. For real-time data, TSUBASA proposes a fast and incremental way of updating a network at interactive speed. TSUBASA is faster than approximate solutions at least one order of magnitude for both historical and real-time data and outperforms a baseline for time-series correlation calculation up to two orders of magnitude.

Questions:

0:54 - Can you introduce your work, describe the problem your paper is aiming to solve and the motivation for doing so?

4:11 - What is the solution you developed? How did you tackle the problem?

6:50 - What is the improvement of TSUBASA over existing work?

8.59 - Are your tools/algorithms publicly available?

10:21 - What is the most interesting lesson or challenge faced whilst working on this topic?

11:51 - What are the future directions for your research?

15:43 - Are there other domains your research can be applied to?

Contact Info:

Email: dracoxu@stanford.edu
Twitter: @DracoyunlongXu

Hosted on Acast. See acast.com/privacy for more information.

Felix S Campbell | Efficient Answering of Historical What-if Queries | #2

Fri, 01 Jul 2022 08:00:25 GMT

Summary:

In this interview Felix discusses "historical what-if queries", a novel type of what-if analysis that determines the effect of a hypothetical change to the transactional history of a database. For example, “how would revenue be affected if we would have charged an additional $6 for shipping?” In his research Felix has developed efficient techniques for answering these historical what-if queries, i.e., determining how a modified history affects the current database state. During the show, Felix talks about reenactment, a replay technique for transactional histories, and how he and his co-authors optimize this process using program and data slicing techniques to determine which updates and what data can be excluded from reenactment without affecting the result.

Questions:

0:42: Can you start off by explaining what are historical what-if queries?

1:56: What is the naive approach to answering these types of questions?

2:47: What are the problems with this naive approach and why is your solution better?

3:45: Tell us about reenactment, how does that work?

4:48: In your paper you mention two additional techniques, data slicing and program slicing, can you tell us more about these?

6:44: How does reenactment, data slicing and program slicing, compare to other techniques in the literature? Where does it improve on the pitfalls of those?

8:00: Are there any commercial DBMSs that provide similar functionality out of the box?

8:57: How did you go about evaluation your solution?

10:40: What are the parameters you varied in your evaluation?

14:11: Where do you see this research being most useful? Who can use this?

15:17: Are the code/toolkit publicly available?

16:15: What is the most interesting aspect of working on what-if queries and more generally in the area of data provenance?

17:36: What do you have planned for future research?

Contact Info:

Email: fcampbell@hawk.iit.edu
LinkedIn

Hosted on Acast. See acast.com/privacy for more information.

Alex Isenko | Where Is My Training Bottleneck? Hidden Trade-Offs in Deep Learning Preprocessing Pipelines | #1

Mon, 27 Jun 2022 08:00:24 GMT

Summary:

Preprocessing pipelines in deep learning aim to provide sufficient data throughput to keep the training processes busy. Maximizing resource utilization is becoming more challenging as the throughput of training processes increases with hardware innovations (e.g., faster GPUs, TPUs, and inter-connects) and advanced parallelization techniques that yield better scalability. At the same time, the amount of training data needed in order to train increasingly complex models is growing. As a consequence of this development, data preprocessing and provisioning are becoming a severe bottleneck in end-to-end deep learning pipelines.

In this interview Alex talks about his in-depth analysis of data preprocessing pipelines from four different machine learning domains. Additionally, he discusses a new perspective on efficiently preparing datasets for end-to-end deep learning pipelines and extract individual trade-offs to optimize throughput, preprocessing time, and storage consumption. Alex and his collaborators have developed an open-source profiling library that can automatically decide on a suitable preprocessing strategy to maximize throughput. By applying their generated insights to real-world use-cases, an increased throughput of 3x to 13x can be obtained compared to an untuned system while keeping the pipeline functionally identical. These findings show the enormous potential of data pipeline tuning.

Questions:

0:36 - Can you explain to our listeners what is a deep learning pipeline?

1:33 - In this pipepline how does data pre-processing become a bottleneck?

5:40 - In the paper you analyse several different domains, can you go into more details about the domains and pipelines?

6:49 - What are the key insights from your analysis?

8:28 - What are the other insights?

13:23 - Your paper introduces PRESTO the opens source profiling library, can you tell us more about that?

15:56 - How does this compare to other tools in the space?

18:46 - Who will find PRESTO useful?

20:13 - What is the most interesting, unexpected, or challenging lesson you encountered whilst working on this topic?

22:10 - What do you have planned for future research?

Contact Info:

Email: alex.isenko@tum.de
LinkedIn

Hosted on Acast. See acast.com/privacy for more information.

Coming Soon | ACM SIGMOD/PODS 2022 | #0

Fri, 03 Jun 2022 19:57:25 GMT

Welcome to Disseminate! The podcast bringing you the cutting edge of Computer Science research in a digestible format. Each series will focus on papers published at a specific Computer Science conference, e.g., SIGMOD, CVPR, so we will cover a wide range of topics from distributed systems to computer vision. Each episode within a series will feature an interview with the author(s) of a paper published at that conference. The podcasts aims to be an alternative source of information for industry practitioners, researchers, and students. The podcast will be of particular use to practitioners as there will be a focus on the practical relevance of research, in an attempt to help bridge the gap between industry and academia. Also, as many interesting ideas/breakthroughs come from the cross pollination of different disciplines within Computer Science, researchers should also find the podcast useful, in addition to it being a source for keeping up with research in their own research area. For students hopefully disseminate will be a useful learning tool.

The first season will focus on the 2022 ACM SIGMOD/PODS International Conference on Management of Data, which is taking place in Philadelphia from Sunday, 12 June to Friday, 17 June. Episodes will start being released in the weeks following the conference.

We look forward to you joining us on this journey!

Hosted on Acast. See acast.com/privacy for more information.