
#562: DuckLake: The Lakehouse That's Just SQL and Parquet
About this episode
Pedro Holanda joined DuckDB in 2018, when it was still a research prototype at CWI. He's the lead DuckLake developer. Guillermo Sanchez Dionis works on DuckLake and the new Quack protocol.
With Quack as the catalog, DuckLake handles 200 transactions a second under heavy contention. No other open table format comes close.
Episode sponsors
Six Feet Up
Talk Python Courses
Links from the show
Pedro Holanda: pedroholanda.org
Guillermo Sanchez: linkedin.com
PhD on progressive indexes: ir.cwi.nl
SQLite: www.sqlite.org
Litestream: litestream.io
boring hardware: talkpython.fm
DuckDB: duckdb.org
episode 491: talkpython.fm
Iceberg: iceberg.apache.org
manifesto: ducklake.select
DuckLake: ducklake.select
spec: ducklake.select
this diagram: blobs.talkpython.fm
Data inlining: ducklake.select
ducklake-dataframe: github.com
Polars course: training.talkpython.fm
CSV parser: duckdb.org
Zero-copy Arrow: duckdb.org
ART index: duckdb.org
async I/O: duckdb.org
v1.0: ducklake.select
Git-like branching: ducklake.select
Watch this episode on YouTube: youtube.com
Episode #562 deep-dive: talkpython.fm/562
Episode transcripts: talkpython.fm
Theme Song: Developer Rap
🥁 Served in a Flask 🎸: talkpython.fm/flasksong
---== Don't be a stranger ==---
YouTube: youtube.com/@talkpython
Bluesky: @talkpython.fm
Mastodon: @[email protected]
X.com: @talkpython
Michael on Bluesky: @mkennedy.codes
Michael on Mastodon: @[email protected]
Michael on X.com: @mkennedy
Get every episode summarized
Each time Talk Python To Me publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
Transcript ready
1,079 searchable segments. Every word is indexed and playable.
Full transcript
Talk Python To Me — #562: DuckLake: The Lakehouse That's Just SQL and Parquet. Machine-transcribed; use the interactive transcript above to jump the player to any line.
How many files does your query read before it reads any data? On some data lakes, you go through JSON and metadata files first, just to learn which parquet files actually matter. Duck Lake asked one SQL question instead. The metadata lives in a real database. The data stays in plain parquet. That's the entire format. Pedro Holanda joined DuckDB in 2018, when it was still a research prototype at CWI. He's the lead duck Lake developer. And Guillermo Sanchez-Doynness works on Duck Lake and the new Quack protocol. With Quack as the catalog, Duck Lake handles 200 transactions per second under heavy contention. No other open table format even comes close. This is TalkPythonemy, Episode 562 recorded August 31st, 2026. Welcome to TalkPythonemy, the number one Python podcast for developers and data scientists.
This is your host, Michael Kennedy. I'm a PSF fellow who's been coding for over 25 years. Let's connect on social media. You'll find me and TalkPython on Macedon, Bluesky, and X. The social links are all in your show notes. You can find over 10 years of past episodes at TalkPython.fm. And if you want to be part of the show, you can join our recording live streams. That's right, we live stream the raw, uncut version of each episode on YouTube. Just visit TalkPython.fm, slash YouTube, to see the schedule of upcoming events. Be sure to subscribe there and press the bell so you'll get notified anytime we're recording. This episode is brought to you by 6 feet up, the Python and AI experts who solve hard software problems. Whether it's scaling an application, driving insights from data, or getting results from AI, 6 feet up helps you move forward faster. See what's possible with 6 feet up? Visit TalkPython.fm, slash 6 feet up. And it's brought to you by us. TalkPython and Python Bites both now have MCP servers.
Point your AI at 10 plus years of Python episodes, transcripts, and show notes. Free like MCP in the nav at talkpython.fm and at Python Bites. Guys, welcome to the show. Awesome to have you here. Yeah, awesome to be here. Yeah, absolutely a pleasure. It is a pleasure. I'm a big fan of databases. I think databases unlock so much potential for software and also data science, especially talking ducty B side of things. But once you really get good with databases, you can just answer questions so well, so quickly and honestly, it kind of makes the other side of programming easy. So I'm interested to see what you all are doing with duct lake, do some refreshers on ducty B and all of these things. Let's roll. I have a little bit. Let's roll. All right. Well, Pedro, let's have you kick us off. Before we get into that, just to give everyone a bit about your background and introduce yourself. All right. So, well, my name is Pedro.
I came to the Netherlands. I think about nine years ago to do my PhD in the database like Dexter's group at CWI. This is also where I met both Hunters and Mark who are the co-creators of ducty B. I actually used to live together with Mark and Hunters was also my co-supervisor. So the connection has been there pre-ducty B in a way. And yeah, I did my PhD at CWI and saw it's like the ends of my PhD. I had already finished it. The book was already being printed. I still have six or nine months to go. So I just came to Mark and Hunters like, hey, you guys have the school research prototype. Can I do something in it? Basically start nerding out about it and be having tons of fun since then. Funny enough, one of the first things I didn't docty B. I think this is 2018 maybe. Like literally the first month that was in was the CW reader. And I came back to it a few years later.
And I'm actually currently working on something that CWI reader as well. So it's my fruit passion, I guess. You were born to work on CSV. Incredible. It seems. It's also known as the unsexiest database problem. But what can I say? Well, I think it's a subset of the unsexiest data science problem. But also the biggest is the data wrangling, clinging, all that kind of stuff. And CSV is a fairly raw format that doesn't actually communicate a lot of stuff about what types it actually intends that column to be and so on. Yeah, the challenge is of course not only being efficient about it, but also it's a pretty crazy lens. I call it like the Wild West of data formats. So being able to actually read these things, it's challenging. And I guess, well, also the main reason I've been invited to here is because I've also been working duct lake for the past year. Oh, yeah, that. I'll dare to say that I'm currently the main duct lake developer
from our team as well for the past year. So I might know a little bit about it. Yeah, that is actually why everybody. I mean, CSVs are interesting. Not giving wrong, but I have a plan of having CSV files replacing parking duct lake. And then that's a big comeback. Oh, yeah, just zipped up. Exactly. Exactly. For rest. CSV. It's just giving a new extension. Nobody will know. Yeah, yeah, exactly. Exactly. So it sounds like you were there from early early days with duct DB. Did it come out of the university there or was it a project after? Yeah, so it came out from CWI. CWI is not actually a university. It's a in fact, surely a research center. So there's also kind of a fun thing. If you're doing a PhD there, you need to also be a free to the editor in university to get a diploma. But you don't really step on the university. You just get the research center. But the story goes that we had a bunch of projects with different companies as well during our PhD.
So I was working on if all the mark was working with Datas too. And these were more like data science projects. And one thing that was very clear to us, and I also had a lot of influencing the art community. And I think that one thing that was clear like to everyone was that everyone was trying to run away from database systems, right? You'd see people using data wrangling tools, data frame tools, things like that. But no one wanted to use a postgres or things like similar to that. So I think like Mark and Hunter's noticed that gap. Notice that like why are people not using database systems and start talking to these data scientists and quickly realize there was a lot of frustration in setting things up, running queries. I remember that during my PhD, anything that I wanted to do with like money to be postgres or taking just like a day to build things from the sources and set up a database system. And we might take it for granted nowadays that you can just like put the BUP in half a second. And in one line you were like querying something but yeah, like 10 years ago or nine years ago,
this was very much not the reality. So they realized the gap and started doing this research database system. Let's say more or more like, yeah, doors academia. And I think at some point they had a realization that a lot of people are taking it quite seriously. And a lot of people were already trying it to assert there's also the funny thing. People try things in production, even though it's not yet to release the same thing with the clicker. We had people using any production before the click 1.0. And I was always like, cool, cool, cool. What could go wrong? So they basically realized. Yeah, when you see things that are exciting, you're like, this is going to solve our problem, right? It's a good sign. You're onto something. No, no, absolutely. Like I think they were super excited about it when having this realization. And same thing for us with the click was like, OK, like people are taking this seriously enough that they're starting to bet their company on it. So that's cool, definitely. I was a little bit scary, but cool.
And yeah, from that on, I think it just took off initially, I think we're just the three of us, like Mark and Hannes as the creators. And I was kind of like the first person to be hired. But then it quickly expanded. I think in the first year where I like six people where now I think 35 maybe. So yeah, it has been quite a journey. Very exciting. It's cool to see it gain interaction. Now quickly here, you also worked on progressive indexes. Is this for your dissertation? If I'm being honest, my plan was to actually work on CSV files during my PhD. I was told that was a solved problem. So I went to some fee else. It was progressive indexes. The basic g is of progressive indexes is that you can create an index while coding the data instead of creating up runs and having the downtime and whatnot. But it's actually quite difficult to implement it in practice in the database system. So in my thesis, I even have a chapter
called Elephant in the Room where I basically did all my work. Now I've pointed out why it's complicated to implement any production. And things that I would personally look into, if I wanted to continue that direction. But it was also like what a fun journey. And then on the, well, you have the image of the website and there's like my little cover. And my little cover is like a tree eating all this database systems and growing from like a little tree to a bigger tree. It is. Have a very cool. Yeah, no. How about you? Welcome to the show. Yes. Thanks, Michael. My background is a little bit different from Pedro. I actually was working as a software and data engineer for quite a while before I joined the TV. And actually, I come from the experience of like running some of these BFF systems. I think Pedro mentioned Monet TV or Postgres, but also like the BigQuery Databricks snowflakes, right? Which is like, let's say a second generation kind of system where the experience is like quite good.
But the cost is quite expensive. And I mean, one of the things actually that promptly to join, I mean, was that that like was released, actually. And I listened to this podcast that announced that like with Hannes and Mark. And yeah, just a lot of it's like really resonated with what I thought was maybe slightly wrong from other open table formats. I just really liked the simplicity of it. I was a fan of Doug D.V. for quite a while now, because if you're interested in like reducing the cost of your stack, it was probably like the best tool out there by far. So then I actually wrote an email to Hannes and Mark directly, say like, hey, I want to work with you guys. Yeah, I initially joined as a Deborah actually, and then got really to work in that like with Pedro. But then as time evolved, I became like this kind of like hybrid role where I was a bit of like product management things. And also I contribute to that clicks sometimes and also
to Quack, which is our new client server protocol. So we're not only embedded anymore. But yeah, so that's the kind of the history of how I go to the back maps. Yeah, if someone reaches out to you and says, I heard you talk about this. And I believe in your mission. Yeah, that's a pretty big endorsement. So I see why they're like, yeah, you have to come work for us. Yeah, for sure. I mean, I literally told them like, I'm quitting my job. When I have one, I work with you. I think it's actually pretty cool, but you're doing it. Who do it? It's done to me. Spark, never. I was already kind of I think done with those technologies because the things like, for example, if you have Databricks as platform, right? Like it's a large thing that for sure, they give you like a lot of like nice services so that you don't have to worry about things. And it's serverless, but it's also so expensive. And I was thinking like there's so much better things out there that you can do that you can run yourself, right? Like if you're interested in these, this is like that click and that maybe feel like foundational things
that you can run any software on. And I think this, this got to be really excited. You know, in software and deploying systems and database, there's the things that just layer and layer until they just get so complicated. And then there's always this alternative. Like, well, what if we just did a simple thing? You know, if we could just make it not so complicated, would it still be work? And if it did, how amazing would that be? So when you say cost, are we talking like operational costs, licensing costs, you know, do I need a large cluster of machines? So there's always a quorum to vote on different things for durability or what do you mean by cost? I mean, Ghost is actually, let's say the service cost that for example, platforms like Databricks or Snowflake, you know, charge you to use infrastructure that they run, right? Because obviously Snowflake and Databricks they run on AWS or Google Cloud, depending on you can even choose your deployment type. But obviously because they provide the service, but they still have to pay for the infrastructure,
the service basically something on top of that infrastructure. And they decide how much it is basically and because they're pretty much the bigger players in the market together with existing Cloud vendors, like GCP, AWS and Azure, they charge quite steep. And I mean, of course, like the service is nice, like you don't have to worry about things, right? You don't have to worry about downtime and all these things. And for sure, they handle the replication and everything that you need. But still though, yeah, if you want control over your stack, this is like a very high level of extraction that they are offering. Yeah. You also mentioned open table formats and that's going to become relevant to our duck lake side of things. So more generally, what are open table formats? Yeah, so I mean, they used to be called open table formats. And then I think people are leaning towards that like this lake house format name, which I'm not sure why. I think it's actually made it because of data rigs
because I think they started calling it the lake house and then everybody's like, yeah, the lake house. But yeah, open table formats are basically, I think Pedro correct me, Hanrog, but it's basically just some metadata on some pocket files that you can query in a transactional manner, right? And you can do operations also, atomic operations on this data as well. The most basic case, right? Is this basically a pointer, like a, sorry, a metadata file with a pointer to a list of files that, you know, all of those files form part of a table, right? And you can also have things like a schema to it, right? Like what data types do you expect from this table? And you can add more complex things like statistics, like, you know, which files contain which data? And you can get obviously a lot more complex than that, but that's sort of the basics, I think. Yeah, I think, I think like the interesting thing of the open table formats is of course, the openness part of it, right?
So they have an specification, the file formats where they're stored, they also have an specification. So by just reading the specification, anyone should be able to implement a reader in writer to that format, right? So the whole idea of this, or maybe the beauty of this is that it should be able to solve the problem where you're locked in into a certain system, right? So you suddenly don't need, like, if you start off your database from Oracle, for example, you're kind of locked in. Like doing an Oracle migration is known to be painful, to be costly, so you're gonna be using, you're gonna be giving Larry ads on a second island to see how I, in the next few years, that's fine. But then if you're actually using open table formats, you kind of solve that problem, right? Like you can literally switch from different engines to different providers, and if you're tired of data breaks, and now you want to use no flake, or now you use, what's your own solution based on ductyby, you can theoretically quickly jump from these, right? And it's also super scalable, because if you're just dealing with file formats,
or like a catalog as well, but your data is in pocket files, it can easily distribute your reads on it. So that's quite some cleverness behind. This portion of talk Python to me is brought to you by six feet up. Let me ask you a question, what's stopping you? Maybe it's an application that won't scale, or an AI initiative that just isn't delivering. That's where six feet up comes in. With deep expertise in Python and AI, they solve hard software problems, modernize platforms, and get teams to market faster. These folks have been doing Python since version one. They know the frameworks and ecosystems like the back their hands. Six feet ups, impact speaks for itself, automated healthcare pipelines for hospitals, helping NASA explore Pluto, building severe weather prediction tools, and applying AI to connect farmers with vital crop data. When the stakes are high and the problems are hard, six feet up is the partner that delivers. See what's possible with six feet up.
Visit talkbython.fm slash six feet up. The link is on the episode page and in your podcast player's show notes. Thanks to six feet up for sponsoring the show. The examples I've seen are like a whole bunch of park hay files on S3 or Azure Blub storage or something like that. And then some other file may be also stored there that you can say, well, let's read the metadata file that describes what files are here. But I mean, your dream could be back. It could be zipped up CSV file stored up there with some metadata, right? So of course, in that lake. So maybe one quick thing is that Iceberg, which is I guess the other main open table formats, it uses files all over, right? So you have your files in the metadata and then your files that actually store the data. In that lake, we'll talk a bit more about this, but the metadata is actually a database system and your file is still for our care files. So there's nothing stopping us from a technical point of view to actually replacing this files of CSV files.
From the decklake spec, that's actually already allowed because we can store in the catalog that these are CSV files and we could theoretically bring the stats necessary. So it is on the cheapable dream. Yes, exactly. The dream of the CSVs are still alive. I don't know, I'm on the CSV kick, but what's the largest CSV you've ever parsed? Personally, I think I've reached close to terabytes on benchmarks. But I've seen people, like, you know, people, they, there were many times in CSV land, there was like, yeah, this sounds like a rather reasonable limitation. No one was ever going to do something. I think I had a limitation, for example, that no one would ever have a line in a CSV file that was over 32 megabytes. And that was kind of nice for me to be able to parallelism. And then, of course, some guy, hey, this CSV is throwing error for me now saying that my line is over 32 megabytes. And like, she's the Scrysman. What are they storing there?
Like a book, a book per line. Yeah. So if I've done a terabytes, I think some people have gone a bit to the crazier. Well, let's jump into our topic here a bit. And I want to work our way sort of do a little bit of talking about ducty B and where I'm a, you talked about Quack. In the early days, ducty B was an embedded database. The most common one of these is SQLite. And Postgres is all the rage, obviously, others as well. But Postgres certainly seems to be taking over a lot of the database side of things. But I think there's still a bit of a research and stuff embedded databases. And I think ducty B is an example of that interest, right? So maybe just you guys give me your thoughts on SQLite and some of these other embedded database ideas. I just wanted to mention that I think Pedro can give more research background on this. But I think from a user perspective, I think the idea of having your database within your
application system was always very attractive to me, right? I think Pedro mentioned this at the beginning. It's very nice to not have to provision a Postgres database or any other type of database on a separate server to run your application. And I think SQLite solved this wonderfully. And of course, when ducty B came by, it was also with the idea to solve this from an analytical database perspective. But of course, ducty B can also do transactions, right? Pedro can talk about it with the keywords. But this also very nice, right? But you can still run transactions in ducty B. But overall, I'm just, I would just be very happy with the idea of being able to embed your database system wherever your application is. Because I think there's a lot of use cases where this is just good enough. And actually, one of the coolest use cases I think that ducty B brings to the table is the fact that you can use a wasn't client in the browser and serve your database over S3, and literally there's no server. And you have an application that can query data from your browser. And I think that's amazing, right? Like it reduces the costs completely.
Because S3 storage is very, very cheap. So that's really incredible. I, you know, there's just so much operational simplicity. If you can make these things work, right? If you make these embedded databases work, right? It's just like, wow, there's nothing to go down. There's no, like, oh, this server needs that security. But then, you know, like just all those things to juggle, right? Just it's a file. It's in process. You had the eliminating like a lot of this complexity of keeping the server for sure. And there are other beauties to it as well, right? Because you suddenly have your database running within your application. So it also can share the same memory space of your application. So you can do a bunch of tricks to avoid copy memory all over. So this is especially interesting, like for data science projects. Because if you're using something like pandas or an ampai, what is an ampai array? Is literally a C array with some makeup on top? What is a duct to be vector is literally an array with some makeup on top.
So you can just change the makeup. And then you can suddenly access the same data with constant cost, right? You don't have to trust from your actual data. That's interesting. So when you do a query, you may be able to just return a piece of the in memory duct. Absolutely. chunk instead of going, okay, ours looks like this. But we're going to copy a million floats over to this thing. And that's called them and then send it back, right? So this is what was also like one of the things that was one of the realizations that database protocols, right? So the way you transfer data from the server to the clients, they're actually quite slow. So this is one of the main frustrations we had seen with the data scientists. It's not only like, oh, it's clumsy to set it up. And you'd like to start a server and create schemas and whatnot. But it's also just to get your data from your ampai or TensorFlow or pandas or whatever you're running into the data space system. And back and forth was super slow. So you completely remove that boundary. And yeah, like you gain all these benefits, which is inspired and very similar to what
SQL either already had. But of course, of the difference that there's a huge focus on analytics. So it's a color format set of a row format. There's a lot of emphasis in compression and vectorized execution and so on. I think probably people out there listening, they're like, wait a minute, columnar row format. What does this mean? Give us a bit more detailer. Absolutely. Well, there's basically two ways that you can store your data. So that is a table. Say that is a table that sells products. So you have products, quantity in your shop and price. One of the ways you can store it like exactly in your memory is continuous row by row. So start first row, then the second row, so on and so forth. However, analytical queries, they are usually of the type of give me the average price of all your products. So in practice, even though you have three columns, you only really want to access one column, right? So on the color format, you actually store both memory and on your storage the columns
first continuously. So if you want to access one column, you don't have to read basically all your table, but you know exactly from your memory or your disk where you have to read this to access that column. And the other nice benefits from that is that lightweight compression is heavily dependent on the proximity of your data, right? So it's much easier, for example, to compress one column of dates. You can maybe do like delta compression quite easily because dates tend to be the same for a long time, and then they just change one by one. That's due to that if it's in a row, right? Because the next value of a row is not going to be a date, it's going to be a double, or it's going to be a string or whatever it is. So this change in formats allows you to access your data much faster with the disadvantage that if you want to update one of your rows, it's a bit more expensive, right? Because again, if you have the row formats and you want to now update the price of your Volkswagen cars or whatever, you just need to fetch that one to Apple and you immediately
can update that. So it's one random access in kind of done. If you are on a color format, you're first going to have to check that one whole column to figure out where are your Volkswagen cars and then go to another part of your date to figure out where the price is and update that value. So it's a bit more costly. And that's why people say that SQLite is good for transactions and transactions in the sense that your updates won't row every now and then. And that could be great for analytics, which is again, this type of query that you breed due checks of your data, but like fewer columns instead of lots of columns. That's a really good explanation. Yeah, I could just see if you want to query, you know, take the average of all of the prices. You've got these rows. You basically got to seek over every row through the entire database to get those things. And if you have a VAR chart sort of thing, then it's even harder because you're not even skipping known links. You got a compute bowl. Okay, where's the price in this particular row? Isn't it?
It's just very different. And maybe just to add the color format also allows you to do vectorized query execution. The basic idea of vectorized query execution is that you can actually have batches of your data being through going through your query plan in one go, right? Unless you have a pipeline breaker, like a join or something like this, but usually you can go through the whole query in one go. And because you have these batches are sufficiently small, they just get cash on your CPU and then you don't have to constantly go through memory to fetch that data anymore while in a type of wise execution, which is like what SQLite does. And it was created more like when memory was small. So you want to like to just have a little bit of your data in memory like from the 90s. You just go and execute through the query plan, tuple by tuple, right? So this calls a bunch of cash misses. You frequently have to go to memory to fetch the next tuple and so on. So that also allows that extra step that makes a huge difference in analytical performance. It's funny.
It used to be things were optimized because memory was expensive in terms of it was scarce, right? That was hard to get enough memory to the handle all the data that you're working with in the 90s. And if people have heard of like third normal form and normalization, all of this is we must not have and we must not waste the memory, right? I mean, other reasons as well. But still we must not waste the memory and then it got kind of cheap. We were able to use it. And now with AI, we're back to memory of scarce again. How about that? Yeah, it's actually for databases, the cycle has started a bit earlier because I think like in the mid 2000s, there was like, oh, we don't need disk. We can have our database as was the e-memory database system, right? So just everything in memory and it's all good. And then you can map like just let the OS handle whatever needs to go to this cover now and then what is going to be good. And the reality is that in these memory has increased drastically, but there's still lots of limitations. If you want to run things on your phone or like in an Arduino or something like this. And that's also the ability to be able to put a lot of efforts in having proper buffer
managers and like all our operators built to disk as well. So even in scenarios that you still have memory constraints, the two will just work. Like it's not going to crash. Maybe like, oh, yeah, there's no memory by because that's also a little bit frustrating. That definitely is. And just the problems people are trying to solve. It's like, yeah, we have 10 terabytes of data that, oh, we didn't need to have that much data typically. That actually one of the parts that I love the most about our website is we have like these series of blog posts that run on exotic hardware. And basically, you know, this can be a Raspberry Pi, but it can also be an iPhone, right? And we try to push it to see where's the limit. There's this very nice image of an iPhone cooling down in a block of ice because we're running the scale factor. I don't know if it's a hundred or even one terabyte on an iPhone 15. Yeah. It's also like a very cool thing about that TV that it truly can run anywhere and that it doesn't matter what constraints it has. It can still run queries, you know, with this out of core processing that has undiability
to spill to these. This portion of talk Python is brought to you by our AI tools. You know that thing where you ask an AI something about Python and it confidently tells you about a library version from 18 months ago. Well, we fix that at least for our shows talk Python and Python bytes both have MCP servers now connect talk Python and your AI can search over 550 episodes, full transcripts, every guest in the entire course catalog connect Python bytes and you get almost 500 episodes of Python news going back to 2026, including every link we've ever put in the show notes. That means you can say things like ask talk Python what Astrol joining open AI means for UV or what has Python bytes said about UV and get real answers with real links, not a hallucination name the show in your prompt in your AI knows exactly where to look. And if you live in the terminal talk Python also has a CLI too one line UV tool install talk dash Python dash CLI and then search episodes transcripts guests and courses without
ever opening a browser. It's open source and it outputs text JSON and markdown so it feeds the AI tools that don't speak MCP yet. But here's the real reason I built it both shows cover around 10 years of Python history. The people the decisions the packages that took over and the ones that quietly didn't this enhanced access is free no account no API key nothing to buy this history should belong to all of us. Visit talkbithon.fm and Python bytes and click the MCP link in the nav bar connect them right now to your agents so that they'll be accessible anytime you need them in the future. There's also some stuff that's been happening to make running these as real back ins for apps better on like the SQLite side we've got the wall or the right head lot log that allows you to have a lot less locking. We've got light stream which will like basically build on top of that and stream constant replication to an s3 bucket. It was a story with ducty B on that kind of stuff.
I mean you talked about clack that sort of ducty B gets to be a little bit more grown up and distributed right. Yeah I mean for sure I think the ducty B experience was always very nice to run on a laptop right or to basically if you weren't wanted to run like an ideal pipeline let's say on an easy to machine like the B was a great tool but I think at least with the original ducty B file formats there was this sort of limitation around you know one writer grabbing the log over the file and then no readers can connect to it and I think like that is not a limitation of course if you're running ducty B your laptop but it is a limitation if you want to share ducty B storage with a brother of the ins. Quack in these is something that solves this problem right but then that means of course that you go into the client server kind of architecture. I think SQL Lite does this a little bit different where even on a single machine with a single
file you can still have writers and readers at the same time but according to me if I've wrong I want to share about these maybe but I don't know what you do in the spot you're not a super light basement so you don't have to. I think so with the right I think with the right head lock it does log rather I think it does allow multiple writers to different parts of the database but there's still some locking. We have the same we also have a right head lock but probably they have a different mechanism that allows for multiple writers at the same time. I mean ducty B does have like these you know within one single process you can still have multiple child connections right into the same database as well I think the problem is if you try to access duct B from a different process that's where the lock in mechanism figures but yeah it is quack it's a good solution for these but it goes more into this like client server recipe again. So if I'm using quack do I host my back end with your mother duck and some sort of cloud situation
do I self host some back end piece to make this possible what's the story? You can host it anywhere you like actually like I mean so some of the nice thing about it being ducty VHS is extremely simple to use right so if you want to try it out you'd go to a terminal you top you to basically type install quack load quack and then the next thing that you do is you call a function called quack serve and that will already speed up a server for you right and then from another terminal client you could just go in connect to this specific address in your local host and then you already have a client server protocol working which you compare it to again traditional systems like post-res is extremely simple and obviously in a more production grade scenario you host these on EC2 and then maybe you need something like a load manager on top of that because also nice thing about quack is it runs on modern HTTP HTTPS so it's also very nice you don't need to reinvent a protocol yourself on top of TCP it's just
plain old HTTP so yeah I think it's relatively simple to still get something up and running on on the cloud and getting to work yeah it's a it's a very nice foundation I think for a client server architecture I like how you all keep it playful as well as you know with quack and ducts and you know it could be now you're going to create a database provider factory and the factory is going to get the provider then the provider is going to you know just like you know I don't know yeah actually very funny story about this another do you bring this up is that it was called quack at the beginning right this was harnesses idea and then someone within the company mentioned like oh mate should we call this a little bit more of a straightforward name like uh you know RPC protocol or something like that and then both me and Gabo which is the main developer in the team we said like no no no no way has to be called quack quack is too good to pass on we cannot pass on these names so good I mean the branding is great you hear quack you like oh that's got to be
ducty be right like yeah no one else is going to take the quack protocol and get away with it all right let's let's talk about duck lake and so yeah let's talk duck lake I think it's a really interesting project it's it's not a direct follow on it's not just a more distributed ducty be right this is a bigger idea yeah so the the basic giz of duck lake is that I think market harness were already looking at data lakes for a while because we had been getting lots of requests for support for iceberg and I think they were already seeing that having the metadata in files was just not as efficient right like you suddenly have to do so many hoops on these files to get to your data the more snapshots you have because maybe researching tons of small files there's more file problem right like we all have heard about this if you've been around data lakes is that
basically if you have lots of small researches like in a streaming fashion for example you create so much of these metadata files that it's hard to get to your actual data so I don't know I guess as database researchers and engineers they were like well why is this not just a database so I actually have it somewhere here oh easy to pick this time this was the new show idea they had so here you can see like all the files right so they're like why can't it just be like a database and there's already like a database here on the top because at the end they needed like a database to get a pointer for the latest metadata file so what if we just put this all in a database and then I made them sign for me in case this was a very successful idea so if anyone wants to buy this as for sale or for the right price so that was the basic jist right because having this whole metadata in a database you can just query with SQL anyone can easily see a schema and write SQL over it so just need specify the scheme of the stables that hold all this information the metadata
and you could you could then have something that's more efficient and and simple right because I think for iceberg you have jzone you have Avro you might have some auto file formats that I'm not just just to start like the metadata so just to be able to read this format to need like three or four different file format readers and then you need to hoop through all of these and with Dark Lake is simply like okay go to the catalog as the catalog which files I have to read to read my table at a certain snapshots you have the list of files and you read them so that there's this implicit effector nice I think maybe the place the the way to make this clear to folks is let's maybe talk through the architecture of these data lakes in general but also the the Dark Lake specific version you know and I know a lot of people have worked with these necessarily worked with databases they worked with APIs and they worked with storage like s3 but the general idea I guess maybe the iceberg original idea was well what if we could just use
all of s3 for scaling like think how scalable that is and we should put a bunch of files all over the place like a bunch of park a files well we need something to gather them together if you're going to a query across them so we'll put a metadata file up there as well but then the problem is well I'm going to do a query so it's a s3 request to the metadata and then you you parse that maybe you got to get some more metadata then you go to the file I mean it's just a lot of back and forth and Peter that's what you're talking about we're like well what if a lot of that back and forth could just be in a database and then you finally get to the data right exactly and I think sure there's something to be said that if you just have your files unless you can theoretically scale this indefinitely right because you could have businear readers three times the same time that's fine but with database systems well I would say that the point you have to scale is much lower a database like that could be can use can make use of small piece of hardware quite drastically but it's also you have distributed database systems right so the only restrictions for the catalog is that
I think it needs to have time step varchar and integer types and its primary keys so any system that has this fork and theoretically be catalog for that like so I do believe that there's also potential for a high first skating there or the case where your metadata is really that big and again this is not about the size of your data right is the size of your metadata all right just here's here the tables here's the elements here they're stored you know it's worth maybe emphasizing that scalability something that's highly scalable doesn't mean it's fast it just means as you add many many more users or amounts of data to it it doesn't get slower but it could have been kind of slow from the start you know what I mean like like 12 back and forth across s3s even for one query is still slow if you don't have a ton of data but maybe it's like consistent as you add a million more right so there's that worth considering it's like performance for one thing versus scale absolutely but as you think the again that from the catalog side this
kind of ability of like it doesn't doesn't necessarily have to suffer like you have no no no yeah absolutely I'm not yeah not not suggesting it does I'm just saying like okay I built a scalable system like great but it was slow from the start so like it does yeah absolutely no I mean all right so basically the architecture looks like this there's three core elements there's storage there's catalog services and there's compute who wants to break that down for people yeah so I mean basically the storage can be as you mentioned before anything that is of the storage like s3 or or blob storage usually you know people go for options that are cloud hosted of course like s3 gcs or Azure blob storage and that's where the parquet files lead in the case of a duct lake and nothing else because we don't have any metadata files then you have the catalog service which you can theoretically run in anything that is big sequel the duct DBE extension of
duct lake right which is basically our implementation of a reader and a writer for duct lake can use can speak to my sequel postgres sequel light and of course duct DBE and quack and because they're not exactly the same right one leaves next to you the other one leaves in any remote server that you did you want it to be and this catalog service is what contains all of the metadata regarding these parquet files that leave you not take storage so basically then the third component which in this case is the compute right and indeed it can be anything it can be spark or it can be duct DBE and of course anything that supports this format and then the only thing that they need to do in order to query for example some tables right is they they make some sequel queries against the catalog service right and the catalog service returns something like a list of files for example and some statistics regarding those files so that the reader knows you know what it needs to read basically and I think this is very nice actually because you just need one query actually in order
to retrieve all of the information that you need to query these parquet files and I think this is the big powerful thing about that lake is that you know we designed it with pretty much iceberg and delta in our heads right where usually there's more than one route trip taking more than one file just to figure out which parquet files you need to read in our case is just once equal query I think that that is basically what collapses the complexity quite a bit for engines that are trying to read this format or write to this format yeah and you can you do database things like a joiner something if you need to right yeah I mean indeed I actually like one of the things that is very interesting about iceberg and delta is that because they grew backwards they started to think about operations in a different way right like you were mentioning for example you know at the beginning universe this right in your parquet files and some metadata that points to them but that's not boiling they wanted to do acid operations on on this data right and then you need to give some sort of guarantees and to keep those guarantees they build this very complex structure of metadata files and
manifest list and manifest files when for us at obesity is just basically you know if you're right into a table right and then yeah you're basically raising another transaction you can just see whether the other transactions succeeded or not right it's just something that is built into your catalog service so for us we didn't have to build any acid it was built into the design basically because the catalog service is a SQL database that already supports acid transactions so yeah I think that that was also an extremely nice thing about the dot click from the get go yeah absolutely no we'll take this give us a sense of the size of these parquet files like I put up I know when I created a date like I've got a bunch of different parquet files I've got a database catalog that will tell me okay if I have this type of query go look at these three but are these five kilobytes five megabytes five gigabytes what are we talking well you you decide the size basically right like you you can say that the target file size is something like 512 megabytes and you know if you're doing
batch operations on your data so batch writing like tag like we'll respect that you write 512 megabytes files with a certain row group size of course if you do very small operations on your data where you do very small rights then you may have if you don't have data in line in on which is a feature that maybe we will discuss a literature you will have very small files and then what happens usually in lake house or open table formats is that you have some sort of compacted functions that the engine offer that allow you to compact these very small files into bigger files so that you know engine is going to scale reading better right because engines don't like to read the files on five kilobytes files they prefer to read you know how fatigue bite file that you can still parallelize by the way but yeah okay that makes a lot of sense I guess the smaller they get the more your pain the cost of the latency round trip than than the actual reading now operationally transferring data is fast on s3 and friends but it can't start to get expensive if you have too much traffic
there what do you recommend running this maybe inside of a data center where you have within the data center traffic to your data storage or what does that look like I think better is the expert on reading from all the other storage yes I think the tip goes setup is that people use like an s3 EC2 kind of instance right so you have in an RDS so in the end you can definitely have this separates I don't I think that's the usual way that people run yeah that's what I would imagine or you know digital ocean servers with spaces or anything it could probably talk to anything that talks s3 I would imagine right yes absolutely yeah okay now back we've got different different ways to set this up the oh before sorry before we move on there was a third piece there's storage there's the catalog and then there's compute what's the story with compute
what kind of stuff in my computing right yeah absolutely theoretically you can use anything that reads spark if files to perform the computation right as Guillermo says in the end if you have a catalog running you just have to reach the correct queries the queries are basically tell you read these files to to answer the query you want you want to answer and you can use any kind of two that is capable of reading these files I think that's most people just use deck TV and there's also like a data fusion extension that's quite evolved already but I'm not sure if there's like a spark proper implementation of the the spec maybe Guillermo knows there is a spark right there I think specifically lead that was developed my mother that because some of their customers were using spark for ETL so they thought that it makes sense to have a spark writer into that lake I say so yeah so the the basic gist is that you can actually use any engine you want as long as you have like
these pieces appropriately builds nice and then you've got some Python code or whatever code that talks to you exactly that and so I have this question go figure it out yeah exactly okay now there's different setups here I would imagine the most common setup is probably post-gress as your catalog and storage is parquet files though a hat tip to compress the svs post post-gress and parquet somewhere but you can also do partly why I jumped on this you could do SQL light as your catalog or you could which is in process and same with the in process ducty B or even where most quack back in that's distributed right like maybe talk about these different options when do you recommend it so I would say that a lot of people also use ducty B as an in process solution to either develop their own data lake on their own machine but I think also a lot of people have been using with their favorites clanker just to store the data also in the local lake house formats but if you're using ducty B specifically it also means that as ducty B
works like the Guillermo said before is that you can only have one writer connected at a time right so you cannot have multiple people connect to the data lake and writing if the ducty B is your catalog so I would say in production what I've seen is that indeed most people use post-gress because it takes that limitation away so you suddenly can operate with it in a more in a way that people are more accolcents to use and so and of course quack is still experimental I think we're going to have the first stable release with ducty B 2.0 like in the month and a half but quack already works with duct lake as well but don't run that in production right now people wait a little bit longer but it's also quite cool because for example when using post-gress one of the problems that we face is that if there's a conflict on a snapshot idea right like you're doing multiple transactions on the same table every time I have the conflicts you have to return to the application from the database system say hey there's a conflict recap later snapshot ideas some
other things send the chords to me again so you have this round tripping and that actually creates quite a big cost on retryos but with quack because it's just ducty B what we can do is that we loads duct lake all the quack server and then you can immediately do the retryos in the server avoiding this process right so to give you an idea if you have post-gress as your catalog I think in a very contentious environment with I don't know like 20 writers you have something like five transactions a second because of this retryos cost but if you're using ducty B and quack is like 200 transactions a second because the retryos are now running on the server and so yeah my expectation is that as time goes on quack will be the defector catalog for that lake maybe something to add to that is that you know I think lake house formats were never designed for transaction on workloads in the first place right but something that we realized quite early to be able to with Pedro is that
you know actually we shoot flag that duct lake is actually better at doing this transaction on workloads because it's one of the selling points right and already with the post-gress setup we were doing quite okay in transaction on workloads versus you know the ducty B also doing transaction on workloads against iceberg and with what Pedro implemented the server side retryos in quack which could theoretically also be done in a way in post-gress we would see if that happens or not it's actually very very powerful to the point where 200 transactions per second in a contention like in a high contention environment is almost unheard of for the other open table format because this open table formats are designed for batch processing right is this we you write once to a table very large chunk of data but they're not really thought for like transaction on loads or streaming data things like that I think the joke was that if you wanted to buy something use your bit scoring it took like 30 seconds because it could only do build transactions and then
iceberg was even slower than that right so yeah it is a bit hard off and just maybe complement what Guillermo said as indeed possible to have something similar to post-gress the main reason I haven't done it yet because of course we can all do the trick of loading duct lake in a post-server however we could theoretically do start procedures but I haven't done a start procedure since college so that's why I've been going to love it. It's really hard to do a store procedure in a CSV. No I honestly for all the the benefits that store procedures have it's just you know in general you have so much more flexibility if you just do SQL and all that on your own but yeah that makes a lot of sense if you could just package it up and go you just call this and the database handles it in process. Yeah so multiple clients works locally or in the cloud I guess for the the non-hosted ones what's the scenario here like either ductyb and process or SQL light is this I'm
working on my machine and I know that the the data is out there or how I share this how do I like wrap it because like you've got the s3 storage with all the stuff that could be shared but then you've got the local SQL or ductyb file. Yeah there's there's a very cool use case for this particular with ductyb in process which people started calling the frozen duct lake and the thing is that to to update a tag lake packed by an in process ductyb catalog you do need to pull the file the ductyb file locally right you can still you know talk to the catalog file locally and then to the updates there write your data to actually s3 directly and then once you've done this for example writing process right which let's say happens once a day it's like a batch workload then you up you upload the the ductyb file to s3 storage and then you can have any amount of read only connections to s3 storage like backed by a ductyb file as a catalog which means that basically you
can have a data lake almost for free right because your catalog database is just a file in s3 which is very powerful so this this works very well if you're doing like let's say one update per day or it's just bad slow the bunch of data and then you can have an arbitrary amount of read only connections to this frozen duct lake yeah this frozen duct lake that's an interesting idea that way it just operationally is easier right like just it runs overnight and the morning download the ductyb file ask all the questions you want yeah I mean it's also like the very nice thing right like you you can literally have anyone asking questions without that without a server catalog running right so you don't need to pay for a post-restubing running or for a quack to be running you can just literally have a ductyb file in s3 storage and then read all of the duct lake from your laptop yeah and operationally it's so much easier yeah so much easier indeed so another interesting optimization that you wrote about is data inlining instead of sending every little request all the way back to the object storage yeah the the basic just here is
again to try to mitigate this small file problem right so there's a mile file problem again is if you have a bunch of liter insertions you generate a bunch of little files and in iceberg you're also going to generate a bunch of metadata files I think it's or files per insertion if I remember correctly but some far from being a nice per specialist and that basically means that if you want to get the information up there one snapshot you have to request all these files and although there's small as you said there's latency playing around and it just adds up and you end up paying quite a big cost on retrieving your data the clique would still have that problem to a small scale right without inline so basically you'd still have a bunch of small data files but you would solve the catalog parts so the idea here was like well we are using a database management system they can start tables what if we start the same table as our data but in our catalog and then store the
inline and they start the actual data that would go to the file in this inline table so that's basically what it is so imagine you have a per key file with like dates and value to columns you go instead of creating that per key file you have if you're using ductive guys or catalog you're going to have a table in your deck to be called something like underscore the clique underscore data table and then you're going to have these two columns plus some extra columns we need to actually run the query so like these snapshots id where that value was asserted the end snapshots if that value that have been deleted while it has not been created to a fire yet and I think we have a third column but I don't remember why hard now that's the basic and then of course at some points if your data grows enough so the idea here is that your insertions are kind of small right like a hundred tables ten tables maybe a thousand tables but it's not really that big and if at some point your table is too big you can just flesh it to disk so you remove the data from your inline table and you create the actual file and then you avoid
this mid situation where you have a bunch of small files and then iceberg tries to solve this issue I think with some streaming tools I'm probably more aware of these things than me but this will again force you to have another two I think it also gets in the way of transactionality or having your snapshots per data file at least because you're going to have to start batching you're the hidden tries to do it in one go and with ductleck you super serve all of these and yeah like I have some numbers of this blog post but of course we just comparing raw iceberg with raw duct like yeah so the numbers is like we got says basically a thousand times faster it's a little bit of a performance improvement yeah it's exactly so it is like a really cool feature in my opinion it is and once you've got this data this local database in place you can use it for things like that right you need the database anyway to have a consistent data like so yeah exactly yeah might as well and then I mean when you say with open some file and Python you get buffering like not every
single byte you write instantly goes to the file because that's a lot slower and it just seems like a much more scaled out one way which is even more important right because the latency of talking to object storage far away absolutely nice all right let's we got a time for a couple more things let's talk about duck lake data frame what is this and so while doing that project with the I had to of course that's a nice work instance to compare with that to be I thought I thought was a bit first trading the experience it took me like a while and then I just had my clanker going over it and I went to do something else and still took my clanker like a while so I start like wondering oh maybe it would be easier to have my clanker implementing you the clank implementation it's to actually get iceberg to work or pie iceberg on this machine of like some glue not glue sorry some the iceberg cataloger forgot the name now I'm working and then the the the oh idea of this experiment is just like this was two months ago so the clanker was not as smart as it is now and I just let it go and
I just allowed it to write it and it came up for something that worked quite well like I didn't reveal the code but like I had nice numbers and I checked the test like over like looks good enough but I also made some fun benchmarks that was like oh what if we just change our scheme off the table a million times was how much faster is this a life's work and I thought it was just funny I mean actually it's funny but it's also like one of the craziest things the the daglick improves over iceberg right because every time you change the schema and iceberg you need to write files as well right and I think that's also the crazy thing about daglick is that why it's not crazy actually it makes a lot of sense but we just do basically get sequel operation and this is just so over time like if you have like all of these different operations that are you know like other table commands for example or stuff like that which actually are not changing data but are growing the metadata then your re-performance we'd decrease because of that and of course I'm changing one million times the schema probably is not a real case scenario but
anyway it's quite interesting. Ah is this one the schema evolution rename I was just like okay I'd have no idea how a cloud font is an interesting metric but okay. I just have a point I think. Yeah what are the things we can do with it yeah very cool but maybe the the more interesting part of this like what I wanted to show or at least to experiment myself is I do believe that compared to the iceberg open table formats the daglick is much easier to implement so that was the idea behind this experiments and since then we've seen some people doing actual implementations there are I think close source implementations from a fireball to think but there's also the people from hot data that've done the one from data fusion there is another implementation that's actually more of Fischo the mine from a Bortjero I think that is the actual data frame daglick so again it was also to show that despite the name a daglick is not only entangled to daglick it's an open table format we want people to also to to write their own readers and writers and to use our own
implementations a reference but also a website has like the spec description so the idea of that experiment was also in line with that position sure that makes a lot of sense when I first saw duck like oh this is duck db but for data lakes how's that working then you know look more into it realize what it was now said don't use this in production when when can you what's what's the path to 1.0 ready for production all these things yeah so I'll say the main like we actually took a a bunch of time to from the click 0.4 I think to 1.0 where decided we're not gonna add in new features we're just gonna focus on fixing bugs and we're only gonna do schema changes if they are fixing bugs and we believe I believe that since April it should be production ready of course as software is there are still issues and we still have bug reports and we still fixing them and we're gonna be releasing the click 1.1 strong I believe most likely in a month and a half of
duck db 2.0 again almost on the same perspective that like it will have some smaller features or some more optimizations but it's mostly focus also on bug fixing so this is really the direction we're going now it's trying to make it's more solids and more compatible with iceberg so we can also read from and export to iceberg so in the sense of being an open table format so also don't want people to be logged in in our own table formats if that makes sense yeah yeah and one thing that is interesting about what Pedro said is that like because we have this specification but also that the duck link implementation of duck db right I think this specification we tried to sort of like free-seat as much as possible and made sure that we only do like new releases of this specification if it's really necessary there's like a breaking change in the in the catalog schema but the extension itself can have like not only bug features but also like performance improvements for example right like I think we've recently seen like some improvements in like the way we
calculate statistics and I think there's a bunch of others so basically the extension can still get better right like the engine can can be faster and do things better anyway so it doesn't mean that that click is not evolving I think it's like the specification evolves slower right as it should be in our opinion and the extension can still get better and better over time now let me just put this out directly it sounds like maybe it is ready for production what's the what's the readiness state of duck lake yeah so I think one of the main things we wanted to achieve for 1.0 before size like all the bug fixing with knots is that one of the main aspects when you run a data lake over a long period of time is that you get more data in it's right so we needed all the check pointing functionality and we've checked pointing as compaction is rewriting data files by adding their deletions is removing orphan files it's removing old files all these things completely ready
and easy to use so I think that's also one of like the main things we ended up spending a lot of time in this production readiness states to be sure that people can just run it and not be completely stationary because their data grew too much right so that was like also one of the main concepts behind this yeah so it is production right I mean for sure there's come I mean there's already a lot of companies that are running this openly and they even base their software on it like I think post hoc is a good example alter table is a good example I think fireball also reworks their their story-changing to use duck lake as well so there's a there's a lot of companies that are they trust duck lake as like yeah you know a thing to trust your software on which is like for me it's like the biggest test up and try to something working yeah and people start to start diving into it that's right to keep the lake analogy going all right let's let's wrap it up with final call to action for people they want to get started but they like you've you convinced them that this is awesome they should be doing it what do you tell
I mean if you want to get started I think the main way of doing it is going to the duck lake website there's gonna be examples I think we have a couple tutorials as well otherwise the clankers are also pretty good these days and it's really easy it's literally like one line and you have a data lake working in one line you can also have a duck lake with quack working to me it's always like a bit mind blowing when like just check the website that's the best way if you find any issues it's always super helpful to us if you open a niche or a github especially with a very reproducible steps like a full on scripts on how to generate the data what exactly was the problem as it's easier as it is for us to reproduce it's easier for us to fix as well and we got people covered like we go through issues we have a way of going through all of them and prioritizing them within our team as yeah I mean I agree with Pedro I think the most important thing is that you try it out for
yourself right and just you whether it works or not for you I think the duck lake can work in multiple different ways right like you can see it as like your pet project to store all of your data which is very a very comfortable way of running it or it can be even the backbone of your software service right which is also where many companies are betting on and so yeah just try it out and let us know what you think right we will be there in the public issue truckers so yeah and uh with that DB 2.0 coming things are also going to get much faster that is true that's like so sorry we now have a acin kio implementers and that will make a huge difference for reading s3 files because you're basically in clog your processing uh so yeah I mean I think this is one of the most exciting times to be a diving into duck lake awesome now finally just you there's all these different ways you can run it with postgres with SQLite with duck DB in process or quack do you recommend people starting one of these I think with duck DB yeah duck DB in process right yeah
like the beam process that's the easiest okay because with postgres is to have to set up the server because that's of course a little bit apart with duck DB is usually one line and you're ready to go yeah it's got all the extensions and it just knows right exactly all right you guys thanks thanks for being on the show and uh sharing this cool project you're working on oh thank you so much for the invites yeah thanks Michael bye this has been another episode of talk path into me thank you to our sponsors be sure to check out what they're offering it really helps support the show thanks again to six feet up the python and ai experts you call for the hardest software problems from scaling applications to simplifying data complexity and unlocking ai outcomes they help you move forward faster see what's possible with six feet up visit talk python.fm slash six feet up and it's brought to you by us talk by thon and python bites both now have mcp servers point your ai at 10 plus years of python episodes transcripts and show notes free like mcp in the nav at talk
by thon.fm and at python bites if you or your team needs to learn python we have over 270 hours of beginner and advanced courses on topics ranging from complete beginners to async code flask jingo htmx and even ellipse best of all there's no subscription in sight browse the catalog at talkbython.fm and if you're not already subscribed to the show on your favorite podcast player what are you waiting for just search for python in your podcast player we should be right at the top if you enjoy that geeky rap song you can download the full track the link is actually in your podcast blur share notes this is your host michael kennedy thank you so much for listening i really appreciate it i'll see you next time
talk by thon and me i think is the norm
More episodes
More from Talk Python To Me

#561: TonIO, a Multi-threaded Async Runtime for Python
Talk Python To Me

#560: Building a Research OS: From Django to 30,000 Samples
Talk Python To Me

#559: 12 Things You Should (and Shouldn't) Do in AWS
Talk Python To Me

#558: Hyper-Personal Software with Python
Talk Python To Me