Skip to content
TrackPodcasts
technologyMar 20, 202612:01

Global Namespace for Precision Medicine: The Data Breakthrough w/ William Baird

Data Unchained

About this episode

Global namespace is becoming a critical strategy for cancer genomics, precision medicine, and life sciences data infrastructure. In this episode of Data Unchained, Molly Presley speaks with William Baird from Guardant Health about managing 160PB of genomic data across HPC, on prem systems, and multi-cloud environments. They break down why 80% of this data is still active, how researchers need faster access without worrying about where data lives, and why global namespace could become a major solution for decentralized data at scale. The conversation also explores vendor flexibility, continuity, standards development, and what it takes to support life sciences workflows when the stakes are incredibly high.


Cyberpunk by jiglr | https://soundcloud.com/jiglrmusic

Music promoted by https://www.free-stock-music.com

Creative Commons Attribution 3.0 Unported License

https://creativecommons.org/licenses/by/3.0/deed.en_US


Hosted on Acast. See acast.com/privacy for more information.

Get every episode summarized

Each time Data Unchained publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.

Email me new episodes

Free for 3 shows. No card needed.

Hosts & guests

Transcript ready

199 searchable segments. Every word is indexed and playable.

Global Namespace for Precision Medicine: The Data Breakthrough w/ William Baird

Data Unchained

0:00
12:01

Full transcript

Data UnchainedGlobal Namespace for Precision Medicine: The Data Breakthrough w/ William Baird. Machine-transcribed; use the interactive transcript above to jump the player to any line.

Hi, I'm Molly Presley, the host of the Data and Chain podcast. And as you can see, we're broadcasting live from a trade show today. We're at Supercomputing 25 in St. Louis. And really excited to be talking this week, not just about high performance, but also about data access and the whole concept of gain access to more data globally. So if you're new to the Data and Chain podcast, let me tell you a little bit about the podcast. We founded this talking about what are some of the challenges of gain access to this decentralized data that's now being created in edge devices, different instruments, different clouds, and maybe needing to be shared with different applications that was originally created for. We're really excited about today's guest because he has driven a thought leadership initiative around some of the challenges in the life sciences community in this exact area. So without further ado, let me introduce William Baird.

William, thank you for joining today. Thank you for having me. So William, before we talk about data and decentralized data, maybe tell us what is your role you're with Garden Health and tell us a little bit about Garden Health as well. Sure. So William Baird, I am the Associate Director for Hyperborn's Computing at Garden Health. We are a liquid biopsy company, and our focus today, and I stress today, is on cancer. So basically, oncology, all its different flavors and forms. We specifically do detection of cancer, and it's various mutations for advanced cancers, or we do monitoring for whether or not how your cancer treatments are progressing, or another one of our products is for early detection. So basically for the screening that we do. So instead of having to go and do more uncomfortable regular checkups, we can actually go off and just do a blood test. Everything to do is based on blood. So it goes through a wet lab process.

It ends up going through a genome sequencer. And then from the genome sequencer, it goes on to our Hyperborn's Computing Environment. And there it produces quite a bit of data. And today, we're now up to curating about 160 petabytes. And all of it is genome. Wow. We have, we either at or past a million patients. And so all of this data, is it actively being used for research? Is it a big archive? Maybe just talk a little bit about how it's being used. It is both used to reproduce the report for the doctor. So the doctor has better results for the patient. It's a precision medicine, so that not all cancers are created equal. Some are more resistant to treatment. Some have required specific treatments. And others are, they change over time. So people need to have a regular checkups and checkouts for those sort of thing. So that is the first step. But when we take that data and use that data for improving our products so

that we can tell doctors better what detected easier, detected more precisely, and so on and so forth. It turns out 80% of our data is actually in play. That we have on-prem day. That's a lot. Well, it's very weird relative to what most people have. Most people have like 5, 10% of their data is actually hot. And the rest of it is cold or cool or at least at most warm. 80% we found in the last six months has been accessed and used in the various analyses that our scientists use. And if I try to push that data off into archive, they pull it back down again. Interesting. So you have personally put a lot of effort and Garden has put a lot of resources into a global namespace initiative. Maybe you could tell us, what is that? Why are you tackling land? Like what's the pain point? And maybe talk a little bit about the initiative itself. So there are kind of three problems our scientists have. One, they need all the data. They need all that data when they need it.

Not when it's available, when you download it from say AWS. Because if something's kept most cheaply in deep glacier, it's going to take 48 hours or more for it to transfer down and into the system. So they don't have a lot of times like I have a request from the FDA. We need to have this data in the next 24 hours. We need this in the next, we have, oh, wait, we found something interesting. And we have a submission that's coming up. We need this other samples that are out there as well. So that was our first pain point. The second pain point is we are diversifying not just on-prem HPC. We're also starting to use the clouds and more than just AWS but others as well. And the scientists moving data management is not what they're great at. They like to produce lots of data, like to hold onto that data, but sorting it, putting it, archiving it, taking care of it is not something they're really good at.

So then they get frustrated as well because they all said, oh, wait, now we have a new environment. We have to rewrite our stuff to where the data actually lives. And so they're very, very stuck by doing the global namespace that we have planned and have been testing and going forward with. We want to have it such that wherever they are, the same data is available everywhere. And it looks the same wherever that particular environment is. So the same file tree will, they get searched in AWS on an instance there as you can on our HPC cluster in the R&D environment, as you can in prod or in often GCP or any other particular cloud, like maybe Azure in the future that we're using. So our intent here is that we make them stop worrying about where the data is and archiving it. Because we'll take that care of it for them.

And are you talking just about the HPC researchers, or is this other types of users of data as well? So today we're working with the HPC users. One of the plans is to share our data with other entities in a curated way so that they can take advantage and use it for their own research and whatnot. As a discussion this morning, data is currency. We happen to have one of the largest, if not the largest repository of cancer genomics data in the world. So very valuable for both human health as well as financially. Exactly. And for being able to find new ways of figuring out health issues, it may be, for example, some of our patients have other conditions as well. And if we can tag, we know that and we can tag that, then other types of research could benefit from this as well. So that's excellent thought leadership as far as traditional HPC architectures

that were designed a little bit differently, but made some of these other challenges difficult. And so adding in this global namespace solves a lot of that. That's great thought leadership. I know you're working to drive this into some standards to make it a little bit more accessible too. Yes, so one of the issues is that today we're producing well over a petabyte a week. Wow. This means that compounded with the problem that this is all hot and then made worse by the fact I have to retain it for 30 years. So in the near future, it's a high potentiality. I'm going to be talking about 21 exabytes of data. Well, that's great and a challenge that we all love. But on the other hand, what if one of my vendors decides they're not going to do this product? I've required, I need to do this anymore. What if I need to have a second replica on a different vendor for in order to make back-end vendor in order to make sure that if there's an outage, like say, what certain cloud providers have had recently?

All very real situations that I am not impacting my cancer patients. Because in many cases, they are very much, this is a life or death for the advanced oncology patients. Yeah, they don't want to wait and they certainly don't want to take it down. It's not a case of wanting to wait. They can't wait. They're literally, in some cases, days away from passing. We don't get our data, the information to them. That's a burden, it's a carry. It's a burden, it's honor. Yeah, exactly. So you had an announcement this week that you're broadening your reach a bit. Tell us just a little bit about your announcement. The intent with this whole announcement and this all this work we've been doing is to build a set of standards. So that if I, with the vendors I have today, we use GPS with Seagate in order to do Seagate Live in order to do this namespace. If we don't, if Seagate decides they don't want to do object storage anymore, or IBM, because everything is very sensitive to price,

if their pricing structure changes, that we can then say, okay, we have other alternatives in the future. And so this set of standards we've been doing is to allow us to ingest the data written by one and be useful in maybe not as quickly, or using all the secret sauce that another vendor has, allow us to ingest it and allow it to run and do continuity of business. So you have a pretty good path at that point. Very good path. And one of the things I want to highlight is this group we've been doing. There's about 34, if I recall, entities that participated with other national labs like Livermore, or there's the various founding members from the storage side, such as, as I mentioned IBM, but also DDN, Weka, and HammerSpace. HammerSpace has been highly helpful in participating and pointing out some of the things that have to be tackled and all of this. Excellent. Well, William, I know you have a busy week.

The show has everybody back to back. Really appreciate you stopping by to take time with us. If people want to learn more about this initiative, do we point them to the Garden Health website? Or is there somewhere else to point them? We do have a GitHub. And you can find the single namespace that we have the presentations from, and the minutes from all our meetings. We're very open about what we're doing. And we are transitioning this over to Oasis, the standards body. And that standards body will be meeting in January. Their call for participation should be going out momentarily. And it should be very much a case that you guys, anyone who's own Oasis member can help shape the standard. What we produce so far is a document that's meant to be a C. Because a lot of the different folks got into the room and had the food fight already to sort out what's what. This allows us to save time to actually moving into a proper standard. OK, so great way for you all to go out and do some of your own research about the initiative. And then I think, William, you have a meeting every three or four months that the approximately the cadence today is the capstone meeting.

OK, great. So we're going to, that particular working group won't move me moving forward. But today is the 18th at 2 p.m. at the National Blues Museum. Anyone who wants to come and participate is more than welcome to. Excellent. Thank you so much for joining the show. Thank you for having me. Thanks for listening to Data Unchained, powered by HammerSpace. To learn more, visit HammerSpace.com. If you have a guest you would like to hear on the show, email me at Molly at HammerSpace.com.

More episodes

More from Data Unchained

View all episodes →