
IPB209: SREs Are Breaking IPv6 — And They Don’t Know It Yet
Get every episode summarized
Each time The Everything Feed - All Packet Pushers Pods publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
About this episode
“Tom Caffeine is out on assignment, so you're stuck with all of us today. We are doing videos, so if you want to catch us on YouTube, you can see what we look like. I don't know if that's helpful or not, but hey, it's probably not helpful.”From the transcript
Get every episode summarized
Each time The Everything Feed - All Packet Pushers Pods publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
Transcript ready
474 searchable segments. Every word is indexed and playable.
Full transcript
The Everything Feed - All Packet Pushers Pods — IPB209: SREs Are Breaking IPv6 — And They Don’t Know It Yet. Machine-transcribed; use the interactive transcript above to jump the player to any line.
Welcome to the IPv6 Buzz, where we dare to dive into the 120-bit address space wormhole. We're glad you joined us today. I'm at Hoily, got Nick Boraleo on. Tom Caffeine is out on assignment, so you're stuck with all of us today. This is my first video. Yeah. Yeah. We are doing videos, so if you want to catch us on YouTube, you can see what we look like. I don't know if that's helpful or not, but hey, it's probably not helpful. They're going to put us right back to audio only after this one. So we had a couple things that we were mulling over, and one of the things that sort of came up was really talking about sort of getting across the last bit of the finish line with IPv6. And the fact that there's a gap that exists in terms of getting to sort of V6 only, even
if you've got your network entirely ready, you've got to IPv6 mostly. You've got V6 running in your data center. You've got all these related components and network and firewall and everything else all dialed in. But that last little bit that you have to crawl across is probably going to be pretty painful because you're going to be dealing with teams that you probably don't have ownership over, any control over, and you're hoping that they're just going to come along for the ride and do the right thing. But that's not going to be the case. And we're going to talk about SRRs and developers, and what their roles are in helping to get V6 across the finish line. And maybe some of the hurdles that you're going to face when talking to them. And they've been sort of absent in both the ITF in terms of contributions and sort of the V6 ops groups and any of the related sort of working groups around what needs to happen specifically for IPv6.
And as a result, I think you're going to find that you're going to see a bunch of technical that related to trying to get these V6 up and operating for services that they may be responsible for maintaining, building, and having operational ownership for. So we thought we'd kick it off of that as a topic. So I guess a slatious thing is maybe the SRE and your developers are sort of breaking your IPv6, but they just don't know it yet. Yeah, I mean, I think that there's probably going to be two diverse paths here, at least in my experience. There's going to be the developers that are just kind of relying on the underlying operating system. They don't care about networking. They're using DNS names for things. So the libraries that they're developing against support IPv6 will just work in IPv6 only. And then you have the other side of that coin, which is we're developing network applications
and we're only considering IPv4 because that's all we've ever done before. And so even in a dual-stacked environment, again, we've talked about this many times how it covers up any type of problems by falling back to the legacy protocol. And once you take that safety net away, like, when none of my stuff is working, well, you don't have that protocol anymore. And so there's the whole process of educating how to essentially backport all of what they've done to support a new backport support forward looking protocol. Two steps forward, three steps back. Something. Yeah, I think one of the challenges becomes that while the networking stack may be available and assuming that you're starting a dual stack and then moving to V6 only data center, that stack may have been available for a long time for them to be able to test and validate
that their V6 is actually working. They probably never took the step to do that. Number one, right? So that's the first hurdle that you have to overcome of saying, like, hey, it's available on the network. You need to test to make sure your application actually functions on V6 only. And so we're going to provide you with V6 only, you know, VLAND and migrate your services over to and confirm that you can function properly. That's a lot of extra work for team, right? And so you're asking them to basically, you know, do a whole set of proof and validation and testing that maybe they never had to do before. So you're going to get resistance on that front. Assuming that they're okay with that and they go ahead and set it up. Let's say they actually run into a problem. They discover that they have a V4 dependency of some type within their code, within their library. Then you're going to deal with the entire how long does it take to actually change code in production to fix that related problem? And are they willing to fix that related problem based off of just the fact that it can't work with IPv6? So there's just sort of two distinct issues.
So maybe we talk through that. Oh, quick. Yeah. I mean, I would even add a third in there as well, which is not necessarily an issue, but it's a, you know, it's a thought process of if I just cannot get, if we as an organization cannot either move away from this thing that requires this legacy support or there's no time to do it for, you know, maybe, maybe your development team has a very structured, you know, process they work through it, you know, in their development pipeline, how do we make this work? And so then there's the, you know, the third pillar of can I do SLB against it? Do I put it in an enclave? Do I, how do I deal with this thing that I can't modernize right now? Right. And I think that that is the, that is the sort of thing that caught me off guard because, you know, when we got so far into the process of removing IPV4, we hadn't run into it.
Like, I don't know how we went this long without really running into this problem at scale. But like, it's like, okay, well, now we need, now we need to well defined, well tested answers. And we have most of that already just doesn't really, it wasn't really put together anywhere. And so we had to write up those things. But yeah, I think that's, that's the third pillar of that, of that. Yeah. So I think, I think if you're talking through enterprise V6 appointments and you've got, you know, tiers of applications, what, what becomes complex is the interdependency between the tiers of applications that are considered, you know, critical to the business. And are those the ones you want to tackle first? It depends on how your organization structured how quickly they want to move. And the risks that are associated. And like you said, Nick, if you don't have a well structured plan to be able to walk someone through and define what those requirements that are actually, you know, actually going
to be, then you, you run into the hurdle of, well, that's not a business requirement. Or that's not code or functional business. So we're not going to do something that puts a risk in those particular areas. So that becomes like the big hurdle you have to get over. Yeah. I mean, it's, it's, I like we were discussing before like this, like when you get to that point, maybe think, like learn from, learn from the pain, right? Don't waste. Don't wait. Don't have this be a surprise. Like try to address these things earlier than maybe you think you need to because they, you know, when they come at the last minute, they're difficult, right? And they're, and they're, they're not small problems either. They're, you know, the application and the SRE layer is, is always been described at least by me as like, this is going to be the hardest part. Right? Even with the, you know, sort of scientific instruments that I have to deal with that are just,
you know, some of them, they're, are literally built into buildings and stuff. Like, okay, those, those would put those over there, right? They're hard to upgrade. But like the application pieces are the, you know, they are the hardest part. Yeah. I would agree with you. And I think, you know, there's, there's some current interesting sort of, you know, draft submissions to the ITF. So Frank Martin submitted a deploying IPv6 data center, one that really sort of targets SREs and software engineers, not necessarily network engineers. And it's really focused on how do we cover building, you know, sort of the right set of frameworks and the right set of components to make it easier for SRE or dev team to figure out how to make use of IPv6 and move to IPv6 only and that they're not risking, you know, object failure, their applications, we buy, buy, by taking it on. And I think that this is sort of the right approach of saying like, hey, there's some standards
or there's some pro standards out there about how to go about doing this process. And that the framework is, you know, is useful because part of the challenge is if you haven't driven that road or that path before, and you don't have technical experience, the likelihood that you want to jump off and go do that stuff is, is, is, you know, you're going to have a higher barrier of, of entry of someone saying like, yeah, let's drop everything and go tackle that to go solve those really problems. The good news is I think for many of them, if they're following sort of standard, you know, development practices and hopefully some standards within your enterprise organization, maybe you're all on Kubernetes or you're all on some given sets of frameworks, maybe you're on OpenShift or something of that nature, right? It becomes a little bit easier that wants those platforms have the right support within them. And once those problems have been solved, it's easier for the developers to be like, well, if that team did it, we should be able to do it too, right? And so you get a little bit of the win together approach in terms of getting across there.
You really need to find that one or two sets of individuals who are advocates who are interested, who know how to tackle these related problems and could go out and prove to everyone like, yeah, this is doable. There's no reason to be upset about it. There's no reason I think this is going to throw your schedule off. And so therefore, you know, you really need to go build bridges with these folks in order to actually get things moving. You don't have to be your best friend. They don't even have, you don't even have to like each other necessarily, but you do need to build bridges around, you know, providing them to support the technical resources they require, providing them, you know, we, you know, I guess technical resources within your team to be able to go solve and work on the problem and work with them, shoulder to shoulder, just sort of fix it. And to help them figure out exactly what's going on because many of them aren't going to be as familiar with you as yourself with doing, you know, wireshark packet captures and doing teapogs and figuring out exactly what's going on on the wire and why their applications
may be behaving the way they are. They really get a looking at code and maybe seeing like how the sockets connecting and what's happening there. But even now with the instructions of these libraries, you won't even see that. They're just going to get a written out air code that says, you know, something and you're hopefully going to be able to, you know, figure that out based off of what they're doing there. But I would say you really have to reach out and work with those teams specifically and take a small area that you can get a win in that isn't necessarily your first prime time production, most critical service. Yeah. I mean, I think, you know, in previous episodes, we've talked about how, you know, it's important to involve the security team like right out of the gate. And I think that one mistake that I've probably made was that I did that, but I did not involve the development team early enough. We brought them in early, but not early enough. And I think that that's a, you know, that's something that I think people should consider, right? Because not just the developers, right, but also the folks that are running the platforms,
the development teams working on, especially them to be honest like that. The containerization pieces, depending on what you're running and what versions of it are, you're running like they don't, they didn't always have the greatest. I could be six support and it definitely wasn't out of the box, right? It wasn't like default on kind of thing. And so, you know, if you're running those kind of platforms that, you know, maybe they're difficult to upgrade for whatever reason, maybe you can't get, you know, it's hard to get maintenance windows or whatever, right? Like it's, it's a chore to do those things. Get those people involved right, right away so that they can start looking at what is and is not supported and start looking at, you know, ways that they can implement the things that you're going to need to then enable the developers, right? So it's a multi-step process, but, you know, those, those are the things that are going to be really, really important, I think.
Yeah. And that was a mistake that I, that I made. I didn't do that. And as I, I didn't do it soon enough. Yeah. I think that's fair. If, you know, that's a self critique I've done the same thing of not, not having those folks involved early enough for them to understand what, what, what is or not being deployed. And I think the other challenge is you need to give them tooling sets to have an understanding of what you're seeing within your environment. So having the monitoring, having the logging, having the capability to show them where you know, both V6 and V4 failures are occurring and how their applications are behaving based off of, you know, that sort of data. So you can walk in with a data driven mindset as opposed to a, we just want to move. And, and, you know, this is, you know, this is something that we think you should do. Sort of conversation is say like, no, we're actually seeing application failures. And the reason why is because we're in a dual stack configuration. We see the following V6 failures that related to your application.
You know, can we, can we talk about why they're failing on the V6 side and how can we get them to be successful and why we're seeing the certain behaviors of V4. And being able to have that conversation from a data driven, this is what's showing up in event logs and we'd like to minimize, you know, failure rates is a much easier conversation to have than we just want to adopt V6 for V6 sake, right? It's almost like communication is important and that all these disciplines are somehow related. Well, and I think this is true too. I think when an application team comes to a network team, it says like, I'm seeing the following errors or failures from my application, you know, happening across your network. Are you seeing something similar? You know, almost every network engineer is going to be like, yeah, let's dig into that. Let's figure out what's going on. But if they come back and say like, I just need you to turn V6 off because we don't like it. That's probably, you're probably going to get a much more negative response from a network team of like, whoa, you know, that's not something we do, right? And kind of if there's a specific business reason or there's a specific need or if it
actually causes a failure, like the times that I've had conversations where we've been like, yeah, we need to turn V6 off is when we've actually had older devices that actually when they saw V6 packets actually fail, their OS couldn't handle the V6 packet and actually cause critical failures in which case we're like, yeah, absolutely. We should be turning V6 off. We think you should be looking or investing to try and replace that piece of equipment eventually. But we understand that that's a critical problem and something that can't happen because this is that particular software application or piece of hardware is critical to the business. We can't continue to do that. And eventually, you have to come to terms of working with your counterpoints around developers and SREs around, you know, hey, these are the standards that we're going to be supporting with in our organization and IPv6 is one of them. And so you need to have your application supporting that. Going in and demanding that that happens overnight and that they put on pause all the
related projects and requirements that they have from their business side is probably not going to be a win. And so I caution everyone that don't go in just demanding that they get working on your set of V6 only capabilities right off the bat. I don't know. They keep you having thoughts there. I mean, I have found that just demanding things in general will elicit a response that you do not want, right? Come in with, you know, don't come in with demands. Maybe don't even come in with solutions. Just come in with use cases. And then ask instead of tell that's my approach. Yeah, I would agree. I tend to agree with you on that one. I mean, there's times when you have to be demanding around certain things because the business needs it. And which case you've got the business to back you up around that demand need requirement. And so it becomes easier conversation. But I agree with you as a general author, I'm coming in with a open mind and coming in
with the use cases and advocating and saying these use cases make sense. I think I saw far more, far better position to come into the conversation with. So anyway, why don't we switch it up? Why don't we switch hats? Why should enterprise admin or SRE or developer care when we walk into room? What's the flip side when we're sitting in the other side of the seat? I mean, the, I mean, I'm not a developer, right? So I can't really answer that. But you can play one on this podcast. And that's how you're telling me once more, I write some crummy go code. I mean, I can, I can tell you from the data center management perspective that the platforms are, I have been told are significantly easier to manage when they're running single stack IPv6. Once you get it to that point because you don't have the spider web of address translation that you get within the Linux bridge or within your EBPF structure within the systems.
And then probably also at the, at the border of whatever it is that you're, you know, that you're building security perimeter. Security perimeter, your PEP, your policy enforcement point. You know, so you basically give everything, everything gets a unique address. And so you can figure out what pod it's in. You can figure out, you know, what the application was maybe, but it's also, you know, much easier because you don't have to protocols. I mean, the same reason the US government wanted to move away from dual stack, right? It's, it's two, you know, data planes that you have to have all things for. And so management, I've been told, and I mean, I've managed network equipment for a long time. That's essentially single stacked. And it is, it's easier. Like it's easier when you don't have both. And it's especially easier when one of them is V6 because everything is unique. Yep. And so there's that. But I mean, I think that from the developer's perspective, you know, you can have, you
know, unless you're writing network applications, my stance has always been, I'm going to rely on the operating system to, you know, I don't care. Like I'm going to use, I'm not going to hard code literals. I'm not going to, you know, I'm not going to do those things. And I just use whatever the underlying operating system has. And it makes life a lot easier now. It gets obviously more complicated when you start developing network applications that need to listen on sockets and things like that. But I mean, again, not a developer, but from my point of view, most of that stuff is being handled by libraries at this point. And most of the libraries support V6 or they just rely on the underlying operating system, which just simplifies everything. Right. Yeah. And having that conversation with your peers and being like, how exactly are you leveraging networking within the applications you develop and to figure out where they fit on that
spectrum that you just identified. Right. Are you hard core down in the kernel actually developing networking related components and stacks in which case you have to be way more fluent in both V4 and V6 in order to be able to do that along with all, obviously, all the operating components of how those services and independent functions work. And you may even need to be familiar with transition technologies. Right. And in order to be able to do that correctly. And then I think the next one is really like if they're relying on everything within either the OS or libraries, then V6 adoption should be pretty transparent to them. And you would hope at that point that they'd be willing to try their application services to figure out exactly where they might have hard code dependencies before literals. Maybe they just don't have a quality record published for a particular service, even though it may even be available on V6. They just didn't know that they needed to put a quality record in there in the V6 address in order to make it available.
And it's just going back and cleaning up your DNS and advertising quality records, right. And things that seem obvious, but maybe we just overlooked as part of the deployment process and or you didn't go over that portion of the checklist again because guess what? Well, it's already working. So it's not something we go and check because that portion is working well, fell back to V4 and the A record work just fine. And so it looks like everything is working exactly as you anticipate. And so those are the sorts of things of going through and building like you mentioned at the beginning, a checklist to work your way through how the applications behave. I think there are many application owners who actually don't build application maps that sort of describe and talk through everything that they anticipate from a flow basis of how the application actually works. And so that's a good lesson for many folks to go through is to actually map that out and have a good understanding of all the related surfaces that touches the dependencies it has, what SOC connections it does.
And that's going to be good information to have when you start moving to zero trust environments anyway, because you're going to need to understand what those dependencies and what applications need to talk to each other. And if you're going to be enforcing zero trust in any of those environments, you're going to need that information anyway. So it's probably a good exercise to go through and do that. Obviously if you're a midsize or a small company, this is probably a lot of extra work beyond beyond what you're already doing. And for many smaller organizations, they may just be relying on SaaS or collaborative related platforms or things of that nature, in which case you are entirely at the mercy of the schedule one time frame of those platforms to support IPv6 in the way that you anticipate or need. So if you have outsourced your ZTNA solution offerings and that platform doesn't have V6 support yet or doesn't have the right dual stack support yet or whatever, you're
going to have to wait till they come up with it. There's just nothing you can do from a time. It doesn't matter what the network team does at that point. You're entirely at the mercy of those providers. And that's going to be true for any of the SaaS related components. So third party sites, website content, anything you're consuming off those platforms if they don't have the six support, it's not like you're going to change it overnight. You can obviously put it through a translator. So you could have control over the capability to put it through a NAT64. But the reality is you can't do that for everyone at home unless you're doing VPN. But if you're doing local internet and a hop off, you have zero control over how that's going to play out. So yeah, I mean, I think there's a couple of things there that are pearls of wisdom. One is you're going to have platforms or software or something. You're going to have something on the software side that doesn't do V6. You need to have a working plan to deal with those for your transition, for your overall
transition plan. What do I do with the things that just don't, they don't have an analog to replace it or we don't have budget to buy it. And we have to have it. Like how do you deal with that? So you need a plan for that. And you need a plan for something else that was groundbreaking that I've already forgotten. Rest assured, it was very important. I had two good points and I forgot all of them. Yeah, that's all good. I think for me, the big challenge of working with SREs and developers is they have a whole set of language and constructs that they work with within their organization that don't match necessarily without network teams operating. And you're going to have to spend a little bit of time and energy to figure out and understand exactly what they're doing with how they talk to each other, how they communicate, how they fix problems. And hopefully, then be able to advocate a little bit about why V6 testing and V6 related
services are going to be important from a scaling and support basis for that organization. So you may be spending a little bit of time on the whiteboard. You may be spending a little bit of time explaining transition technologies. You may be explaining a little bit about where the internet is at today because not everyone knows it. Greater than 50% of the internet is running V6 today. And obviously that's only going to increase over time. And so I think those points are sometimes lost because people still see V6 as they I heard about this 20 years ago phase as opposed to where it's at today. And I think that's part of the challenge that you're going to deal with. I mean, that's even a challenge even in the networking community. It's like, why am I paying attention to this? And this is something I need to know. The other mistake that I've seen that's been made many times for both developers and other teams. This is the assumption that IPv6 behaves the same way as IPv4. And so you're going to have to spend a little bit of time doing education about how the difference in the protocol, what's lack is all about where DHEPV6 is different than DHEP,
just how multicast is used, the lack of broadcast, the lack of network. Network numbering and subnetting is just different. Prefixes are different. And how do you think about the address space is different. And fundamentals about how some of the protocol works for application developers, the flow label may be really important in terms of building entropy in for how their applications are actually being balanced across data center links, things of that nature that are good conversations that you should actually be having with those teams to sort of explain, like, hey, don't signal thread your Oracle SQL query and not put any entropy in. You're going to overload this part of our least spying configuration or whatever needs to happen. And those are pure technical conversations that are worth still having with those teams, and being able to make use of the entropy in the flow label might be useful for those that are actually building applications in that area.
So those are good things to sort of bring up as topics with your counterparts. I don't know. Is there anything else you think we want to cover before? Before we wrap this? No, I still don't remember that other point that I had, but I mean, I think that the, you know, we have a couple of themes, right? Like communication is real important. Yeah. Like being able to talk through the other groups, getting them involved. And I think the human communication of these six communication. Yeah, human communication is important. Interpersonal, which, you know, some IT people don't like doing that. And that's fine, but like, it still has to get done. And, you know, the, the whole like, make this an hour problem kind of thing, like give you, give the folks that are going to do the work, some sort of invested, you know, interest in the successful outcome of whatever it is you're doing is, I think, very understated. Like, we don't talk about that enough.
I think that's really, really important. Well, and in this case, you're going to be a vested interest in their outcome of having success. So it's going to be you contributing in on that. So yes, it's going to take us to the time. It's going to be extra effort on your part to make that happen. So yeah, 100%. Cool. All right. Let's wrap it and say like, hey, we can pretty much say that the last bit of the finish line is really going to be around developers and estuaries getting the applications to work on V6. So once you get that all deployed as a network engineer, you're going to have this last little bit hurdle. And hopefully maybe some of those devices is useful in terms of helping you to get them to actually get the applications working the way you need to. All right. Well, thanks for joining us for this episode of I-PV6 Buzz. If you got feedback or a follow-up on this topic, send us a message at packupusures.net such as FU, where the FU stands for follow-up. We'd love to hear from you and continue the conversation. Also at packupusures.net, you'll find a range of other deep dive technical podcasts for IT Pros.
So definitely go check out what's available there. So so long until next time, we'll see you on the Internet. The IPv6 Internet, that is.
More episodes
More from The Everything Feed - All Packet Pushers Pods

HN844: Multipath Reliable Connection (MRC)
The Everything Feed - All Packet Pushers Pods

TNO075: What Does Network Operations Look Like in 2030? (Sponsored)
The Everything Feed - All Packet Pushers Pods

N4N065: Well Actually 4: Multicast, MPLS, and More
The Everything Feed - All Packet Pushers Pods

NAN132: The AI-Augmented Engineer
The Everything Feed - All Packet Pushers Pods