Skip to content
TrackPodcasts
technologyOct 2, 20261:07:24

HN844: Multipath Reliable Connection (MRC)

Get every episode summarized

Each time The Everything Feed - All Packet Pushers Pods publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.

Email me new episodes

Free for 3 shows. No card needed.

About this episode

“I am Ethan Banks with Drew, Conray Murray. You can find us on LinkedIn or the Packet Pusher's Community Slack Group and we would love to connect with you.”From the transcript
Today’s show puts the “heavy” in Heavy Networking with a discussion about Multipath Reliable Connection (MRC). MRC is about getting an even distribution of RDMA traffic across equal cost multipath links in massive data centers, and doing so over plain old lossy Ethernet. Ethan, Drew, and guests discuss how packet spraying, resilience to link and... Read more »

Hosts & guests

Transcript ready

294 searchable segments. Every word is indexed and playable.

HN844: Multipath Reliable Connection (MRC)

The Everything Feed - All Packet Pushers Pods

0:00
1:07:24

Full transcript

The Everything Feed - All Packet Pushers Pods — HN844: Multipath Reliable Connection (MRC). Machine-transcribed; use the interactive transcript above to jump the player to any line.

Welcome to Heavy Networking, the flagship podcast from the Packet Pusher's Network of Fine Technical Content to make you better at your job in IT. I am Ethan Banks with Drew, Conray Murray. You can find us on LinkedIn or the Packet Pusher's Community Slack Group and we would love to connect with you. Bonser Itentials Flow AI delivers agentic operations for infrastructure, meaning easily build AI agents that actually work the way engineers need them to, governed, deterministic and built for production. Add intelligent automations to your network operations without the usual AI chaos, find out more at itential.com slash flow AI, that is itential.com slash flow AI. On today's episode, we're putting the heavy and heavy networking today with a discussion about multi-path reliable connection. MRC is about getting an even distribution of RDMA traffic across equal cost multi-path links and massive data centers and doing so over plain old Lossy Ethernet. And we're going to get into the use cases shortly, but you already know why we're doing this. You know, our guest today are RIP so hand Distinguished Engineer, Eric Davis Master Engineer and Eric Spada,

Technical Director and Distinguished Engineer. All three of these gentlemen, happy to work for Broadcom, but this is not a sponsored episode. We're just here to talk about the relatively new multi-path reliable connection protocol. Well, RIP, I think I'm going to send the first question off to you. This is going to be a brand new topic for a lot of folks. So just give us the 10,000 foot overview in just a sentence or two. What is multi-path reliable connection? So multi-path reliable connection is a specialized transport for AI workloads. And by that, I mean, it's an open extension to the Rocky V2 protocol that exists today. We add a package spring, we add resilience to Lincoln fabric failures, we add congestion control while preserving the application interface so that customers can get the benefits of a transfer that is engineered for doing well in AI with AI workloads, but don't have to change their applications or their network stacks.

So where did MRC come from? What's the history because there are other efforts going on to address these kinds of issues? So the MRC consortium was a collaboration with Microsoft OpenAI Intel AMD and Nvidia. So it was primarily driven initially by Microsoft and by OpenAI. They wanted to get AI as we know moving very quickly. So they wanted to get something as quickly as they could that solved some of the key problems that RIP distressed primarily the package spring and multi-pathing was a key. Enableer for them to build some of the scale networks that they're talking about. And they wanted to do it in a rapid way as quickly as possible. So they brought together some of the key Nick vendors to kind of collaborate together to come up with a minimal viable product to meet their primarily training or cloud needs.

So that's kind of the history has been there's been now has been great is being deployed in actual networks right now. It's not a theoretical program. So that's kind of the history where it came from. Well, let's explain the problem in more detail because you talk about package spring and so on. I assume we're talking about equal cost multi-path and making even distribution at QCMP has been a challenge. I assume that's the big one we're getting at here is using all those links and some kind of if you look at the way that if you look at the way that Rocky is typically deployed today. There's a single path from source of destination files and it can use ECMP but ECMP typically will provide reordering behavior in the network unless you nail up the flows to a single path. So what package spring allows you to do is to utilize with it within a single message multiple paths network efficiently. So if you look at it from a from a protocol perspective message comes in it's broken up in the packets very similar to where Rocky is on today.

And those packets are delivered across multiple paths in the network and the out of order data placement feature that we added as part of and Marcee allows you to recover that reordering at the other side. So really what it allows you to do is to kind of spray the marbles across a whole bunch of paths to get a more uniform distribution of the traffic across the address and network. And generally that's driven by the requirements of these AI training workloads where you're trying to minimize things like packet drops packet retransmission. If you look at the training workloads their workloads are very typically very versi and which they want very high bandwidth from node to node as quickly as possible. So it's a really quick start up time. So you know so it's unlike a standard Ethernet application or an internet application where you have kind of the summation of large flows can get you to good performance in this case.

There's usually a very small number of very high bandwidth flows and being able to treat them well as an aggregate is a key. Different share to be able to get the flow completion time of the collective or wherever operation they're trying to achieve is is is critical. Yeah, yeah. One thing that I would add though is is with RC QPs that what we call like legacy rocky today. When you have a large collective going on with a bunch of you know a bunch of peers till latency is a killer rocky right here the application is only give you as fast as that that's lowest. That's low as flow right so MRC with the spring of the network across the network. And you just get a huge advantage for performance. Okay, so we have to talk about a few different things here we got with a tail latency in the second but before that the issue of parallel links and very high bandwidth in these data centers we are talking about the problem of trying to move data into GPU that is part of a cluster doing a massive mathematical operation.

In the issue of data needing to arrive at all the GPUs all the data needing to write that all the GPUs appropriately so that the calculations can continue if there is delay you are idling all of the GPUs that are in that cluster until everyone's got the data that they need to proceed do I understand that right. That's absolutely correct correct the P the P 99.5 telly and see is a key determination of what the application system is actually really if you look at the collective there's an all all or even other ones that are going on. By definition you need all of the complete in order to to and for for the operation at the network level took at the application loads of complete. Yeah, yeah. So tail latency then is we're waiting for the slowest member of the cluster to complete operations and so that and that slows everybody else down that's the tail we're talking about. Yep and the tail latency will determine the sort of the job state machine right so everybody sort of waiting until the tail is done and then they can move forward.

The other thing that we should point out is that it's not just about tail latency it's also about when the collective is going on is to make effective uses of your links without overwhelming anyone given receiver. So in order to make these applications simple and scalable there is no global schedule right so when a collective starts up given knows knows how many how many other nodes it needs to send data to and it just sends them it's not talking to the other notes now think about it from the point of view of the receiver. So receiver sitting around in all of a sudden 10x the bandwidth shows up you're going to get drops you're going to get delays in the network and in MRC a large part of what we did was to just congestion control element that allows us to do to basically back off receivers so that everybody is able to communicate without requiring a lot of application level coordination and scheduling. Okay and in addition to that the retransmission of a drop that says all handled in hardware without the application involvement at all so it just sends the request in and and really as good network and folks we try to make sure the right thing happens in this case is to make sure that anything that was dropped to be transmitted and that we can control the flows from from end to end.

So the major use for this today is is is AI we talked about that that is an imagining AI that is doing the training of models so very large data sets that are being used to train these models and it takes the jobs can take anywhere from days weeks or even months depending on the size of the data set right. That's correct. The critical thing here no idle GPUs that are just wasting time takes longer to train the models longer for that model to come to market of burn and up power you're wasting time so this has MRC is instrumental in making sure that we are optimizing AI model training is that that's all true. The other point to make is that there's been a lot of work done on allowing the job to continue operation if things happen in the fabric that are inefficient so in particular if a link goes down if a node goes down if a link flaps for example we put a lot of effort into MRC to make sure that at a node level the nick that sending out these these packets can detect.

These sorts of events can start to exclude those paths and then has has a way of getting getting this this information back when the path comes up so that it can begin to include it. The big thing that that sort of we have been getting from our primary customers is that they no longer have to think about the reliability of the job at a minute to minute level which was a big problem for them before right so it's it's a self healing fabric is is the. Yeah I think that's really important like what we're saying is one thing to take away from MRC is the application can always be making forward progress it never blocks it never stalls because of a single connection on a single path it will always continue to make progress. We're going to talk about this in detail we're going to nail all this a few more intro questions just to help everybody getting that getting that okay okay because you guys are getting ahead of us. So one another qualifying question here we're talking about RDMA operations design to stand it so is MRC useful for anything other than transporting RDMA right operations specifically across the best effort ethernet fabric.

So MRC was designed to be kind of a minimal super minimum subset of rocky and the big thing to to make that efficient the only two operations that are supported are right and right with the yet and this aligns with how many of the training stacks are implemented today if you look at Nickola Riggle or any of the other stars CCL libraries as we'd like to call them almost all of them are heavily dominated by right and right with immediate a lot of protocols. So we're literally using we're literally using MRC as a transport to carry memory operations and writing data into memory across the network fabric that's what's happening here just like rocky yeah yeah yeah and and it's also important to note that you know just like rocky there's no seepian involvement right so so so the hardware is doing all of the memory placement it the CPU only knows when a right immediate ends up on the other end. Otherwise it's completely one side dead that the recipient isn't involved in the in the loop at all and of course the reason for this is we want to try and save CPU cycles when we do these operations.

So you said MRC is a subset of rocky and just again for base lining rocky is RDMA over converged ethernet so is MRC like a rocky replacement and am I replacing anything else in my network fabric with this. So I wouldn't say so much or a replacement as it's augmenting right rocky so think of it as sitting aside rocky next that implement MRC will probably also implement RC right so does the standard ethernet capability or sorry standard rocky capability. And so it's augmenting that the other aspects that it augment is that standard rocky runs over pfc we we do best effort ethernet or you know sometimes it's called loss ethernet so so what that allows us to do is it allows us to not have had a line blocking and allows scaling to happen a little bit more. We have a bunch of extended capabilities in MRC that provide better performance for example we have packet trimming and we have the ability to do probing of the network and congestion control and all of this is completely compatible with all the existing traffic that's running in your network.

You may need to configure it in certain ways that they don't they don't interfere with each other or you may not depending on what your application in use case looks like but think a bit more as extending what's there rather than ending up replacing things. If I'm doing a ground up design and I'm considering in the independent I could say not going to use ethernet instead and leverage MRC to give me a kind of a result I would have gotten within a band said. I mean I think if you look in the industry there's really only one incentive and number vendor there are several rocky vendors that have a decent to rocky so I think from a least from our perspective if you look at the industry you know ethernet is becoming obviously even more ubiquitous and we wanted to bring these type of capabilities to kind of an open ethernet ecosystem which has been proven to be very resilient and and robust over you know. So several decades. Well how does MRC relate to all three ethernet then we've talked to the ultra ethernet consortium guys rep you're in on that conversation I think that we have that with J mats yes I'll start and then you know Eric and Eric can can can kind of chime in so I look at MRC as maybe I'll treat that right so it's a limited subset of all three to net for a work loads of today uses a lot of the design principles and even implementation principles.

Our very similar we borrow a lot if you look at the spec it it basically will reference the ultra ethernet spec all over the place and there are multiple reasons for that one was the same sort of companies that was contributing to both efforts the other was we were looking at MRC as a as an initial design point to get our implementations out there and then eventually make them complete ultra ethernet so think of it as bridging the gap between the old world which is standard RC to the absolute newer. The absolute newer which will be ultra ethernet MRC sits in the in the middle and is sort of like you know holding hands across both of these these different design points right. I know I so reply you know that so I mean so you know when the when the group was built again almost all everyone was in the this group that was in UEC we'd already spent an enormous amount of effort I can just control on the sacking infrastructure. The out of water placed in all all the and trimming all the basal features and we just as a group decided that you know we've already agreed to all the stuff so we don't have to you know reintroduce different concepts we already do it we know how it works there's a basic set of simulation infrastructure in place particular around the congestion control and sacking so we're just going to to borrow this and really most of MRC is still is at the semantic layer which is the same.

So it's a deal with how do you deal with the different protocols and the multi passing and adding those different features so it really is complimentary to you see from our from from from a bit from a picture perspective. Yeah I'll add on to that you know UEC is a is a complete clean slate protocol from the ground up I mean it has some very advanced features beyond AI does a lot of HPC workload stuff. And MRC you know if you're a vendor and you do it or you have an R neck today it's going to be a heavy lift to evolve into ultra ethernet well you look at MRC what we did is we iterated on our DMA you know our CQP today so it's much easier to jump to MRC and get a Bible product out in the market versus jumping to UEC so so yeah MRC is a stopgap solution that's the way I look at it. So you you think ultra ethernet is going to have all the features of MRC at some point and eventually it is it is right now it's like I see look at 1.0 it has all the features that are in if you look at the AI based profile which is defined in in UEC it is basically a superset of MRC.

Now as Eric said it's it's implementation is very different but if you look at it from a feature set perspective from a customer perspective it it is MRC plus a couple of other features so it's a complete superset go ahead. I was probably I did this out but you know MRC does have a bunch of stuff that isn't in in that ultra ethernet right so that I guess the question guys is does is MRC replaced by UEC at some point by I'm sorry by UET at some point or does that mean well MRC have a long run here. Well it's a good question right so what we expect to see is that over time ultra ethernet will replace MRC and the reason that will happen is for for various well there are various reasons one is there's an ecosystem around open ultra ethernet that will form right and that that gives you that economy of scale etc but also the workloads are changing right over time and like you're hearing now about inference and inference is very important.

Different in terms of traffic patterns and what will end up happening is either you have to object either you continue to develop MRC not so sure we want to do that or you move to an ultra ethernet ecosystem that gives you all the capabilities that MRC gives you plus more but has but MRC has played its its role in being the initial point of think of it as the canary in the coal mine where people have deployed it they have got good feedback they have. So many changes that have been fed back into the ultra ethernet design point and just making everything better so MRC will have done its job from the perspective of enabling ultra ethernet long term. Okay let's jump into how MRC works given to some of the nuts and bolts of it for folks that are familiar with with with TCP folks that are familiar with the ECMP and they're used to load bouncing they're used to how TCP does it's acknowledgments and knows when traffic is lost and how to react to that and so on some of these concepts are going to translate over so if you're a network and you never heard of MRC this is different what's going on but it's going to be very similar to other things that you're you're familiar with.

So let's let's talk about how ECMP traditionally load balances traffic across members that are in a link bundle and then use that as a as a foundation to describe how MRC would perform that same task. Yeah so let's talk about ECMP right so right now in the TCP world how do I load balance my flows across the network what I do is I set up my flows and I make sure that the UDP source port of these flows is not the same right so from again so from a given source to given destination in terms of layer three I am making sure that even though the source and destination IP address is the same I'm perturbing the UDP source port and by perturbing the UDP source port I'm able to spread my flows across networks now what what we're doing today in TCP is that for a given flow that UDP source port is exactly the same for all the packets for that flow MRC takes at a step further it makes that a per packet decision so same flow right flow number one but each each each packet has a has a different UDP source port number the existing ECMP hashing still works as as intended but what happens is rather than on a flow level on a packet level things get spread all over the network.

Okay so so we're dealing with the problem here of hashing a flow based on source port source and destination for some combination how we're doing that hashing with the switch and if that information is the same the entire flow is going to get has to the same link member which in our world here of training AI data center models are training AI models in our AI data center. These are huge flows we can't want an entire flow to be hash to one link you're going to bury that one link and have uneven load distribution across the other links so we're changing the game okay we're going to has the same way the same same process that we're familiar with here we're going to change the source port for every packet so now in theory every packet is going to be hash to a different member of that ECMP bundle and we're going to end up with much more even load distribution yes that's correct and as a byproduct of that will happen. This is at the receiving end you have a little bit more logic so that the receiving Nick is able to work out for every packet that comes in off of the wire which flow does this belong to what is the state of the packets that were preceding this packet should I be giving a notification should I be telling the remote and that I I I I there's a packet in the middle that's lost so we move some of that that that complexity into the Nick endpoint we keep the switch exactly the same and we provide better performance overall.

Which we'd have to bake all that additional logic in because if we didn't know any better if there was no MRC logic all of a sudden we've got a bunch of different flows is how it would appear to a traditional network all got different source port and different on all these different packets these are all different flows but in fact the members of the same flow and so. MRC on the receiving end has to know how to reconstruct that flow put it all back together know that everybody is a member of the same of the same flow okay so what let's talk about that what's actually happening then with how are we how are we identifying that all these packets that we change the source port value on our in fact members of the same flow. Okay so the way we're doing that is basically we're using an upper layer transport header so think of it we have from the perspective of we have a connection ID and this connection ID is independent of the layer three or the layer two concepts and we're basically keeping some amount of aggregate state for every connection right so this is the the at the q pair level which is the construct we use and along with that we're keeping the same flow in the same way we're using the same flow as the other ones that we're using.

And along with that we're keeping a bunch of receiver state so you know in particular which packets I'm expecting expecting each packet has its own sequence number which packets are are are are missing etc we're we're willing that information back to the to the. To the request to the article the request is doing the same thing it's keeping its own bit map of packets that are out there in flight which I mean receive which ones haven't and the two are reconciling with each other the the request is using information that is coming back from the responder to do either retransmit or to make decisions such as black listing a path and the request is using the same information to. Place data in memory as it comes in and then sending notifications to the upper layer software if needed. So you are still doing what I would sort of expect with TCP and maintaining some communication between center and receiver but you've just made it more efficient and doing it packet by packet as opposed to flow by flow.

It's it's yes packet by packet information so recall that even with TCP you get acts right in the acts accumulative acts we have similar sorts of concepts but on a on a per flow level what you're doing is you're just sending more information back about the state of the receiver to the to the request. And this feels like this is a lot of technology recall is what rip is referring to so what we effectively do is take the packets of a message is param network perspective they really have no idea that they're part of a message just to network is unaware. So just use packets it deals with packets that's and then at the receive side based upon the message the message ID that's that's carried in the packet and the key number the state is reconstructed and we can make sure that we got all the packets after they've sprayed and we can write them to the host out of water based upon the place of information in the packets.

It feels like a ton of state to keep track of though what where is that it's actually very similar to what is being done for for it's very similar was being done for rocky now one of the things that's different is is because it's out of order to look at rocky are see there's an expected sequence number that so you only have to store one sequence number at the receives side there's implied a bit map that you need to make sure that score each of the packets as a receive so that them receive. That water you can make sure that the packets of a message have been received before you effectively you complete that message from a protocol perspective so that's really the big hardware protocol difference between the two standard rocky and and and and and MRC so there is a bunch of state but it's not it's not untrodden territory we we we not do yes as we look at any rocky implementation there's a fair amount of cute pair state already I mean if you look at it there's

you know unfortunately quite a bit this adds effectively a packet delivery context which is you know most in most implementations of that map which scores the packets that I've been received so that you know that you know that so that you can make sure that you've received all the packets at the and then the other thing just to add to that is that the amount of state is in the order of your round trip time and the round trip time the on understate per connection is in the order of the round trip time per connection which is pretty short in these data centers so you're looking at a loaded RTT of 10 to 15 microseconds in sort of worst case because you know everything needs to happen fast and then the other point just to just to be clear is that we anticipate that we are going to be able to do that. We've been paid to this right when we did the MRC protocol so there's a bunch of stuff in the spec where an implementation can trade off the amount of state for the amount of the number of in flight packets and it can dynamically adjust the state per flow by telling by by the by the receiver receiver telling the center. Hey you can now send more in flight you can send less in flight all of that is available for for implementations that are really simple have a small number of connections you could imagine like a

you know very fixed sort of set of state parameters for implementation that are really much more complicated going over many years this could be much more dynamic implementation to make better use of your available resources. We mentioned selective acts so let's talk about that mechanism here I'm a sender I am sending traffic it is being sprayed about across a bunch of paths because we're changing and as far as our conversations gone so far we're changing our UDP source port number to something unique. That's going to allow me to distribute those packets across my ECMP group. They're received on the other end and then the other end receives those and will send selective acts sacks back to the center communicating a variety of things in the MRC context on the one I assume it's going to basically I got all the packets and I got them at a timely fashion is that first is that so what else you can do so big picture the way to think about this is you've got a less.

Vegetable window and that left edge means I've received everything older than the sequins summer so that's in MRC terminology is called a cumulative acknowledgement. And then additionally a sack message can provide a bit map of up to 64 packets that will selectively say hey by the way I have not seen these packets in this fit. So that is communicating I got everything but but but 16 28 and yeah exactly those in shock now there's a bunch of sack rules at the receiver to send out to determine because you don't send a sack in every packet you basically have a set of rules up to when you actually send a sack message when the six sack messages is back to receiver there's a bunch of processing to figure out okay am I actually going to retransmit these because I think they're lost or the same thing. Do I set a timer on them the sketch is a little vague on this but it allows the receiver to communicate to the transmitter that there are holes in the bit map these are the packets I have not seen yet and depending upon your six of the sender you can say hey I can do a fast retransmission which is which is described in the spec I can do other things to kind of just send the packets that are missing out of the bit map unlike traditional rocky where once you go out of water.

You get a net message that goes back everything goes back to the expected sequence number and you re send everything and when the network delays are when I see the delay bandwidth product is bigger there's a lot more stuff flight so the expensive you know dealing with of recovering using go back and is is significant higher. And unlike TCP we're not there's no sliding window here we're not slowing down because all the sun going on right in the network window is really just described in the packet arrival heuristics the rate is being controlled by the con just for a while and there's a variety of different pieces that are being used for that primarily the at get our TT across the past. But generally speaking if a sender gets a message from a receiver saying I've missed x number of packets the sender doesn't automatically retransmit it can run through you know different decision trees to the yeah I mean yes exactly and if you look at the at the at the spec there's also another heuristic that's provided in the sack message which is the out of order count.

And this basically tells you the difference between the left edge and right edge so this tells you how many bits are set to tell you is your new mistake of how many packets were out of one so under conditions where you have a lot of out of orderness you know it's likely that you're just so that you're just haven't seen them yet. But if there's not a lot of out of order at the receiver and you're missing packet to almost certain those have been lost in a network and you can expect them. Okay so it's not necessarily like sending probes down links to be like oh I think this one is just a little slow so please wait it's using other. Yes it's inferring from others network state that whether or not you retransmit okay yeah there's a whole bunch of signals that are provided in every acknowledgement back right so you get the RTT you get ECN marks you get your bit map you get your out of order window there's like a bunch of stuff that implementations can pick and choose. Which signals they care about in order to make the decisions on when to retransmit.

This quick break is courtesy of sponsor potential now if you follow network automation you have seen AI change how automation is done automation is not just about sources of truth and infrastructure is code and orchestrating scripts and playbooks and repositories and testing and yeah it is still about all those things. But AI adds even more capabilities and so then the question is how do I fit AI into my toolbox because you don't you don't just start throwing data into an LLM and expect to get back these production quality results that's not how that works. NetObsings require that AI behaves like an engineer and that it's bound by the security we use to govern any technology that touches our beloved routers and switches. It's all about that it's deterministic so for the same input you're going to get the same output it's bounded by security constraints and it is a gentick now and agenda is interesting a gentick AI is a set of agents each with a capability that can be dispatched to gather information or perform a task think of each agent like a specialized tool.

So flow AI can securely dispatch agents in response to network events or because you asked at a plane English prompt and then once dispatched they can bring back a deterministic result and then recommend next steps and then you control what happens after that you can approve the action or deny it or refine it whatever is appropriate in the moment. Flow AI is technology that makes you a more efficient and capable network engineer find out more at itential.com slash flow AI they have a white paper there about flow AI that is worth reading and you don't have to give it your contact information to read it again that's itential.com slash flow AI and please tell them packet pushers sent you. Does do those signals are those also what the sender is examining to determine if a particular path was is bad or it should not be using the we haven't said the word entropy value yet but but that's the. So the answer to peace comes down to your old ECMP and if you remember some of the older implementations the hashing was terrible so what people would do is they would essentially randomize the source port in order to get better distribution in the hash algorithm and that's not true in any of the modern switchings but historically.

If you had or entropy in that label you would get very poor low bound decisions so in this case we re re use that in really entropy in this context is a path through the network and you know so there is a built in part of all and there's an existing there's a you know required example is in uses ECN to skip a path and also use the end explicit congestion notification that sent by. Someone intermediate long way I have congestion here you need to slow down i'm letting you know so what it means is that there's the there's the four path in reverse path so per you know ECN standard if any the packets along the if any of the switches along the path have. Our above the ECN threshold they will mark the packet and marking this essentially sets as particular DAC key point and the ECN ECT bits in the in the in the header as the receive side what we do is we will reflect that this path was marked in the sack message.

So that gives the sender an idea that hey the path to the destination was congested and maybe I should make a decision on my load balancer to to to to essentially. Rebounce and most out we don't mean some engineers are going to be pick up like like my five box not a yes i'm not a question of the question in this case is specific typically hardware firmware within the Nick that is choosing which of the past to the destination that packets should go on that's the that's the load balancing decision so in the specification there's a state machine that basically says that if you can see and you do a skip. And really what this means is that you just skip it once through your shuffle of the past because congestion is a temporary thing that's a temporary again even for now. Yes even very good implementations the load balancing decision is never perfect so i'm as you will as you know basically you get some depth in the network that's asymmetric what you want to be able to do is just stop using that path one round that will essentially add a slight bubble in that.

Interface and kind of re re balance sure. The passing of network this is a really important point MRC does not know the network apology. It has a pool of entropy values UDP source port in our current example that it is rotating through knowing that the way hashing works the odds are as it changes different numbers it's going to get a fairly even low distribution but as you said Eric it's not going to be perfect and so occasionally we're going to get signals back like ECN telling us that links congested okay i'm going to skip that for now because it doesn't really know the physical hop to hop to hop what's actually going on it's just rotating a number and. Hoping presuming that the behavior will result in even that's a hundred percent correct is a really good observation and one of the other things is to be clear is the bubble and some network in cases is a bad thing bubble in this case is just using that stop using a path they would not introduce a bubble in networks you still maintain full line rate between the pieces you're just using one of the other pass i'm just moving traffic from one path to another in order to to rebalance the traffic.

We briefly mentioned intermediary switches they can note congestion and mark packets and that gets sent on to the receiver who can make a decision about whether I need to tell a center to back off do intermediary switches perform any other functions in MRC or they basically just ship and stuff as fast as so the requirement in this requirement came from u.c. as well when all switch vendors in u.c. got together we just said okay the baseline is a standard e to that switch that we're going to get a lot of information about. Support ECM market that's what we're going to basically have as a requirement okay as in addition to that we specified that packet trimming which is the ability to instead of ECM marking you would basically trim the packet which means make it smaller and send it at high priority to the receiver is an optional port portion of both the u.c. standard implementation is also optional port portion of of the MRC. Wait a term the packet meaning what I cut the payload off she has so this is the one that i'm sending a 4k packet so the track packet trimming would do is it would send 64 bytes or some smaller amount of that it was sent just that to the destination so you can see is a switch I freed up a lot of offering resources and at the reset and at the receiver.

I have an explicit notification that this was dropped I don't have to guess that this might have been dropped to kind of a strats it were singling to the receiver yeah that this packet was dropped this packet maybe more resource and by the way instead of the single bit signal which is which is the edge is the ECM mark is okay this is the actual packet that would have been dropped and hope and here's the you know the typically the first portions of it and from that you can you know you can schedule for your transmission and I'm going to go to the case you effectively will sit will knock it which is to send a packet back to the receiver that will effectively tell the to the sender that hey you should retracement this because he was explicitly dropped in a network. Now normally what you would need to do in this case is wait for time outs and other pieces but in this case it's an explicit notification and you know there's a toll bunch of papers and stuff on this on the on the performance improvements you get free to use these sub detectings.

So those are the only two real features in the in the fabric that's other than that are based on requirement acquired from from use from a MRC perspective. So most of this I guess we can call it state is being monitored by the sender and the receiver and intermediate switches are just doing congestion marking. Yeah, the intermediate switch and that MRC aware at all yeah yeah like like Eric said I mean you you can just take a standard switch that exists today that does ECMP and it will work with MRC. Okay. Yeah. Okay. All right. We've talked about hash based entropy values that rotating UDP source number as the I think the introduction introductory way to kind of get the concepts across here how packet spring is going how packets are being reordered the fact that MRC is known to the switches but it's to the sender and the receiver on either end they know about MRC and how to reconstruct the flows and some of how they're doing that. Well, all right, there's a couple of other ways that we can deal with spraying packets across our ECMP bundles there's structured EV entropy values and then there's SRV6 microsegments.

Okay. What let's talk about these other options. Right. So I think we started with the ECMP version which obviously supported and all the switches today about any modern switch of support this that the structure EV and the microsegents allow you to either specify explicitly portions of path of the entire path. If we go back to early days and networking there was this concept called source routing where you would define the path of the network at the source and say, okay, I wanted to take precisely these hops. It was a young engineer Eric. It was known to do source routing and the thing was you got to turn that off with the security breach don't think that there's all kind of right exactly. So what this allows you to do is depending upon how sophisticated you want to be your network, you can effectively engineer like we said with ECMP even with the best even with the best session, there are collisions. You can construct both structured EVs and SRV6 pass that are 100% non-overliving if that's your desire.

Jan, you would need knowledge of network topology to pull that off. You absolutely need knowledge and network topology and also you need you know at a minimum. So depending if you're using structured EVs, you need some type of ACL engine in order to process the tag and obviously if you go all the way to SRV6, you need an SRV6 capable switch to deal with the microsegents processing. But this allows you to like I said, explicitly engineer the path from source to destination with the case of SRV6 and with structured EVs typically designed to spray to the root and then you have a single path down from the group is typically what that is a typical deployment methodology. So I need a controller don't I for that. You need definitely need knowledge of the network. You don't necessarily need a controller, but you need a knowledge of the policy and one of the things that is unlike many networks that is that if you look at the AI training networks and most of the networks, they're deployed in one chop as in you know exactly you're bringing a building up this building has 64,000 GPUs.

This is exact it's it's a very well known topology and that topologies known you to sometimes years before it's built. So that's not the radical to know it's you know, you know, you know, you know, you know, you're actually right and that's not typical in most networks, but these AI networks that want to do path engineering, they know very well. And in fact, in some cases, they've chosen the topology based on being able to control some of these heuristics. But what is the advantage to this path engineering over just spraying packets across so two things a is like I said, there's zero overlapping paths 100%. So that's the distribution. Yeah, now you still the next still has to make a decision. So you know, there's there's still there's still some ambiguity there and most next, you know, they're not perfect there either. But the past and the networks are 100% known and can be engineered to be you know, non congested or whatever you're looking for.

The other piece is that, you know, if you do control past each of the GPUs, you can control effectively exactly how your the network is is is is presented based by essentially giving a subset of past each of the GPUs or each of the notes in the network. It can become increasingly sophisticated depending upon what you're again, this gives you 100% control on precisely the path that goes to the net. But if there is an issue, if there is for whatever reason, congestion or delay in an SR V6 mode and I've traffic engineer my paths, can I use a different path? Is that so it just sort of like stuck in general. This the source Nick is both balancing across the number of paths. Those paths in this case, this map them to different like a said stacks in the network. And just like we had before, you can effectively you load balance across those. So there's ECN marked, you did the skip, you all of the same processing you would do really from a part of all perspective is nothing new.

But one of the nice things about this as well for all of these technologies is that if you detect the path is failed, that failure can be mapped out in the directly in the data path, which is kind of a new concept for MRC. That is, I was going to say the advantage of doing that is of course that your convergence when a path fails to detect that it's failed and be able to use a different path is in the order of one RTT right because you're not relying on the switch fabric, the routine protocol to do the conversion. So you get much less downtime. That's a huge advantage for a fixed fail. Very fast reaction then to the link going bad as opposed to like waiting for bgp or whatever you did the center of the work is to detect that failure, you reconverge and go from there. Now you know very quickly off this is bad. You said one round trip times we're talking in order. It's you can depend upon what you want to risk you want to put on it, but it's it's one or it's low number of RTT times, which is orders of magnitude faster than any routing portable is going to converge.

And you know the key thing is it's also done like I said in the data path, the you know the control of the application is completely unaware of that. I went instead of you know I'm starting using six for pass. I'm not using 63 because this one was we thought was bad as long as that we put it as long as there's still enough diversity in the network. You want to see any reduction and bandwidth between the two notes. If I were to be using a controller that in this structured EV or a microsit deployments what role does it play because it can't be the one that's detecting a bad path and reacting and said that would take way too long. What role does the controller actually. So you know different customers have different views, but I think the way I look at it is the controller is responsible for giving you the possible good pass between two destinations. Now that may be different than the topology because thinking things out for maintenance. There may be you know there might be problems in the network.

So the controller would come in periodically and the period is probably long. So okay these are the past that are are good in the network with physical perspective. So then each of the each of the nicks in the network will use these past to create connections to the destination. And then if a link goes down or laser goes bad or congestion all this will be detected in the data path and then we'll stop using that and then there is a set of APIs that we can that that are go up to the control plane. And there's a couple of different things you can do you can you know take another good path and replace it or you could you know go to network management there's a bunch of customers specific you know things you can do when you've detected persistently bad path. But I just want to we've talked about controller and I think typically when people think about controller means I'm often hunting up to the controller to get a decision about a path but that's not really what's happening.

That's not the case here. Okay. No, we went to a lot of effort in MRC to make sure that the data plane is completely independent of the controller for various reasons. One is it simplifies the implementation overall system level. Two is we wanted the fast reaction times to be able to deal with congested and with the down door or or flapping hands. And the third is that from overall resiliency perspective it provides a better overall resiliency model when you don't have a piece of software, I either controller in. Right. So we've talked a lot about the nix and it sounds like most of the intelligence is residing in the Nick. What's the role of the ASIC in an MRC environment. And do I need to specialize as a stick to operate in this environment. So I think Eric talked about this before and that typically what we're seeing is that MRC would be added to an existing our Nick.

So if you have a rocky Nick today. You would add MRC to this. And if you look at the product, I use it me the end user or the far know the Nick vendor. Yes, okay. Our Nick vendor would be the would would add this as a feature to their to their to their either their firmware or hardware implementation depending upon the Nick and specific. That's been done by somebody already. Is that true? I've been done by number different vendor by many people already so you know all the vendors that are involved in MRC. So AMD brought come Intel and video have implementations as far as I know. And the since the spec has been public we have anecdotal evidence that there are start ups and other vendors that are looking to implement this too. And then on the switch side as far as the ASIC goes again, it depends on what entropy value method we're using.

There's nothing special about UDP source for it. If we're using structured EV there's a I think it's a 32 bit feel that's got to get parsed you can decode the hot by hot path which you would do via access control list entries that can parse that and then it's likely true that you're using ACL engine. And the structure TV is overlaid on top of the full label and on top of the UDP source value. So these are typical fields that are already parsed by most ACL engines and most of which is we know about. So those were the available to the ACL engine and then you know there's definitely some specific processing you need to be able to select the port based on that. And again, most of these are most of this which is support to type of things. And again, also look at the structured EV as a halfway point between ECMP and SRV 6 right. So all the major news which is are doing SRV 6 because it's being used in other use cases.

It's very popular in 6G for example and is a lot of standard bread and butter there. And when we designed MRC we were sort of the timeframe was spanning these two right. So we had the capability to do ECMP we wanted some sort of structure TV. We put that in and sort of the world was moving towards SRV 6 as the deployment models for our customers. And so both options are available for you for source. So when we started this kind of a history piece the as the structured was first because when we started this SRV 6 was kind of unknown. And then I think midway through the development process of the spec there was an open source implementation in SANA. And that really jumped start the the discussions like okay this is now a real protocol. I can really go get an open source implementation of this implemented on a variety of switches.

And also there was you know some discussion within the group about having particularly with that remote data centers being able to have finer grain control from the route switch out to the end points was a desire. So at that point you know there was also some discussion about this is you know standard base the structure TV is you know is obviously a is something dedicated to MRC. So those I think you're too big your choices between that structure TV is slightly more efficient because the SMT the micruses take another out the V6 et cetera. But these are workloads you know most of the traffic is large packets so that one red is minimal. Okay what about APIs is that something that I should be keeping an eye on or what what are my options. Yes yes so part of the OCP project that we have with the spec is there's also a GitHub repo where we have defined all of the APIs necessary from the application.

To create MRC QPs and drive those as well as the controller managing eb and in the nick itself right. So the one thing that we did with the application APIs is we modeled them very similar to what you see in and with IB verbs today. Okay and what it is. So you've got create QP modify QP a very very common set of attributes that you do with each of those and if you were familiar with verbs program today you would be able to look at our MRC APIs and go oh yeah I can do this. And just hit the ground run it. So if I built an application stack thinking about Rocky V2 you're saying I can just kind of make a few API tweaks and plug into MRC. That's right. It just a couple of couple of line changes for simplicity and then there's a whole another aspect of these APIs and that's for this controller right so if you want to manage your evs your structured paths on your SRV 6 whether that how you how you not only how you manage on.

But they could be generated on the nick they could be explicitly defined by the controller controller also has APIs to send out of them probes like if you want to probe a specific path it can do that. There's also a whole event mechanism so as we talked about you know we've when we're the hardware is actually probing in it could skip evs right so when it skips in ev we have an event mechanism where we can actually send an asynchronous certification up to the control. Or saying hey you know this ev we see a problem on this and then you that controller can proactively you know figure out what's going on there right so it's a very rich set of APIs and I would also call out in this GitHub repo that we have. There's a there's like a full programmers guide right so the it talks about like all of the different ev programming all the different application programming a very step by step basis and how you would do this right and it's. It's very cohesive and it's it's broad and you know as we talked about all the companies chip in on these APIs and it's it was a it was a monumental effort there was a good two years it was about it was about 50 50 to be quite honest with you.

So I guess a ton of time on the API's one of the other things is that if you look in the GitHub repository there is essentially worked examples on how do you structure these srv6 ecmp I mean there's stop there's read me files in there that will go through you know very detailed okay if I want to do this this is a set of APIs that we did to call. It's not in the same. Yeah, you know another thing that's in there is there's a command line application that I wrote that we can actually experiment with evs so like so I have a network I have a three tier network I want to have 64 pass or 128 pass between you know myself in the spine. And you can actually break it down this tool actually break down all the ev generation it'll show you bit with and like how you might want to program your switches you know what bits you have to look at it'll do srv6 and all that so it's a really good guideline to see what's capable on these APIs and then how to move forward. Yeah, okay, the other thing talking about a PAS that comes to me we didn't really talk about telemetry what can I get telemetry it sounds like a lot of these decisions are being made by the next at very like non human speeds so what do I as a network engineer or from an observability or monitoring perspective doing.

And what can you see to me so so there's two sorts of telemetry there's a telemetry that you'll get as a facet of the implementation of the Nick right so a lot of our next counters that they're that they're updating with every center received packet. And that information is is fed and all the implementations that are doing MRC will have MRC specific counters for sure and then beyond that you have telemetry that as Eric talked about come from the controller perspective so the controller has this capability called eb bench which is I think the most important aspect of of MRC which is that when. And a packet times out or you get a congestion mark you get a notification actually just when a packet times out we drop the other one you get you get notifications that are sent to software that it can aggregate and then this these these these aggregate telemetry events can be sent to the controller that can be used by the controller to either do an update or you know some sort of further processing for example troping the path that telemetry is is is is is standing.

And then finally from the perspective of the overall end to end system right so you're going to have a bunch of next a bunch of switches we have made sure that our we have documented things well enough and and there's a set of third party tools that are now being written by some of our partners some of which are open source that will allow. Different networking engineers to be able to pick up and do things like wire wire sharp traces etc that's all all the time as well okay. Are there you mentioned GitHub so it sounds like I can actually start taking a look at things but other other places or ways I can explore MRC are there organizations that I could talk to or learn from yeah there so i'll talk a little bit and i'll let the other guys talk so. So from the point of view of what is MRC right there's a OCP spec it's a very dense complicated spec I would suggest keep that as secondary reading first reading is there's a four page paper which i'm sure you'll be able to like to on the website yeah which was in hot interconnects this year and there's a presentation that Eric spot and I did on that as well so if you want something else go read that but the four page paper is a very succinct network engine.

Near perspective of what MRC is so so go go look at those four pages then the specs a second way of getting information there's another paper which is a bigger paper written by opening high on which they talk about deploying MRC at scale and what they've seen in the field and how it behaves that will give you a good idea of the capabilities of this thing at runtime and then forth and I mean this quite literally. Eric Eric and I would be very happy to field questions right so we'll give you our context just let us know if you have questions and we're happy to respond yeah we'll have your LinkedIn's we've got that here and we'll make sure those are so open the show. Final question then we've talked about MRC and you have positioned MRC is sort of a bridge or a stopgap to get capabilities out but moving toward uec and uet so is there anything next for MRC are you guys continuing to add features or make improvements what what happens with MRC going forward.

So so again right now with MRC the you know it's just out there we've got a couple of implementations we're sort of fine feelings as implementations we've got a couple of customers there big customers and they're getting the stuff out our bandwidth is being consumed with getting this thing deployed implemented. You know seeing it in in in the field and then using the feedback we're getting we expect to either. To now implementations for this I will be one aspect but we definitely expect to use that feedback to feed into either modified version of MRC as as things you know are changing but. More likely we'll also pull that stuff into the uec because remember I said MRC is a stopgap it's not be all an end all and so. The ideas to use MRC as a proving ground for the uec ideas and improve uec further as a as a function of it.

Okay so we could see MRC sort of. Refolded into uec in the future yeah yeah that's the right way to get on it okay that makes sense so it's not like if I'm spending a lot of time implementing and learning MRC it's a waste of effort for the one thing I would tell your listeners is. If you want to get into uec without reading the 700 big spec understand MRC right like you'll have got. I can't buy it 90% of the way there from from the point of view of the things that matter right okay let's go to advice. All right well you mentioned a bunch of resources will have those links in the show notes that accompany this podcast will also have your LinkedIn so people want to reach out to you with questions thank you all of you for for being here thank you for the work you put into MRC I know. AI data center networking can seem like sort of rarefied but it's fascinating stuff and I sort of wonder if some of these capabilities will. You know make their way down into more standard enterprise data center networks or maybe more enterprises are going to be someday building their own you know AI training fabrics so.

So I'm a maybe I'll let me just add a little bit yeah please which is that. This is my opinion my personal opinion but I am seeing a lot of. Use cases for our mix or RDMA in the data center right like as line rich increase and then we want to make more effective use of our CPUs. RDMA technologies are becoming important okay so while we are presented here is like one x one end of the scale which is extreme performance right like multi path and crazy things happening with probing and entry values and SRV 6 now for different use cases a subset of this stuff will be useful but. The world is definitely moving towards a configuration where hard risk fast processing capabilities incredibly fast for the GPU or whatever and then to feed that you're going to need and still the resources whether storage whether it's secondary compute whatever it is so you can imagine that the fabric that connects all of this would have to make use of all the capabilities that we are providing with MRC today for other use cases as well so to the list nurse I would say.

You would do very well and you would be ahead of the curve by making sure that you understand these these capabilities you understand them at a sort of principle level and you you can reason about them that's putting your head of the pack. That's fantastic yeah like the perspective well again thank you rip and thank you Eric and Eric for spending the time with us and for going into so much detail really appreciate it Ethan banks had potentially a fiber cut or something else that accidentally kicked me out of the. Accidentally kicked him off we didn't we didn't send him away that was some internet issues but he would says farewell i'm Drew Conry Murray from pack of persons you can follow some LinkedIn you can join the pack of persons community slack with thousands of others check out pack of persons that net to figure out what how to do that send up a follow-up message discover all the other content we offer for free we do respect your privacy by the way we aren't tracking you there's no log in required so. Coming up on heavy networking we've got episodes planned with meter and selector as well as discussions about BMP and net claw and again gentlemen thank you for putting the heavy and heavy networking for today show and until the next one just remember that too much networking would never be enough.

More episodes

More from The Everything Feed - All Packet Pushers Pods

View all episodes →