Learn more
Don't trust your CMDB? Try IP Fabric's ServiceNow integration, available in the ServiceNow marketplace!
read more

Continuous network assurance through operational intent verification

Performing root cause analysis is hard enough, but setting continuous network verification for the specific cause of the issue can be impractical or downright impossible. Find out how you can proactively increase network reliability using continuous verification of operational intent of the network infrastructure.

Transcript

Thank you once again for joining this webinar. My name is Pawel Bykov. I'm co founder of IP Fabric. And, thank you for joining this, webinar about, continuous network assurance, through operational intent verification. So, network infrastructure complexity is constantly growing.

And as the need arises for, advanced level of, continuous verification, which is, goes above and beyond, what current monitoring solutions can provide. Today, I wanted to walk you through a, use case, that, where we can see how to not only properly analyze, root cause and how complex the cause of an issue might be, but, most importantly, how to, ensure that the, the root cause is no longer present in the network and that, through continuous assurance, we are improving the design and improving the SLAs by, ensuring that, our relevant redundancies are present in the network and how to achieve this increase in, in the service availability, even though the outages of the hardware, will will keep happening, as they are as they happen currently. So the specific example that I wanted to show you is in case of an outage that happens in a critical site where the goal of operations is to recover the failing hardware or failing device as as quickly as possible. So there is really no time for root cause analysis, no time to collect relevant information. So this is where, IP Fabric solution comes in.

As monitoring is lacking in-depth, IP fabric takes a full and complete snapshot of, whole network state. So in this case, from the, time when the, issue was happening, operators took a snapshot of, a failing network to for engineers to be able to analyze, root cause later on. We also took, a snapshot at a different time, which we know that the network was working because it was verified by the user. So in the initial situation, we have a case where, when comparing the baseline network snapshot to the, snapshot when the, outage has happened, we saw that there were devices that were no longer available in, in one of the critical sites. So the service outage was suspected, and therefore root cause is warranted.

But, because there was no time to troubleshoot and only selected clients had access to the services at that site, and they weren't available even for user testing or to assist with troubleshooting. The only thing that ops could do is to recover the hardware or recover the failing hardware and and and take the snapshot for us to, analyze later on. So what we could do, we also could look at the specific, changes in the, routing topology and see what, what has changed on the on that specific site from the routing point of view, and from connectivity as well. By visualizing this data, you can see the specific device where which has lost the relevant irrelevant connections. But, in this case, it's not telling us much because, the connections are missing.

But we want to ensure we want to know, if the outage that has happened has actually caused a service outage and not only a device outage. So for that, we need to take a deep dive into data and perform a root cause analysis. We need to ensure that, if indeed it was an outage of only of a single device and nothing was impacted or if, it was an outage that actually, caused an impact to our SLAs. To that end, we can use, we can use end to end path simulations, which are available in IP fabric platform based on our network model of collected data. So we know that, certain users were able are supposed to have access to the, to the services beyond in in in the data center, beyond the point where the outage has happened.

So we could take one of those, users. And even though, users were never present for, to assist with troubleshooting or assist with this root cause analysis, by running by running network simulations, we can actually see, the specific behavior of, of the network at that time. We can see that, this the the path status is that it is, it is failing. So indeed it was not working at the time, when the snapshot was taken. We can also look at the, verified time and see, that the the path is actually passing.

So, indeed, there was a change in the path behavior, and, all the services were likely unavailable. To analyze the issue further, we can display not only path, but, actually, the relevant, the relevant topology, as well where we would look at, how the 2, the source and destinations are communicating with each other and ensure that, the network is behaving as expected. So this is the premise of intent based analysis where our intent is for clients to be able to reach the servers, and we need to ensure that this intent is actually, satisfied. So with this, intent expressed as, a source and destination path, we can, we can save this view as root cause analysis of an outage that happened on, 9th May. And, we can see that, indeed, during the verified time and during the baseline time, the the status of the path is passing, and we can see the individual decisions that are made, during the, during the past mapping through the forwarding state tables of of the network.

So going to the, outage snapshot, we can see that exactly same path is actually failing. We can look at why this path is failing and if indeed this is something that has caused a service outage. So the general information that the path is failing, and the the result of the failure is because of a routing lookup failure. So it says that the route is not available, at that on that particular route. Therefore, we can go into the, into the routing table, which, as IP fabric collects information about every single routing table, from every single device, we can analyze routing tables holistically and ask which routers were able to route to that to a particular IP address during the verified time.

And we are interested in prefixes that are, longer than 16 that are not, really summaries, but just, pure, unsummarized prefixes. We can see that during verified time, there are a number of, external and internal, prefixes. We can exclude BGP or connected routes. But for now, let's, focus. Let's look only at the OSPF protocols or, connected, connected routes at this time.

So during the verified time, there were a number of routes present. However, during the outage, we can see that no routes were present, that only a single connected route which originated the network, was was present in all of the routing tables throughout, throughout the network. So that means that even though the route did exist, none of the network resources were able to reach this, to bridge the servers, definitely causing a service outage. So even though locally service service monitoring and load balancers did not report an outage because it was monitored locally, and we couldn't monitor those, services remotely because of, tight security restrictions. We could see from the snapshot, from the detailed information, that routes were actually missing.

So we know that some something, has gone terribly wrong with, with how our redundancy is set up or how our, you know, network is set up in this case. So we can get go back to the diagrams, load our particular view, and look at how the path is, functioning during the verified state. So looking at the path as a whole, we can see that there are number of hops. Some of the, devices are load balancing the traffic, such as in the case of, this router, where we can seek from the forwarding perspective the inbound links and outbound links are, outbound links are redundant. Here we can see that, there are some links that are not utilized.

So there is probably, issues with load balancing. However, we want to see where are really problems with redundancy. For this for this option, we have a single points of failure and non redundant links check. So if we are looking at single points of failure, on that particular diagram, we can see that, even though the links are not used, they are, they can be used as backup, and therefore, there is no single point of failure. However, with this, with some of the devices, we can see that actually there are, single points of failure in our network.

So, actually, this is this was one of the failing devices, during the, during the outage. And, of course, it's normal that, when device fails, we want to bring it up back up as soon as possible. However, in this case, or generally, when when the device, fails from a pair, we don't expect the whole pair to fail. We expect to have a pair of devices to increase our redundancy, to increase our availability. But in this case, we can see that on this particular path, the forwarding topology between this particular source and destination pair, from the source to this, destination, server that's, connected to the, to a particular switch and uses, firewall as a default gateway.

That this cluster, the way it is set up, the forwarding actually makes the, the network behavior non redundant instead of providing more redundancy. So instead of having a a cluster that increases redundancy, the whole cluster fails when one of the member fails because, for some reason, routing is set up in a way that it has to hairpin through a, through another box. It's really up to the, box administrator to, to create a better design. Of course, we can based on this finding, we can suggest suggest to look for a particular way how to, how to fix, this issue, for example, by recommending that protocols are set up in a way which, which provide redundancy. For example, if we will show how the individual routing protocols are set up in that particular site and turn off everything else that's, that we are not interested in.

We can also turn off the, redundancy and, how routings are set up. We can see that, actually, in this case, the firewall set up, in a way where a cluster is connected via OSPF, via 2 links, but there are BGP neighbors only to one of the, to one of the devices. And when tracing, when when plotting a path, we could see that it had to go, had to be used as a hairpin. So definitely, this setup is optimal. And, what we need to do now is ensure that this problem never happens again.

So we need to set up continuous verification that will ensure that the root cause will no longer be present in the network. And if it is present, we want to be, alerted about it. So, one of the things that, we can do is, again, going back to the, to the specific, end to end path that we have set, what we can do is we can, actually, create a, path verification check and ensure that this, this path, is, is never, is never failing, that the, that the path is always, passing. So, for that, we want this, path to, actually work. And when we expect this path to work, it is permanently entered into a path verification stable.

And for this particular path that we have just entered, we expect the the the state, the this path to work, and we expect, the result not to be, not to be a form of error. So if we would look at how this check would look even though we have just set up this check, if we would look at how, this the resulting behavior would be or alerts would be during the during the outage, we could see that actually, this path would, would have a result of forwarding failure, because we wouldn't be able to reach this particular network, at that particular time. So this this is, this will alert us if a service is actually out even though no user or no monitoring system is actually, is actually, watching it. However, we also would like to prevent this outage from ever happening, again. So to that, we need some form of, additional, like, smarter testing, which we can, utilize the community communicative, routing table that's available in the system of all of the routes, that are in the system and set up a verification check.

So let's set up a continuous verification of path redundancy in, in the data center. So how we want to interpret this redundancy? Again, through the operational intent. So our operational intent is for path to be redundant. That means that every, in in, to express this intent, we would use, next hop next hops.

So, to to generalize the intent of redundant, redundant paths, where, load balancing is happening. We can, say that the net next account has to be has to be larger, larger than, 1. So, vice versa, if next hop count is actually, lower than 2, then this is something that we want to watch out for. Additionally, we want to, restrict this check, only for, data center site. So site should contain, DC.

And, also, we want to, restrict it to a particular route. So we are not looking at a specific prefix. We are just looking at routes that can actually route to, the particular IP address, or a range of servers in which we are interested in. We also don't want to consider any, external routes or, or routes that that are connected. And, we want to ensure that we are checking only the right internal gateway protocol.

So in this case, that would be, OSPF. We also it it should be equal. Oh, no. It should begin with o. So we can actually use regex and say that it actually begins with an o.

We also don't want to consider summaries. So, just in case, we want to, limit the length of, this check to be higher than, than 16. And, yeah, that's pretty much, what we want to check for. So let's, let's click on the preview button. So, in this case, we can see that during root cause during the type of the root cause, there weren't there weren't, routes at all to that network so that we have, verified that.

However, when the network is functioning as expected, it's still not up to standards in a sense that there are, less than 2 next ops to those particular routes. So we can create this rule and, essentially, have a continuous verification where for each for each snapshot that will be taken in the future or that has been taken, in the past, we can actually see if there are any, redundancy issues. And if they pop up, we can proactively solve them, actually preventing the, the outage. So that is the premise of, continuous network assurance, through intent based verification where we express our intent either in path form or in a structured representation of our, redundancy intent. And then running the, these checks to, ensure that the network complies, with our, design in or operational intent.

And if we prevent these issues, and fix, fix this non redundancy issue before the, before the outage, we actually, are able to increase, service SLA and effectively prevent outages from happening, ever again because, of course, hardware will always fail. Of course, we will always have, various, various outages. But by ensuring proper design and ensuring that we don't have situations of, of, non redundant links and, single points of failure, like in this case. We will then, prevent a failure of a single box affecting a whole slew of services that are actually beyond beyond these firewalls or, that are actually being serviced by these by these boxes. So, yeah, the the goal of proper network engineering and design is not to eliminate all types of outages.

It is to provide, high higher level of resiliency, high better, network design to be able to live with those outages that are happening on a daily basis and ensuring that those outages do not do not affect overall SLAs or overall KPIs. With that, I wanted to, thank you for your attention. I hope that this webinar has been, useful, for you. If you would like to know more about our platform, please visit, ipfabric.io, to request a free trial to see, how it can help you with, providing better visibility, topology mapping, network insurance, and network simulation, for your, enterprise network infrastructure. Thank you.

Webinar notes

Episode Title:

Continuous network assurance through operational intent verification

Topics:

  • Intent Verification
  • Network Assurance
  • Network Infrastructure

Our hosts

Pavel Bykov

Pavel Bykov

CEO