Register for webinar
Modernize change management for automation and audit readiness.
read more

Preventing Service Outages through Root Cause Analysis

Having the best monitoring system ensures that an outage is always detected and reacted to as fast as possible. But monitoring systems do not understand complex network infrastructure relationships, and why a component failure would cause a service outage.

Transcript

Hello and welcome to IP Fabric Webinar about preventing, service outages through root cause analysis. Today, we are going to look at the common scenarios of the outages and, how can IP fabric prevent the outages, in your environment. When looking at the common, outage scenario, we can have a situation when a node has gone offline. So outages of individual nodes, in the network are rather common. We can never get rid of them.

Nodes will always fail. It's just the weight of way of life. What we do need to have in our network is sufficient redundancy to ensure that individual node outage does not cause a larger service outage or that we are able to be resilient to a number of outage scenarios. Also, we need to have ability to detect the outage, which is usually done by a monitoring system. So, in this scenario, when a node has gone offline, monitoring system has quickly detected an outage and, operators, were able to quickly recover the failed node.

In this case, they were able to do it in, within 15 minutes, so the outage did not last more than 15 minutes. However, the problem was that the node outage, caused, a service outage as well. Now all of the engineering teams involved, they are adamant that the service outage the service is actually redundant and that a single node, should not have caused a, service outage and should not have, impacted production. So this is one of the common scenarios when, an outage happens, when the goal of operators is to, recover, the failure as soon as possible, but then there is no time to actually go in-depth and analyze the problem. In this case, users really disagree that there was no outage.

So how can we solve this, issue? So if we look at how this scenario plays out, if, there is an outage in a critical site, of course, the monitoring systems detect the outage very quickly. There is no time to go in-depth in what has caused an outage or why it's, why why there is a problem. What operators need to do is to recover the service as quickly as possible. Of course, it's one thing to detect a note outage, but a whole other thing to, to analyze it because monitoring systems, they can tell you if, a device is, running or, if it's not running, but they cannot tell you in-depth information about the, about the service, about the network state that's needed for, network engineers.

So what we need to do is to perform a root cause analysis of what has actually happened. So looking at the, at at this scenario and, how, IP Fabric platform can help you with, with this scenario is, first of all, when an outage happens and operators need to recover the service as quickly as possible, what can be done is, the platform can be used to take a snapshot during the outage. So in this case, we did take a snapshot during, during the outage. We, called it we named this snapshot, that there was an outage and root cause analysis is needed. We can actually, look at what has happened, in that time by comparing, our baseline snapshot with the snapshot when the outage has happened and see what has, been unavailable, what was the problem at that time, by look at by looking at, what was, what was missing.

So there is there is a device that was missing. We can see that it did indeed fail and it was not available at that time, but that was already provided by the, by the monitoring system. So where, IP fabric platform value is is in providing you in-depth information about about the outage. In this case, if we would look at, the service that's provided by by the failed device, so when users complained that actually the service was unavailable, during the outage time, we can look like look how the service looks like during our baseline. So during the baseline, users, as depicted, by the icon on the left here, were able to, reach, the servers throughout our infrastructure during the baseline snapshot.

But during the outage, if we rerun those, network simulations, switch the snapshot, to the one taken during the outage, we can see that the, that the path was not available from the user to the, to the destination servers. So when performing root cause analysis without a platform such as IP fabric, you would need to simulate the scenario again. Basically, schedule a change window for, Saturday, decide, replicate the scenario that happened during the outage and start collecting, this the state information from the, or, start extensive testing and and see exactly how the network behaved during an outage. Thanks to, IP Fabric platform, ability to collect all of the network state data, we cannot only look at one scenarios, we can look at multiple scenarios from any, from any user, and test how the service behaved during this outage and compare it to our baseline and see that during the baseline, this everything worked. But during the outage, this the, there was a problem, and, actually, the service was unavailable.

So not only the individual node was not, reachable, but also the the service itself was not available to the user. So that means the outage of an individual node actually caused a production outage. An application was not, was not available. We can also see, the specific reason of why why this path was not available. So if we will look at, this last hop device, we can see that there was a routing lookup for a specific destination, but, it just wasn't, it just wasn't available.

We can look at in more detail, but by looking at the, routing tables. So if we would look at the at the overall routing tables and look, at the situation during the during the baseline, we can look up, that route for a for a particular destination, eliminate the summaries, and we can see that the route to this destination was present throughout the network. So it's distributed here using, through OSPF, through BGP. It's in a lot of places. And, actually, you can see that that route to that specific route was, at, 287, different routing tables.

So that was the situation during the baseline. But during the outage, we can say that the only route present was, was a connected route at at only one site and only on a single device. So, that's a device that's, hosts that specific, you know, that specific network, And we can see, that this is a, Juniper, SRX firewall, that's hosting that network. But, the route was not available in any of the other sites, so it wasn't, propagated properly. We can, look at that particular site and see that, even though the route was present there, somehow it was not being propagated.

Using the snapshot capability, we can start performing root cause analysis without the need to actually schedule a change window or simulate the outage in our real network because we actually had a had an outage, and we collected all of the state data, in a snapshot. So we can use that snapshot information to look at how the site, looks normally, and we can look at, how the site looked like during the outage. We can see that during the outage, there was definitely a difference in how the site is, what devices are reachable at the site, but we can also look at the baseline side diagram and ignoring the protocols that do not establish connectivity. So turning off, layer layer 3 and or turning off layer 1, maybe turn off even layer 2 and just looking at the layer 3 connectivity. So looking at layer 3 connectivity, we can see that from the routings point of view, the site is actually not redundant, And we can check that by looking at, in the options, for single points of failure and, and non redundant links.

Of course, here, there is a problem that there is no routing session that would go from router 5 to firewall 9 and from router 6 to firewall 10. So there is a number of, issues here where security team set up a redundant firewall cluster, but they forgot first of all to connect the cluster to the, to the routing infrastructure and to redundantly connect the, the primary master device. So then in that case when, when a device goes down, if the firewall 9 would go down, there would be, there would be a complete service outage. But looking at the baseline or looking at the outage itself, we can see that firewall 9 was not down, and there was a, a subnet that was, that that was hosted on it, and it, so it should have been, reachable. So then there is a question of why, why this subnet was not being propagated during the outage, because, it's clearly connected to the firewall.

It's clearly there, but it's not, but it was not reachable. So for that, we can go back to the simulation capabilities and, again, simulate a user path or a path from a user to the server. And look during the baseline what could be the problem. So having the single points of failure and not redundant links turned on, we can look at this specific path and see, okay. So here that we have a a red connection where, there is only single routing link from this site.

So I guess this site is not redundant. That was not, part of our problem, so we can ignore that. However, we do have, an issue here that these devices are not redundant. So we already know that because we looked at the site at and we saw that the devices were not redundantly connected. But what we are interested in to find out now is why because, why only, firewall 10 failed, firewall 9 stained the flight, but this path was actually not available and was not functioning.

So for that, we can simply click on the firewall 9 and look at how the path, looks like. So when we are looking at the routing, we can see, just by hovering by hovering the line, we can see the individual path components light up on the diagram here in green and we can see that the path is coming from the left from r 6 and r 5 and is forwarded to firewall 9, which does not forward it directly to the destination, but is actually forwarding it to its second cluster. So that's definitely a peculiar situation and peculiar setup of the firewall rules. It could be because security engineers wanted to set up a firewall cascade. And we see that firewall 9 sends the traffic to firewall 10, and then, fw10 sends the traffic back to the, f w 9, which then forwards, the traffic over layer 3 to the user directly because of over IP.

That's the next hop. And over over layer 2, it actually has to pass through one switch then through another switch and then finally to the to the last destination So from from this analysis, we can see that actually this firewall setup is, very wrong in in on multiple levels. First of all, there is no, routing redundancy, from, from connect, routing connection standpoint. And second of all, there is no, the traffic passes through fw10 before being diverted back to the firewall 9. Using the results of this firewall and of of this outage or from this root cause analysis, we can say that for this critical site, for the, for this data center, we want all of the paths to be redundant here.

So to express this, we can say that paths are created by routing information base, by routing, and, that we want all of the next hops to the to this particular network to be, to be at least at least redundant to have at least 2 paths. So to express that, we can go look at the routing itself. So we can look at the routing and set up a verification check where routing to this network to to this particular IP address, so any network that can route to this particular IP address at the Ostrava, DC site, which is not a summary, or we can specify it because we know that there is a specific routing protocol. So in this case, it's ospf, and, we can also look at that, look at which which routing protocol is being used at the site by turning on the routing protocols and and sees exactly, where individual, routing protocols, routing protocols are used. There is a BGP session between the between the devices, but, the primary protocol that's being used here is, OSPF.

So we can limit this, lookup or this rule also to to OSPF. And, we can set up a a rule, which I've done previously here, which would check the next hop counts for a particular, for for routing to a particular destination at a particular site for a particular protocol. And if it's less than 2, then we would get a warning. So, tests like these are easy to set up, and we can do an example here, routing redundancy verification, where we would look at the particular route that if a route to to a particular, IP address. And let's check here.

If a route is to to a a particular particular IP address, has next hop count less than 2. And, protocol, contains OSPF, then we want to be, warned, about about this, situation. So we'll just let, well, this calculation to, to finish, and we can see all of the, issues that actually, that exist in the network and where we need, to fix them. So in case such an issue comes up again, we know, we would be warned beforehand that actually there is a missing redundancy in our network that we need to fix. And in the meantime, we can schedule, a change window to fix to fix this situation and actually fix the, the firewall setup, so it it would become properly redundant.

So this this type of analysis is very time consuming, and, usually should be done during the day, not during the outage. It's not possible to conduct such an in-depth analysis, during the outage because during the outage, what we're interested in is to bring, bring the service back up as soon as possible. But by doing everything very fast, we don't have time to analyze, anything and to analyze it in-depth. We would need a lab. We would need to recreate the scenario, recreate the outage, which would be much more time consuming than simply having all of these informations at hand at all times, by taking a snapshot of the network, during the outage and analyzing it in-depth.

In addition, by finding the issue, we were able to look at the, specific reasons for that outage that there was, insufficient redundancy in a critical site. And we then can set up a rule that would check the, the redundancy setup and, ensure that if there is such a case anywhere else in the network, it can be detected and found out before the outage actually happens. And then we can schedule of would fix, before it actually affects production. So this was an example of how, IP Fabric platform can be used to perform root cause analysis, and to help, to, analyze the network, to find those hidden risks that could cause an outage. Otherwise, the platform provides, a number of benefits to network engineers, since the the platform was made, from practical experience to, assist network engineers throughout, their net network life cycle activities throughout their work by discovering all of the network infrastructure devices, mapping the, the topology, enabling to run network simulations, and, collecting and structuring all of the network state and protocol data to, provide, continuous network assurance.

Lastly, the snapshots, provide full network history where a network engineer can go back in time, analyze how the network looked like before, compare it to how it looks like now, and, improve the network that way. So so, if you have any questions, just type them in. Otherwise, I hope that, the session was useful to you. And, you can visit ipfabric.io for more information.

Webinar notes

Episode Title:

Preventing Service Outages through Root Cause Analysis

Topics:

  • Service Outages
  • Root Cause Analysis
  • Complex Network Infrastructure

Our hosts

Pavel Bykov

Pavel Bykov

CEO