Welcome to the KuppingerCole Analyst Chat. I'm your host. My name is Matthias Reinwarth. I'm analyst and advisor with KuppingerCole Analysts. My guests today, it's a plural, I have two guests. I have with me Martin Kuppinger. He is the principal analyst of KuppingerCole, and Alexei Balaganski, he's a lead analyst in cybersecurity.
Hi, Alexei. Hi, Martin. Hello. Welcome. Great to have you, and when I have two guests, it's usually for a good reason, and it is again, I want to start with a small anecdote. I'm old enough to remember back to 1992 or early 1993, when in the company I was working at at that time, a friend or a colleague came in and said, I want to show you the future of the internet, and he showed us the World Wide Web.
By the end of that day, we had created one of the first 10 commercial web servers in Germany that was a list at Freie Universität Berlin, where you could look up all these internet services, commercial ones, and of course, academic ones. By the end of the day, we were really blown away by the idea that there is a decentralized failover-capable, secure, democratic infrastructure around for spreading information, also for doing business, and we were really, really looking forward to a bright future of decentralized and democratized information provision.
Fast forward 25 plus years, so 32 years, I think, we want to talk about the internet and its status for today. And we want to talk about something that happened on the 18th of November, that was an outage of a platform, of a security and a delivery platform called Cloudflare.
Alexei, what happened? Well, it all went down unexpectedly, as usual. I think the first feeling was that half of the entire internet was suddenly unavailable. And the worst feeling for an IT specialist, I guess, is that you cannot do anything. It's somebody else's problem. It's a global scale problem. And not only you cannot fix it yourself, you cannot even find out what's going on. So basically, we had to wait for hours and refreshing the Cloudflare status page and seeing that, yes, they are working on the investigation and bringing back the services slowly.
And of course, in the end, we found out it was another, quote-unquote, honest mistake of somebody pushing out an incorrect configuration or something like that. Basically, it's definitely not the first time a company does something like that. It's not even the first time for Cloudflare or an oops like that. And all that scale and impact. The issue is basically, what can we do next time? Apparently, some vendors just don't learn from their own mistakes. Can we make them or should we do something, should we prepare and be ready for such outages ourselves and can we even?
I guess this is exactly what we are going to discuss today. If we look at the architecture, way back in 1993, we had three servers, PC, SEO, Unix, to be precise. And if I switched one off, then that portion of the internet would go down.
Today, we have the architectural possibility that the outage of one service takes half of the internet down. That would have been impossible way back then. So what changed from an architectural perspective? And maybe that is the lever to use how we can look at that, Martin. Yeah. So I think, when we go back, even before 1992, before the World Wide Web, I think the basic idea when we go back in the anti-history was to build a resilient communication infrastructure, a resilient communication network.
It was the whole idea behind everything from ARPANET and others, starting, whatever, in the 1960s or so, if I'm right, which really took a long time then to become what the internet and the web, et cetera, were. And this is the foundation. There are definitely elements when you go to the, so to speak, core of the internet, which are still built for that resilience. You can have multiple DNS servers, et cetera, and the traffic finds its path, in a sense. So at the end, the underlying infrastructure still follows this principle.
What has changed is that we have more and more overlays, one thing, like content delivery networks, and we have more and more very widely used services that can potentially fail. We have seen these failures virtually everywhere. We have seen CDNs like Fastly having their issues. We have seen email security solutions like Proofpoint. We have seen Cloudflare. We have seen AWS, certain microservices, soft services, et cetera, being off for a while. And a while, basically, when we look at our usage and consumption at a global level, a while is everything above a few seconds.
But we have had these outages in the hours up to days range, sometimes even mostly several hours, happening every now and then. And this is impacting business. Right. And if we look at these platforms that are obviously built on top of the regular internet infrastructure, they are providing redundancy, failover, segmentation, as long as they're working. And as you said, Alexei, once we use these services and they're working all the time, most of the time pretty well, but if they go away, we're potentially lost and not able to do anything.
So this consolidation or concentration or lock-in, is this really something that we should reconsider, Alexei? Well, you just used the very E word for this whole discussion, platform. Everybody nowadays, especially in the security business, everybody wants to have a platform for one simple reason. If you have just one universal tool, which combines and unifies multiple functionalities and reporting capabilities and whatnot, it's obviously much easier and cheaper both to run and to use as a consumer, the customer.
So of course, companies like CloudFlare and Palo Alto Networks, and even though all the big cloud hyperscalers, of course, they push this notion that the platform is always better than an independent set of tools. Even if each of those tools is best of breed on its own, it's then your responsibility and a lot of work to basically juggle all those tools and make sure that they work together. So they will happily take this burden off you for a price.
And they promise, of course, to have an easy, universal experience where you only need a credit card to unlock the entire digital business, if you will. But then again, the same applies to them as service providers, too. If it's really a platform designed by the lowest bidder, if you will, if it's an architecture, well, the only reason to make it integrated is to reduce cost.
Of course, if some tiny part of this platform fails, you just have this whole domino effect. If basically the entirety of your international cloud architecture depends on a single DNS service in one of the regions, obviously it will never survive a nuclear blast as it was originally intended by DARPA. It won't even survive an intern unplugging their own cable from a network, right? It goes down because of a simple mistake. And then again, half of the Internet goes down with it. This is exactly the direct consequence of platformization of the entire Internet, if you will. Yeah.
And I think there are two aspects maybe to bring in. The one is vendors, platform providers potentially could do a better job in increasing their own availability and resilience against failure. And failure happens. The other is, I think, the equation a lot of customers use for making decisions might be somewhat incomplete because it looks at, I think what customers frequently don't factor in sufficiently is the risk and cost of outages, of failure, and also the lock-in risk. I think these are different aspects.
When I look at the first one, let's take something where something leads to updates being installed on a large number of endpoints and bring these endpoints down. Something by the way, me as a Windows user from basically more or less day one, I remember these Microsoft patches in the past, which brought down a lot of Windows systems. We haven't seen these major things, I've seen whatever, 15 or 20 years ago, we haven't seen them for a very, very long time. Hopefully this remains as it is, but I think you can improve.
And I think it's very important to learn from failure, but I think it's also important that vendors put in, I would say, controls to understand whether changes have an impact to stop changes to roll back quickly if needed. So if I take rolling out an update and then you observe first in Australia, oh, there are issues, but it runs all over the globe, through all the time zones, hour by hour, and you can't stop it, why not having something in which is basically a kill switch? Something where the update component, the update service asks central components first, can I run this update now or not?
If you collect telemetry data, you then have an option for a kill switch, easy to implement, really not rocket science at all, things like that. I think we need to also think a bit beyond. So I think that would be one call to the vendors, looks like you want to say something before I move to the other side, the end user side, Alexei? I just wanted to say, what you are talking about is an easy question you could ask yourself because you already know the answer, if you will, but unfortunately not everybody in the world is as responsible and forward thinking. A lot of people just, well, why not?
Because nobody ever thought about it until it's too late. But the problem is not that, I mean, we are all human, the world is human after all. The problem is that a lot of people continue not thinking about this even after the first issue.
You know, we are analysts and we analyze software services for decades. Maybe even before doing as an IT freelance journalist, I think since 35 years, I look at software and it was always interesting to see that there were mistakes that were made by many recurring and could have been easy to fix. So I remember back in the days, I had a lot of Windows NT software and Windows NT services installed. And in the early days, a lot of them failed coming from US or UK software vendors. I installed them on a German version of NT and it didn't work because they were looking for the group domain admins.
But in German, it says domain administrator. So the text string is different, but the GID below it was the same. So they just would have done it right following the APIs and it would have worked. And I saw plenty of solutions until they learned in some sense.
And yes, I think it's an important thing that vendors, and for many areas, step back and think about what can we do differently. So really also a bit of a shift left thinking, maybe thinking about what is really a different way to tackle it. By the way, this is also something analysts are there for. I think we won't fix everything, we won't have a great idea for everything, but I think we are extremely good sparring or discussion partners to think about maybe from a broader perspective of how you could fix this. So just reach out to us because this is where the experts are there for.
The other thing I think is clearly on the end user side. And I think we massively, or the customer side, we massively tend to underestimate price. We have to pay for some of the things like platformization, like the super powerful hyperscalers. There are always consequences. There's no such thing as a free lunch. So platformization comes at a cost. It comes at a promise and a cost.
I think one of the things we probably will discuss more in the future, but I think when you, for instance, count the number of services of hyperscalers and other cloud providers, you can come to the conclusion that some of the well-known hyperscalers are by far the best because they have the most services. But what happens if you use all of these services? You hardly can then run a multi-cloud strategy anymore. You hardly can move to a sovereign cloud if you want or need to do that because you have a lock-in.
And I think what everyone, what you don't do well overall when we do product decisions is thinking about how do I get out of it again? This is one of the most fundamental questions. So what can go wrong? What happens if this fails and how do we get out? Ask these questions.
You know, this reminds me of a really old joke about a guy who just found a new job at a firefighter station and basically his friends asked him, how is it going? How are you enjoying your job so far?
He said, you know, it's awesome. I get to just sit on my couch all day long, eating sandwiches, watching TV, would love to do it the whole life. But you know, every time there is a fire, I'm seriously considering quitting. And I think this kind of perfectly describes the whole situation in the IT industry nowadays. I guess it's all in this whole, it's a cultural thing, if you will. Like we have this tinkerer culture in the whole IT industry and some people just have never grown up from that stage. They love to play with things, nowadays are very expensive and kind of cloud scale things as well.
But they never think about the real responsibilities of an adult, consequences of their actions If something goes down, somebody has to pay for it. And if it's not you, who then? Until our industry grows up, well, we will be facing this challenges over and over again. And I think it's also, basically, this is also one of the reasons why most organizations have a sew of tools in cybersecurity. Because okay, something goes wrong, I need a tool to fix it without thinking too much about could I do this in a different way? Maybe I have tools, I just need to use them the right way.
How do I handle these tools, et cetera? I think overall, and I think this holds true for everything. If you go for a content delivery network, it's very clear that helps you in delivering your content. If it goes down, you're in trouble. So what is your exit strategy or what is your failover strategy? This is the question which isn't asked frequently enough. And there might be consequences.
If you say, I don't want to rely on a single hyperscaler, I go for a multi-cloud strategy and I want to leave the door open for having serenity, at least in certain critical areas, then it means you can't utilize all the services. And then it means, okay, then I need to pay a price for it. And I think we must become more honest to ourselves as decision makers to ask these questions and to think about the fallback strategies, the failover strategies, all the stuff we need to have in place, depending on the criticality and the risk of a service.
Some of the services we may say, we don't care much if they're out for a couple of hours. For other services, it's just business critical that they run. So if your business depends on selling stuff over the internet and your servers are not available for a couple of hours, then it costs you effectively revenue. One of the consequences of this platformization is if you, for example, go in, let's say CloudFlare, kind of emit to all of their services completely, it's not easy to switch off it in terms of a crisis. Like if it goes down, everything goes down.
If for example, your DNS is hosted on CloudFlare as well, you will never be able to switch to a different CDN, even if you have one. So do you have to kind of build another independent overlay on top of all those overlays on top of the internet, or do we have to come up with some completely new 40-year-old networking patterns which we forgot how to use in the cloud times?
Yeah, I think that the first question is making up your mind, can I stand these scenarios? And I think we have sufficient examples of what has happened already. And we can extrapolate that other situations will occur. There's a likeliness that something will happen also in the not too distant future that takes longer than a couple of hours. And then the question is, which are the services you need to have up and running and where you say, okay, if the risk is too high because I can't easily build a failover, then you need to take a different route, probably.
You have to think even deeper than that, because you have to consider your entire digital supply chain too. If you are, as you just mentioned, an online store, and if you are worried about your CDN going down, fine, you can build a second one or a third one. But what if your credit card processor goes down as well, because it was also dependent on CloudFlare? What if your identity provider is down as well? What if your whole security stack is unavailable because it was exclusively supposed to be managed through the cloud, and the cloud is down?
Where do you draw the line before reinventing the entirety of your existing infrastructure? Yeah, I think it's a tricky question. I think the first thing is to be aware of what can go wrong. If it's a manual process for certain things, and yes, if you say, I don't build on a platform, then you have to pay the price somewhere else. It's not that saying, I just don't use them doesn't solve the challenge. It is about understanding what it means if you do something. I think this is the main point, and then making a really informed decision. I think we can transform this to a much simpler scenario.
I know a lot of organizations that ended up with EntraID as their primary system for user authentication on the endpoints, because someone made a decision to go for Office 365 or Microsoft 365. The decision wasn't about EntraID, the decision was about Microsoft 365, which in consequence meant EntraID was there, which triggered our decisions, which can be good or bad.
I like EntraID, to be honest, but it must not happen that this comes in through the back door, without understanding in a decision made for a workplace, having a decision that has a huge impact on your ecosystem or identity security. It's clearly something you can handle, but it has an impact, and doing decisions without thinking about the consequences, the full consequences, I think is important. One of the most overlooked things is at the end of the day, lock-in. What does it mean from a lock-in when you do that or that or that? Can you get out of it?
Relatively simple to get away from your office solution. Way more tricky to get away from your ERP system. Having this in mind and understanding what you can do helps you also to develop better resilience strategies. I think this is, at the end, a bit of the starting point, not to assume that there's something that doesn't come at a cost. You always pay a price for that, and if it's convenient, then there's also a price to pay.
If I play the devil's advocate, I would say that – or I could say, I never would do that, of course – that these platforms introduce systemic risk for organizations, and systemic risk should be something that should be prevented. In all these episodes that I did in this podcast, almost every third ended up with the idea of proper risk management, understanding and evaluating and mitigating risk, understanding what that actually means.
Of course, these vendors are able and capable of understanding what went wrong. Of course, it's documented, everybody knows, and how to fix this issue. But the question is – and this is the much more tricky question, as you said, Martin – it's really to understand what the user organization, the end user of these systems, platforms, or only one platform can actually do to mitigate the risks of being out of business for six hours, for 12 hours, for 18 hours. And these customers could be hospitals, could be governmental organizations, could be military, could be anything.
So the question is really, how do I deal with risks? So this is way back then, and so we've come full circle. It's us understanding what we need to do better. There's an established methodology, which is called business impact analysis, BIA. Our advisory team does this. And it's interesting to see the results of it, because every customer learns about it, learns a lot of things when doing this, and understanding which are the things that are at risk. And I think that is something every organization must do to understand which are the really business-critical things.
And if you don't map it to your IT infrastructure and understand what happens, there could be a situation where you say, okay, with the one approach, I might have a higher risk of smaller outages, latency, et cetera. With the other, this fixes a lot of this. I'm always available, high performance. But if it's out for a longer time, this can go beyond the time I can stand a failure. So sometimes the one long incident is more problematic than sort of a continuous weakness or the other way around. And understanding this is important.
I think at the end, it's about doing this more properly, more thoughtful, and as I've said, also really understanding the lock-in risk. And as I've said, the most important question when you purchase a software is, how can I get rid of it again? You need to factor in the access strategy when you actually consider purchasing something.
Alexei, final thoughts. This is not a light ending for an easy episode, but what are your thoughts at the end of this discussion?
Well, to put it somewhat bluntly, if I may, I think that we have just collective as the entire Western society, not just IT industry or whatever. We forgot how to do risk management somehow. I don't know. Maybe it's because people just kind of got used to, you know, that electricity is always coming out of the wall. There's always food in the local supermarkets and stuff like that, and the ramps are driving around and cars don't brake suddenly. We are not exposed to risk on a daily basis, as opposed to a lot of countries out there in the world.
So we just forgot that risk management is a thing, and it's a critical survival skill, if you will, both for physical and digital businesses. And there are rules, as Martin mentioned, there are rules, methodologies of doing it properly.
I mean, at the very basic level, I would say, risk management boils down to, first of all, understanding your own business domain, because nobody can tell you how to do your business properly. But on the other hand, you have to do some mathematics on top of that. And of course, there are some basic rules, you know, like there are rules for handling guns, for example, which were created by people who experience those issues on a daily basis. And now you don't have to shoot yourself in the foot to learn how to use guns.
The same should be applied to clouds, if you will, and everything on the internet. You have to be able to learn from other people's mistakes, you have to repeat all those stupid things. I think it's interesting, you know, I think we tend to be somewhat naive sometimes. One of my favorite examples, and I wasn't a fan of trust in time, delivery logistics way ahead of the pandemic and all the supply chain disruptions that occurred then. Because at the end, this was a classical mathematical or a failure in the equation.
So there was the advantage of lower capital binding, so lower sort of bone capital and having only smaller storage units, et cetera. All true, but I think every strike already showed, okay, that may disrupt the supply chains. And it was very clear that there could be larger disruptions of that. And I think we have a bit of a tendency to ignore the risk side of things and look more at the positive side of things.
Ideally, economic professors should have done it properly when bringing up the idea of trust in time. Obviously, they didn't do it that well. It's true in other areas as well.
I think, you know, automotive industry right now suffers from availability of chips. And at the end of the day, it has been clear for whatever, 15, 20 years at least, that the main differentiator in an automobile is not the fuel-powered engine anymore, but what is in the chip. So that should have been a center of attention of every C-level leader in the automotive industry to build resilience and autonomy and sovereignty, in that sense, for this part of the vehicle. And so I think it's something we can transfer to a lot of other areas in the economy, and we need to get better here. Yes.
So the summary for today is that we send all back with some homework, rethinking their risk posture and understanding what the actual risks are. And that is based on their own business context, and then try to find better solutions to processes that are in place. Would that be a good summary? And that would involve this event as well? I would say we cannot do the first step for you. We as like IT analysts, like coping your call, because again, only you, your peers in your industry know the exact risks, the impacts and probabilities.
But as soon as you have figured out those numbers and you identified your top priorities, this is where experts like ourselves, if you will, come into play. This is why you should start talking around, finding the right people and understanding. You should learn from other people's mistakes. You don't have to do the same. You don't have to step on the same rake over and over again.
And well, the events like our EIC, for example, are the right place to talk to those people, both successful and formerly unsuccessful in their businesses. But at least they've learned their lessons and they will be happy to share their experiences. Final words, Martin?
Well, I think we've said enough. Just take risk earnest. True. Okay. Thank you very much, Martin. Thank you very much, Alexei, for being my guests today. And that was, yeah, a tough discussion, but sometimes there are no easy solutions and we can even support you in finding not so easy solutions. So thank you again, Alexei and Martin, and see you in the next episode sometimes. And we are getting close to Christmas. Have a happy new year. This is the final episode that you're listening to for 2025. But don't be too happy. I'll be back early next year.
So there will be more analyst chats in January. So thank you again and looking forward to meeting you again. Bye-bye. Bye.