RSS Amplifier

Techie Who Writes · Sep 14, 2025

An OpenStack Conversation

0
Sign in to vote or save

Abby Nduta · Techie Who Writes

A little bit earlier in the year, I asked the OpenStack community on LinkedIn what kind of OpenStack content they were curious about.

I wanted to feature community voices too. Frank reached out and said that he wanted to share his experience using OpenStack during the 5G Testbeds and Trials Programme, funded by the UK government.

In this podcast-format, we explore:

  • how the team ended up choosing OpenStack,

  • some use cases in the trial like distributed concerts and telesurgery,

  • how the importance of latency varies from one scenario to another,

  • and why Frank thinks coding is an important skill to have as a network engineer.

Agnes (00:00)

You can start maybe by telling us a bit about yourself and the kind of work you do. Yeah.

Frank Sardis (00:06)

Yeah, sure. So my name is Frank Sardis. I used to be a researcher at some point in time, a few years back. And I was mainly looking into 5G telecommunications and cloud computing and how those two work together.

How you combine the two technologies, and effectively how do you enable use cases that require ultra low latency, or ultra reliable communications.

And so in my time as a researcher, I focused a lot on these areas and developed some demonstrators and use cases, and eventually some live demos and events where we showcased the capabilities. And at the center of it all was OpenStack.

2017, 2018 and 2019, those three years are when I was kind of really working hands-on with this technology and I was responsible for all the operations.

Agnes (01:14)

Yeah, that sounds really interesting. I have read a little about the whole 5G adoption in the UK. So how is the adoption looking like right now?

Frank Sardis (01:27)

It's growing, it's going well and now we're looking at 5G advanced and we're looking at slowly the 6G standardization, right, we're entering that phase.

Although I haven't been keeping up to speed with telecoms, again, my background being mostly on the actual network and the cloud part of it. And then you have all the radio and the network functions and so on that run the radio and run all the subscriber services and so on, which is a completely different part of the picture.

But ultimately, the common denominator in all of this is that, you know, they have to run somewhere and typically they will run on some kind of private cloud platform, public cloud platform. And at the time when I was involved in these, OpenStack was sort of the weapon of choice at the time.

Agnes (02:24)

Yeah, so how did you get involved in the, you called it the test testbeds and trials program?

Frank Sardis (02:30)

Yes, that was the UK funded program. So that was a UK government funded program. the concept behind it was to build a test bed that will showcase the capabilities of 5G and therefore accelerate the adoption and generate knowledge and hopefully create business opportunities as well.

That was the general idea of it. So it was a collaboration between three universities. It was King's College London, which is where I was. We had the University of Bristol and University of Surrey. And then from our side with King's College, we had a collaboration with Ericsson who provided the prototype 5G equipment for us.

Agnes (03:19)

Interesting, interesting. So this began in a very academic setting, I may call it that.

Frank Sardis (03:25)

Yeah, it was strictly research. It was the very early days of 5G. Nothing was sort of concluded. All right? And everything we're working with was experimental, pre-release.

And we also then had our own sort of developments in the lab where we had completely experimental vendor agnostic solutions for things like function splits and testing. How you can do function splits on the radio. At what different levels you can do them? And how can you run those on a virtualization platform?

And do they work? Do they not work? Because some of them had absolutely insane requirements. Let's say, when it comes to latency and jitter, down to microseconds, for instance. So some of them naturally were not really feasible. But some of them that were a bit more kind of relaxed, they could work. So it was a matter of finding the sweet spot.

And then on top of that, we had all the other supporting technologies. How would you do things like video broadcasts? How would you transfer data for haptic communication? At some point we tried VR streaming, we tried robotics, telesurgery and so on and so forth. So we had various applications running in the cloud to support these use cases. Mainly things that would process data packets and they would instruct robots or machinery or give you haptic feedback and so on, depending on what was happening on the other side.

So it was a platform that we used to research both the connectivity itself as well as the technologies that support the use case.

Agnes (05:17)

Yeah, that sounds really interesting. You mentioned distributed concerts, and you mentioned latency. Maybe explain to someone who doesn't know what distributed concerts are how important it was for you to have them and what was the lowest latency you could possibly allow.

Frank Sardis (05:29)

Yeah.

Okay, so that's a trick question. Let's take it from the... So the music concerts was one of those kind of, let's say, events that caught more attention.

We did...a few of them in fact, some small scale, some bigger scale. The principle was the same. Some of them were between different countries. The very first one we did was between UK and Germany. Then after that we did within the UK, couple of locations. And then eventually we had a big one, which was, I want to say three locations in the UK.

I don't remember exactly where they were. There was one in London, obviously, which is where I was stationed. There was one, of course, in Bristol and I think Birmingham was the third one, which was like the bigger concert we had with this band effectively distributed in different places. So in London, we had a singer pianist performing and then in the other locations we will have the band which had supporting singers backing vocals and then drummers, guitars, bass, all sorts of other instruments.

So the first challenge is how do you connect all these locations? You can get fiber to connect all these locations. Then the question becomes, how do you ensure that the latency is consistent between the locations? Right? And then that you sort of tackle with SDN, and you kind of optimize things which we'll get to a bit later on.

And ultimately, to get to your question about the latency, so what really matters was not so much the network latency as in I'm pinging something, right? Because that's going to return microseconds, which is fine.

But the real challenge was, and that applies to robotics as well or VR, was the total end-to-end latency. So when you're dealing with audio and video, you have to take into account the latency of encoding and decoding at each end.

Frank Sardis (08:00)

When it comes down to video, and if it's VR and you're interacting with things, you have to take into account the latency on the monitor itself. Usually LCD displays and OLED displays and all that, come with like 8 millisecond latency or higher or thereabouts anyway.

So you have to account like the total latency, right? And depending on what we were doing, the latency would fluctuate. But for audio, one thing we did is for audio, wanted consistency in the music events.

So the audio latency, including sort of codec, let's say delays, was about 30 milliseconds. And for the video, we were up to the hundreds, 150, 160 or so. And then depending on what resolution we would do, if it was HD or 4K, that could grow to 300, 400 milliseconds. So there was a noticeable kind of difference between the audio and the video.

But if you kept it low as in HD at the time, there wasn't much difference, it wasn't noticeable. So you had very, good quality of video and very, very good quality of audio and they were synchronized.

Agnes (09:21)

Wow, that's interesting. I can imagine attending a concert in like five locations. The instrumentalists are in different locations from the background vocalist, from the lead singer, and having one seamless experience. I think that would be really cool to see.

Frank Sardis (09:39)

Yeah, I from the audience's perspective, you would have essentially the main singer in front of you. So the way we run it is the venue in London where we had the singer and pianist is where we also had the audience. And then the rest of the musicians were effectively in a studio environment, right? So they were streaming from studios in these locations.

And so the audience was just in one place. It wasn't like an audience in different locations that each one would see different band members. But yeah, it's entirely possible if you wanted to do that sort of thing, you could have audience at every location because every location had a video feed and audio feed so they could hear each other play and see each other as well.

Agnes (10:29)

Great. So I'm wondering, what kind of thinking went into this? What kind of tech we need for this, and where does OpenStack play into all of this?

Frank Sardis (10:41)

So the selection of OpenStack came long before any of this. I was running different projects at the time. One of them had to do with virtualization and software-defined networks for energy, renewable energy.

And it had to do with how do you guarantee reliability. How do you guarantee quality of service? So from early on, I had sort of settled on ⁓ the usual kind of we had Linux in the lab and then OpenStack sort of became a natural choice at the time, mainly because our partners were also kind of working on OpenStack and everyone was familiar with the platform.

So it was...the natural choice there. Now, from there on, yeah, we already had that, but once we got into the 5G Testbeds and Trials Programme, we took everything we had in the lab, threw it out, brought in entirely new equipment, much more high-performance equipment in terms of servers, CPUs and memories, hard disks and so on.

And at some point, if I remember correctly, that small data center we had, had over two terabytes of memory across the nodes and over 2,500 cores in total. So it was a decently sized, let's say edge cloud, if you will, or on-prem cloud.

And essentially, that's the back story of it because the idea was once we got into the 5G test results and trials, we knew we had to go for really high performance.

We knew that we had some events coming up, and we needed demonstrators and we had to go outside the lab effectively.

So the entire design started from pretty much down to the basics of what fiber do we need? Just reinstall fiber, bring all the new equipment in, the data center, what cooling do we need for the data center, what power do we need for the data center, from the ground up design. And it all had to happen within six, seven months is all we had.

Frank Sardis (13:05)

⁓ So all the equipment came in and then again, what would you choose there? Well, we were familiar with OpenStack already. So my choice was let's do OpenStack again. And we knew that OpenStack is not going to be a problem from previous experience with testing in other projects. But also because largely when it comes to performance of these functions that we were hosting in OpenStack, it's not OpenStack itself that...

Frank Sardis (13:33)

deals with it. It's all the networking and the underlying hardware that need to perform and the OpenStack is there as the middle layer to manage the VMs and deploy things or destroy things when you don’t need them anymore. And so yeah, it was the natural choice as I said before.

Agnes (13:43)

Yeah, that sounds interesting. Sounds like it was an easy sell to the rest of the team.

Frank Sardis (13:58)

Yeah, it had all the features we needed. In fact, it had more features than we ever needed. And towards the end of the testbeds life, I started implementing more and more things as well in OpenStack. So at some point, we were able to have live migrations of virtual machines between nodes for failover purposes.

But then I was able to combine that with other things. Again, being a research setting, we were able to demonstrate that you could move a VM from one node to another based on a specific user's location. And if that node was attached to a location and the user moved to another location, then that VM that belonged to that user could move to a node that is closer to their location, still managed by OpenStack.

And at some point I even had one of the OpenStack nodes running at my place at home over VPN. So when I would come at home, I could just move a VM over the internet to my home computer. But yeah, that was just an experimental setup.

Agnes (15:07)

Wow, that sounds incredibly complex, but so much fun to implement. Yeah, I know OpenStack has a lot of services. There's a whole lot of networking, compute. What were the main services that you worked with the most?

Frank Sardis (15:12)

Yeah.

⁓ the core ones, right? So we'd have like, you can't avoid Keystone. You need that for identity, right? Then you have all the usual Neutron, Nova, you know, for the scheduling of the VMs, the networking, the storage one, which I forget what it's called now. And then we have...

Agnes (15:40)

There are several for storage.

Frank Sardis (15:47)

We had glance obviously again for the images to deploy the VMs and I had Horizon as well just to have a pretty GUI type all the time in the terminal. And that was it. That was the core services really. It was a very, very lean deployment. A lot of the optimization was again, the underlying hardware like bias configurations, Linux kernel configuration as well.

How to do with things like huge pages, CPU, core isolation, things like that. Which again, I can go into detail if you want.

Agnes (16:30)

Yeah, I think that's fine. I would imagine the OpenStack community knows the different services that they use. Yeah.

So I'm wondering, as you are setting up all this hardware and managing it with OpenStack and just trying out all these things, one of the other things that you mentioned was telesurgery.

That sounds very risky. know, there's life involved. So what were some of the failover or recovery options that you thought about as you were working on that?

Frank Sardis (17:08)

Yeah, okay.

Yeah. So we didn't try it on anything alive. When we were experimenting with it, we were experimenting with robots developed by PhD students in the medical school. So they would provide, I don't remember what they were called, but they were kind of...

It was like a snake. It was effectively a catheter that could move. It was very flexible. And so we were controlling these with various devices. We could control it effectively if we wanted with a keyboard and give it instructions of where to go.

Eventually, what we did is we had haptic gloves, some prototype haptic gloves that we procured and then we were able to use the haptic gloves with hand gestures, control the robot and then the great advantage of it is one of our PhD students at the time developed this solution where if the robot found an obstacle or if the robot detected an anomaly you would receive vibrations in your fingers to indicate that, you know, there's an obstacle there or there's something that shouldn't be there.

The idea behind it being like a catheter, this soft robot, was that it would go inside your body and sort of like in a minimally invasive way feel let's say your kidney and see if there's a tumor there.

So as it passed over the tissue, if there was a tumor there, it was slightly harder, it would then transmit that back to your fingers so you could sense that there's something there. And that was the idea of it.

So we had various ways of detecting that. We had a sample of a silicone implant that we put a plastic bead inside and then the robot would go over it and detect if the bead is there or not and would transmit it to our fingers, then we could use our fingers to move it around. in all that, so on one side, if you can imagine, we had a doctor, let's say, right, the operator where we would wear the gloves and had the laptop and that would connect to the infrastructure. It would transmit oral gestures and all the packet data effectively going into the cloud, into OpenStack where we had a VM process the data and then send it to the other side to the robot where the robot would then translate those motions into, sorry, those gestures into actual movement and where it needed to go. And then in turn, would transmit back data to us so that we can feel whether there's an obstacle or anything else.

Agnes (19:58)

Wow, so you had to work with the actual doctors who are PhD students at the time.

Frank Sardis (20:05)

Yeah. Yeah.

Agnes (20:07)

That sounds interesting. So, I feel like what you guys did in the team was we know OpenStack. Let's push it and see how far we can push it. Were there any surprises, you know, as you implemented, you're like, we didn't know this could happen. We didn't know this kind of capacity was there.

Frank Sardis (20:25)

Oh yeah. I mean, it had multiple layers, right? And there were kind of human errors, 99 % of the time it was human errors in the sense that there would be a typo in the configuration file somewhere that when you type it, you don't notice it. And then you go there. Initially it looks fine. And then you start deploying things and errors all over the place. Why did this happen?

So you go into the error loads, you're to figure out what it is, then you go to your settings like, there it is. There's a typo there. That's what Trigger did. So you had mistakes like that, then you had more serious design mistakes initially

And because of the time constraints, I approached it in a very kind of naive manner. So the network was very simplistic. It was just kind of a flat topology, everything essentially in the same broadcast domain.

Well, turns out not a great idea, right? So when you're going for high performance, you want to isolate traffic as much as possible. You want to start implementing traffic control rules with SDM controller. So eventually I had to rebuild the entire network and we had dedicated physical networks for all the OpenStack overlay networks for Neutron.

Then we had dedicated physical networks for SSH and management of the nodes. So none of that traffic is going to be on the actual network with the user data, right? And then we had things like the *API's open stack* for management, various other things that again, they would go into their own independent networks. And everything was 40 gigabit.

Some of the networks that didn't need it, they could be a gigabit, some of them could be 10 gigabit. In terms of redundancy, everything was...all the links were kind of aggregated, so we could have failures, two links at minimum, sometimes four links paired together and in some cases, we also had redundant paths as well. So if a whole switch would go down, then there was an alternative path that we could use. Thankfully, we never had to go into that sort of failure mode.

Agnes (22:55)

Hmm, interesting. So there's a joke that the human is always the weakest part of every system. Yeah, so what would you do differently if you were to do this again?

Frank Sardis (23:04)

Yeah

⁓ If we're talking specifically about the OpenStack deployment, I would most certainly take the time in the beginning to design it properly because we were sort of two months into the project when I realized I needed to redesign this network and make it a bit more, you know, let's say professional, a bit more serious for this kind of use cases, right?

This became apparent especially when our partners started bringing in their equipment that needed to attach to the testbed and then you had all the prototype equipment, you had experimental equipment, and other people were coming in who wanted to try use cases.

Everything was just flat, right, so there was a risk of security as well. Why would you have everything flat where everything can access everything, right? Everything can talk to each other if they wanted, so we started isolating things at the time and being more careful but also then, kind of having that fine-grain segmentation so that we can control the traffic effectively.

I would definitely, definitely, definitely save myself a lot of stress if I had done it properly from the beginning. So that's the main thing I would say. The other one would be, I'd probably, let's say, do more coding because during the entire program, my role was to design the testbed, build the testbed, make sure the testbed runs, then help everyone bring their stuff in and make it run.

And a lot of the time, because I knew the capabilities of the testbed, I knew also what people wanted and I was trying to bridge the gap if there was a gap, but that effectively meant that I wasn't there participating with the development of the applications, right? So I didn't write any code for the robots.

I was sitting with the people doing these things and they would tell me I need lower latency here. Can we optimize that? So at which point, you you start looking at what's happening on the network. What's the traffic like, how many packets per second, so on and so forth. And you start looking at window sizes. You're looking at frame sizes, you're looking at empty use and all these things. And you're trying to optimize it at that point, to make the use case work really as best as it can.

And then the other thing that's interesting, of course, many use cases had contradicting requirements. Like video seemed to work very well if we could have a large empty use or a whole frame could fit in one packet and send it out. But then other use cases wanted smaller packets just to maintain that low latency and keep sending packets.

So we had that, that challenge as well. And there was a bit of reconfiguring things. Then we had, I remember one big surprise was a few days before one of the big, big, big demo events. I don't remember which one it was because we ran several, but we connected to the test beds with our partners, Bristle and Surrey.

Nothing was communicating. And the reason for that is because I had set my MTU to something like 8,000. I was using jumbo frames, particularly because I was testing video and then everyone else wasn't. So immediately that was a problem. And we're just scratching our heads, what is going on here? Why is this happening? And I realized, okay, well, wrong MTU sizes. And then once you get to that...

That means now, change everything in VMs, OpenStack, the hosts, the nodes, and the switches, everything, all the interface. So that was okay. Now let's go back manually, change everything. So not fun, but still quite an experience.

Agnes (27:24)

So I'm trying to understand why you said coding would have been important. Would it have helped with like automating? Like if you're making a change, like the MTU change, just make it in one place, not 10 or 15.

Frank Sardis (27:38)

That was definitely one aspect of it. The other aspect of it was infrastructure as code, which I started doing again later on towards the end of life of the test bed. I started using heat a lot more to automate the deployment of virtual machines because we had, know, wasn't, it was rarely just one VM to be deployed. was usually a bundle of virtual machines that had to be deployed and each one was doing a specific function.

Sometimes we wanted to deploy a whole cluster of virtual machines that would host Kubernetes and then would have containers. Initially, because of the time constraint, during the early life of the test bed, all that was done manually because it was a one-off thing, but also because the requirements came in and somebody told you, you need to have that ready in a few days around the first test.

You just wanted to start deploying as quickly as possible, get it up and then ask people to test it, see if it works and then go back and make adjustments as required to help everyone. Oftentimes, yeah, I would work with the developers. I would work with people that did the robot programming. I would work with people that did the function splits for 5G, VNFs, but yeah, I wasn't the one writing the code for the most part.

I wrote code for my sort of own pet project on the side, which was a little robot. It was a little Rover robot where I had all the AI, all the intelligence running on OpenStack in the cloud. And then the robot was simply sending signals and receiving instructions from the AI running in the cloud. And then we were able to play with the latencies and the bandwidth to see how those would affect the performance of the robot as it was operating.

And you would see things, for instance, if the robot goes slowly and you have a high latency, so when it received the signal that there's an obstacle, it would transmit that signal back to OpenStack where we had the VM to process that signal and basically tell it to stop so it doesn't crash. So if it was moving slowly and you had a high latency,

It would transmit the signal. The latency was high, sure. But by the time it got to the cloud, got processed, and then sent back to the robot, you had sufficient time to actually make it stop. But if you still had that high latency and the robot was traveling really, really quickly, by the time it got the signal from OpenStack to tell it to stop, it would have already crashed because it would have covered much longer distance at the time. So there was this kind of correlation between the speed of the movement and the latency.

And we did a demo as well, eventually, where I was able, depending on the speed that you would tell the robot to move, it would also send a signal to the SDM controllers to prioritize the traffic accordingly to ensure that no packets are lost or we get lower latency, especially if there is other traffic on the network that could conflict and, you know, kind of...contend for resources.

Agnes (30:57)

So what I'm hearing you say is when you're designing, especially complex architectures, you would rather spend some time designing for scale, even if you're not there yet. But on the other hand, there's the issue of time, you know, you had seven months, you couldn't spend three months just saying we're designing, we're designing for scale. So how would you work that balance if you were to do it again today?

Frank Sardis (31:24)

It's again going to be dictated by the timeline, and sort of budget, and the number of people you have, and how quickly does it need to be delivered. How complex is it? What is the future plan for it? Is it a long-term thing? Is it the one-off thing?

There is no point in building something to scale if you know that it's only going to be within that very specific use case and then it's going to be discarded and repurposed for something else. So you need to choose your battles, but if you do know that there's a long-term plan for whatever it is you're building, then I would say measure twice, cut once.

Take the time, and even if this is an important trade-off basically, but a very important decision, even if you have this time constraint where somebody may tell you well you need to have let's say the prototype the the very basic thing running in three four months six months whatever, right?

Then you're thinking, okay, I'm in a hurry, I need to do it as quickly as possible, maybe I need to cut corners. And then that sacrifices some of the long-term because then once you get out of that, you have to go back and say, well, okay, the prototype works, but it wasn't really designed for scale, now I have to redesign everything.

So you have to choose your battles there very carefully and say that, you know what, okay, maybe it is seven, eight months, but I think it's worth taking the time now to get it right.

And if you can't it, like if you can't meet the deadline, maybe you can try to extend it. If you cannot extend it, then choose specifically what design you want to have in place. And even if it's not necessarily a perfect design, then at least get the core components in place.

So for instance, in my case, I could have easily done a proper segmentation of the network, proper design of the network, and then still bringing all the nodes and all the experimental equipment and attach it like in a flat way, but then all I had to do was redesign, sorry, no, reattach it to different places on the network and redo the routing.

But the way I started, everything was flat by design at the hardware level, right? In terms of connectivity. So there was nowhere I could go. So this is one mistake you can make, right? And then you say, OK, well, maybe design the network properly, but don't worry so much about where everything attaches for now. Get it running.

And then once you go through that first phase, the seven, eight months, the proof of concept that everything is there, then unplug everything and then put it in the right places. But at least you will have the foundation there in terms of design and everything will be waiting in the right place to be connected.

Agnes (34:35)

So it's a trade off of time, of resources, and let's just get this working. Once it works, we have proofed the concepts, then we can think about scale. Interesting. I've also heard you talk about when you thought about latency within like a video streaming context, vis-a-vis a different context. So does it mean that latency is more of a use case thing? It's not a one size fits all kind of thing.

Frank Sardis (35:04)

Absolutely. yeah, absolutely. So a very good example of that is when we were experimenting with cloud gaming, effectively streaming games from OpenStack, we had some nodes that had GPUs, so we were able to stream video games. And you have slow video games like strategy video games, and then you have very fast-paced video games, so racing games, shooting games, that sort of thing.

And the latency is vastly different, right? One can survive in terms of like streaming the images, one can survive with 30 milliseconds, the other one doesn't, it needs like 10-15 for example, right?

And once you start bringing VR into the mix, then it's not so much about the interaction of the player with the game, it's also the human factor and the way our brain works because if you move your head around, but then the VR image is slow to react, you will start feeling sick. You get nausea, right?

So that becomes then kind of a hard requirement and it doesn't depend on anything. Doesn't depend on whether or not you're playing a slow game or a fast game in VR. It's more of a, you know, I need it to run. Otherwise people will not feel well. User experience will be terrible anyway.

So yeah.

Agnes (36:36)

Wow. So I'm wondering, after all this work that you did with the test bed trials, are there actual companies or businesses or organizations that you can point to and say, yeah, this is the direct result of the work that we did?

Frank Sardis (36:55)

A direct result of the work we did there…there is, I'm aware that there is something happening with telesurgery. It's been developing for a few years, but again, I have been out of the game since 2018 now. No, sorry, I'm lying, 2019, end of 2019.

So I haven't really followed up on the research and all the latest developments, but I am aware there is something happening with the tele-surgery and it's being explored in the bigger context of 5G and how, like to go back to one of your earlier questions, how do you provide reliable connectivity if a link goes down, right?

So you can have 5G as the failover link and 5G can give you, depending on the setting, but it can give you very low latency as well. And that's how we did these events as well.

With the distributed music concerts and the robotics, the 5G connectivity was able to provide low latency for it. Because we had access to 4G as well, and with 4G, it wasn't even close. It was almost double the latency.

Having said that, it was also a matter of how do you optimize 5G for it? So if you try to do the same thing with just your standard public 5G connectivity now, it's not going to work.

You need to actually change the schedulers on the radio and make sure you get the right traffic prioritization and so on. The network slicing is another term that people are talking about and it's part of this story. So I know this is happening with sort of telesurgery.

And then from there on the telecom side, the direction the industry is going now with edge clouds and virtualizing the network functions, this has happened already with 5G. But now the talk of having AI, for instance, and AI agents at the edge to manage the radio functions, all these will have an underlying cloud platform. And it's...effectively like the edge clouds will be OpenStack or Kubernetes or a combination of sorts.

Any other work now that comes to mind?

I know there is talk about autonomous vehicles, right? And again, you have this sort of vehicle to infrastructure connectivity where you would have edge clouds and the vehicles in an area, self-driving vehicles in the area, would be able to share information with each other via the infrastructure, via that edge cloud. And that edge cloud would actually send an emergency message or instructions that there's a blockage here or avoid an obstacle coming up ahead and so on. So you have use cases like that.

But it's a world of opportunities out there when it comes to Edge Cloud and its applications.

Agnes (40:09)

Okay, so I've been seeing a lot of people talking about the de facto tech stack for OpenStack is really Linux, Kubernetes, and OpenStack. What would you have to say about that?

Frank Sardis (40:23)

Yes, I mean, it's what I did. It's what I did as well. And one of the interesting things that I wasn't really able to take advantage of at the time is the standard way of using everything I described, the standard way of deploying things was we had OpenStack deployed and I basically deployed all the OpenStack components manually at the time.

So just, you know...all the installation of packages, configuration, everything was done manually. And then we would have Kubernetes running in VMs as needed in OpenStack. But then later on, you were able to deploy OpenStack using Kubernetes. So it will be Kubernetes on the bare metal to deploy OpenStack components.

That I didn't get a chance to play with, but I think that is great because you can do so many things and it can improve the reliability of OpenStack components as well. It has a lot of advantages.

So I totally agree. I think they complement each other. There's kind of a tendency now where people go say to Kubernetes and they use it as sort of their infrastructure manager and they deploy applications. But OpenStack brings you that infrastructure as a service kind of functionality that Kubernetes does not natively have. And it gives you isolation at the kernel level as well.

If you want to, and not if you want to be, it's part of the design but, if you need it, it's there so this is normally how I would do things as well unless I know there is a very specific kind of use case where I don't need infrastructure service and I can just go bare metal Kubernetes.

Agnes (42:13)

Yeah, so right now AI is making everyone worried. It's coming for jobs, know, and we're trying to analyze what will I be doing the next five years. For a professional with the kind of skill set that you have, networking, Kubernetes, OpenStack. Are you worried? Are you worried about AI?

Frank Sardis (42:34)

I'm not worried about AI, I'm optimistic, very very optimistic about AI, I like some of the things that I'm seeing. On the social impact front there are concerns, very very understandable concerns, very justified concerns as well.

Because essentially it can behave as an assistant, it can function as an assistant that everyone has, and therefore now you need fewer hands on deck, let's say, to do a given task. So I totally understand that.

On the technology side...in terms of what it can do for technology and again the use cases go from AI to manage the network, self-healing of networks and so on and so forth.

So the opportunities are absolutely endless, and they can have a massive impact if you have this sort of self-healing networks and self-managed networks.

One other thing that I personally like is the intent-based networking that you can do.

You can, with SDN, and it was a thing before AI, but you could just give it an intent and say I want high quality of service between this application and that backend service.

And then the SDN controller will compile everything and figure out what's the best path and everything. So now you can have an AI in the loop to figure that out. So you don't need to kind of hard code it in the old fashioned kind of symbolic AI as we would. You can have the…generative model there or some sort of machine learning there that will figure out how to do things. But it needs to be supervised

Agnes (44:23)

If you were to advise your younger self, what would you tell them?

Frank Sardis (44:28)

That's a good one.

More programming, mainly because I got to the point where...the networks basically became software defined. So the old fashioned network engineer was less relevant and it was more about your programming skills.

Cause now you could go in and start programming the STM controller to do various things. And...programming skills become very, very useful there because knowing the network stack and how everything works is one thing, but then if you cannot program an SDN controller, you have to rely on someone else to do it, which is what I was doing a lot of the time as well.

But yeah, it would have been great if I had more hands-on programming for like...a lot of these things and not only just for the use cases, but on the SDN controllers and so on, all the other supporting technologies we have.

And therefore, I would say, get a bit more into programming, even if you are considering yourself to be a sort of network and infrastructure person, programming skills are sort of mandatory at this point.

It's not just for infrastructure as code, which is basically configuration files, but it's because the underlying infrastructure now, things like OpenStack, things like SDN controllers. If you want to get to that low level of building things, then you need to have a programming skill, or you need to find good programmers and then tell them exactly what to do.

For me, it was a good experience. I would have loved a bit more hands-on involvement with coding. And I would definitely advise myself earlier on to just get a little bit deeper into programming.

Agnes (46:39)

Yeah, that's an interesting perspective. Yeah, that's a really, really interesting perspective. So yeah, so maybe the last question, it's not a question, just a comment. If people wanted to follow you or learn more about the test beds and your work, OpenStack, where should they talk to you?

Frank Sardis (46:59)

LinkedIn, it's, it's everything is there. That's the main hub. Yeah. It's everything on LinkedIn.

Agnes (47:03)

Great. Sounds like you had a very interesting experience with the test beds, OpenStack, just building out the networks. Yes. So thank you so much for taking the time to come and share your experience and hopefully show the world and other people who are starting now that you can actually build real world things with OpenStack and open source technologies. Of course, we know we run on open source, but you kind of leave this in the background.

Yeah, so it is really, really interesting to learn about all the things that you worked on, and yeah, so I hope this will be that people will find valuable.

Frank Sardis (47:53)

Yeah.

Agnes (47:55)

Alright then, I don't know if there's anything from your end.

Frank Sardis (48:01)

No, nothing, to add. Like I said, I've been out of the game for past few years, but I do keep an eye whenever I can on OpenStack and I see it has progressed so much more and so many things have changed. And I believe now it's going to be an even easier thing to get into and simpler, let's say, but also more powerful, have more features and more functionality.

I would, if anyone wants to learn how a cloud actually works and what's happening in the background, how do things communicate, how a VM gets created, how the configuration is passed to the VM, how the overlay network's working, how's the routing done.

I mean, OpenStack is going to give you full access to all that knowledge. If you spend time with it and you learn how it works, then you start learning a lot about how clouds work in general.

Agnes (49:08)

Yeah, that's true. That's true. I'm curious, did you ever contribute to OpenStack?

Frank Sardis (49:15)

You mean code? No, no, no, we did not contribute code to OpenStack.

We did, however, make minor changes here and there. Simple things like if you wanted to have some additional information from the CLI.

Then we would go and add a few lines of code that would, you when you run the command, it would also give you something additional. This is something that a lot of people did, a lot of vendors did, that use OpenStack or generally any other open source platform. You'll see that they kind of extend the functionality with their own function.

So, maybe if you get the vanilla open source version of OpenStack and you type a command, you're going to get a certain output. But then if you get a vendor specifically, let's say version, they type the same command, they're going to get additional information that you don't.

And that's basically modifying a little bit the code to make the appropriate calls in the APIs to give you additional information. So very light modifications of that sort, but we didn't go deep in the sense of, let's write a new scheduler or anything like that.

Agnes (50:28)

Great. So you did extend the functionality.

Frank Sardis (50:32)

Yeah, I wouldn't even call it that, but yeah, if you wish to call it that.

Agnes (50:39)

Okay, yeah, so it's been really, really great to learn from your experience and I hope that more people will also learn and try OpenStack out and use it even in their real world use cases.

Frank Sardis (50:55)

Yeah, absolutely. It's a great to many kind of on-prem solutions. So it's definitely a good candidate for on-prem solutions if you want your own private cloud. So highly recommend it. Yeah.

Agnes (50:57)

All right,

Awesome. Just one thing that came to mind just before we finish. The whole cloud repatriation movement. What's your thoughts on it?

Frank Sardis (51:22)

In my experience, it depends on what you do. So what I mean by that, there is definitely places where you have the performance requirements, the scale requirements, the operational requirements to be using a public cloud.

But then there are cases where you're better off running your own sovereign cloud on-prem and have complete independence. It could be cheaper as well, depending on what you're doing.

If I want to be training machine learning and I have the space and I have already the infrastructure, maybe if I don't have the infrastructure, maybe over the time that I'm going to be running the simulations or machine learning, training models and so on and so forth, maybe the cloud cost will justify me just buying the equipment and doing it independently, right?

So it's not, it's a question where basically the answer depends on what you're actually doing. I mean, there's a whole...like the whole concept of FinOps behind it, which is like, know, money wise, it like, how do I reduce my cloud costs? How do I make intelligent decisions about how I deploy things in the cloud?

And then at some point you sort of ask yourself the question of, you know, should I be in the cloud or should I be on-prem and independent based on what I'm doing and based on, you know, the costs and any other.

There's regulatory as well. There's a lot of factors, right? So yeah, I don't think there is a good answer that captures the details of it. But yes, a lot of people do realize that maybe, you know, we're better off on-prem.

Agnes (53:27)

Okay, then. So thank you so much for sharing your experience.

Frank Sardis (53:32)

Thank you. Very good to meet you. Yeah.

Agnes (53:34)

Yeah, thank you so much.

Thank you so much Frank. Okay, bye.

Frank Sardis (53:36)

Thank you, guys. Take care. Bye-bye.

No posts

Read the original on techiewhowrites.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.