A featured contribution from Leadership Perspectives: a curated forum reserved for leaders nominated by our subscribers and vetted by the Telecom Business Review Advisory Board.

Blizzard Entertainment

Catalina Cicioiu, Sr. Manager, Production Network and Data Center Engineering

From Firefighting to Engineering: How Great Infrastructure Leaders Build Teams that Can Outgrow the Problems they were Hired to Solve

Catalina Cicioiu

Catalina Cicioiu

Infrastructure Leadership Architect

Infrastructure is successful when it becomes invisible and most visible when it fails.

An outrage, a system performance issue, a network capacity problem or an unexpected failure can instantly turn an intentionally planned day into firefighting chaos.

The engineer who knows the system best gets called. The incident bridge opens. People start troubleshooting. And eventually, someone saves the day. We celebrate these moments and rightly so. But there is a danger in building a culture around them. A team that repeatedly saves production may look exceptional. A team that systematically eliminates the reasons production needs saving is exceptional.

Over the years, working across financial services and gaming, building high-performance systems and low-latency networks, operating large-scale infrastructures, I've come to believe that one of the most important responsibilities of an infrastructure leader is not simply to solve difficult technical problems. It is to build a team that eventually outgrows the problems it was hired to solve. These three core principles have been guiding me over the years and every team I worked with will recognize them clearly.

You get to be a hero, but only once.

Every leader knows the feeling when the subject matter expert enters the incident chat. It is the feeling of constant relief because you know the production outage is going to be soon resolved. You have your hero in the room. But your impact as a leader will be determined by what you do afterwards.

If only one person understands a critical system, knows how to recover it or can navigate its undocumented complexity, that person has become a single point of failure. This is where you start leading. Documentation becomes no-negotiable. So does automation. So does pairing engineers on difficult problems and deliberately giving others opportunities to operate systems outside their comfort zone.

A great infrastructure leader makes sure their team learns and matures. That might mean redesigning an architecture, automating a manual process, improving observability, eliminating a single point of failure or simply making sure the knowledge no longer lives in one person's head.

The Tech Debt Bucket

Each infrastructure has a story to tell and some have many skeletons to hide. Systems are built over years, dependencies deepen organically and once new and modern architectures become legacy systems. Then come the decisions made for the problem of today: the access-list line that solved an outage once, configured in isolation by someone who long left the company, that ends up being someone else’s two-day troubleshooting problem when deploying a new service. Then come the shortcuts, the cut-corner engineering decisions, the deferred upgrades, all justifiable when taken in isolation, but contributing to the hidden danger of feeding the technical debt monster.

Infrastructure is successful when it becomes invisible and most visible when it fails.

I learned the same lesson over and over again: The shortcut you take today will become tomorrow’s incident. The one-time patch will become next year’s rework project. Reading “The Phoenix Project” in my early team lead days helped me visualize the consequences of not being intentional about the Tech Debt Bucket.

The biggest danger begins when technical debt becomes invisible, when teams normalize workarounds, accept recurring incidents or continue investing in temporary fixes because there is never enough time to address the underlying problem. A great infrastructure leader makes technical debt visible, measurable and part of the engineering roadmap. Not every piece of debt needs to be eliminated immediately, but every significant piece needs an owner, a risk assessment and a decision. Technical debt is not a failure of engineering; unmanaged technical debt is a failure of leadership.

The Focus Week

Every team operating a live infrastructure environment, can easily become trapped in operational work: changes, escalations, troubleshooting, maintenances and then doing it all over again. But people don't grow simply because they have more work. They grow when they have ownership of meaningful problems.

You can help your team break out from the routine, designing intentional time to reset. These are your focus weeks. Think of them as intentionally and meaningfully planned breaks from the day-to-day, allowing new and creative pieces of work to take over. When we first did this as a team, I got to see the energy shifting, the excitement growing, the multitude of ideas, imagining new tools that could evolve into real engineering solutions. Some engineers got to surprise themselves with the prototype they built by the end of the week. But besides the tangible results, it allowed space for the team to come together, look deeper at the underlying problem and create an engineering solution that will solve and improve the day-to-day operations for the long run. 

I feel very fortunate to have had the chance to learn from great mentors, who helped me sediment these three core principles, supporting my ultimate goal as an infrastructure leader: to create an organization that can continuously improve and build teams capable of facing problems we haven't anticipated yet.

Moving knowledge from people's heads into systems. Moving decisions from individuals into clear principles. Moving repetitive work into automation. Moving operational lessons into architectural improvements. And moving engineers from executing tasks to owning outcomes.

And that, to me, is the difference between firefighting and engineering. Instead of asking: How do we get through this? we are asking: How do we make sure we're better prepared for whatever comes next? And great infrastructure leadership goes one step further: How do I build a team that can answer that question without me?

The articles from these contributors are based on their personal expertise and viewpoints, and do not necessarily reflect the opinions of their employers or affiliated organizations.