Thank you for Subscribing to Telecom Business Review Weekly Brief

From Firefighting to Engineering: How Great Infrastructure Leaders Build Teams that Can Outgrow the Problems they were Hired to Solve


Infrastructure is successful when it becomes invisible and most visible when it fails.
An outrage, a system performance issue, a network capacity problem or an unexpected failure can instantly turn an intentionally planned day into firefighting chaos. The engineer who knows the system best gets called. The incident bridge opens. People start troubleshooting. And eventually, someone saves the day. We celebrate these moments and rightly so. But there is a danger in building a culture around them. A team that repeatedly saves production may look exceptional. A team that systematically eliminates the reasons production needs saving is exceptional. Over the years, working across financial services and gaming, building high-performance systems and low-latency networks, operating large-scale infrastructures, I've come to believe that one of the most important responsibilities of an infrastructure leader is not simply to solve difficult technical problems. It is to build a team that eventually outgrows the problems it was hired to solve. These three core principles have been guiding me over the years and every team I worked with will recognize them clearly. You get to be a hero, but only once. Every leader knows the feeling when the subject matter expert enters the incident chat. It is the feeling of constant relief because you know the production outage is going to be soon resolved. You have your hero in the room. But your impact as a leader will be determined by what you do afterwards. If only one person understands a critical system, knows how to recover it or can navigate its undocumented complexity, that person has become a single point of failure. This is where you start leading. Documentation becomes no-negotiable. So does automation. So does pairing engineers on difficult problems and deliberately giving others opportunities to operate systems outside their comfort zone. A great infrastructure leader makes sure their team learns and matures. That might mean redesigning an architecture, automating a manual process, improving observability, eliminating a single point of failure or simply making sure the knowledge no longer lives in one person's head. The Tech Debt Bucket Each infrastructure has a story to tell and some have many skeletons to hide. Systems are built over years, dependencies deepen organically and once new and modern architectures become legacy systems. Then come the decisions made for the problem of today: the access-list line that solved an outage once, configured in isolation by someone who long left the company, that ends up being someone else’s two-day troubleshooting problem when deploying a new service. Then come the shortcuts, the cut-corner engineering decisions, the deferred upgrades, all justifiable when taken in isolation, but contributing to the hidden danger of feeding the technical debt monster.Infrastructure is successful when it becomes invisible and most visible when it fails.