
Site Reliability Engineering Professional
Job Description
About the role
Our ambition is to create brilliant, converged services over the best network, aligning across OR&SM (Operational Resilience and Service Management), Digital Operations, Fixed Networks and Mobile Operations. The Service Now Platform has a unique position in BT by helping all customer facing units and ensuring outstanding customer experience is delivered through our IT, Fixed and Mobile networks, across all our products.
The SRE role will create, build, and maintain automated deployment pipelines, implement intelligent monitoring and event management solutions, and leverage AIOps capabilities to detect anomalies, anticipate failures, and accelerate incident resolution. Working closely with development, platform engineering, security, and operations teams, the engineer will drive continuous improvement across the software delivery lifecycle, enabling faster and safer releases while maintaining high service reliability.
Key responsibilities include managing CI/CD pipelines, automating infrastructure and operational processes, improving observability through metrics, logs and traces, developing self-healing capabilities, analysing system performance, helping incident and problem management processes, and championing reliability engineering best practices. The role also contributes to cloud platform optimisation, DevSecOps adoption, and the implementation of service level objectives (SLOs), service level indicators (SLIs), and error budgets.
What you’ll be doing
Actively lead the ServiceNow platforms across BT, proactive monitoring of system health and performance resolution of system defects, delivery of service requests
Capable for full ownership, innovation, and continuous improvement of operational processes
Constantly engage with service teams to develop intelligence and maintain a high technical standard and availability of the platform. Fully covering progressive dev/ops methods of working.
Decision Making - Gathers information, and analyses different scenarios, assesses alternative resolutions and reaches a decision.
Proactively lead on the end-to-end health of the IT services, by using the proactive tools to monitor customer journeys
Manage incidents and planned outages within a regulated environment
Design and implement automation for monitoring, remediation, deployments, and operational tasks
Build and maintain ServiceNow with CI/CD pipelines and enterprise tools
Support DevOps and SRE adoption across teams
Deliver on AIOPs initiatives to automate incident resolution and to proactively resolve platform issues
Essential Skills / Experience
Hold ITIL Foundation or similar certification and confidently display knowledge of incident management practice.
Understand how incident management connects with change, problem, and other ITIL practices.
Convey clearly through professional verbal and written updates during incidents and coordination activities.
Understand diagnostic tools like NetBrain, TSNA, Ansible, ThousandEyes etc. for troubleshooting help.
Prioritise tasks effectively and collaborate well with internal teams and external partners.
Demonstrate team management skills that helps engagement, performance, and operational stability.
Maintain a strong customer‑focused mindset to ensure helpful, timely, and professional service experiences.
Demonstrate soft skills like proactivity, ownership and flexibility as part of professional maturity.
Strong experience in SRE, production help, or platform reliability roles
Our Package
BT Group is the UK’s leading communications group and the holding company behind some of the country’s most recognised brands – including BT, EE, Openreach and Plusnet. Our purpose is as simple as it is ambitious: we connect for good. Our customers include consumers, small, medium and large businesses, public sector organisations and other communications providers.
BT Group’s role is about setting direction, unlocking value and creating the conditions for our brands and businesses to thrive.
Having come through the most capital-intensive phase of our fibre investment, our focus now is on what comes next – simplifying how we operate, using technology and AI to work smarter, and organising ourselves to serve customers better and grow sustainably. Group teams shape strategy, policy, brand, capital allocation and transformation, helping the whole organisation perform at its best.
We have a singular culture that unites all our people: we are customer-first challengers, who are committed, clear and connected. These behaviours unite us as one team to deliver for our colleagues, our customers, our stakeholders and the country. Joining BT Group means working at the heart of a business that matters to the UK, with the opportunity to shape decisions, influence outcomes and help set the future course of one of the country’s most important companies.
Other jobs you might like
Site Reliability Engineering Professional
Bengaluru, IN
Site Reliability Engineer
Aalborg, DK
Working at BT Group

3 office days / week

A little flex time

