Location:- Bengaluru
Last date of application:- 22-May-2026
|
About the role |
|
As a Site Reliability Engineer (SRE) within the Network Operations team, BTI International, you will be responsible for ensuring the reliability, resilience and performance of our Global Platforms including Global Fabric. You will collaborate closely with Engineering, Product and ASG teams to embed SRE principles such as automation, observability and proactive incident reduction into day‑to‑day operations. By improving how we monitor, maintain and evolve our services, you will help reduce risk, improve service quality and increase operational efficiency. Through this role, you will support BTI International’s strategy by enabling stable, secure and scalable platforms that support business growth, accelerate delivery of new capabilities, and protect customer experience. |
|
What you will be doing (Role Accountabilities) |
|
|
What you’ll need to succeed (Skills & Experience) |
|
|
BT Group’s Behaviours |
|
|
Customer First |
Prioritize customer needs in every decision and action. |
|
Challengers |
Challenge the status quo and bring innovative ideas to life. |
|
Committed |
Own outcomes and deliver with integrity. |
|
Clear |
Communicate openly and simply, ensuring alignment. |
|
Connected |
Collaborate across teams to achieve shared goals. |
About the role
The Site Reliability Engineering Specialist independently executes activities that help ensures BT is in the best position to deliver the service performance, reliability and availability that internal and external customers expect, through enabling cross-team engineering discussions to achieve scalable, measurable, fault-tolerant, and cost-effective cloud services.
What you’ll be doing
1. Executes the implementation of new software development life cycle automation tools, frameworks, and code pipelines (continuous integration/continuous delivery pipelines whilst executing best practices with a focus on the re-use of application code, demonstrates consistent software delivery practices and produces continuous integration/continuous delivery platform solutions using Amazon Web Services cloud, infrastructure as code (IaC), GitOps, and container technologies
2. Coordinates a diverse team and creates the initial test schedule to deliver all aspects of testing to time, budget and quality targets, ensuring producing outlines of solutions and defining depth of testing required
3. Executes the implementation of automation technologies to ensure repeatability, eliminating toil, reducing mean time to detection and resolution and repair services
4. Proactively identifies and manages risk through regular assessment and diligent execution of controls and mitigations, proactively raising any concerns
5. Leads scale testing to measure, tune and optimise system performance
6. Executes metric/monitoring analysis that creates stability, security, and performance improvements
7. Designs, analyses, develops and troubleshoots highly-distributed large-scale production systems spanning on-prem and cloud-based hosting
8. Executes approaches that scale systems sustainably through mechanisms like automation and evolves systems by pushing for changes that improve reliability and velocity
9. Writes and delivers infrastructure as code software to improve the availability, scalability, latency, and efficiency of services
10. Implements robust monitoring and alerting systems and performs root cause analysis and post-mortems with an eye towards future prevention
11. Inspects queue and support processing to ensure early warning of support issues
12. Executes retrospective and preventive actions after each high severity production incident
13. Analyses complex systems from a reliability and resilience perspective and identifies sources of instability in distributed systems
14. Champions, continuously develops and shares with team knowledge on emerging trends and changes in site reliability engineering best practices and industry standards
15. Mentors other site reliability engineers, helping to improve the team's abilities by acting as a technical resource