Sr. Lead - Service Delivery - Observability & DevOps Platforms
Software Engineering
Bengaluru, Karnataka, India · Pune, Maharashtra, India
About Northern Trust
As a global leader in innovative wealth management, asset servicing, asset management and banking services, Northern Trust (Nasdaq: NTRS) is proud to guide the world’s most successful individuals, families, corporations and institutions.
Since 1889, we have aligned our efforts with our three guiding Principles That Endure: Service, Expertise, and Integrity. Together, they reflect the three cornerstones of business conduct which we strive to instill in our employees, whom we call partners, and to provide to our clients and the communities we serve worldwide.
With more than 135 years of financial experience and over 24,000 partners, we serve the world’s most sophisticated clients using leading technology and exceptional service.
Position : Sr. Lead - Service Delivery - Observability & DevOps Platforms
We are seeking an experienced and highly motivated Service Delivery Sr, Lead to lead the operational governance, service delivery, risk management, customer engagement, and continuous improvement of enterprise Observability and Developer Platform services.
This role serves as the critical bridge between Engineering, Operations Support, Infrastructure Teams, Application Teams, Security, Vendors, and Business Stakeholders, ensuring reliable, secure, scalable, and compliant delivery of strategic platform services.
The Service Delivery Manager will be accountable for the operational success, service quality, risk posture, vendor management, capacity planning, and modernization initiatives supporting enterprise observability solutions, event management platforms, and software delivery toolchains.
The ideal candidate will possess extensive experience operating within large-scale, highly regulated environments such as Financial Services, Manufacturing, Healthcare, or similarly compliance-driven organizations.
Scope of Responsibility
Observability & Monitoring Platforms
Dynatrace
Elastic / ELK
Azure Log Analytics
SCOM (Microsoft System Center Operations Manager)
ServiceNow Event Management
Enterprise Alerting Platforms
Developer Platforms & Productivity Tools
GitHub Enterprise
GitHub Actions & GitHub Runners
Azure DevOps
Sonatype Nexus Repository
CI/CD Platforms
Enterprise Desktop Software & Developer Tooling
Key Responsibilities
Service Delivery & Operational Leadership
Own end-to-end service delivery for Observability, Monitoring, Event Management, and Developer Tooling platforms.
Ensure operational excellence and stability of services supporting critical enterprise workloads.
Maintain high-performing operational support functions for global "Run the Business" activities.
Establish and monitor service level objectives, KPIs, SLAs, OLAs, and operational scorecards.
Drive service maturity, operational effectiveness, and customer satisfaction across supported platforms.
Ensure operational procedures, governance controls, and service management processes are consistently followed.
Engineering & Operations Partnership
Act as the primary liaison between Engineering Teams, Infrastructure Teams, Operations Support, Security, Vendors, and Application Owners.
Facilitate seamless transition of projects, upgrades, and new capabilities into production support.
Drive alignment between strategic engineering initiatives and operational support requirements.
Ensure operational readiness reviews are conducted prior to production releases.
Champion supportability, observability, resiliency, and operational excellence during platform modernization initiatives.
Major Incident & Escalation Management
Own and lead enterprise-wide major incident management processes.
Serve as the escalation owner during high-priority incidents affecting production services.
Coordinate engineering, infrastructure, support, vendor, and business stakeholders during service disruptions.
Ensure effective executive communication throughout incident lifecycles.
Lead post-incident reviews, Root Cause Analysis (RCA), and corrective action planning.
Drive permanent resolution of recurring operational issues through Problem Management practices.
Risk, Compliance & Governance
Ensure compliance with enterprise security, regulatory, and operational standards.
Identify, assess, track, and mitigate operational risks across services.
Support internal audits, regulatory examinations, compliance reviews, and risk assessments.
Maintain service governance frameworks, controls documentation, and operational procedures.
Partner closely with Risk, Security, Compliance, and Audit organizations to address findings and remediation activities.
Ensure service delivery aligns with organizational governance requirements and regulatory obligations.
Capacity, Availability & Performance Management
Own capacity planning processes for observability and platform services.
Ensure future demand from business growth, projects, and modernization initiatives is incorporated into capacity plans.
Monitor platform utilization, service performance, and infrastructure health.
Develop forecasting models and operational dashboards to support strategic planning.
Drive continuous optimization of availability, performance, scalability, and cost efficiency.
Vendor & Third-Party Management
Manage strategic relationships with vendors and service providers including Microsoft, Dynatrace, Elastic, GitHub, Sonatype, and managed service partners.
Conduct service review meetings covering:
Service Performance
Risk & Compliance
Security
Financial Management
Continuous Improvement
SLA Adherence
Ensure vendors meet contractual obligations and agreed service levels.
Manage vendor escalations, service improvement plans, and remediation programs.
Evaluate vendor performance from both operational and financial perspectives.
Service Improvement & Platform Modernization
Identify opportunities to improve service reliability, efficiency, usability, and customer experience.
Lead service improvement programs across monitoring, alerting, event management, and CI/CD ecosystems.
Drive adoption of automation, self-service capabilities, observability best practices, and platform engineering principles.
Develop and execute Service Improvement Plans (SIPs).
Ensure actions are tracked through completion with measurable business outcomes.
Promote automation-first and reliability engineering approaches throughout service operations.
Reporting & Executive Communications
Provide regular and accurate service performance reporting to leadership.
Deliver executive dashboards covering:
Availability
Reliability
Capacity
Risk
Compliance
Customer Satisfaction
Operational Trends
Present service health reviews to senior leadership and governance forums.
Communicate service impacts, operational risks, and strategic recommendations to stakeholders.
Required Qualifications
Bachelor's Degree in Computer Science, Information Technology, Engineering, or related discipline.
12+ years of experience in IT Infrastructure Operations, Service Delivery, Platform Operations, Engineering Operations, or Production Support.
Extensive experience in highly regulated enterprise environments.
Proven experience managing enterprise observability and monitoring ecosystems.
Strong understanding of:
Dynatrace
SCOM
Elastic / ELK
Azure Log Analytics
ServiceNow Event Management
GitHub Enterprise
Azure DevOps
Nexus Repository Manager
Strong understanding of ITIL disciplines including:
Incident Management
Problem Management
Change Management
Capacity Management
Availability Management
Service Level Management
Experience managing major incidents and critical production escalations.
Experience managing vendor contracts and third-party service providers.
Excellent communication, stakeholder management, negotiation, and leadership skills.
Preferred Qualifications
ITIL Foundation / ITIL Managing Professional Certification.
Experience in Financial Services, Manufacturing, Healthcare, or other regulated industries.
Understanding of Site Reliability Engineering (SRE) principles.
Experience with ServiceNow ITSM.
Exposure Cloud Technologies (Azure, AWS).
Working knowledge of:
GitHub Actions
Azure DevOps
Infrastructure as Code
Ansible
Python
PowerShell
Automation Frameworks
Familiarity with Disaster Recovery, Business Continuity, Backup, Storage, Virtualization, Linux, Windows, and Citrix technologies.
Key Competencies
Leadership
Strategic Thinking
Executive Communication
Stakeholder Management
Vendor Management
Team Leadership
Decision Making
Service Management
ITIL
Service Governance
Major Incident Management
Capacity Management
Performance Management
Risk Management
Technical Domain Knowledge
Observability
Monitoring
Event Management
AIOps
DevOps Platforms
CI/CD Tooling
Platform Operations
Working with Us
As a Northern Trust partner, you will be part of a flexible and collaborative work culture, which has a strong history of financial strength and stability. Movement within the organization is encouraged, senior leaders are accessible, and you can take pride in working for a company committed to an inclusive workplace and assisting the communities we serve.
Philanthropy is deeply rooted in Northern Trust’s history and is an essential element of our culture. Employees around the world give their time and talent to work for the greater good of their communities.
Reasonable Accommodation
Northern Trust is committed to working with and providing adjustments to individuals with health conditions and disabilities. If you need a reasonable accommodation for any part of the employment process, please email our HR Service Center at MyHRHelp@ntrs.com, or alternatively you can discuss your individual requirements with the recruiter you are working with.
About Our Pune Office
The Northern Trust Pune office, established in 2016, is now home to over 3,000 employees. The office handles various functions, including Operations for Asset Servicing and Wealth Management, as well as delivering critical technology solutions that support business operations across the globe.
Our Pune team takes our commitment to service to heart. In 2024, they volunteered more than 10,000+ hours into the communities where they live and work. Learn more.