Major Incident Manager (24*7 environment)
Hyderabad, Telangana, India
Role overview:
We are currently seeking an experienced Major Incident Manager to lead Major Incident Management, Root Cause Analysis, SLA Compliance, Reporting and Team Leadership across our Infrastructure and Cloud Service Assurance function, working with global teams across UK, EU and Canada.
Ownership of Strategic Work
Leading key strategic incident response and continuous improvement initiatives for the team, which could involve technologies or domains outside your immediate experience.
Taking accountability for the resolution of high-impact, business-critical major incidents end-to-end.
Owning incident outcomes even when root cause sits outside your core technical expertise.
Operating in Ambiguity
Able to work effectively with high levels of ambiguity and incomplete information during live, high-severity incidents.
Applying broad, structured and creative thinking to solve complex, multi-party incident problems.
Acting as a thought leader on incident management practice, not just an administrator of process.
Demonstrating strong self-starting behaviour, identifying problems and improvement opportunities without direction.
Key Duties & Responsibilities
Take command of major incidents across Keyloop, chairing bridge calls and driving them to resolution at pace — not simply monitoring status, but actively directing resolver teams and third-party suppliers.
Own business-critical escalations end-to-end, from identification through to resolution and stakeholder sign-off, ensuring SLA compliance and managing breach risk proactively.
Line-manage the MI Co-ordinator team, including the on-call/shift model and 24x7 coverage, setting standards for incident handling, communication and RCA quality.
Lead root cause analysis for all major incidents across Keyloop and third-party providers, feeding Problem Management and driving continuous improvement across people, process and tooling.
Prepare and drive leadership and executive updates during and after major incidents, using AI assistants (Copilot/Claude) to draft meeting outcomes and summarise bridge calls efficiently.
Create and maintain operational dashboards for real-time monitoring and major incident reporting.
Enforce ITIL-aligned incident, major incident and problem management processes, keeping all systems current with accurate status, actions and timelines.
Coach, develop and manage the performance of the MI Co-ordinator team, and act as an escalation and decision point for incident, process and supplier issues.
Undertake other duties as reasonably required in support of Service Assurance objectives.
Essential skillsets:
Have a minimum of 8+ years' experience in IT including demonstrable Major Incident leadership.
Have a proven track record of leading high-severity major incidents and influencing outcomes on live bridge calls.
Have advanced proficiency in Jira, preferably JSM, and strong command of the ITIL framework and incident management best practice.
Using AI in day to day work — hands-on use of AI assistants (Copilot/Claude) for drafting meeting outcomes, leadership updates, dashboard creation and metrics/trend analysis.
Have excellent english written and verbal communication; calm, decisive and credible during high-severity escalations.
Be a diplomatic, professional communicator across all levels of stakeholder, with a hands-on, strong-ownership approach.
Clear and confident communicator in English, able to collaborate effectively across time zones and cultures.
Have strong working knowledge of application integration, AWS and other cloud services, middleware, data centres, networks, compute and databases.
Have experience leading or mentoring a small team.
Be willing to work in shifts to support 24x7 major incident coverage.