NebiusNebius

Incident Manager

Added 4 months ago

Description

The role 

As an Incident Manager, you will be responsible for managing Incident and Problem processes across our data center infrastructure. You will coordinate between internal teams and external partners to ensure quick and effective resolution of hardware, firmware, and infrastructure-related issues.

You are someone who enjoys diving deep into processes, continuously improving and automating workflows to deliver a more reliable and efficient service.

Your responsibilities will include: 

  • Lead Incident and Problem Management for all hardware, firmware, and IT operational issues that impact services.
  • Coordinate between the IT Infrastructure team, Engineering, and external vendors/contractors to resolve issues and optimize communication.
  • Manage and maintain the Knowledge Base, encouraging contributions to service operation documentation.
  • Drive process improvements and follow up on corrective actions.
  • Identify opportunities for automation to prevent incidents and detect issues at an early stage.

You’re welcome to work in our office in Amsterdam, The Netherlands. 

We expect you to have: 

  • Proven experience in Incident and Problem Management.
  • Strong understanding of data center operations, server, and network equipment.
  • Hands-on experience in process design and implementation.
  • Technical background in troubleshooting data center IT hardware.
  • Excellent communication and coordination skills.
  • Ability to create and deliver presentations, charts, diagrams, and reports.
  • Strong data analysis and scripting skills for identifying deviations and trends.
  • High proficiency in spoken and written English.
  • Proactive, detail-oriented, and responsible approach to work.

It will be an added bonus if you have: 

  • ITIL or equivalent service management certification.
  • Driving license (category B).

Company

Nebius provides an AI-focused cloud platform enabling scalable GPU clusters (from single GPU to thousands of NVIDIA GPUs) with pre-configured drivers, InfiniBand networking, and orchestrators like Kubernetes or Slurm. It offers fully managed services (MLflow, PostgreSQL, Apache Spark), cloud-native tooling (Terraform, API, CLI), ready-to-go solutions, and expert support. Nebius also runs data centers and is active in AI research collaborations and open-source AI ecosystem examples (vLLM, CRISPR-GPT references) and has partnerships with NVIDIA as Reference Platform Cloud Partner.

See more incident manager jobs