Senior AI Operations Engineer

Senior AI Operations Engineer

Barcelona

Categoría: DevOps

Salario: ~

Empresa: AstraZeneca

Estudios:

Experiencia: 3-5 años

Tipo de contrato: Indefinido

Jornada: Jornada completa

Fecha: 23/08/2026

Descripción: Do you have expertise in, and passion for, AI-powered operational excellence? Would you like to apply your expertise to impact a company that follows the science and turns ideas into life changing medicines? If so, AstraZeneca might be the one for you!

ABOUT ASTRAZENECA

AstraZeneca is a global, innovation-driven BioPharmaceutical business that focuses on the discovery, development and commercialisation of prescription medicines for some of the world´s most serious diseases. But we´re more than one of the world´s leading pharmaceutical companies.

At AstraZeneca we´re dedicated to being a Great Place to Work. Where you are empowered to push the boundaries of science and unleash your entrepreneurial spirit. There´s no better place to make a difference to medicine, patients and society. An inclusive culture that champions diversity and collaboration. Always committed to lifelong learning, growth and development.

ABOUT OUR ENTERPRISE AI PLATFORMS & SERVICES TEAM

A place to do important work. We connect across the whole business to power each function to better influence patient outcomes and improve their lives. Impactful and valuable, this is where you come to raise your profile and do good for others. Play an increasingly crucial role in driving disruptive transformation on our journey to becoming a digital and data-led enterprise. Unleash the power of our latest innovations in data, machine learning and technology to turn complex information into life-changing and practical insights. Work in synergy with leading experts in our specialist communities. Here we bring the brightest minds to bear, with access to cutting-edge techniques and the opportunity to be part of novel solutions. It´s up to us to drive the outcomes forward, inventing and building, expanding our knowledge to identify the next opportunity. An inclusive team, we bring together diverse areas - different functions as well as external partners. Pooling from an unrivalled source of knowledge, we share, learn and challenge. It powers us to decode business needs and apply our technical know-how to add greater value. Rise to the challenge of shaping the future of an evolving business in the technology space.

ABOUT THE ROLE

The Enterprise AI Platforms & Services Team are responsible for building and running the platforms, tooling and infrastructure that powers AstraZeneca´s ambition to use AI in every step of the value chain, from discovering new compounds to patient safety systems.

We are looking for a Senior AI Operations Engineer to be a hands-on technical leader within our AI/ML platform operations function. Reporting to the AI for Operations Lead, you will be the primary executor and technical driver across our growing platform estate, built on Azure and AWS. The ideal candidate will have deep, current experience in site reliability engineering and will combine strong individual contribution with mentoring and technical leadership of L1 and L2 engineers.

You will operate within a three-tier operations structure (L1 Runbook Operators, L2 Site Reliability Engineers, L3 Product Engineering interface) and be the senior hands-on practitioner who builds the automation, designs the observability, and drives the continuous improvement flywheel day-to-day. You will be a key contributor to AI-augmented operations - building and refining AI-powered tools for runbook querying, incident pattern analysis, and automated diagnosis.

This is not a traditional support engineer role. This is a senior technical position for someone who sees operations as an engineering discipline and who can design, build, and instrument systems to the highest standard while coaching others to do the same.

KEY ACCOUNTABILITIES

Technical Execution and Design

- Design, build, and maintain the centralised observability layer - instrumentation, dashboards, alerting rules, and telemetry pipelines using platforms such as Datadog, New Relic, Grafana, or Splunk

- Architect and implement AI-augmented operations tooling: conversational runbook interfaces, AI-driven incident analysis, and pattern recognition systems

- Lead complex incident response, perform root-cause analysis, and produce actionable post-mortems

- Contribute patches and instrumentation to product engineering codebases, ensuring they are architecturally sound and address root causes

- Contribute to vendor evaluation to assess and recommend appropriate technologies and automation tooling.

Automation and Continuous Improvement

- Build and maintain the automation that powers the continuous improvement cycle: every incident results in a runbook, an automation, or a patch

- Own the technical execution of automation investments prioritised by frequency, resolution time, and blast radius

- Proactively identify and eliminate toil, maintaining the standard of no more than 50% of SRE time on reactive work

- Develop and maintain runbooks, ensuring they are accurate, current, and progressively automated

Operational Readiness and Platform Onboarding

- Implement the Operational Readiness Gate - assess platforms against defined criteria before they transition from product engineering to operations ownership

- Shape platform operability during development - contributing instrumentation, health checks, and observability hooks before platforms enter BAU

- Provide technical input to platform architecture reviews from an operational perspective

Mentoring and Team Contribution

- Mentor and coach L1 operators and junior L2 SREs, building their technical capability and engineering mindset

- Identify operators with engineering aptitude and support their development pathway from L1 to L2

- Contribute to the operations review by preparing flywheel metrics: L1 resolution rate, repeat incident rate, automation coverage, and toil budget compliance

- Foster a culture where SREs are engineers whose product is operational excellence

Stakeholder Collaboration

- Partner with product engineering teams on post-mortems, handover assessments, and operational obligation delivery

- Represent operational requirements in technical design discussions

- Communicate incident findings, operational risks, and improvement proposals clearly to the Operations Lead and wider team

CANDIDATE KNOWLEDGE, SKILLS AND EXPERIENCE

Essential

- BSc/MSc degree in Computer Science or related quantitative or analytical field

- Significant hands-on experience as a Site Reliability Engineer or platform operations engineer at scale - you build and run systems, not just direct others

- Strong expertise in observability platforms (Datadog, New Relic, Grafana, Splunk, or equivalent) including dashboard design, alerting strategies, and telemetry pipeline implementation

- Strong working knowledge of OpenTelemetry, distributed tracing, and structured logging standards

- Proven track record of designing and implementing automation that materially reduces operational toil - strong scripting skills in Python and Bash, with the ability to build robust tooling beyond one-off scripts

- Experience assessing platform readiness and contributing to operational handover processes

- Strong hands-on skills with cloud infrastructure (Azure and/or AWS) including container orchestration, serverless architectures, and managed services

- Experience running or contributing significantly to post-mortem processes and translating findings into preventive engineering work

- Ability to implement precise technical solutions - you can instrument a system, configure meaningful alerts, and build automation that intervenes at the right point

- Experience mentoring junior engineers and contributing to team development

Desirable

- Experience applying AI/ML to operational challenges - intelligent alerting, automated diagnosis, predictive incident detection, or conversational operations interfaces

- Familiarity with the AstraZeneca technology estate or regulated pharmaceutical environments

- Experience operating platforms that serve AI/ML workloads (LLM inference, model serving, data pipelines)

- ITIL, SRE, or operational excellence certifications or equivalent practical frameworks

- Experience working within multi-tier support structures with clear escalation paths

- Infrastructure-as-code expertise (Terraform, CloudFormation)

- Awareness to GxP and audit trail awareness.

Personal Qualities

- An engineering mindset applied to operations - you see every repeated manual task as an automation opportunity

- Comfort operating at the boundary between deep technical work and collaborative team contribution

- Creative, collaborative, and resilient

- Strong communicator who can explain technical complexity clearly to both engineers and leadership

- A bias toward systems thinking - you address root causes, not symptoms

- Self-directed with the ability to manage competing priorities across multiple platforms

When we put unexpected teams in the same room, we unleash bold thinking with the power to inspire life-changing medicines. In-person working gives us the platform we need to connect, work at pace and challenge perceptions. That´s why we work, on average, a minimum of three days per week from the office. But that doesn´t mean we´re not flexible. We balance the expectation of being in the office while respecting individual flexibility. Join us in our unique and ambitious world.

Publicado 24-08-2026


Síguenos:

Servicios:

Empleo Público:

Ofertas empleo en BARCELONA relacionadas

Cloud Operations Engineer Azure y AWS

Cloud Operations Engineer Azure y AWSBarcelona (Híbrido)Categoría: DevOps Salario: ~ Empresa: SeremEstudios: FP

Publicado: 24-08-2026

Field Service Engineer

Field Service EngineerBarcelonaCategoría: Electrónica - Ingenieros/Industria Salario: ~ Empresa: Michael PageEs

Publicado: 24-08-2026

SAP MM y SD Senior Consultant

SAP MM y SD Senior ConsultantBarcelona (Híbrido)Categoría: Consultor Salario: ~ Empresa: knowmad moodEstudios:

Publicado: 24-08-2026

Últimas Noticias de Empleo

Semana de sacudidas: el hantavirus pone a prueba al Estado, Trump pierde en el Supremo y el Real Madrid estalla por dentro

Semana de sacudidas: el hantavirus pone a prueba al Estado, Trump pierde en el Supremo y el Real Madrid estalla por dentro

Semana de sacudidas: el hantavirus pone a prueba al Estado, Trump pierde en el Supremo y el Real Madrid estalla por dentro

El crucero del hantavirus sacude la política española mientras Europa negocia, los mercados celebran y el mundo no da tregua

El crucero del hantavirus sacude la política española mientras Europa negocia, los mercados celebran y el mundo no da tregua

El crucero del hantavirus sacude la política española mientras Europa negocia, los mercados celebran y el mundo no da tregua

Del hantavirus en alta mar a la mayor crisis energética de la historia: España y el mundo ante múltiples frentes

Del hantavirus en alta mar a la mayor crisis energética de la historia: España y el mundo ante múltiples frentes

Del hantavirus en alta mar a la mayor crisis energética de la historia: España y el mundo ante múltiples frentes. Resumen noticias 5 de mayo de 2026

Ormuz al borde del colapso, un crucero en cuarentena y la mayor indemnización médica de España: una jornada de crisis simultáneas

Ormuz al borde del colapso, un crucero en cuarentena y la mayor indemnización médica de España: una jornada de crisis simultáneas

La jornada del lunes estuvo marcada por una sucesión de crisis simultáneas que mantienen en vilo al mundo, con el estrecho de Ormuz como principal foco geopolítico. Emiratos Árabes Unidos denunció ataques de Irán con misiles y drones que provocaron un gran incendio en la Zona Industrial Petrolera del emirato de Fuyaira, con tres trabajadores […]

El poder del correo electrónico

El poder del correo electrónico

Hace unos 3 años creé una aplicación de gestión orienta a su uso en ayuntamientos, que les permitiera gestionar el alquiler de instalaciones deportivas o no deportivas (pistas de padel, pabellones, piscinas, centros de conferencias, etc). La tengo instalada en varios ayuntamientos, con una curva de crecimientos exponencial en cuanto a resultados. La mayoría de […]

El nuevo control horario en España: el fin de la improvisación empresarial

El nuevo control horario en España: el fin de la improvisación empresarial

España se prepara para uno de los mayores cambios en materia laboral de los últimos años. El nuevo sistema de fichaje y control horario, previsto para consolidarse a lo largo de 2026, no es una simple actualización normativa: es un endurecimiento real del control sobre las empresas. Quien piense que esto será “más de lo […]

;