Jobs›Mumbai

EY - GDS Consulting - AIA - Gen AI - Manager

EY · Mumbai
PayPay not listed
WhereMumbaiMaharashtra
TypeFull time
Posted17 Sep22 days ago, via SimplyHired
Skills they list19 named
FiddlerAzureCloud architectureDoctoral degreeComputer ScienceMCPKubernetesMaster's degreeBachelor's degreeDistributed systemsAPIsScalabilityData scienceAILeadershipCommunication skillsPythonData ScienceAnalytics
About this job
Location
Mumbai
Other locations
Anywhere in Country
Salary
Competitive
Date
Sep 17, 2026

Job description

Requisition ID
1723811

At EY, you’ll have the chance to build a career as unique as you are, with the global scale, support, inclusive culture and technology to become the best version of you. And we’re counting on your unique voice and perspective to help EY become even better, too. Join us and build an exceptional experience for yourself, and a better working world for all.

Position
Manager
Department
Technology Consulting, AI&A
Title
Enterprise Observability – AI Data Engineer – Manager
Educational qualification
BTech/Masters/PhD

The opportunity

We are seeking an experienced AI Data Engineer Manager with 8–12+ years of professional experience and deep expertise in AI engineering, enterprise observability, telemetry, traceability, and operational monitoring of AI systems. The ideal candidate should be passionate about building reliable, scalable, and governable AI platforms with strong experience in Agentic AI implementations, observability architectures, monitoring frameworks, and enterprise AI operations.

The individual will lead the design, development, implementation, and optimization of observability solutions for AI, Generative AI, Agentic AI, and Agentic RAG systems while collaborating with cross-functional teams, clients, and stakeholders to improve reliability, transparency, performance, and operational readiness across enterprise AI platforms.

Important note: This role requires professionally delivered AI experience in business environments. Personal projects, certifications, hackathons, and tutorial-based work may support an application, but do not replace the required core experience.

Application guidance: Candidates should be able to clearly explain at least two relevant AI implementations, including the business use case, their direct contribution, the technical approach, and the outcome delivered.

Your key responsibilities

Lead the design and implementation of enterprise-scale observability frameworks for AI, GenAI, Agentic AI, and Agentic RAG systems.

Define enterprise standards for telemetry, tracing, monitoring, logging, and operational governance of AI solutions.

Design and implement traceability frameworks to track agent reasoning, tool usage, retrieval paths, model interactions, and workflow execution.

Build end-to-end observability solutions for AI agents, retrieval systems, APIs, and distributed AI applications.

Design and implement telemetry collection pipelines to monitor model performance, latency, cost, quality, reliability, and user interactions.

Develop monitoring and evaluation frameworks for Agentic AI workflows using LangGraph, LangChain, MCP integrations, and Azure AI services.

Build production-grade observability services, APIs, and monitoring components using Python and FastAPI.

Design operational dashboards, alerting frameworks, and real-time monitoring solutions for AI workloads.

Establish evaluation, benchmarking, experimentation, and continuous improvement processes for enterprise AI systems.

Implement traceability and governance controls to support auditing, compliance, security, and Responsible AI requirements.

Deploy and manage scalable AI observability solutions on Azure Kubernetes Service (AKS).

Integrate telemetry data from AI workloads, enterprise applications, APIs, databases, and distributed systems.

Design data pipelines for collection, processing, and enrichment of observability and monitoring data.

Collaborate with AI engineers, platform teams, data engineers, and business stakeholders to improve operational visibility and system reliability.

Create accelerators, reusable observability frameworks, monitoring standards, and operational best practices.

Stay current with advancements in Agentic AI, observability, telemetry platforms, distributed tracing, AI evaluation frameworks, and emerging technologies.

Skills and Attributes

Professional Experience

8–12+ years of total experience, including 4+ years of directly relevant experience in AI engineering, observability, telemetry, distributed systems monitoring, or AI operations.

2+ years of people, technical, or delivery leadership experience.

Proven experience delivering enterprise-scale monitoring, observability, or operational intelligence platforms.

Strong depth in at least 3 of the following areas, with direct ownership in at least 2

AI observability

Agentic AI implementations

Monitoring and telemetry

Traceability and governance

Enterprise AI platforms

Distributed systems monitoring

AI evaluation frameworks

Cloud-native architecture

AI operations and reliability engineering

Data engineering and operational analytics

Strong ability to translate operational and governance requirements into scalable technical solutions.

Comfortable working across engineering, platform, business, risk, and governance stakeholders.

Educational Background

Bachelor’s/master’s degree in computer science, Data Science, Engineering, or related field.

Technical Skills

Deep expertise in Agentic AI, Agentic AI Implementation, and Agentic RAG Systems.

Strong understanding of observability principles, telemetry collection, traceability frameworks, monitoring architectures, and operational analytics.

Experience establishing end-to-end observability for AI applications, AI agents, retrieval systems, and enterprise AI platforms.

Experience working with distributed tracing, telemetry pipelines, logging frameworks, and monitoring solutions.

Strong understanding of AI evaluation, reliability engineering, performance measurement, quality monitoring, and operational excellence.

Hands-on experience implementing observability frameworks using LangChain, LangGraph, and Model Context Protocol (MCP).

Experience monitoring agent workflows, tool chains, retrieval paths, model interactions, and multi-agent systems.

Strong experience with Microsoft Azure AI Pla

Never pay to get work. If a listing asks for a fee, it is a scam. The ten signs →

Apply on SimplyHired
Opens simplyhired.co.in in a new tab