Jobs›Platform Engineer, Pune

Platform Engineer

innovistors · Pune
Pay₹20L–45La year, as listed
WherePuneMaharashtra
TypeFull time
Posted29 Sep7 days ago, via SimplyHired
Kaam checked
No fee, deposit or pay-to-apply signs
Night shifts
No vehicle or licence needed
Skills they list4 named
KubernetesDNSLinuxAI
About this job

The AI Platform Operations Engineer is part of the 24×7 operations team supporting the Central AI Kitchen platform and its AI services.

The role is responsible for continuous monitoring, first-level incident detection and troubleshooting, execution of approved recovery procedures, and timely escalation to the appropriate platform, infrastructure and application support teams.

The platform consists of AI services running on Red Hat OpenShift, hosted on on-prem infrastructure. The engineer will operate primarily using established dashboards, alerts, runbooks and GitOps-based recovery procedures.

This is an entry-level operations role designed for engineers who have foundational Linux, networking and Kubernetes/OpenShift knowledge and are interested in developing deeper platform engineering and SRE capabilities.

Key Responsibilities

24×7 Platform Monitoring

Monitor platform and application dashboards, alerts and operational mailboxes.

Identify availability, infrastructure, application and service degradation events.

Acknowledge alerts promptly and determine initial severity based on established procedures.

Maintain accurate shift handover and operational records.

First-Level Incident Troubleshooting

Perform basic network connectivity checks including ping, DNS resolution and endpoint connectivity.

Check OpenShift/Kubernetes resource status including nodes, pods, deployments and services.

Review basic application and platform logs to identify common failure conditions.

Check infrastructure and service health using approved dashboards and operational tools.

Collect relevant diagnostic information before escalation.

Incident Recovery

Execute documented recovery procedures and operational runbooks.

Restart or redeploy affected workloads using approved GitOps processes.

Verify service recovery through dashboards, health checks and application endpoints.

Escalate when recovery procedures are unsuccessful or when an incident falls outside the approved operating scope.

Incident Coordination & Escalation

Create and maintain incident tickets with accurate timestamps, symptoms, actions and observations.

Engage the appropriate application, platform, infrastructure or network teams based on established escalation procedures.

Provide clear status updates during active incidents.

Support incident bridges by providing operational information and executing actions requested by L2/L3 engineers.

Ensure effective handover of unresolved incidents between shifts.

Operational Procedures

Follow established Standard Operating Procedures (SOPs), runbooks and change-management processes.

Document newly encountered symptoms and successful troubleshooting steps.

Highlight recurring alerts or operational problems to senior platform engineers.

Participate in operational drills and recovery exercises.

Pay
₹2,000,000.00 - ₹4,500,000.00 per year
Work Location
In person

Never pay to get work. If a listing asks for a fee, it is a scam. The ten signs →

Apply on SimplyHired
Opens simplyhired.co.in in a new tab