Senior DevOps Engineer – AI Platforms
Unico Connect
5 - 10 years
Mumbai
Posted: 28/05/2026
Job Description
Senior DevOps Engineer AI Platforms
AI Platform Infrastructure, Agent Operations & Cloud DevOps
Unico Connect|Mumbai (On-site)|Full-time|58 years
Unico Connect is an AI-first product engineering agency building custom software, mobile apps, and AI solutions for clients across the USA, Europe, UK, Australia, Singapore, and India. We are hiring a Senior DevOps Engineer to own the platform infrastructure layer of the AI systems we design and ship, spanning multi-tenant AI platforms, agent orchestration runtimes, AI tool integration server deployments, and the API gateway and SDK infrastructure that connects these layers to end clients.
The mandatory requirement for this role is hands-on production experience deploying and operating infrastructure for an AI or LLM application. You will work directly with AI/ML engineers and product engineers to build and sustain the platform backbone: Kubernetes-based services, CI/CD pipelines for agent workflows and SDK releases, multi-tenant tenant isolation, observability across platform layers, and cloud cost management at the per-tenant level. A typical week involves standing up new service components, reviewing deployment pipelines, resolving reliability issues in running AI services, and translating infrastructure decisions into clear trade-off explanations for non-technical stakeholders.
Platform infrastructure and cloud architecture: Design and operate multi-tenant AI platform infrastructure on AWS, covering EKS cluster management, VPC architecture, IAM, KMS, S3, RDS, ElastiCache, and API Gateway. Establish the infrastructure patterns that scale across SMB, enterprise, and regulated client segments (government, Big 4) without per-client configuration sprawl.
AI tool integration server deployment and operations: Deploy, version, and operate AI tool integration servers in containerised environments. Manage server lifecycle, routing, versioning, and inter-service communication so that AI agents can reliably invoke tools across tenants without cross-contamination or latency spikes.
Agent orchestration infrastructure: Provision and maintain the runtime layer on which tool agents (TA) execute, including scheduling, concurrency management, queue depth control, and graceful failure handling. Work with AI engineers using LangGraph, CrewAI, or AutoGen to translate orchestration requirements into stable, observable infrastructure.
API gateway and SDK infrastructure: Own the gateway layer that routes client SDK calls through to the platform: rate limiting, tenant-based routing, authentication token validation, versioning, and graceful degradation. Ensure SDK-facing endpoints meet latency and availability SLOs agreed with clients.
CI/CD for AI-native applications: Build and maintain deployment pipelines for AI services, agent workflows, and SDK releases using GitHub Actions, ArgoCD, or equivalent. Implement promotion gates that include model evaluation checks, smoke tests, and rollback triggers before production traffic is shifted.
Monitoring, observability, and AI SRE: Define and instrument SLIs and SLOs for platform components: p95 latency per service, agent task success rate, AI service error rate, and per-tenant throughput. Build dashboards and alerting in Datadog, Grafana, or CloudWatch. Own runbooks and on-call response for platform-level incidents.
Security and tenant isolation: Implement and maintain multi-tenant isolation at the network, IAM, and data layer. Apply security controls aligned with OWASP, SOC 2, HIPAA, GDPR, and DPDP Act requirements as client contracts demand. Run secrets management via AWS Secrets Manager or HashiCorp Vault. Participate in client security reviews where infrastructure evidence is required.
FinOps and cost management: Instrument per-tenant cost attribution across compute, inference, storage, and egress. Establish budgets, set anomaly alerts, and produce monthly cost reports that allow product and commercial teams to price and plan accurately. Identify and act on optimisation opportunities without degrading reliability.
Infrastructure as code and environment management: Author and maintain Terraform modules for all platform infrastructure. Manage environment parity (dev, staging, production) and enforce drift detection. Document infrastructure decisions in architecture decision records (ADRs).
Architecture, mentorship, and client engagement: Evaluate managed services versus self-hosted options with documented cost, scale, and security reasoning. Contribute to technical proposals and scoping exercises. Mentor mid-level engineers through code and design reviews.
Hands-on production DevOps or platform engineering for an AI or LLM application (mandatory). Must have personally deployed and operated infrastructure that serves a live AI or LLM workload for an external client or internal product. Infrastructure-adjacent roles (e.g., DevOps for web apps only, or ML experimentation environments) do not qualify.
5+ years of DevOps or platform engineering experience, with at least 2 years in an environment deploying AI or LLM services to production.
AWS at depth: EKS, EC2, ECS, API Gateway, Lambda, S3, RDS, ElastiCache, CloudWatch, IAM, KMS, and Secrets Manager. Equivalent Azure or GCP experience considered with willingness to work in AWS.
Container and IaC stack: Docker, Kubernetes (cluster setup, Helm, network policies, autoscaling), Terraform, and at least one CI/CD platform (GitHub Actions, ArgoCD, GitLab CI, or Jenkins).
API gateway and service mesh experience: Hands-on with AWS API Gateway, Kong, Envoy, Istio, or equivalent. Able to configure routing rules, rate limiting, authentication middleware, and observability hooks at the gateway layer.
Familiarity with AI platform tooling at the infrastructure level: Comfortable reading and deploying services built on LangChain, LangGraph, LlamaIndex, or similar agent frameworks. Does not need to write AI code, but must understand how agent services consume compute and network resources so infrastructure is sized and secured correctly.
Working knowledge of cloud security controls and compliance frameworks, including IAM design, network segmentation, secrets management, audit logging, and at least one compliance standard (SOC 2, HIPAA, GDPR, or DPDP Act).
Clear written and verbal communication. Able to explain infrastructure trade-offs to AI engineers, product owners, and client stakeholders in their language, and to produce clear documentation without hand-holding.
Nice to have: AWS Certified DevOps Engineer Professional or CKA (Certified Kubernetes Administrator); multi-tenant SaaS platform experience; FinOps tooling hands-on (AWS Cost Explorer, Kubecost, Infracost); familiarity with agent orchestration frameworks (LangGraph, CrewAI, AutoGen) at deployment level; prior agency or consulting experience.
Services you might be interested in
We Search & Apply Jobs for You!
Our team scans through 1000s of opportunities and applies to roles best suited to your profile
Save 100+ hours and focus on what matters - cracking interviews and landing offers.
