About the role
Meesho is hiring an Engineering Manager to lead its AI Platform team, which serves 1M+ real-time deep-learning model inferences per second (scaling 3x+ on sale days) across Search, Recommendations, Personalized Ranking, Logistics, Fraud Detection, and Image Match. You'll architect and scale the AI platform, drive inference optimization, and lead a team of AI engineers.
What you'll do
Lead, mentor, and grow a team of AI engineers, setting technical direction and owning delivery
Architect and scale Meesho's AI platform: cross-region inference, GPU fleet management, distributed training
Drive inference optimization across the stack - GPU kernel tuning, quantization, memory/IO optimization
Optimize open-weight models at both model and inference-engine level
Partner with Product, Data Science, and Platform teams to turn AI capabilities into production impact
What we're looking for
9++ years of relevant experience
9+ years of software engineering experience, including 2+ years managing engineers
Strong hands-on experience with modern LLM inference stacks (TensorRT-LLM, vLLM, SGLang) and production low-latency model serving
Depth in inference optimization: GPU kernel tuning, quantization, speculative decoding, KV-cache/memory optimization
Experience with distributed training frameworks (PyTorch FSDP, DeepSpeed, Megatron, or Ray)
Experience running GPU fleets in production (Kubernetes/GKE, GPU scheduling, multi-region deployment)
Proficiency in Python
systems-level fluency (C++/Go/Rust) for performance-critical paths
Nice to have
Open-source contributions to inference engines, training frameworks, or ML infra toolingExperience managing GPU cost/efficiency (FinOps) for a large fleetTrack record building platforms for high-scale consumer productsFamiliarity with observability and reliability for ML systems (SLOs, autoscaling)