About the role
Meesho is hiring an SDE III to design and implement scalable, fault-tolerant data pipelines (batch and streaming) using Apache Spark, Flink, and Kafka. You'll lead development of data platforms and reusable frameworks serving multiple teams, with deep enough Spark expertise to modify or extend the open-source codebase where needed.
What you'll do
Design and implement scalable, fault-tolerant data pipelines using Spark, Flink, and Kafka
Lead design and development of data platforms and reusable frameworks
Build and optimize data models and schemas for operational and analytical workloads
Develop streaming solutions using Apache Flink and Spark Structured Streaming
Ensure data quality, consistency, and governance through validation frameworks and access controls
What we're looking for
5 - 8+ years of relevant experience
5-8 years of professional experience in software/data engineering with a focus on distributed data systems
Strong programming skills in Java, Scala, or Python, with SQL expertise
2+ years hands-on with big data systems (Kafka, Spark/EMR/Dataproc, Hive, Delta Lake, Presto/Trino, Airflow)
Experience implementing and tuning Spark/Delta Lake/Presto at terabyte-scale or beyond
Strong understanding of Apache Spark internals (Catalyst, Tungsten, shuffle)
Nice to have
Contributions to open-source projects in the big data ecosystem (Spark, Kafka, Hive, Airflow)Familiarity with OLAP data cubes and BI/reporting tools (Tableau, Power BI, Superset, Looker)Exposure to backend technologies (RxJava, Spring Boot, microservices)