Chennai · Staff/Principal
Applicants who checked fit first are 3.1× more likely to hear back
Your score for this role already exists
ASAI compared this JD against 41 signals - skills, seniority, domain, stack overlap etc. Add a resume and it unlocks in about 30 seconds.
No credit card · 1 tap with Google
MEDIUM
enGen Global is reviewing applications at a steady pace. Expect a standard response time as they evaluate the current pool.
First 72 hours
Window passed - posted 45d agoEarly applicants get seen before the pile builds.
Not a repost
The first time we've seen this listing - it hasn't been closed and reopened.
You almost certainly match several of these already. Unlock your skill map to see the matches, the gaps, and what to fix first.
Lead PySpark Cloud Data Engineer:
enGen Global is an emerging global healthcare partner that delivers strategic innovation, expertise, and flexibility to its healthcare partners. Being a US healthcare conglomerate captive, we have direct access to deeper insights that help us accelerate our learning process and keeps us ahead of the curve. Thryve delivers next-generation solutions that enable our healthcare partners to provide positive experiences to their consumers.
Our global collaborative of healthcare, operations, and IT experts creates innovative and sustainable processes for our clients, which keeps the ever-evolving consumers engaged and assists them in managing the future of their healthcare better. We recognize that our people are our strength and the diverse talents they bring to our global workforce are directly linked to our success. Thryve is an equal opportunity employer and places a high value on integrity, diversity, and inclusion in the organization. We do not discriminate based on any protected attribute. For more information about the organization, please visit www.thryvedigital.com
Role Summary:
This job involves understand the overall requirement of the enterprise Data need and design & develop the robust, highly scalable & resilient data pipeline using Pyspark, Dataproc, open source table formats and other Google services. Job will involve extensive interfacing & co-ordination with other senior tech folks (across India & US) and lead the design & development of the data ingestion pipelines
Essential Responsibilities
• Design, develop, and implement highly scalable, reliable, and performant data pipelines using PySpark, Dataproc, and other Google cloud-native technologies
• Create pipelines for data lakehouse architectures utilizing Apache Iceberg for open table formats and BigLake Metastore for unified metadata management
• Develop and maintain data transformation logic using Pyspark and / or dbt to create clean, consistent, and production-ready data models
• Work extensively with Google Cloud Platform (GCP) services, including Dataproc, BigQuery, Cloud Storage, and other relevant data services.
• Optimize data pipelines and queries for performance, efficiency, and cost-effectiveness.
• Implement and enforce data quality checks, monitoring, and governance best practices to ensure data integrity and reliability
• Provide technical guidance and mentorship to junior data engineers, fostering a culture of continuous learning and improvement
• Create and maintain comprehensive technical documentation for data pipelines, data models, and platform architecture
• Provide regular updates on the tasks, status and risks to project manager
The experience we are looking to add to our team
Required
• Bachelor’s degree or higher from a reputed university
• 8 to 12 years total experience with majority of that experience related to building high performant data pipelines – batch and streaming
• Strong Hands on experience / expert-level proficiency in PySpark for data processing, transformation, and analysis
• Extensive experience in implementing large scale data ingestion and curation solutions
• Hands-on experience with cloud-based data platforms, preferably Google Cloud Platform (GCP) and services like Dataproc, BigQuery, and Cloud Storage
• Knowledge in Apache Iceberg for open table formats
• Advanced SQL skills for data querying, manipulation, and optimization
• Proficiency with Git and collaborative development workflows
• Excellent analytical and problem-solving skills with a keen attention to detail
• Strong communication and interpersonal skills, with the ability to explain complex technical concepts to both technical and non-technical audiences
Good to have
• Expertise in Google Cloud services
• Proven experience with Apache Iceberg for open table formats and BigLake Metastore for unified metadata management
• Strong experience with dbt (data build tool) for data modeling, transformation, and orchestration
• Experience in Data Governance and Data quality tools like Atlan, Monte Carlo etc.
• Experience in processing streaming data using Kafka / Pub-Sub
• Healthcare industry experience
• Experience in Agile
Free · no signup
Daily job drops, skill trends and free resources - posted straight to the group. Leave any time.
No spam. Just jobs and resources.
Why people use ASAI
Scored, not searched. Every role ranked against your actual profile.
Alerts as often as hourly. Reach new roles while the pile is still small.
Skill gaps, spelled out. See exactly which requirements you don't meet yet.
Verified jobs, only. Say no to ghost jobs. Your time deserves respect.
More Data Engineer roles in Chennai
See allKeep browsing