Hyderabad · Director
Applicants who checked fit first are 3.1× more likely to hear back
Your score for this role already exists
ASAI compared this JD against 41 signals - skills, seniority, domain, stack overlap etc. Add a resume and it unlocks in about 30 seconds.
No credit card · 1 tap with Google
MEDIUM
Egen is reviewing applications at a steady pace. Expect a standard response time as they evaluate the current pool.
First 72 hours
Window passed - posted 7d agoEarly applicants get seen before the pile builds.
Not a repost
The first time we've seen this listing - it hasn't been closed and reopened.
You almost certainly match several of these already. Unlock your skill map to see the matches, the gaps, and what to fix first.
Job Overview:
We are looking for a skilled and motivated Lead Data Engineer with strong experience in Python programming, PySpark, Databricks and Google Cloud Platform (GCP) to join our data engineering team. The ideal candidate will be responsible for requirements gathering, designing, architecting the solution, developing, and maintaining robust and scalable ETL (Extract, Transform, Load) & ELT data pipelines. The role involves working with customers directly, gathering requirements, discovery phase, designing, architecting the solution, using various GCP services, implementing data transformations, data ingestion, data quality, and consistency across systems, and post post-delivery support.
Experience Level:
10 to 16 years of relevant IT experience
Design, develop, test, and maintain scalable ETL data pipelines using Python, PySpark, Databricks & GCP / Azure.
Architect the enterprise solutions with various technologies like GCP, Azure, Databricks, PySpark and Spark SQL.
Work extensively on Google Cloud Platform (GCP) services such as:
Dataflow for real-time and batch data processing
Cloud Functions for lightweight serverless compute
BigQuery for data warehousing and analytics
Cloud Composer for orchestration of data workflows (on Apache Airflow)
Google Cloud Storage (GCS) for managing data at scale
IAM for access control and security
Cloud Run for containerized applications
Develop production-grade Databricks notebooks and workflows.
Build data transformation pipelines using PySpark and Spark SQL.
Implement Delta Lake architecture.
Design Bronze, Silver, and Gold data layers using the Medallion Architecture.
Implement Databricks Workflows/Jobs and dependency management.
Tune Spark jobs for large-scale data processing.
Optimize cluster configuration and compute utilization.
Implement appropriate partitioning, caching, and file-size optimization strategies.
Perform data ingestion from various sources and apply transformation and cleansing logic to ensure high-quality data delivery.
Implement and enforce data quality checks, validation rules, and monitoring.
Collaborate with data scientists, analysts, and other engineering teams to understand data needs and deliver efficient data solutions.
Manage version control using GitHub and participate in CI/CD pipeline deployments for data projects.
Write complex SQL queries for data extraction and validation from relational databases such as SQL Server, Oracle, or PostgreSQL.
Document pipeline designs, data flow diagrams, and operational support procedures.
10+ years of hands-on experience in Python for backend or data engineering projects.
Strong understanding and working experience with GCP cloud services (especially Dataflow, BigQuery, Cloud Functions, Cloud Composer, etc.).
Working experience with Azure Data Factory (ADF), Azure Databricks, Azure Data Lake Storage Gen2 (ADLS).
Solid understanding of data pipeline architecture, data integration, and transformation techniques.
Experience in working with version control systems like GitHub and knowledge of CI/CD practices.
Experience in Apache Spark, Kafka, Redis, Fast APIs, Airflow, GCP Composer DAGs.
Strong experience in SQL with at least one enterprise database (SQL Server, Oracle, PostgreSQL, etc.).
Experience with PySpark is required.
Experience in data migrations from on-premise data sources to Cloud platforms.
Good to Have (Optional Skills):
Experience with AWS services.
Additional Details:
Excellent problem-solving and analytical skills.
Strong communication skills and ability to collaborate in a team environment.
Education:
● Bachelor's degree in Computer Science, a related field, or equivalent experience.
Free · no signup
Daily job drops, skill trends and free resources - posted straight to the group. Leave any time.
No spam. Just jobs and resources.
Why people use ASAI
Scored, not searched. Every role ranked against your actual profile.
Alerts as often as hourly. Reach new roles while the pile is still small.
Skill gaps, spelled out. See exactly which requirements you don't meet yet.
Verified jobs, only. Say no to ghost jobs. Your time deserves respect.
More Data Engineer roles in Hyderabad
See allKeep browsing