Mumbai · Staff/Principal
Applicants who checked fit first are 3.1× more likely to hear back
Your score for this role already exists
ASAI compared this JD against 41 signals - skills, seniority, domain, stack overlap etc. Add a resume and it unlocks in about 30 seconds.
No credit card · 1 tap with Google
LOW
Gdtc is showing limited hiring activity lately. Expect slight delayed response.
First 72 hours
Window passed - posted 991d agoEarly applicants get seen before the pile builds.
Not a repost
The first time we've seen this listing - it hasn't been closed and reopened.
You almost certainly match several of these already. Unlock your skill map to see the matches, the gaps, and what to fix first.
Location: Mumbai
Experience: 4yrsto 8yrs
Technologies / Skills: Advanced SQL, Python and associated librarieslike Pandas, Numpy etc., Pyspark , Shell scripting, DataModelling, Big data, Hadoop, Hive, ETL pipelines.
Responsibilities:
• Proven successin communicating with users, other technical teams, and senior management to collect requirements, describe data modeling decisions and develop data engineering strategy.
• Ability to work with business ownersto define key businessrequirements and convert to user stories with required technical specifications.
• Communicate results and businessimpacts of insight initiatives to key stakeholders to collaboratively solve business problems.
• Working closely with the overall Enterprise Data & Analytics Architect and Engineering practice leads to ensure adherence with the best practices and design principles.
• Assures quality, security and compliance requirements are met forsupported area.
• Design and create fault-tolerance data pipelinesrunning on cluster
• Excellent communication skills with the ability to influence client business and IT teams
• Should have design data engineering solutions end to end. Ability to come up with scalable and modular solutions
Required Qualification:
• 3+ years of hands-on experience Designing and developing Data Pipelinesfor Data Ingestion or Transformation using Python (PySpark)/Spark SQL in AWS cloud
• Experience in design and development of data pipelines and processing of data at scale.
• Advanced experience in writing and optimizing efficient SQL queries with Python and Hive handling Large Data Sets in Big-Data Environments
• Experience in debugging, tunning and optimizing PySpark data pipelines
• Should have implemented concepts and have good knowledge of Pyspark data frames, joins, caching, memory management, partitioning, parallelism etc.
• Understanding of Spark UI, Event Timelines, DAG, Spark config parameters, in order to tune the long running data pipelines.
• Experience working in Agile implementations
• Experience with building data pipelinesin streaming and batch mode.
• Experience with Git and CI/CD pipelines to deploy cloud applications
• Good knowledge of designing Hive tables with partitioning for performance.
Desired Qualification:
• Experience in data modelling
• Hands on creating workflows on any Scheduling Tool like Autosys, CA Workload Automation
• Proficiency in using SDKsfor interacting with native AWS services
• Strong understanding of concepts of ETL, ELT and data modeling.
Free · no signup
Daily job drops, skill trends and free resources - posted straight to the group. Leave any time.
No spam. Just jobs and resources.
Why people use ASAI
Scored, not searched. Every role ranked against your actual profile.
Alerts as often as hourly. Reach new roles while the pile is still small.
Skill gaps, spelled out. See exactly which requirements you don't meet yet.
Verified jobs, only. Say no to ghost jobs. Your time deserves respect.
More Data Engineer roles in Mumbai
See allKeep browsing