Job In A Week – get a job in a weekJob In A Week – get a job in a week
JobsInternshipsCompaniesSalariesResumePricingPost a job
Sign in
Jobs / Backend developer jobs in India / Backend developer jobs in Bengaluru /
S

Data Engineer – PySpark, Hadoop, Hive & Big Data Pipeline Development

Synechron · Bengaluru

Posted
today
Experience
6–10 yrs
Pay
Not stated
Role
Backend developer
SparkStakeholder managementTroubleshootingSQLPythonAgileProblem solving
Apply on Synechron's siteOpens the company's own careers page.

About the job

Job Summary Synechron is seeking a Data Engineer with strong expertise in PySpark and the Cloudera Data Platform (CDP) to design, develop, and maintain scalable data pipelines. The role is responsible for building reliable data ingestion and transformation processes, working with large-scale datasets, and supporting distributed data-processing environments. The Data Engineer will contribute to business objectives by improving data availability, quality, integrity, processing efficiency, and operational reliability. The successful candidate will collaborate with cross-functional teams to deliver data-driven solutions that support business and technology requirements.

Software Requirements

Required Software Skills

- PySpark: Strong hands-on experience developing, optimizing, and maintaining production-grade data pipelines.

- Cloudera Data Platform (CDP): Strong practical experience working with CDP environments and related data-processing capabilities.

- Apache Spark: Hands-on experience with Spark processing, job execution, performance tuning, and troubleshooting.

- Hadoop: Experience working with Hadoop-based distributed data ecosystems.

- Hive: Experience developing and optimizing Hive queries, tables, and data-processing workflows.

- HDFS: Experience managing and processing data stored in the Hadoop Distributed File System.

- ETL tools and processes: Practical experience designing data ingestion, transformation, validation, and loading workflows.

- Data warehousing technologies: Experience working with data warehouse concepts, structures, and processing patterns.

- Cloud and distributed data environments: Experience with cloud-based or distributed data-processing platforms.

- Version experience: Experience with the versions of PySpark, Spark, Hadoop, Hive, HDFS, CDP, and associated tools adopted by the assigned Synechron project; ability to work with current supported releases and understand version-related compatibility considerations.

Preferred Software Skills

- Experience with cloud-native data platforms and managed distributed-processing services.

- Familiarity with workflow orchestration, data- pipeline monitoring, scheduling, and alerting tools.

- Experience with source-control, code-review, build, deployment, and automated testing tools.

- Familiarity with data-quality, metadata-management, lineage, and observability tools.

- Experience working with containerized or infrastructure-as-code environments.

- Exposure to streaming, real-time processing, or event-driven data platforms. Equivalent tools may be considered where they provide comparable capabilities.

Overall Responsibilities

- Design and develop scalable data pipelines using PySpark and Cloudera Data Platform (CDP) .

- Build, optimize, and maintain data ingestion and transformation processes for large-scale datasets.

- Develop reusable, maintainable, and testable data-processing components.

- Implement data validation and quality checks to support data accuracy, consistency, integrity, and availability.

- Work with Hadoop, Hive, HDFS, Spark, ETL processes, and related big data technologies.

- Apply data warehousing and big data architecture principles when designing data solutions.

- Collaborate with cross-functional teams to understand data requirements and deliver data-driven solutions.

- Monitor data workflows, processing jobs, resource utilization, and pipeline performance.

- Investigate and resolve data-quality issues, job failures, performance bottlenecks, and operational incidents.

- Optimize Spark jobs, PySpark code, queries, data layouts, partitioning, and resource usage where required.

- Document data pipelines, transformation logic, dependencies, operational procedures, and known limitations.

- Support deployment, release, testing, and production-readiness activities for data solutions.

- Contribute to improvements in data engineering standards, development practices, monitoring, and operational support.

- Consider sustainable engineering practices by reducing unnecessary data movement, optimizing compute and storage usage, reusing pipeline components, and supporting maintainable solutions.

- Deliver data pipelines that meet agreed requirements for quality, performance, reliability, security, and availability.

- Escalate material risks, dependencies, data-quality concerns, and delivery constraints through appropriate channels.

Technical Skills (By Category)

Programming Languages

Essential

- Strong hands-on experience with Python , particularly for PySpark-based data processing.

- Experience writing maintainable, modular, testable, and performance-conscious data-processing code.

- Ability to use SQL for data querying, validation, transformation, and analysis.

Preferred

- Familiarity with Scala or another language used in distributed data processing.

- Experience writing shell scripts for data operations, job execution, monitoring, or automation. Databases/Data Management

Essential

- Hands-on experience with Hive, HDFS, Hadoop, and data warehousing concepts .

- Understanding of structured, semi-structured, and large-scale distributed data.

- Experience designing data ingestion, transformation, storage, and retrieval processes.

- Knowledge of data quality, data integrity, validation, partitioning, schema management, and data availability.

- Understanding of ETL processes and big data architecture.

Preferred

- Experience with relational and non-relational databases.

- Familiarity with data lakes, lakehouse architectures, metadata management, data lineage, or data cataloguing.

- Experience working with incremental loads, change-data processing, historical data, and schema evolution.

- Exposure to batch and streaming data-processing patterns.

Cloud Technologies

Essential

- Experience working in cloud or distributed data-processing environments.

- Understanding of cloud-related considerations such as scalability, storage, compute usage, availability, access controls, and operational monitoring.

- Ability to design or support data pipelines that run reliably in distributed environments.

Preferred

- Experience with cloud-native data platforms, managed Spark services, object storage, or cloud-based data warehouses.

- Familiarity with cloud monitoring, logging, automation, and cost-optimization practices.

- Experience migrating or modernizing Hadoop or CDP workloads in cloud environments.

Frameworks and Libraries

Essential

- Strong experience with PySpark and Apache Spark .

- Experience using Spark DataFrame, SQL, transformation, action, partitioning, and optimization capabilities.

- Ability to select appropriate Spark processing patterns for scalability, reliability, and performance.

- Experience working with ETL frameworks or reusable data-pipeline components.

Preferred

- Familiarity with structured streaming or other distributed streaming frameworks.

- Experience with data-validation, testing, serialization, or data-format libraries.

- Exposure to reusable pipeline frameworks and configuration-driven processing.

Development Tools and Methodologies

Essential

- Experience developing, testing, deploying, monitoring, and supporting data pipelines across development and production environments.

- Familiarity with source control, code reviews, defect tracking, and controlled release practices.

Apply now

Do you fit this job?

Create your profile and every open job, this one included, gets a fit score for you.

More jobs like this

  • Backend developer jobs in Bengaluru
  • All jobs in Bengaluru
  • All jobs at Synechron

Similar jobs

All backend developer jobs →
O

Staff Software Engineer, Platform & Infrastructure

Okta
Bengaluru5+ yrsBackend developer
CI/CDKubernetesGoPythonAWSGCP
Posted todayApply
R

Software Engineer - Cloud: Azure

Rubrik
Bengaluru0–2 yrsBackend developerEarly career (0–2 yrs)
AzurePythonJavaMicroservicesDockerKubernetes
Posted todayApply
AN

Software Engineer — Test Automation

Arista Networks
Bengaluru3–6 yrsBackend developer
CI/CDGoREST APIsTroubleshootingManual testingPython
Posted todayApply
PS

Member of Technical Staff, Backend Engineering- Portworx Enterprise

Pure Storage
Bengaluru7+ yrsBackend developer
MicroservicesGoProduct managementStakeholder managementTroubleshooting
Posted todayApply
ZT

Software Engineer Professional II

Zebra Technologies
Bengaluru5+ yrsBackend developer
Stakeholder managementGitProblem solvingTroubleshootingCI/CDNLP
Posted todayApply
A

Software Development Engineer II, EU INTech Partner Growth Experience

Amazon
Bengaluru3+ yrsBackend developer
JavaC#NLP
Posted todayApply

Jobs by role

  • Backend developer jobs in India
  • Full-stack developer jobs in India
  • Program / project manager jobs in India
  • Data scientist jobs in India
  • Data analyst jobs in India
  • Consultant jobs in India
  • Product manager jobs in India
  • Strategy & operations jobs in India
  • QA / Test engineer jobs in India
  • DevOps engineer jobs in India
  • Frontend developer jobs in India
  • Tech support jobs in India
  • Research analyst jobs in India
  • Mobile developer jobs in India

Jobs by city

  • Jobs in Bengaluru
  • Jobs in Hyderabad
  • Jobs in Gurugram
  • Jobs in Pune
  • Jobs in Chennai
  • Jobs in Mumbai
  • Jobs in Noida
  • Jobs in Delhi
  • Jobs in Coimbatore
  • Jobs in Ahmedabad
  • Jobs in Kolkata
  • Jobs in Chandigarh

Popular searches

  • Backend developer jobs in Bengaluru
  • Backend developer jobs in Hyderabad
  • Full-stack developer jobs in Bengaluru
  • Program / project manager jobs in Bengaluru
  • Data scientist jobs in Bengaluru
  • Product manager jobs in Bengaluru
  • Backend developer jobs in Chennai
  • Data analyst jobs in Bengaluru
  • Backend developer jobs in Pune
  • Consultant jobs in Bengaluru
  • QA / Test engineer jobs in Bengaluru
  • Backend developer jobs in Gurugram
  • Full-stack developer jobs in Hyderabad
  • DevOps engineer jobs in Bengaluru
  • Frontend developer jobs in Bengaluru
  • Program / project manager jobs in Hyderabad
  • Strategy & operations jobs in Bengaluru
  • Data analyst jobs in Hyderabad

JobInAWeek

  • How to get a job in a week
  • Career guides
  • Post a job (free)
  • Create your profile
  • Internships & fresher jobs
  • Companies hiring
  • Salary check
  • Swipe jobs
  • Resume builder
  • Pricing
  • Privacy
  • Terms
  • Refunds & cancellation
  • Contact
  • About our job crawler
Job In A Week – Skills today. Job tomorrow.Job In A Week – Skills today. Job tomorrow.

Job In A Week (jobinaweek.com) helps you get a job in a week: fresh jobs from India's top companies, matched to your profile. Jobs come from company careers pages and link to the company's own apply page. Updated every two hours. © 2026 Job In A Week.