Loading page...
Loading page...
Join freshers getting daily off-campus drives, internships & Remote jobs
Loading job details...
Corning Incorporated
Experience / Eligibility
B.E / B.Tech (Freshers / 0-1 Yr)
Salary
Not Disclosed / As per Industry Standards
Location
Pune, Maharashtra, India
As a data engineer for our advanced analytics platforms, your main responsibilities will be:
Design, test, deploy and maintain production big-data ingestion pipelines using established frameworks, patterns of practice, agile software development and CI/CD practices, working closely with the Principal Software Engineer – Data Ingestion
Work with cross-organizational data source teams to define data ingestion requirements for structured, unstructured and semi-structured data, pilot their implementation, ensure the data source teams accept the resulting landed data as valid
Define and implement automated validation and profiling capabilities needed to ensure reliable data delivery, using agile software development and CI/CD practices
Work with data source teams, domain experts and data scientists to define data cleansing and data enrichment requirements for landed data
Implement data cleansing and enrichment code using established patterns of practice General - Corning (L4)
Work with data source teams, domain experts and data scientists to validate landed, cleansed and enriched data, using agile software development and CI/CD practices, while ensuring that the final datasets are directly usable by them without additional processing effort
Actively participate in code reviews and technical information sharing with your team members and the broader software engineering community at Corning
Stay up to date with industry standards and technological advancements that will improve the quality, productivity and performance of your work.
Provide support in a DevOps environment to monitor tokens, jobs and overall system performance. Qualifications
Bachelor's degree in computer science, engineering, mathematics, or a related technical discipline
Understand concepts of big data engineering, developing and maintaining ETL and ELT pipelines for data warehousing, on-premise and cloud data lake environments
Demonstrated production programming proficiency in at least one modern JVM language such as Java, Scala or Kotlin, as well as an interpreted declarative programming language such as Python
Entry level experience with AWS platform services, including AWS S3 & EC2, Data Migration Services (DMS), RDS, EMR, RedShift, Lambda, DynamoDB, CloudWatch, CloudTrail
Hands-on technical familiarity with Apache Spark architecture, S3, parquet and Delta Lake architecture, technologies and tools
Basic proficiency with both traditional relational and polyglot persistence technologies
Familiarity with agile software development & continuous integration + continuous deployment methodologies along with supporting tools such as Git (Gitlab), Jira, Terraform, New Relic
Familiarity with notebook environments including JupyterHub
Able to Communicate well with users, other technical team to collect requirements, describe data modeling decisions.