Skills required
Technologies this role asks for
Prepare for this role
Practice and Learn are coming soon for Python, Machine Learning — browse articles meanwhile, then apply.
Job description
What you’ll do in this role
Job Overview
Citi is hiring a Big Data Engineer - Python and Spark for its Data/Information Management function in Pune. This is a hybrid, full-time role within the Decision Management job family and focuses on working with large-scale data, data processing, software systems, and data management technologies.
The role is described as a trainee professional position requiring knowledge of relevant processes, procedures, and systems. The position involves analyzing factual information, resolving technical problems, improving data and compute processes, and communicating solutions clearly to stakeholders.
Key Responsibilities
- Work with large and multiple datasets, data warehouses, and programs used to retrieve and process data.
- Develop and work with data solutions using Python, Scala, SQL, and relevant big data technologies.
- Work with unstructured and undocumented code and contribute to code redesign and improvements for costly compute and data processes.
- Apply development standards while improving existing data and software processes.
- Perform data preprocessing and application engineering activities.
- Work with batch and real-time processing requirements for computationally intensive software systems.
- Apply supervised and unsupervised machine learning techniques where relevant to assigned work.
- Support data ingestion, ETL activities, data modeling, and related data management processes.
- Work with data management, data governance, data security, and regulatory practices.
- Identify and solve complex business problems and present solutions to management in a structured and understandable manner.
- Communicate information concisely and work effectively with diverse audiences and delivery teams.
Required Skills
- Experience with big data systems including Hive, Hadoop, and Spark.
- Hands-on programming experience with Python and Scala.
- Strong experience with SQL.
- Experience with Unix Scripting.
- Ability to work with large datasets and data warehouses and retrieve data using relevant programs and coding.
- Understanding of data preprocessing and application engineering.
- Understanding of supervised and unsupervised machine learning techniques.
- Knowledge of data management, governance, security, and regulatory practices.
- Strong problem-solving, analytical, communication, and presentation skills.
- High attention to detail.
Preferred and Additional Skills
- Exposure to data ingestion and ETL tools such as Talend.
- Exposure to modeling tools and performance management tooling such as Pepper Data.
- Exposure to the Cloudera stack.
- Previous related experience is preferred.
- Experience working in onsite and offsite delivery models.
Education
The requirements mention a Master’s or Engineering Degree for the role. The education section also states that a Bachelor’s or University degree, or equivalent experience, is accepted. No specific engineering branch is stated.
Experience
The qualifications section specifies 0 to 2 years of experience in big data systems and related technologies. It also separately mentions at least 3 years of experience designing software systems with intensive computational needs across real-time and batch processes. These requirements are both stated in the supplied job description.
Technologies
- Python – Programming language required for big data and application work.
- Scala – Programming language mentioned for big data programming.
- SQL – Strong SQL experience is required for working with data and data warehouses.
- Apache Spark – Big data processing technology required for the role.
- Apache Hadoop – Big data system experience is required.
- Apache Hive – Experience with Hive is explicitly required.
- Unix – Unix scripting experience is required.
- Machine Learning – Understanding of supervised and unsupervised techniques is required.
- Extract Transform Load (ETL) – ETL exposure is listed among the additional requirements.
Work Mode & Location
This is a hybrid full-time position based in Pune, Maharashtra, India. The role is part of Citi’s Decision Management job family and Data/Information Management function.
Who Should Apply
This opportunity is intended for candidates with an engineering or university degree and experience or knowledge relevant to big data systems, data engineering, programming, SQL, data processing, and data management. Candidates should be comfortable working with large datasets, existing codebases, and complex data or software problems.
How to Apply
Apply now through SoftoJobs to explore this opportunity.
Eligibility
Education, passing batch and experience
Education
Any Engineering Degree — All Branches
Experience
0 - 2 years (Freshers welcome)
Similar Fresher Jobs
More openings you may like