This role will be responsible for building and maintaining the software infrastructure that enables computation over large data sets.
- Create, optimize and maintain optimal data collection, flow and pipeline for data science POCs and industrialized solutions
- Assemble, cleanse, integrate and transform large, complex data sets into model ready data
- Identify, design, and implement internal process improvements, including automating manual processes, optimizing data delivery, re-designing infrastructure for greater scalability, etc.
- Build the infrastructure required for optimal extraction, transformation, and loading of data from a wide variety of data sources using SQL and Cloudera ‘big data’ technologies
- Build data cleansing, management and visualization tools that provide actionable in-sights into data features that might be relevant for POCs and industrialized solutions
- Create data tools for Data Science and other analytics team members that assist them in building and optimizing POCs and industrialized solutions
- Work with stakeholders including Data Scientists, Data Architects, commercial and other stakeholders to assist with data-related technical issues and support their data infrastructure needs
- Work with data and analytics experts to strive for greater functionality in our data systems
