Bachelor’s, Master’s, or PhD in Computer Science, Machine Learning, Applied Statistics, Mathematics, Engineering, Physics, or other technical fields
2+ years of professional experience applying machine learning and data mining techniques to solve real-world problems with substantial data sets
Programming experience (focus on machine learning): SQL and Python’s Data Science stack; good knowledge of at least one big data framework (e.g., PySpark, Hive, Hadoop) is a plus. R, SPSS, and SAS are considered nice-to-have
Strong understanding of machine learning methods and experience applying them to complex, data-rich environments
Ability to prototype and deploy statistical and machine learning algorithms, and translate analytical outputs into data-driven solutions
Experience deploying ML/AI technologies into production or applied business environments is a plus
While we advocate using the right tech for the right task, we often leverage: Python, PySpark, the PyData stack, SQL, Airflow, Databricks, Kedro (our open-source data pipelining framework), Dask/RAPIDS, Docker, Kubernetes, and cloud solutions such as AWS, GCP, and Azure
Familiarity with Generative AI (GenAI) and agentic systems is a strong plus
Excellent time management skills to handle responsibilities in a complex and largely autonomous environment
Willingness to travel
Strong communication skills, both verbal and written, in English with the ability to adapt your style to different audiences and seniority levels