Responsibilities
- Develop Java, Scala and CUDA/C++ libraries to accelerate DataFrames and I/O operations on common file formats such as Parquet, ORC and JSON
- Enable interoperability with table formats such as Apache Iceberg and Delta Lake, and metastores such as Unity Catalog
- Work with open source communities to enhance libraries like NVIDIA cuDF, CCCL and UCX through technical discussion and code contributions
- Collaborate with distributed systems teams to craft solutions to distributed processing problems challenges at large scale
- Provide recommendations and feedback to teams regarding decisions surrounding topics such as infrastructure, continuous integration and testing strategy
- Build, test and optimize CUDA/C++ libraries across different platforms
Requirements
- BS, MS, or PhD in Computer Science, Computer Engineering, or closely related field (or equivalent experience)
- 15+ years of work experience in software development
- Outstanding technical skills in designing and implementing high-quality distributed systems
- Excellent programming skills in C++, Java, and/or Scala
- Ability to work with teams across organizational boundaries and geographies
- Familiarity with the open source data platform ecosystem (Apache Spark, Velox, Presto, Apache Arrow, Apache DataFusion, etc.). Meaningful contributions to the OSS community a plus.
- Highly motivated with strong interpersonal skills
- Database query optimization is a strong plus
Core Competencies
Demonstrates expertise in developing high-quality distributed systems using Java, Scala, and CUDA/C++. Proficient in optimizing libraries for data processing and contributing to open source projects within the data platform ecosystem.
#J-18808-Ljbffr