Unknown Company

VLM Data Science Expert

san ramon, ca • Posted 6 days ago
Onsite Full Time IT & Technology

Overview

At CitiusTech , we constantly strive to solve the industry\'s greatest challenges with technology, creativity, and agility. With over 8,500 healthcare technology professionals worldwide, CitiusTech powers healthcare digital innovation, business transformation, and industry-wide convergence for over 140 organizations through next-generation technologies, solutions, and products. We aim to accelerate the transition to a human-first, sustainable, and digital healthcare ecosystem with the world\'s leading Healthcare and life sciences organizations and our partners.

Here is an opportunity for you to make a difference and collaborate with global leaders to shape the future of healthcare and positively impact human lives.

Our vision: To inspire new possibilities for the health ecosystem with technology and human ingenuity.

Base pay range

$160,000.00/yr - $200,000.00/yr

Direct message the job poster from CitiusTech

To learn more about CitiusTech, visit

What is in it for you?

If you\'re a Senior Data Scientist with a strong background in Vision-Language Models (VLMs), this is a chance to lead the charge in building smart, scalable multimodal AI solutions. We’re looking for someone who’s worked hands-on with cutting-edge frameworks like VILA, Isaac, and VSS—and who knows how to take models from concept to production in real-world settings. If you\'ve got experience in healthcare, especially with medical devices, that\'s a big plus. You\'ll be diving into the latest VLM techniques and deploying them on cloud platforms like AWS, helping shape the future of AI in a meaningful, impactful way.

Key Responsibilities

  • Design, train, and deploy efficient Vision-Language Models (e.g., VILA, Isaac Sim) for multimodal applications including image captioning, visual search, and document understanding, pose understanding, pose comparison.
  • Develop and manage Digital Twin frameworks using AWS IoT TwinMaker, SiteWise, and Greengrass to simulate and optimize real-world systems.
  • Develop Digital Avatars using AWS services integrated with 3D rendering engines, animation pipelines, and real-time data feeds.
  • Explore cost-effective methods such as knowledge distillation, modal-adaptive pruning, and LoRA fine-tuning to optimize training and inference.
  • Implement scalable pipelines for training/testing VLMs on cloud platforms (AWS services such as SageMaker, Bedrock, Rekognition, Comprehend, and Textract.)

NVIDIA Platforms

  • Should develop a blend of technical expertise, tool proficiency, and domain- specific knowledge on below NVIDIA Platforms:
  • NeMo Framework: Training and scaling VLMs across thousands of GPUs.
  • DeepStream SDK: Integrates pose models like TRTPose and OpenPose, Real-time video analytics and multi-stream processing.

Multimodal AI Solutions

  • Develop solutions that integrate vision and language capabilities for applications like image-text matching, visual question answering (VQA), and document data extraction.
  • Leverage interleaved image-text datasets and advanced techniques (e.g., cross-attention layers) to enhance model performance.

Image Processing and Computer Vision

  • Develop solutions that integrate Vision based deep learning models for applications like live video streaming integration and processing, object detection, image segmentation, pose Estimation, Object Tracking and Image Classification and defect detection on medical Xray images
  • Knowledge of real-time video analytics, multi-camera tracking, and object detection.
  • Training and testing the deep learning models on customized data

Healthcare Domain Expertise (Nice to Have)

  • While it’s not a must, having experience in the healthcare space—especially with medical imaging, motion detection, or patient monitoring—can be a big advantage.
  • You’ll be applying Vision-Language Models to use cases like analyzing scans, detecting positioning and movement, and making precise measurements.
  • If you\'re familiar with healthcare standards and know how to handle sensitive data responsibly, that’s a definite plus.
  • Evaluate trade-offs between model size, performance, and cost using techniques like elastic visual encoders or lightweight architectures.
  • Benchmark different VLMs (e.g., GPT-4V, Claude 3.5, Nova Lite) for accuracy, speed, and cost-effectiveness on specific tasks.
  • Benchmarking on GPU vs CPU
  • Collaborate with cross-functional teams including engineers and domain experts to define project requirements.
  • Mentor junior team members and provide technical leadership on complex projects.

Experience

  • 10+ Years

Location

  • San Ramon, CA or Milwaukee, WI (Onsite)

Qualifications

  • Education: Master’s or Ph.D. in Computer Science, Data Science, Machine Learning, or a related field.

Experience

  • Minimum of 10+ years of experience in Machine Learning or Data Science roles with a focus on Vision-Language Models.
  • Proven expertise in deploying production-grade multimodal AI solutions.
  • Experience in self driving cars and self navigating robots.

Technical Skills

  • Proficiency in Python and ML frameworks (e.g., PyTorch, TensorFlow).
  • Hands-on experience with VLMs such as VILA, Isaac Sim, or VSS.
  • Familiarity with cloud platforms like AWS SageMaker or Azure ML Studio for scalable AI deployment.
  • CUDA, cuDNN

Domain Knowledge — A Valuable Bonus

It’s helpful if you’ve got a solid grasp of medical datasets, especially imaging data, and an understanding of healthcare regulations.

Knowing how to navigate the complexities of clinical data and compliance can really elevate your impact in this role

Soft Skills

  • Strong problem-solving skills with the ability to optimize models for real-world constraints.
  • Excellent communication skills to explain technical concepts to diverse stakeholders.

Preferred Technologies

  • Multimodal Techniques: Cross-attention layers, interleaved image-text datasets
  • MLOps Tools: Docker, MLflow

Life at CitiusTech

We focus on building highly motivated engineering teams and thought leaders with an entrepreneurial mindset, centered on our core values of Passion, Respect, Openness, Unity, and Depth (PROUD) of knowledge . Our success lies in creating a fun, transparent, non-hierarchical, diverse work culture that focuses on continuous learning and work-life balance.

Rated by our employees as the ‘Great Place to Work ’ for’ according to the Great Place to Work survey. We offer you a comprehensive set of benefits to ensure that you have a long and rewarding career with us.

This posting complies with applicable pay transparency laws in states such as California, Colorado, New York, New Jersey, Illinois, and others.

Equal Opportunity Employer . We celebrate diversity and are committed to creating an inclusive environment for all employees. All qualified applicants will receive consideration for employment without regard to race, color, religion, gender, gender identity or expression, sexual orientation, national origin, genetics, disability, age, or veteran status

Our EVP

Be You Be Awesome is our EVP and it reflects our continuing efforts to create CitiusTech as a great place to work where our employees can thrive, both personally and professionally. It encompasses the unique benefits and opportunities we offer to support your growth, well-being, and success throughout your journey with us and beyond. Together with our clients, we are solving some of the greatest healthcare challenges and positively impacting human lives. Welcome to the world of Faster Growth, Higher Learning, and Stronger Impact.

Join CitiusTech. Be You. Be Awesome.

To learn more about CitiusTech, visit

#J-18808-Ljbffr
Back to Job Search