- Evaluate and select models from major and emerging providers using rigorous benchmarks and LLM-as-judge frameworks.
- Apply statistical analysis to make model decisions defensible.
AI Evaluation & Development:
- Design and maintain offline evaluation sets, judge calibration, and regression benchmarks.
- Build classification and fine-tuned models for routing, categorization, and detection.
- Develop agentic AI products and prompt/context engineering strategies.
Technical Leadership & Research:
- Partner with data engineering, software engineering, and product teams.
- Investigate model hallucinations, drift, and prompt sensitivity from first principles.
- Maintain high engineering standards for code quality, testing, and documentation.
Jobgether
A production agentic AI platform that builds and deploys advanced machine learning systems. The team is collaborative and values innovation, offering a remote work environment with high autonomy.
#J-18808-LjbffrSenior ML Specialist (Fully Remote) in workfromhome at Unknown Company
This position is listed as full time and able to be worked remotely.