Unknown Company

Llama Developer (Generative AI / LLM Engineer)

ca • Posted 3 days ago
Onsite Full Time General

Job TitleBuild AI applications using Llama (Llama 3 / Llama Stack / Llama API / local LLM inference).Fine-tune and evaluate Llama models on proprietary and domain-specific datasets.Implement Retrieval-Augmented Generation (RAG) pipelines using vector databases.Develop conversational agents, copilots, or knowledge assistants for business workflows.Optimize model performance via quantization, prompt engineering, and latency reduction.Integrate LLM capabilities into back-end services, microservices, APIs, or cloud platforms.Ensure compliance, safety, and responsible-AI standards for generated content.Collaborate with Data Science, MLOps, and Product teams to deploy scalable AI products.Required Skills & Experience3–8+ years of software engineering or machine-learning experience.Proven experience with Llama models (self-hosted or via Meta API).Proficiency in Python (FastAPI, LangChain, LlamaIndex, Hugging Face ecosystem).Experience with vector databases (FAISS, Pinecone, Weaviate, ChromaDB).Strong understanding of prompt engineering, model fine-tuning, LoRA / QLoRA.Hands-on experience with GPU computing, PyTorch, Docker, Kubernetes.Familiarity with MLOps practices: CI/CD for ML, model monitoring, logging.Preferred / Nice to HaveExperience deploying models on AWS / GCP / Azure / on-prem GPU clusters.Knowledge of RAG architectures, knowledge graphs, and document parsing pipelines.Understanding of model safety, hallucination mitigation, red-team testing.Experience with llama.cpp, vLLM, Ollama, or NVIDIA Triton.Contributions to open-source LLMs or AI frameworks.

Back to Job Search