Infrastructure EngineerHOAi is a fast-growing startup revolutionizing the community association management industry. Our AI workforce platform integrates machine learning technology to streamline labor-heavy processes, eliminating inefficiencies and driving scalability. With rapid growth in the AI space, we are pushing boundaries to redefine industry standards. HOAi is the leading AI solution for the community association management industry, enabling organizations to deploy AI agents that function like experienced managers.
These AI agents go beyond traditional AI by proactively executing complex, multi-step processes with human-like reasoning—working autonomously, 24/7, across your entire operation. This transformation optimizes labor costs, enables growth without additional hires, and ensures faster, higher-quality service for residents and board members. HOAi was acquired by Vantaca in the fall of 2024. Vantaca just achieved unicorn status with a $1.25B valuation, so it's safe to say we're past the "scrappy startup phase." We're not just building a successful company – we're building the category-defining platform that will transform how an entire industry operates.
Here's the reality of our trajectory:Growing 100% year-over-yearOur AI product (HOAi) went from $0 to millions in monthsBacked by Cove Hill Partners and JMI Private Equity6M+ doors on our platform, displacing legacy systemsThe AI Infrastructure Engineer at HOAi is responsible for scaling and maintaining the infrastructure that powers our AI-driven products and services. This role sits at the intersection of infrastructure engineering, machine learning operations, and product development, ensuring our AI systems operate with exceptional reliability, performance, and efficiency. The ideal candidate is someone who gets excited about making AI systems fundamentally faster and more scalable. You'll work directly with our engineering and product teams to build the foundational infrastructure that enables HOAi to deliver the most advanced AI product in the community association management industry.Infrastructure Ownership : Design, build, and maintain the cloud architecture, model serving infrastructure, and ML pipelines that power HOAi's productsPerformance Optimization : Profile and optimize AI workloads to achieve sub-second inference latency while managing costs effectivelyScalability & Reliability : Build auto-scaling systems, implement robust failover mechanisms, and ensure 99.99% uptime for mission-critical AI servicesMLOps Excellence : Develop and maintain CI/CD pipelines for model deployment, monitoring, and versioning across development and production environmentsDeveloper Enablement : Create tooling and infrastructure that allows product engineers to deploy AI features quickly and safelySecurity & Compliance : Implement security best practices and ensure compliance requirements are met across all AI infrastructureInfrastructure uptime and reliabilityAI inference latency (p95, p99) and throughput metricsInfrastructure cost efficiency and optimization (cost per inference, GPU utilization)Time to deploy new models and workflows (deployment velocity)Developer satisfaction and productivity using AI infrastructure toolsSystem observability and incident response timePerformance & ScalabilityProfile and optimize database queries, API endpoints, and ML inference pipelinesImplement caching strategies, connection pooling, and distributed systems for scaleMonitor and optimize GPU utilization, memory usage, and compute costsDesign load balancing and auto-scaling policies for variable AI workloadsBuild disaster recovery systems with redundancyMLOps & DeploymentBuild and maintain CI/CD pipelines specifically for model deploymentImplement model versioning, A/B testing infrastructure, and rollout mechanismsCreate automated testing frameworks for model quality and performance regressionDevelop infrastructure for model monitoring, drift detection, and retraining workflowsManage experiment tracking and model registry systemsObservability & ReliabilityImplement comprehensive monitoring, logging, and alerting across the AI stackRefine dashboards for real-time visibility into system health and performanceConduct post-mortems and implement reliability improvementsDesign circuit breakers, retry logic, and graceful degradation for critical servicesSecurity & ComplianceRefine security best practices for AI infrastructure and data handlingEnsure compliance with data privacy regulations and industry standardsManage credentials and access control across infrastructureSupport security audits and vulnerability assessmentsCollaboration & DocumentationWork closely with Product & Engineering team to understand infrastructure needs and to enable fast, safe feature deploymentDocument infrastructure architecture, runbooks, and operational proceduresMentor team members on infrastructure best practices and toolingContribute to technical strategy and architectural decisionsRequired Experience8+ years of experience in infrastructure engineering, DevOps, or SREStrong cloud platform expertiseExperience building and maintaining deployment pipelinesExperience with PostgreSQL, Redis, or other production databasesExperience with APM tools, metrics, logging, and alertingFamiliarity with vector databases, model serving frameworks and cross-system observability and traceabilityManaging and optimizing GPU workReal-time inference with low-latency serving infrastructureLLM deploymentTrack record of achieving 10x performance improvementsSkills & CompetenciesAble to debug complex distributed systems and find root causesObsessed with latency, throughput, and resource efficiencyDefaults to automating repetitive tasks and building scalable solutionsUnderstands security implications and implements best practicesAble to explain complex technical concepts clearlyWorks effectively across teams and functionsTakes initiative to identify and solve problems before they become criticalComfortable with ambiguity and changing priorities in a fast-moving startupSupporting A/B deployment strategyMindset & ApproachExtreme ownership: Takes full responsibility for outcomes, not just inputsCustomer-focused: Understands how infrastructure decisions impact end-user experienceData-driven: Makes decisions based on metrics and evidenceContinuous learner: Stays current with evolving technologies and best practicesQuality-focused: Builds systems that are reliable, maintainable, and elegantVelocity-minded: Balances speed with quality; ships incrementally and iteratesBar raiser: Sets and maintains high standards; elevates team performance and output qualityStrong delegator: Empowers others effectively; distributes work based on strengths and growth opportunitiesAlways Growing: Likes change and enjoys finding new ways to improve their knowledge and the product.
Always ready to learn quickly, helping themselves and the team grow.Win as a Team: Builds trust and works together by making sure everyone communicates well. Actively involved in daily work, working closely with the team, listening to their ideas, and celebrating successes together.Accountability Starts with Me: Notices problems and takes personal action to solve them.Unwavering Commitment to Customer Experience: Regularly talks to customers, taking personal responsibility to understand what they need, address concerns, and make their experience better with improved Vantaca processes.Innovate Boldly: We challenge the status quo and push boundaries to create meaningful change. We act with urgency and purpose, knowing that innovation drives our success.Shoot for Impossible and Make