Unknown Company

Principal AI Infrastructure Engineer (Part-time -> Full time)

boston, ma • Posted 1 weeks ago
Onsite Full Time Networks & Systems

Engagement Type: Fractional (Advisory) to start, opportunity to move into full time role following initial engagement.

Compensation: $250–$500 per hour (depending on experience)

Time Commitment: 5–10 hours per week to start

Initial Duration: 4–12 weeks

Location: United States only. Boston area preferred (within a few hours’ drive) for occasional in-person meetings with founder.

About Us

We are a bootstrapped AI company building a high-throughput research and intelligence engine for the public-sector market.

We ingest and analyze public records at scale to identify government agencies entering active buying cycles for our clients’ solutions.

  • $100K revenue in first 6 months
  • 80%+ retention rate
  • Annual agreements with brand-name companies
  • Projecting $500K revenue this year
  • On track for profitability
  • Lean, senior team

Our initial internal AI research platform was built by one senior engineer and has successfully supported early customer growth.

Now, as customer volume increases and use cases diversify, we are adding senior talent and need to architect the next generation of our internal and external tools to support 100x current capacity.

The Role

We are seeking a Principal-level AI Infrastructure Engineer to:

  • Review and pressure-test our current architecture
  • Design a next-generation, LLM-agnostic system capable of 100x scale
  • Help guide and support implementation of that architecture

This is an engineering and systems role — not a management position.

There is a clear opportunity to evolve into a full-time lead engineer role after the initial engagement for the right person.

Engagement Phases

Phase 1 – Architecture & Codebase Review

  • Review current system architecture and codebase
  • Evaluate LLM usage patterns and token efficiency
  • Assess API orchestration, rate limiting, batching, queuing, and retry logic
  • Identify bottlenecks, fragility points, and scaling risks
  • Deliver a structured architectural assessment

Phase 2 – Next-Generation Architecture Design (100x Scale)

  • Design a scalable, LLM-agnostic AI architecture
  • Plan for 100x current throughput
  • Architect for:
    • Token and inference cost control
    • Provider abstraction (closed + open models)
    • Resilience and fallback routing
    • Distributed job orchestration
    • High-concurrency environments
    • Advise on local vs hosted inference strategy
    • Evaluate GPU cost, latency, and inference tradeoffs

Phase 3 – Implementation Support

  • Guide implementation of the new architecture
  • Review critical technical decisions during buildPressure-test scaling assumptions
  • Help prevent structural technical debt

Required Experience

  • Built and scaled LLM-agnostic systems
  • Scaled AI or API-heavy systems under real production load
  • Experience operating at billion-token-per-day scale (or comparable throughput environments)Deep expertise in rate limits, retries, batching, queuing, and distributed failure modes
  • Designed token-efficient architectures
  • Worked with both closed-model providers and open-source models
  • Deployed models locally or within controlled infrastructure
  • Evaluated GPU cost, latency, and inference tradeoffs

Preferred Background

  • Former CTO, Principal Engineer, or Staff Engineer
  • Experience at a VC-backed startup with a successful outcome or a major public technology company
  • History of scaling AI-native or API-intensive systems
  • Comfortable collaborating closely with a technically involved founder and senior engineer
  • Systems-oriented, pragmatic, and product-aware

Why This Is Interesting

  • Strong early product-market fit
  • Real production workload and scaling pressure
  • High ownership and architectural influence
  • Lean team with meaningful upside
  • Clear path to deeper involvement for the right person

#J-18808-Ljbffr
Back to Job Search