Engagement Type: Fractional (Advisory) to start, opportunity to move into full time role following initial engagement.
Compensation: $250–$500 per hour (depending on experience)
Time Commitment: 5–10 hours per week to start
Initial Duration: 4–12 weeks
Location: United States only. Boston area preferred (within a few hours’ drive) for occasional in-person meetings with founder.
About Us
We are a bootstrapped AI company building a high-throughput research and intelligence engine for the public-sector market.
We ingest and analyze public records at scale to identify government agencies entering active buying cycles for our clients’ solutions.
- $100K revenue in first 6 months
- 80%+ retention rate
- Annual agreements with brand-name companies
- Projecting $500K revenue this year
- On track for profitability
- Lean, senior team
Our initial internal AI research platform was built by one senior engineer and has successfully supported early customer growth.
Now, as customer volume increases and use cases diversify, we are adding senior talent and need to architect the next generation of our internal and external tools to support 100x current capacity.
The Role
We are seeking a Principal-level AI Infrastructure Engineer to:
- Review and pressure-test our current architecture
- Design a next-generation, LLM-agnostic system capable of 100x scale
- Help guide and support implementation of that architecture
This is an engineering and systems role — not a management position.
There is a clear opportunity to evolve into a full-time lead engineer role after the initial engagement for the right person.
Engagement Phases
Phase 1 – Architecture & Codebase Review
- Review current system architecture and codebase
- Evaluate LLM usage patterns and token efficiency
- Assess API orchestration, rate limiting, batching, queuing, and retry logic
- Identify bottlenecks, fragility points, and scaling risks
- Deliver a structured architectural assessment
Phase 2 – Next-Generation Architecture Design (100x Scale)
- Design a scalable, LLM-agnostic AI architecture
- Plan for 100x current throughput
- Architect for:
- Token and inference cost control
- Provider abstraction (closed + open models)
- Resilience and fallback routing
- Distributed job orchestration
- High-concurrency environments
- Advise on local vs hosted inference strategy
- Evaluate GPU cost, latency, and inference tradeoffs
Phase 3 – Implementation Support
- Guide implementation of the new architecture
- Review critical technical decisions during buildPressure-test scaling assumptions
- Help prevent structural technical debt
Required Experience
- Built and scaled LLM-agnostic systems
- Scaled AI or API-heavy systems under real production load
- Experience operating at billion-token-per-day scale (or comparable throughput environments)Deep expertise in rate limits, retries, batching, queuing, and distributed failure modes
- Designed token-efficient architectures
- Worked with both closed-model providers and open-source models
- Deployed models locally or within controlled infrastructure
- Evaluated GPU cost, latency, and inference tradeoffs
Preferred Background
- Former CTO, Principal Engineer, or Staff Engineer
- Experience at a VC-backed startup with a successful outcome or a major public technology company
- History of scaling AI-native or API-intensive systems
- Comfortable collaborating closely with a technically involved founder and senior engineer
- Systems-oriented, pragmatic, and product-aware
Why This Is Interesting
- Strong early product-market fit
- Real production workload and scaling pressure
- High ownership and architectural influence
- Lean team with meaningful upside
- Clear path to deeper involvement for the right person