AI Architect (Voice AI)
Neurons Lab
We are seeking an AI Architect to lead the technical vision and delivery of a real-time voice copilot for the largest US network of in-home veterinary hospice and end-of-life care. This role is backed by a major US private-equity sponsor and involves spearheading the transition from a validated proof of concept to full production. The copilot, which we have developed, listens to live calls with pet families and extracts appointment and clinical details in real time, populating the client's scheduling system via a Chrome extension. A parallel initiative, the Vet Visit Copilot, delivers AI-generated pre-visit briefings to veterinarians through Amazon SES. The production phase is slated to begin in September 2026, with a multi-month timeline and a strong likelihood of extension. The commitment is a minimum of 0.5 FTE, scaling toward 1.0 FTE as production demands grow. The current architect is transitioning to another strategic project but will remain available at 0.15–0.2 FTE for supervision and knowledge transfer, ensuring a smooth handover. Key responsibilities include owning the end-to-end architecture and delivery of the voice copilot, meeting stringent benchmarks for latency, accuracy, concurrency, and cost. You will drive the reduction of P95 latency from ~6 seconds toward ~2 seconds, eliminate occasional ~1-minute lags in specific field processing, and manage model A/B testing (currently Claude Haiku vs. GPT Luna) using golden-set evaluations for phonetic name and email accuracy. You will also oversee Langfuse-based observability, cost tracking per live call, and optimization strategies. Production hardening is essential, including support for 5–10+ concurrent calls, strict data isolation, monitoring, alerting, and safe rollback procedures. You will ship features end-to-end, such as the SES email briefing service, while maintaining a demo fallback to ensure live sessions never fail. In this client-facing role, you will lead technical discussions with a meticulous client, while VCCs test edge cases and expect production-grade quality. You will present concrete system behavior backed by data, as this account values evidence over presentations. You will also manage scope by tying every feedback item to the SOW and routing roadmap items—like learning loops and persistent memory—to future phases. All client-facing materials must pass ADM review, and internal discussions stay internal. As the team lead, you will guide an AI Engineer and the pod, setting tasks, reviewing output, and removing blockers. You will absorb knowledge from the outgoing architect and conduct knowledge-transfer sessions to eliminate single points of failure. You will also provide estimates and architecture options to the account team when requested. Required skills include hands-on experience with real-time voice pipelines (streaming STT, turn handling, low-latency LLM inference), LLM engineering (prompt design, structured extraction, guardrails, A/B evaluation), observability and evals (Langfuse or similar, golden datasets, latency/accuracy/cost dashboards), and AWS services (Bedrock, serverless patterns, SES). You should have strong Python skills and enough TypeScript/Chrome-extension knowledge to own the integration. Excellent spoken and written English is a must for engaging with demanding US executives. Ideal candidates will have a background in contact-center or agent-assist metrics (handle time, cost per call, concurrency) and production LLM operations (load testing, data isolation, incident handling). Experience in empathy-sensitive domains like healthcare, veterinary, or insurance, and PE-sponsored rollouts is a plus. We are looking for someone with 6+ years in hands-on AI/ML engineering, including strong recent LLM production experience. You must have shipped at least one real-time or speech product to real users (e.g., agent assist, voice bot, live transcription copilot) and be able to demo real artifacts at the interview. A proven track record of reducing P95 latency and fixing concurrency issues in live systems is essential. Consulting or client-facing seniority, with the ability to remain calm and precise under detailed UAT scrutiny, is required. Experience with Chrome extensions, telephony/streaming stacks (Amazon Connect, Twilio, LiveKit), and Langfuse in production is a plus, as is familiarity with US clients and Eastern-time overlap.
- Início
- 1 de setembro de 2026
- Duração
- multi-month
- Publicado
- 27 de agosto de 2026
A carregar obras relevantes...