You will own the backbone that runs thousands of concurrent AI-powered phone calls inside bank-grade on-premise environments. Not just keeping pods alive, you architect distributed systems that handle real-time voice, scale STT / LLM / TTS inference across customer GPU clusters, integrate with enterprise telephony (Cisco CUBE, Genesys, Asterisk), and deploy behind the firewalls of largest financial institutions. Your work decides whether our platform answers a bank's rush-hour traffic or leaves customers on dead air.
• Scale GPU inference infrastructure for our STT, TTS, and LLM models across multiple customer environments (H100 / H200, NVLink, Triton or vLLM).
• Integrate with telephony: Asterisk, SIP trunks, Cisco CUBE, Genesys, WebRTC. SIP header parsing (X-Genesys-*), direction routing, warm transfers, DTMF.
• Own reliability: Splunk SIEM forwarding, Langfuse and Grafana observability, incident playbooks for bank-grade 24/7 SLAs.
• Security and compliance: RBAC, pentest remediation, KVKK and BDDK compliance patterns, pod security policies.
• Scale with growth: we onboard a new bank or insurer every quarter. Each is a new on-prem environment with its own constraints.
• Spot flaws early. We are building new architecture for a regulated industry. You help us see what needs to be solved next.
Interesting Problems to Own
• On-prem meets streaming. Most voice AI stacks assume cloud. We run the same stack inside banks with zero internet egress. Novel problems in image delivery, model updates, secrets rotation.
• Bank-scale concurrency. A single campaign can put millions of customers on the line the same afternoon. Queueing, graceful degradation, GPU-aware autoscaling are yours to design.
• Legacy-meets-new telephony. Cisco CUBE, Asterisk, Genesys, SIP, WebRTC. You wrangle old-school protocols alongside modern streaming stacks.
What Makes You a Great Fit
• 3+ years building and scaling distributed systems. Deep Kubernetes / OpenShift knowledge. AWS / GCP helpful.
• Fundamentals plus. You can sketch how a SIP INVITE flows through a proxy, explain a K8s GPU scheduler, or tell us the obscure thing you fell asleep reading last night.
• Real-time systems experience. Low-latency streaming or inference. Voice / video is a big plus.
• On-prem mentality. You have shipped software into environments you did not fully control: Turkish bank, European healthcare, US regulated finance, anything similar.
• Startup hats. You have worked where problems find you before process does.
• Opinionated without alienating. Opinions drive progress, but you find compromises with customers and teammates.