12,000$
We are building our own AI technology stack and a portfolio of AI-driven assets: from consumer AI products and smart devices to on-prem LLM clusters and foundational layers (memory layer, RAM/VRAM optimization, orchestration, AI software for infrastructure and compute).
This includes R&D in local, distributed, and orbital data centers and distributed computing.
Goal: join the world's top 50 AI companies with a combined asset valuation of $10B+ by 2032.
We are an OKR-driven company: measurable outcomes, speed, transparency, and ownership.
We hire A-players: autonomous, fast, strong execution, high accountability.
You lead a small team that builds and runs the bare-metal GPU clusters behind our self-hosted LLMs. This is a hands-on Lead DevOps role: you write Terraform, debug GPU driver and infrastructure issues, and carry the pager, while also assigning work, reviewing your team's changes, and unblocking them day to day. You report directly to the Head of Engineering. You shape the infra strategy for your area yourself and bring it to the Head of Engineering for sign-off — then you own execution: the build, the incidents, and the cost per GPU-hour.
Our compute runs primarily on our own hardware: racked GPU servers, our own network fabric, our own power and cooling planning. AWS and GCP exist as a hybrid layer for burst and failover, not the default environment. Think in terms of rack space and GPU supply lead times, not just API calls.
Cost is a standing responsibility, not a cleanup task. You own the unit economics of the clusters you run — cost per GPU-hour, cost per token, utilization — and the call between buying more hardware, tightening scheduling, or bursting to cloud.
The load is high. English is the working language, written and spoken.
Bare-metal and GPU infrastructure
Build, rack, and run on-prem GPU clusters as the primary compute platform: hardware provisioning, network fabric, power and cooling capacity
Own GPU infrastructure at a deep level — scheduling, utilization, memory allocation, failure modes, firmware and driver lifecycle
Keep hybrid cloud (AWS, GCP) wired in for burst and failover, and know when it's the right call instead of the fallback
Plan physical capacity two to three quarters out: rack space, power draw, network throughput, GPU procurement lead times
LLM as a production service
Operate and evolve self-hosted open-weight LLMs on the team's clusters: inference and fine-tuning pipelines, day-2 operations
Keep RAG and model-serving paths inside SLO under production traffic
Track latency, throughput, cost, and memory as one operating picture
Cost optimization and FinOps
Own the unit economics of the infrastructure you run: cost per GPU-hour, cost per token/request, cluster utilization
Build and maintain cost visibility — tagging, chargeback or showback, dashboards — across on-prem and cloud spend
Make and defend the capex-vs-opex call (buy more hardware vs. burst to cloud vs. optimize scheduling) with numbers
Report infrastructure budget status and cost trends to the Head of Engineering on a regular cadence; flag overruns before they land
Find and reclaim idle or underutilized GPU capacity as a standing practice
Delivery, GitOps, and team leadership
Lead a small team (2-4 engineers): assign work, review Terraform and Kubernetes changes, unblock people, and stay in the codebase yourself
Own Kubernetes, Terraform, and ArgoCD as the default change path for your clusters
Build GitLab CI and custom pipelines so changes are reviewable, repeatable, and reversible
Keep Linux, networking, and performance work at the depth the cluster actually needs
Reliability, observability, and security
Run Prometheus, Grafana, and OpenTelemetry against SLOs and error budgets for the systems you ownSit in the on-call rotation; lead incident response and postmortems for your team's surface
Implement IAM, RBAC, and secrets management (Vault / KMS / SSM) for your clusters, inside the security model the Head of Engineering sets
Hard requirement: direct, hands-on experience building and operating your own data center or owned hardware infrastructure — bare-metal server provisioning, racking, network fabric, and GPU infrastructure, with deep reasoning about compute and memory under load. Not only managed cloud; candidates without real bare-metal/on-prem ownership will not be considered
6+ years in DevOps, Infrastructure, or Platform Engineering
Experience owning infrastructure cost directly: capacity planning tied to unit economics, cost tagging/visibility, or FinOps practice in a prior role
Production experience operating self-hosted open-weight LLMs on local GPU clusters: inference, fine-tuning, day-2 operations
Deep Linux (internals, networking, performance), Docker, and production-scale Kubernetes
Terraform as a default; ArgoCD or equivalent GitOps tooling in real production use
GitLab CI and custom pipelines you've designed and run in production
Working AWS or GCP experience for hybrid and burst scenarios
Prometheus, Grafana, OpenTelemetry; SLO / error budgets; incident response and postmortems
Experience leading or managing a team of engineers, while staying hands-on — this is a working lead, not a pure people-manager seat
English B2, written and spoken; Russian, written and spoken
STEM background or equivalent practical experience
Hugging Face, LoRA / QLoRA fine-tuning in production
Vector databases in production paths: Qdrant, Pinecone, or Weaviate
Ansible or Pulumi in real use
Enterprise-level AWS or GCPSpecialization in AI/GPU-dense infrastructure specifically — high-density GPU racks, power and cooling planning built around AI/ML workloads, beyond generic bare-metal ops
Bare-metal servers, GPU infrastructure, on-prem data center operations
Self-hosted LLM inference and fine-tuning; Hugging Face, LoRA / QLoRA, RAG pipelines
Linux, Docker, Kubernetes (production-scale)
Terraform (required), Ansible / Pulumi
GitOps: ArgoCD (required)
CI/CD: GitLab CI, custom pipelines
Hybrid cloud: AWS and GCP (burst and failover, not primary)
Vector databases: Qdrant / Pinecone / Weaviate
Prometheus, Grafana, OpenTelemetry
Vault / KMS / SSM, IAM, RBAC
Cost per GPU-hour and cost per token are tracked and hold flat or trend down as workload grows
Cluster utilization stays above the agreed threshold; idle capacity gets caught and reclaimed inside an agreed window
Infrastructure budget reporting to the Head of Engineering is current and accurate — no cost surprises
Self-hosted LLM inference and fine-tuning run as production services with explicit latency, throughput, cost, and memory targets
Platform changes go through Terraform and ArgoCD; rollback is a practiced path, not a theory
SLOs and error budgets exist for your team's surface; incidents produce postmortems that change the system
Your team ships against its roadmap commitments, and the engineers you lead grow in scope
remote; sync windows by agreement (CET and other zones as needed). Fully remote regardless of where the physical data center sits — on-site racking/cabling is handled by local remote-hands, you drive the work remotely.
Location: Europe or UAE Reporting line Head of Engineering
Compensation Up to 12000 USDT total (base fix + shipment bonuses)
Equity: Company equity option, per level
Network: Work with top venture funds from Silicon Valley and a strong professional team
AI exposure: Portfolio of AI products in production; regular exposure to leading AI solutions shipped to production
Published on: 9/18/2026
Unimatch Lab
AI-native venture studio for the post-singularity world. We build products for financial, physical and mental health.
Please let Unimatch Lab know you found this job on Wantapply.com. It helps us to get more jobs on our site. Thanks!