01. ABOUT THE COMPANY
Dizzaract is a product-driven company operating at the intersection of gaming, digital platforms, and AI. We build and scale multiple products, including FAR Labs & Gamed — each exploring a different space, yet united by a shared approach: moving fast, staying curious, and focusing on things that people actually use.
We operate as a collaborative, non-hierarchical team where ideas are valued based on their impact, not their origin, and where AI is embedded across everything we build, from infrastructure to product decisions.
02. ABOUT THE ROLE
FAR Labs is building a distributed inference platform designed to serve large language models efficiently across diverse hardware. Serving quality — latency, throughput, memory efficiency, cost, and model quality — sits at the core of the product.
As our systems grow in complexity and scale, reliability becomes an engineering discipline of its own. We are looking for an SRE Lead who can take end-to-end ownership of infrastructure reliability, observability, automation, and production engineering across FAR Labs.
This is a hands-on technical leadership role. You will own the reliability, scalability, and operational maturity of our production environment while defining the SRE standards used across engineering.
You will work closely with Backend, Platform, Product, Security, and Engineering leadership to ensure our systems remain reliable as the business scales.
We are looking for someone who can move comfortably between architecture, hands-on engineering, incident response, automation, and technical leadership.
03. WHAT YOU'LL OWN
Infrastructure & Reliability
Own the reliability, availability, scalability, and performance of FAR Labs production infrastructure.
Design and evolve resilient infrastructure supporting growing and distributed workloads.
Identify infrastructure bottlenecks, single points of failure, and operational risks.
Establish capacity planning, backup, recovery, and disaster-recovery practices.
Drive improvements in infrastructure performance, stability, and cost efficiency.
Own the architecture, operation, and optimization of Kubernetes environments.
Improve cluster reliability, networking, resource utilization, deployment strategies, and workload management.
Build and maintain infrastructure through Terraform, Helm, Ansible, or equivalent IaC tooling.
Ensure infrastructure is automated, reproducible, version-controlled, and scalable.
Work deeply with Linux, Docker/containerd, networking, storage, and cloud-native infrastructure.
CI/CD Automation
Design and improve production-grade CI/CD pipelines.
Automate infrastructure provisioning, deployments, configuration, testing, and operational workflows.
Improve deployment safety through automated validation, rollback mechanisms, and appropriate release controls.
Reduce manual operational work through automation.
Observability
Own the observability strategy across FAR Labs.
Build monitoring, logging, tracing, dashboards, and alerting infrastructure.
Work with Prometheus, Grafana, VictoriaMetrics, ELK/EFK, or equivalent technologies.
Establish visibility into infrastructure health, application health, performance, and failures.
Improve alert quality and reduce unnecessary operational noise.
SRE Standards
Define and implement SLIs, SLOs, SLAs, and error budgets.
Establish reliability standards for critical services.
Partner with engineering teams to make reliability part of the development lifecycle.
Use operational data to identify recurring problems and prioritize improvements.
Incident Management
Lead technical response to critical production incidents.
Establish incident management and escalation processes.
Lead RCA and blameless post-mortems.
Translate incidents into concrete engineering improvements.
Participate in and help structure the on-call rotation.
Security & Operational Readiness
Embed infrastructure and cloud security best practices.
Strengthen access management, secrets management, network security, CI/CD security, and runtime security.
Support security, compliance, and audit-readiness requirements where relevant.
Set the technical direction for SRE and infrastructure engineering.
Establish engineering standards and best practices.
Review infrastructure architecture and major technical changes.
Mentor engineers and increase infrastructure expertise across teams.
Remain hands-on with implementation and troubleshooting.
04. REQUIREMENTS
7+ years of experience in SRE, DevOps, Platform Engineering, Infrastructure Engineering, or similar roles.
Strong experience owning production infrastructure at scale.
Deep hands-on expertise with Kubernetes.
Strong experience with Terraform and Infrastructure as Code.
Experience designing and operating production-grade CI/CD pipelines.
Strong knowledge of Linux, networking, DNS, load balancing, storage, containers, and distributed systems.
Experience with Prometheus, Grafana, VictoriaMetrics, ELK/EFK, or similar observability stacks.
Strong scripting/programming skills in Python, Go, Bash, or similar.
Practical experience with SLIs, SLOs, SLAs, alerting, and error budgets.
Experience leading production incidents, RCA, and post-mortems.
Strong understanding of cloud-native and infrastructure security.
Experience making architectural decisions and driving technical standards.
Strong ownership, communication, and problem-solving skills.
05. NICE TO HAVE
Experience supporting AI/ML or HPC workloads.
Experience with GPU infrastructure, scheduling, resource management, or performance optimization.
Experience operating large-scale distributed systems.
Multi-cloud, hybrid-cloud, or self-managed infrastructure experience.
Experience optimizing infrastructure costs and capacity.
Experience with service mesh, distributed tracing, or advanced Kubernetes networking.
Previous experience as an SRE Lead, Platform Lead, Infrastructure Lead, or Tech Lead.
06. WHAT WE OFFER
Real technical ownership of site reliability, infrastructure, and production operations across FAR Labs.
The opportunity to define the SRE strategy, standards, and technical direction as our platform scales.
Hands-on work with technically challenging problems across Kubernetes, distributed systems, cloud infrastructure, networking, observability, and automation.
The opportunity to build and evolve high-availability, scalable production infrastructure rather than simply maintain existing systems.
Direct influence over architecture, SLOs, incident management, CI/CD, observability, and infrastructure performance.
A small, senior engineering environment where you can set technical standards, mentor engineers, and see your decisions translate directly into platform and product outcomes.
Fast execution, low bureaucracy, and a highly collaborative, idea-driven team.
Competitive salary with performance-based incentives.
24 days annual leave, plus public holidays.
Health insurance.
Modern office in Yas Creative Hub.
Continuous learning through real-world problem solving — not just theory.
The opportunity to influence both technical architecture and product direction.
A diverse, open-minded team where ideas are genuinely heard.
Published on: 9/19/2026
Farcana
Farcana is a tournament of Stars. It was established to test the capabilities of Axelar, a compound found on Mars that can bind to humans and give them amazing powers at a high cost. In the year 2101, most major issues of humanity are solved not through wars or political conflict but through powerful, enhanced individuals battling them out on the Farcana Arena. Each one of them has personal goals, dreams, and ideals, which only the winners will get to implement in the world.
Please let Farcana know you found this job on Wantapply.com. It helps us to get more jobs on our site. Thanks!