Senior Elasticsearch Engineer
We're looking for a Senior Elasticsearch Engineer who can own the full lifecycle of our search and analytics data platform: capacity planning, cluster architecture, performance tuning, incident response, migration strategy, and operational excellence. You'll be the single point of deep expertise across all Elasticsearch and OpenSearch clusters at Chess.com.
This is not a monitoring-from-dashboards role. You'll be hands-on with cluster internals, write ILM/ISM policies, push infrastructure changes through GitOps, and make real-time decisions about replica allocation when a cluster goes red.
What you'll do
Incident Response & Reliability
Shard allocation strategy for write-heavy data streams at high throughput (millions of documents per minute)
Disk watermark management, retention policy tuning, and rollover orchestration for high-volume indices
Performance optimization and I/O tuning on bare-metal nodes
Write queue analysis, thread pool diagnostics, and shard rebalancing under load
Incident Response & Reliability
On-call ownership for Elasticsearch-related incidents: cluster health degradation, node loss, disk pressure, shard imbalance, and write rejection cascades
Real-time cluster triage and cross-team coordination during production incidents
Post-mortem authoring and systemic reliability improvements
Snapshot and disaster recovery management across clusters
Migration & Strategy
Elasticsearch-to-OpenSearch migration analysis and execution, including compatibility evaluation across ILM/ISM, security models, and plugin ecosystems
Version upgrade planning and rolling restart orchestration with zero-downtime requirements
End-to-end new cluster provisioning and onboarding
Cross-Team Enablement
Advise engineering teams on index design, mapping strategy, retention policies, and query optimization
Manage Kibana and OpenSearch Dashboards access and configuration for internal consumers
Define and maintain workload priority tiers across clusters
Preferred Skills
7+ years operating Elasticsearch at scale (multi-TB clusters, dozens of nodes, high write throughput)
Deep understanding of Elasticsearch internals: segment merging, translog, shard allocation, and cluster state management
Production experience with ECK (Elastic Cloud on Kubernetes) or equivalent operator-based deployments
Proficiency with Kubernetes operations for stateful workloads (StatefulSets, persistent storage, resource management)
Hands-on Linux systems administration with a focus on storage and I/O performance
Experience managing both Elasticsearch and OpenSearch in production, including an informed opinion on their respective trade-offs
Incident command experience: ability to diagnose and mitigate cluster emergencies under pressure while communicating clearly to stakeholders
Git-based infrastructure management (GitOps): Helm charts, ArgoCD/Flux, infrastructure-as-code for cluster configuration
Fluency with the Elastic stack APIs: cluster administration, index templates, data streams, ILM policies, snapshot/restore
Bonus Experience
OpenSearch ISM policies and security plugin (fine-grained access control)
GCS or S3 snapshot repository configuration and cross-cluster replication
Grafana + Prometheus monitoring for Elasticsearch metrics
Kibana Discover, Dev Tools, and data view management at scale
Java internals relevant to Elasticsearch JVM tuning (heap sizing, GC tuning, circuit breakers)
Vault integration for secrets management in Kubernetes-deployed search clusters
Fluentd/Fluent Bit log pipeline configuration feeding OpenSearch
Hardware selection experience for search-optimized server configurations
Python or scripting for operational analysis and automation
What Makes This Role Special
Full autonomy. You are the Elasticsearch authority. You make the architecture calls, set the priorities, and own the outcomes.
Real scale. Hundreds of terabytes of data, billions of documents, millions of daily queries. The problems here don't exist at smaller companies.
Bare metal. No managed Elastic Cloud. You're operating directly on the hardware. This is hands-on engineering.
Strategic impact. Your decisions on ES vs. OpenSearch migration, cluster topology, and capacity planning directly affect product capabilities and infrastructure costs.
Small team, big trust. Chess.com runs lean. You won't be buried in process or approvals. Ship changes, fix problems, improve systems.
About the Opportunity
This is a full-time opportunity. We are 100% remote (work from anywhere!)
You can learn more about us here: How Chess.com Virtual Team Works Together
About Chess.com
Published on: 7/21/2026

Chess.com
Chess.com is one of the largest gaming sites in the world and the #1 platform for playing, learning, and enjoying chess.
Please let Chess.com know you found this job on Wantapply.com. It helps us to get more jobs on our site. Thanks!
Unlock access with Plus