Lead Site Reliability Engineer (Performance & Scalability) | Contract | Remote US

1 week, 6 days ago
Contract
Lead
DevOps and Infrastructure
Tech Holding

Tech Holding

Tech Holding: California's #1 website design company offering full-service technology consulting with expertise in software management, AI, and security.

Internet Software & Services
51-250
Founded 2016

Description

  • Establish performance, throughput, latency, and capacity baselines for critical platform workflows.
  • Define and maintain SLOs, error budgets, performance budgets, dashboards, alerts, and reliability thresholds.
  • Instrument and analyze the full request path across services, infrastructure, databases, networking, caches, queues, DNS, and third-party dependencies.
  • Identify system bottlenecks and lead cross-functional remediation efforts with engineering teams.
  • Build capacity models that show current limits, emerging constraints, and the cost of additional scale.
  • Lead load, stress, soak, spike, failure, and recovery testing in representative environments.
  • Develop realistic demand scenarios for major customers, partnerships, pilots, and high-volume events.
  • Drive architecture hardening, graceful degradation, dependency-failure planning, and resilience improvements.
  • Partner with Test Automation and Scalability Engineering on automated performance testing, regression coverage, and production release gates.
  • Own technical readiness assessments for major pilots, partnerships, and production launches.
  • Create operational runbooks for scale-up events, incidents, rollback, recovery, and dependency failures.
  • Lead performance and reliability investigations during incidents and incorporate lessons into future engineering work.
  • Communicate infrastructure cost, performance, and reliability tradeoffs to engineering and executive leadership.
  • Recommend capacity and reliability investments before they become production constraints.

Requirements

  • Significant experience in Site Reliability Engineering, performance engineering, platform engineering, distributed systems, or a closely related discipline.
  • Experience supporting production systems with meaningful scale, traffic, latency, or availability requirements.
  • Deep understanding of observability, performance analysis, capacity planning, and reliability engineering.
  • Strong hands-on experience with cloud infrastructure and production distributed systems.
  • Deep knowledge of databases, networking, caching, queueing, compute, storage, and common distributed-system failure modes.
  • Experience defining and operating against SLOs, SLIs, error budgets, and production reliability metrics.
  • Hands-on experience performing load, stress, soak, scalability, and resilience testing.
  • Ability to profile systems, diagnose bottlenecks, tune architecture, and work directly with engineering teams to implement improvements.
  • Experience designing for graceful degradation, dependency failures, recovery, and high-demand scenarios.
  • Strong incident management and root-cause analysis experience.
  • Ability to translate technical performance and reliability risks into clear business implications for senior leadership.
  • Strong judgment around when systems need optimization versus when added complexity is premature.
  • Experience operating high-scale SaaS, identity, DNS, registry, infrastructure, or other highly distributed platforms (preferred).
  • Experience creating capacity-cost models and forecasting infrastructure requirements (preferred).
  • Experience building performance and reliability gates into CI/CD pipelines (preferred).
  • Experience preparing platforms for major traffic increases from enterprise customers or strategic partnerships (preferred).
  • Experience leading reliability or performance initiatives across multiple engineering teams (preferred).
  • Applicants must be authorized to work for any employer in the U.S.; no visa sponsorship is available.

Benefits

  • Remote work opportunity (#LI-Remote).
  • Contract employment type.
  • Equal Opportunity Employer commitment.
  • Inclusive workplace across all backgrounds and experiences.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Reliability Engineer

Sapsol Technologies 51-250 Internet Software & Services

A medical device company is seeking a Reliability Engineer to develop and execute reliability requirements, testing, and risk analyses for IVD products in a highly regulated environment.

MATLAB Python R
2 months, 3 weeks ago

Blockchain Site Reliability Engineer

InfStones 51-250 Internet Software & Services

InfStones is hiring a remote Blockchain Site Reliability Engineer in Dallas to ensure the reliability, availability, and performance of its blockchain node infrastructure.

Docker Ethereum Go Grafana JavaScript Kubernetes Linux Prometheus Python Rust Solana
4 months, 2 weeks ago

Senior Site Reliability Engineer

Intuition Machines 51-250 Life Sciences Tools & Services

Intuition Machines is hiring a Senior Site Reliability Engineer to support its internet-scale AI/ML security products, with a focus on improving the performance, availability, security, and cost efficiency of systems serving millions of users.

C++ CI/CD Cloudflare Cybersecurity Go JavaScript Kubernetes Load Balancing Machine Learning Python Rust
4 months, 2 weeks ago

DevOps Engineer (Cloud) - Freelance

Lingaro 5K-10K IT Services

Build and evolve an Azure-based monitoring and observability platform that improves system reliability, data quality, operational insight, automation, and cloud cost management across supply chain systems.

Ansible Azure Bash CI/CD Databricks Docker GitHub Actions Grafana Kafka Kubernetes Linux Power BI PowerShell Prometheus Python SonarQube Terraform Windows Server
9 hours, 33 minutes ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers