Site Reliability Engineer, Tech Lead

13 hours, 8 minutes ago
Contract
Lead
DevOps and Infrastructure
Loadsmart

Loadsmart

Loadsmart is a logistics solutions provider automating freight transportation with innovative technology to move more efficiently.

Air Freight & Logistics
251-1K
$346M raised

Description

  • Design, deploy, and operate critical systems while balancing reliability, cost, and agility.
  • Lead reliability initiatives across multiple engineering squads.
  • Build, maintain, and take ownership of internal platform and software infrastructure projects.
  • Define and maintain platform Service Level Agreements and Objectives.
  • Collect operational metrics and connect them to business impact.
  • Troubleshoot production issues and conduct root-cause analyses.
  • Provide infrastructure support during off-hours as needed.
  • Collaborate with engineering teams and internal stakeholders on technical projects.
  • Support code and specification reviews while giving and receiving constructive feedback.
  • Apply emerging AI tooling, including AI-assisted coding, LLMs, and agentic workflows, to improve reliability and operations.

Requirements

  • 1–3 years of experience leading reliability work across multiple engineering squads.
  • 5+ years of experience in cloud computing, SRE, or DevOps.
  • Experience collaborating with stakeholders across multiple engineering teams.
  • Strong project management, delegation, mentoring, and leadership skills.
  • Fluent written and spoken English for daily collaboration with international teams.
  • Strong understanding of software engineering principles and system internals.
  • In-depth knowledge of modern networking and operating systems.
  • Proficiency with AWS, cloud environments, Docker, Kubernetes, containers, DevOps practices, testing, and CI/CD pipelines.
  • Experience with Terraform, Ansible, Chef, or similar infrastructure automation tools.
  • Production troubleshooting and systems engineering experience in UNIX/Linux environments.
  • Experience with monitoring, alerting, incident management, and scripting languages such as Python or Bash.
  • PostgreSQL and database administration experience is preferred.
  • Familiarity with AI agents, agentic workflows, LLMs, and MCP servers or gateways is preferred.

Benefits

  • Remote work from anywhere in Brazil.
  • Competitive base salary.
  • Competitive equity package.
  • Unlimited PTO and sick days.
  • Equal opportunity workplace committed to diversity and inclusion.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Incident management / reliability / SRE Evaluator

Weekday 11-50 Construction & Engineering

An independent contractor Evaluator will remotely assess AI-generated documents, spreadsheets, and presentations for accuracy, rigor, and quality using incident management, reliability, and SRE expertise.

1 day, 13 hours ago

Senior Site Reliability Engineer

Sports Academy Education Services

Texas Sports Academy is seeking a part-time Senior Site Reliability Engineer consultant to audit, improve, and scale the infrastructure supporting its AI-first K–12 school.

AWS CI/CD Datadog Grafana Prometheus
1 week, 5 days ago

Reliability Engineer

Sapsol Technologies 51-250 Internet Software & Services

A medical device company is seeking a Reliability Engineer to develop and execute reliability requirements, testing, and risk analyses for IVD products in a highly regulated environment.

MATLAB Python R
3 months, 1 week ago

Blockchain Site Reliability Engineer

InfStones 51-250 Internet Software & Services

InfStones is hiring a remote Blockchain Site Reliability Engineer in Dallas to ensure the reliability, availability, and performance of its blockchain node infrastructure.

Docker Ethereum Go Grafana JavaScript Kubernetes Linux Prometheus Python Rust Solana
5 months ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers