Beauty For All Industries (BFA)

Beauty For All Industries (BFA)

Beauty For All Industries (BFA), also known as Beauty for All, is a leading beauty innovation platform that integrates technology and community to make beauty accessible to everyone. The company is dedicated to inspiring self-expression and operates at the intersection of technology and beauty, utilizing extensive member data to create personalized shopping experiences. BFA serves over 30 million subscribers and generates over $1 billion in annual revenue. Headquartered in San Mateo, California, with additional offices in New York, Miami, Santa Monica, and Argentina, BFA is a digitally native company that functions as a beauty subscription powerhouse and brand incubator. Its portfolio includes notable brands such as IPSY, the largest beauty subscription service in the world, BoxyCharm, Madeby Collective, and Refreshments. BFA focuses on inclusivity and aims to create a welcoming community for all, while also attracting entrepreneurial thinkers and creative problem solvers to drive its mission forward.

information technology & services
51-200
Founded 2020
$333M raised

Description

  • Build and maintain Datadog observability, including dashboards, monitors, APM, log pipelines, and low-noise alerts.
  • Define and track SLIs, SLOs, and error budgets with service owners.
  • Participate in on-call rotations and support incident triage, prioritization, escalation, and resolution.
  • Lead incident response documentation, status updates, handoffs, and ownership transfers.
  • Manage alerting and escalation through Opsgenie and incident communication through Slack.
  • Conduct blameless post-incident reviews, identify root causes, and track preventive actions.
  • Automate operational toil through scripts, tooling, self-healing, and automated remediation.
  • Use AI tools such as Claude and Cursor to accelerate debugging, runbook creation, RCA drafting, and automation.
  • Improve reliability across AWS and third-party services including Netlify, CommerceTools, Auth0, and Contentful.
  • Contribute to CI/CD reliability, deployment safety, infrastructure-as-code, runbooks, and triage workflows.

Requirements

  • Experience with observability and monitoring tools, ideally Datadog; Grafana, Prometheus, New Relic, or CloudWatch experience also applies.
  • Experience participating in on-call rotations and incident response, including triage, prioritization, escalation, and post-incident reviews.
  • Working knowledge of SRE principles, including SLIs, SLOs, error budgets, and toil reduction.
  • Working knowledge of AWS and distributed, microservice, and API-gateway architectures.
  • Scripting and automation skills in Python, Bash, or similar languages, with the ability to read and reason about code.
  • Familiarity with CI/CD pipelines and infrastructure-as-code tools such as Terraform.
  • Strong written and verbal communication during incidents and in RCA documentation.
  • Ability to collaborate across distributed, multi-time-zone teams.
  • Comfort using AI tools for debugging, documentation, RCAs, and reliability workflows.
  • Preferred: experience with high-traffic consumer or e-commerce platforms, Opsgenie or PagerDuty, CommerceTools, Auth0, or Amplitude.
  • Must work remotely from Mexico or Colombia while covering the PST time zone.
  • Resume/CV must be submitted in English.

Benefits

  • Competitive salary paid in USD.
  • Paid time off.
  • Work-from-home flexibility.
  • Remote-first work environment with virtual activities and company-wide offsites.
  • Professional development and learning sessions.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Site Reliability Engineer, Tech Lead

Loadsmart 251-1K Air Freight & Logistics

Loadsmart is hiring a remote SRE Tech Lead in Brazil to build and operate its internal engineering platform, improve reliability, and enable safe, dependable applications across engineering teams.

Ansible AWS Bash Chef CI/CD Docker Kubernetes PostgreSQL Python Terraform
1 week, 4 days ago

Incident management / reliability / SRE Evaluator

Weekday 11-50 Construction & Engineering

An independent contractor Evaluator will remotely assess AI-generated documents, spreadsheets, and presentations for accuracy, rigor, and quality using incident management, reliability, and SRE expertise.

1 week, 5 days ago

Reliability Engineer

Sapsol Technologies 51-250 Internet Software & Services

A medical device company is seeking a Reliability Engineer to develop and execute reliability requirements, testing, and risk analyses for IVD products in a highly regulated environment.

MATLAB Python R
3 months, 3 weeks ago

Blockchain Site Reliability Engineer

InfStones 51-250 Internet Software & Services

InfStones is hiring a remote Blockchain Site Reliability Engineer in Dallas to ensure the reliability, availability, and performance of its blockchain node infrastructure.

Docker Ethereum Go Grafana JavaScript Kubernetes Linux Prometheus Python Rust Solana
5 months, 2 weeks ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers