Margo Bank

Margo Bank

Unlock excellence with MARGO Consulting: where ambition, expertise, and innovation drive the most complex tech challenges.

Professional Services
$8M raised

Description

  • Build and support large-scale AI infrastructure with monitoring, diagnosis, and remediation of production incidents.
  • Troubleshoot high-impact production issues in collaboration with other engineering teams.
  • Participate in an on-call rotation to handle incidents and ensure service continuity.
  • Implement and maintain observability solutions to monitor AI infrastructure and application health.
  • Contribute to AI infrastructure lifecycle management across different environments and countries.
  • Promote and apply best practices for stability, resiliency, scalability, and security.
  • Maintain clear technical documentation for tools and procedures.
  • Contribute to the evolution of systems and tools based on production feedback.
  • Collaborate closely with development teams to ensure infrastructure readiness.
  • Participate in team rituals and knowledge-sharing initiatives.

Requirements

  • Experience with Go or Python.
  • Strong scripting skills in Bash and Python.
  • Hands-on experience with Linux systems, especially Ubuntu/Debian.
  • Preferred hands-on experience with GPU and HPC infrastructure.
  • Knowledge of networking concepts such as VLAN/LAN, TCP/IP, DNS, BGP, load-balancing, and IPv6.
  • Familiarity with monitoring and logging tools such as Prometheus, Grafana, and Elastic.
  • Comfort with Infrastructure-as-Code tools such as Ansible, Salt, and AWX.
  • Experience managing relational databases, especially MariaDB.
  • Understanding of CI/CD pipelines, especially GitLab.
  • Comfortable communicating in English, both written and spoken.
  • Proactive and solution-oriented mindset.
  • Passion for automation and continuous improvement.
  • Strong collaboration and communication skills.
  • Ability to work independently and as part of a team.
  • Willingness to mentor others and share knowledge.

Benefits

  • Remote work arrangement.
  • Permanent contract or B2B contract option.
  • Hourly rate of 200 zł - 250 zł.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

SRE Engineer Contractor

Beauty For All Industries (BFA) 51-200 information technology & services

IPSY is seeking a remote Site Reliability Engineer Contractor in Mexico or Colombia, covering the PST time zone, to improve the availability, resilience, and operational reliability of its beauty membership platform.

Amplitude AWS Bash CI/CD Contentful Datadog Grafana Microservices Netlify New Relic OpsGenie PagerDuty Prometheus Python Terraform
1 week, 5 days ago

Site Reliability Engineer, Tech Lead

Loadsmart 251-1K Air Freight & Logistics

Loadsmart is hiring a remote SRE Tech Lead in Brazil to build and operate its internal engineering platform, improve reliability, and enable safe, dependable applications across engineering teams.

Ansible AWS Bash Chef CI/CD Docker Kubernetes PostgreSQL Python Terraform
3 weeks, 2 days ago

Incident management / reliability / SRE Evaluator

Weekday 11-50 Construction & Engineering

An independent contractor Evaluator will remotely assess AI-generated documents, spreadsheets, and presentations for accuracy, rigor, and quality using incident management, reliability, and SRE expertise.

3 weeks, 3 days ago

Reliability Engineer

Sapsol Technologies 51-250 Internet Software & Services

A medical device company is seeking a Reliability Engineer to develop and execute reliability requirements, testing, and risk analyses for IVD products in a highly regulated environment.

MATLAB Python R
4 months ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers