Smart Working

Smart Working

Smart Working is a company that specializes in software development outsourcing and staff augmentation. They offer nearshore software development services, outsourcing solutions, and staff augmentation with a focus on providing highly skilled Indian de...

Internet Software & Services

Description

  • Architect and build scalable data pipelines and infrastructure to support AI and product systems.
  • Design and maintain data ingestion, transformation, and storage architectures for operational and AI workloads.
  • Develop and manage batch and real-time data pipelines.
  • Build and optimize systems for vector search, retrieval, and machine learning data pipelines.
  • Ensure data reliability, security, and governance across the platform.
  • Collaborate with AI, backend, product, and leadership teams to support training, inference, and product features.
  • Implement monitoring, observability, and data quality frameworks.
  • Optimize the performance of large-scale datasets and query systems.
  • Contribute to technical architecture decisions and long-term data strategy.
  • Act as the founding data hire and help define the culture, standards, and hiring bar for the growing data function.

Requirements

  • 7+ years of professional experience, with most of that in dedicated data engineering roles.
  • Strong experience designing and building data pipelines and distributed data systems.
  • Experience with relational databases, with PostgreSQL preferred; MySQL or similar acceptable.
  • Experience working with NoSQL databases.
  • Experience with vector databases used in modern AI systems.
  • Strong programming experience in Python.
  • Demonstrated ability to make and justify architectural decisions.
  • Experience building scalable backend systems.
  • Experience designing data models and storage architectures.
  • Strong understanding of data processing performance and optimization.
  • Experience with Apache Spark, Apache Airflow, Kafka, and Elasticsearch or OpenSearch is highly desirable.
  • Experience with PostgreSQL, MongoDB, and vector databases such as Qdrant, Milvus, or pgvector is highly desirable.
  • Experience with Python data-processing libraries such as Pandas or Polars is highly desirable.
  • Experience working on AI or machine learning platforms (nice to have).
  • Familiarity with stream processing and event-driven architectures (nice to have).
  • Experience with cloud infrastructure such as GCP, AWS, or Azure (nice to have).
  • Experience working in high-growth startups or early-stage companies (nice to have).

Benefits

  • Remote-first work environment.
  • Full-time, long-term role.
  • Opportunity to join a genuine community focused on growth and well-being.
  • Work with outstanding global teams and products.
  • Exposure to a high-rated workplace recognized on Glassdoor.

Interested in this position?

Apply directly on the company website

Apply Now

Similar Roles

Senior Fabric Data Engineer

Spirit Omega 11-50 human resources

Spirit Omega is hiring a Senior Fabric Data Engineer to modernize and migrate learning-platform data into the Cornerstone Galaxy ecosystem using Microsoft Fabric and Azure-based data pipelines.

Agile Apache Spark Azure CI/CD Databricks Feature Engineering Git Python SQL
3 days ago

Data Engineer

CodeRoad 51-250 Internet Software & Services

Coderoad is hiring a Data Engineer to build and improve cloud-based data pipelines and analytics infrastructure for real-world software projects.

Apache Spark Databricks GCP MySQL Python Scala
3 days, 23 hours ago

Sr. Data Engineer

Teachable 51-250 Internet Software & Services

Teachable is hiring a Senior Data Engineering leader to guide data architecture and pipelines for a remote Brazil-based team supporting company-wide decision-making across engineering, product, finance, and business functions.

Apache Airflow AWS dbt Kafka Terraform
1 week, 2 days ago

Data Engineer (Snowflake & GCP)

Coderio 51-250 Internet Software & Services

Coderio is hiring a Data Engineer to join its data engineering team and lead the design and optimization of cloud-native data pipelines connecting Snowflake and Google Cloud Platform.

GCP Python Snowflake
2 weeks, 2 days ago

You're on a roll! Sign up now to keep applying.

Sign Up

Already have an account? Log in

Used by 14,729+ remote workers