Job-ID33079160Reference26-00621Remote100% RemoteConexess Group is aiding a large healthcare client in their search for a Staff Machine Learning Engineer in a remote capacity. This is a direct hire opportunity with a competitive compensation package.
Position Overview
Operating with full autonomy, you'll establish MLOps excellence, mentor engineering teams, and drive strategic technical decisions that directly impact patient outcomes. This role requires a seasoned engineer who can balance innovation with production reliability while building systems that handle healthcare's most sensitive and complex data.
Key ResponsibilitiesArchitect and maintain enterprise-grade ML infrastructure, including model versioning, automated testing frameworks, containerization strategies, CI/CD pipelines, and comprehensive monitoring systems for model performance, data quality, and drift detection.Drive MLOps strategy and standards across the organization. Mentor data scientists and engineers on production best practices, system design, and scalable architecture patterns.Own end to end journey from model development through production deployment, including real-time and batch inference systems, A/B testing frameworks, and automated retraining pipelines.Collaborate with clinical leaders, product teams, and data scientists to translate complex healthcare requirements into robust, scalable ML solutions. Present technical strategies to executive stakeholders.Build fault-tolerant, compliant systems that meet healthcare security and privacy standards. Establish SLAs, incident response protocols, and disaster recovery procedures for mission-critical ML services.Evaluate and integrate cutting-edge MLOps tools and practices. Design systems that scale growth while reducing operational overhead and improving model iteration velocity.Required QualificationsBachelor's degree in computer science, engineering, or related field required; master's degree preferredMinimum of ten (10) years in software engineering with five (5) years focused on ML infrastructure, MLOps, or production ML systems and Python development with strong software engineering fundamentals and three (3) years architecting and deploying production ML systems on cloud platforms (Azure preferred)Proven track record building and scaling ML platforms from the ground upHealthcare or regulated industry experience strongly preferredExpert-level proficiency with MLOps tooling (MLflow, Kubeflow, SageMaker, Azure ML, etc.)Deep experience with containerization (Docker, Kubernetes), orchestration tools (Airflow, Prefect), and infrastructure-as-code (Terraform, ARM templates)Advanced knowledge of CI/CD systems, automated testing strategies, and GitOps workflowsData engineering skills: SQL, Spark/PySpark, Databricks, data pipeline optimizationExpertise in model monitoring, observability, feature stores, and experiment tracking at scaleProduction experience with both batch and real-time inference architecturesUnderstanding of healthcare data standards (FHIR, HL7, claims data) is a plusDemonstrated ability to influence technical direction and mentor senior engineersProven communication skills with ability to distill complex technical concepts for diverse audiencesTrack record of driving consensus on architectural decisions across multiple stakeholdersSystems thinking skills with focus on reliability, scalability, and maintainability preferredUnderstanding of security, compliance, and privacy requirements in healthcare (HIPAA) preferred#LI-Remote
#LI-CB2