ironquill.tech/board

$ cat jobs/healthcare-data-scientist-cherokee-federal-66772a17920a.json

Healthcare Data Scientist

cherokee federal·US·United States·mid
pythonsqlazurepandas
Apply on himalayas → Get AI match score →
ATA is seeking a Data Scientist to support data pipeline development, validation, and analysis within a cloud-based Health IT data platform. This role is hands-on and delivery-focused, with an emphasis on building reliable, reproducible data workflows using SQL, Python, and PySpark in an Azure Synapse environment. A core expectation of this role is the ability to work across the full data lifecycle, from ingestion through transformation to final dataset delivery, while maintaining data quality and traceability. The ideal candidate is comfortable debugging data issues end-to-end, understands how data structure and join logic impact outputs, and applies disciplined validation and documentation practices. This role also supports exploratory data analysis and the development of derived datasets to enable analytics and downstream use cases. The position will work extensively with healthcare data originating from EHR systems and interface feeds, including HL7 v2 and FHIR data, clinical terminology, and source-to-target data mappings. Key Responsibilities: Data Pipeline Development Develop, maintain, and optimize data pipelines using SQL, Python (pandas), and PySpark. ATA, LLC | 752 Walker Road | Suite D | Great Falls, VA 22066 | www.ata-llc.com Advanced Technology Applications Execute and manage notebook-based workflows within Azure Synapse, including debugging and documentation. Process and transform structured and semi-structured data in formats such as CSV, JSON/NDJSON, and Parquet. Work within ETL/ELT pipelines across raw, curated, and production data layers. Ingest, profile, map, and transform healthcare data from EHR systems and interface feeds while preserving source lineage and clinical context Data Quality and Validation Perform structured data validation, including row counts, null checks, duplicate detection, schema validation, and allowed value enforcement. Identify and resolve data quality issues such as schema drift, inconsistencies, and transformation error

Similar remote roles

Machine Learning Engineer
Corpay · Worldwide · mid
Software Engineer
nimble gravity · LATAM · mid
Data Engineer
CIMSOLUTIONS · Worldwide · mid
Machine Learning Engineer
Data Wizards · Worldwide · mid
AI/ML ENGINEER
Recordly · Worldwide · mid
Bioinformatics Engineer II
tempus · US · mid
11920543 - Analista de Dados e Compliance Pleno
mtp metodos e tecnologia brasil · LATAM · mid
Data Engineer
assistrx · US · mid