MT
Data Scientist
Marathon TS, Inc.
Location not stated
Remote
2 days ago
- AI
- calibration
- MLflow
- Databricks
- XGBoost
- LightGBM
- SciPy
- Python
- SQL
- FedRAMP
- NIST
- CMMC
- Machine Learning
- Core ML
- Unity Catalog
2 days ago
Responsibilities
Build risk-scoring models over synthetic tabular data, engineering features from curated medallion-layer tables.
Build anomaly and outlier detection to surface irregularities in records and process data.
Build optimization models for prioritization, routing, and resource allocation.
Validate honestly — calibration, discrimination, stability, explainability. A correctly characterized model matters more than a flattering headline metric.
Package deliverables as jobs and Asset Bundles, tracked in MLflow, and document assumptions, limitations, and what must be revalidated against real data post-ATO.
Required Qualifications
U.S. citizenship and active T5/SSBI federally adjudicated clearance required.
Hands-on Databricks.
Feature engineering on tabular and time-series data — encoding, aggregation, leakage prevention, and selection grounded in domain reasoning rather than automated search alone.
Supervised learning on tabular data: gradient boosting (XGBoost/LightGBM), regularized regression, and the judgment to know when the simpler model is the right answer.
Model calibration and evaluation under class imbalance — you can explain why AUC alone is insufficient for a risk score.
Anomaly detection: isolation forests, autoencoders, statistical process control, or comparable — with a clear account of how you validated detections without labels.
Optimization: LP/MIP or heuristic methods (OR-Tools, Pyomo, SciPy, or equivalent) applied to a real allocation or prioritization problem.
Explainability (SHAP or comparable) in a decision-support context.
Privacy-preserving synthetic data generation from CUI, PII, or comparably restricted source data — relational tabular data with distributional fidelity, cross-column correlations, referential integrity, and preservation of the rare-event structure that anomaly detection and risk scoring depend on. Includes an understanding of re-identification risk.
Strong Python, SQL, and Spark.
Government or defense contracting experience.
Preferred Qualifications
Modeling on federal investigative, vetting, fraud, or insider-threat data.
Direct experience with FedRAMP, NIST 800-171, CMMC L2, or CUI handling.
Familiarity with LLM/GenAI workflows — useful for collaboration with a peer document-intelligence workstream, but secondary to the core ML skill set.
H2O (Driverless AI, H2O-3).
MLflow, Databricks Asset Bundles, Unity Catalog.
Fairness / adverse-impact analysis in a regulated or decision-support setting.
Soft Skills
Self-directed execution against a fixed milestone with minimal oversight.
Honest reporting of model behavior — comfortable stating what synthetic-data performance does and does not establish about real-world accuracy.
Collaboration across technical and non-technical teams.
Clear documentation and active knowledge transfer.
Data Scientist · Marathon TS, Inc.