
Internship – End-of-Studies Research Project
- Airflow
- Machine Learning
- Python
- scikit-learn
- Linux
- Jupyter
Context & Motivation
- Reducing the environmental footprint of large-scale storage systems
- Improving capacity planning and thermal management
- Enabling proactive power optimization in production environments
Modern distributed storage infrastructures are made up of thousands of hard disk drives (HDDs) operating continuously. While compute and network energy costs are increasingly well understood, HDD/SSD power consumption under real-world workloads remains difficult to measure and predict accurately. Understanding and modeling this consumption is critical for:
To date, no robust, interpretable multi-parameter model exists that can accurately estimate disk energy consumption from observable parameters in a production environment. This internship aims to fill that gap.
Your Mission
- Drive temperature
- Fan speed and airflow
- Read/write throughput and IOPS
- Idle vs. active state transitions
- Calibrate and validate the model
- Quantify model accuracy across workload types
- Identify parameters with the greatest predictive value
As part of our R&D team, you will design, build, and validate a multi-parameter model for estimating HDD/SSD power consumption. Your work will include:
1. State of the Art
Review existing literature and open-source projects related to HDD/SSD power modeling, thermal behavior, and energy measurement in storage systems (including tools such as PowerAPI).
2. Physical Modeling
Develop a physics-based multi-parameter model of temperature and power draw, capturing relationships between:
3. Interpretable Machine Learning Model
Design a multi-parameter ML model (e.g., gradient boosting, linear regression with feature engineering) that is both accurate and interpretable — enabling engineers to understand which parameters drive consumption under different workload profiles.
4. Measurement & ValidationInstrument real drives in a controlled laboratory environment and under
production-representative workloads to collect ground-truth measurements. Use this data to:
Expected Outcomes & Valorization
- Depending on results, the work may be valorized through:
- A scientific publication or technical white paper
- Integration into the open-source PowerAPI project
- Direct integration into Scality's internal monitoring and capacity planning tooling
Technical Stack
- Python (primary language for data collection, modeling, and analysis)
- scikit-learn (ML modeling and evaluation)
- PowerAPI ecosystem (https://powerapi.org/)
- Linux system tooling for hardware instrumentation (smartctl, lm-sensors, etc.)
- Jupyter Notebooks for exploratory analysis and result visualization
Candidate Profile
- Final-year student in a Master's program or Engineering school (Bac+5)
- Strong interest in physical modeling and/or applied machine learning
- Comfortable working with real hardware and experimental data
- Autonomous, curious, and rigorous in your approach to problem-solving
- Able to communicate results clearly in written and spoken English
Work Environment
- You will join a multicultural R&D team with colleagues across France, the US, and Asia. English is our primary working language. You will benefit from:
- Mentorship from senior R&D engineers with expertise in distributed systems and performance engineering
- Access to real production-grade storage hardware and laboratory infrastructure
- A high-trust environment where your findings will directly influence engineering decision
Why Join Scality?
- Work on a concrete research problem with real-world industrial impact
- Contribute to open-source energy-efficiency tooling used beyond Scality
- Be part of a team building infrastructure trusted by Fortune 500 companies
- Potential to publish research or continue as a full-time engineer after the internship
Internship – End-of-Studies Research Project · Scality