
Training that turns data teams into practitioners of applied machine learning and modern data platform engineering, grounded in real business decisions.
What the program covers
Modules cover data platform architecture, pipeline engineering, applied machine learning, and responsible AI deployment. Teams work with anonymized real-world datasets rather than toy examples.
Recent outcomes
- A retail analytics team built a demand-forecasting model that is now in production.
- A healthcare client's data team redesigned its pipeline architecture, cutting report latency from days to hours.
- An internal team shipped a first responsible-AI governance checklist adopted company-wide.
Technologies covered
- Python
- Pandas
- scikit-learn
- PyTorch
- Airflow
- Spark
- MLflow
Program agenda
Data platform (Day 1)
- Data platform architecture
- Modeling & data warehousing
- Data quality
- Governance
Pipeline engineering (Day 2)
- Orchestration with Airflow
- Distributed processing (Spark)
- Batch & streaming pipelines
- Data testing
Applied machine learning (Days 3–4)
- Data prep & feature engineering
- Supervised models
- Evaluation & validation
- Production deployment (MLOps)
Responsible AI (Day 5)
- Bias & fairness
- Explainability
- Regulatory framework (AI Act)
- Governance checklist
Deliverables
- Working data pipeline
- Deployed ML model
- AI governance checklist
- Certificate of completion
Commitments
- Group of 6 to 10 people
- Prerequisites: basic Python & statistics
- Duration: 5 days
- Anonymized real-world datasets provided
Equipment & materials
- Jupyter environment provided
- GPU access for ML modules
- Prepared datasets
- Course materials + notebooks