Fix the data before the AI
Every project plan contains hundreds of guesses. How long will this task take? Who needs to finish first? In most large projects, those guesses still come from expert judgement.
For my master’s thesis at Bosch, I wanted to know whether machine learning could do better. I trained four models on 4,000+ tasks from two real semiconductor projects and compared how well they predicted task durations.
What worked
The two ensemble models, Random Forest and Gradient Boosting, clearly won. The best one explained over 85% of the variation in task durations. The neural network, which most papers recommend, performed worst.
Why
Not because neural networks are bad, but because the results depend heavily on the quantity and quality of the data. Real project data is rarely as complete and consistent as research data, and ensemble models cope with that better.
Before jumping on the AI bandwagon, fix the data foundation.
What I would improve next
- Improve the data set itself: complete, clean records.
- Add resources as a feature, not just tasks and durations.
- Connect data between planning processes such as schedule, risk and resources.
- Include real-time factors, so schedules become more resilient.
Good data makes AI useful. Bad data makes it confidently wrong.