Experiment Design for ML

Hard
practical-experience Netflix Google

When comparing two ML models in an offline experiment, why is it insufficient to compare them on a single train-test split?

Multiple Choice

Correct!
Incorrect — try again next time!