databricks / spark-deep-learning

Deep Learning Pipelines for Apache Spark
https://databricks.github.io/spark-deep-learning
Apache License 2.0
1.99k stars 494 forks source link

Horovod Runner is stuck. Not passing through the first epoch after start training. #250

Open camposwalacy opened 3 weeks ago

camposwalacy commented 3 weeks ago

Hello, folks!

I am using HorovodRunner within Databricks runtime LTS 14.2 ML with Tensorflow 14.0 through sparkdl. My data is in TFRecords format, and this issue started to happen after 25th June. I migrated my workload to Unity Catalog. I am debugging on my side if there is something that might have changed, but I couldn't find a way to fix this yet.