spark模型运行时无法连接摸个excutors异常org.apache.spark.shuffle.FetchFailedException: Failed to connect to xxxx/xx.xx.xx.xx:xxxx

error:org.apache.spark.shuffle.FetchFailedException: Failed to connect to xxxx/xx.xx.xx.xx:xxxx

定位来定位去与防火墙等无关。反复查看日志:

2019-09-30 11:00:46,521 | WARN | [dispatcher-event-loop-50] | Lost task 5.0 in stage 1.2 (TID 24441, dggsafe0321-cm, executor 7): ExecutorLostFailure (executor 7 exited caused by one of the running tasks) Reason: Container killed by YARN for exceeding memory limits. 4.6 GB of 4.5 GB physical memory used. Consider boosting spark.yarn.executor.memoryOverhead. | org.apache.spark.internal.Logging$class.logWarning(Logging.scala:66)
2019-09-30 11:00:46,521 | INFO | [dag-scheduler-event-loop] | Resubmitted ShuffleMapTask(6, 25830), so marking it as still running | org.apache.spark.internal.Logging$class.logInfo(Logging.scala:54)
2019-09-30 11:00:46,522 | WARN | [dispatcher-event-loop-50] | Lost task 4.0 in stage 1.2 (TID 24440, dggsafe0321-cm, executor 7): ExecutorLostFailure (executor 7 exited caused by one of the running tasks) Reason: Container killed by YARN for exceeding memory limits. 4.6 GB of 4.5 GB physical memory used. Consider boosting spark.yarn.executor.memoryOverhead. | org.apache.spark.internal.Logging$class.logWarning(Logging.scala:66)
2019-09-30 11:00:46,522 | INFO | [dag-scheduler-event-loop] | Resubmitted ShuffleMapTask(6, 15603), so marking it as still running | org.apache.spark.internal.Logging$class.logInfo(Logging.scala:54)

发现节点内存溢出,导致节假死,导致节点无法访问,扩展相应执行内存重启就行。

--driver-memory 4g --executor-memory 6g 

 

posted @ 2019-09-30 17:44  ~清风煮酒~  阅读(1209)  评论(0编辑  收藏  举报