基于改进YOLOv8的煤矿井下巡检机器人目标检测方法研究

    Target detection method of coal mine underground inspection robot based on improved YOLOv8

    • 摘要: 为解决传统煤矿井下巡检机器人存在的巡检效率低下、识别精度较低等问题,提出了基于改进的YOLOv8煤矿井下巡检机器人目标检测方法。以YOLOv8为基础,针对煤矿井下光照不均匀、尘雾干扰、目标多尺度及遮挡等复杂环境特点,引入卷积注意力模块(Convolutional Block Attention Module,CBAM)和加权双向特征金字塔网络(Weighted Bidirectional Feature Pyramid Network,BiFPN)对模型进行优化。在YOLOv8的颈部网络中嵌入CBAM注意力模块,在上采样层输出端、特征拼接操作末端及多尺度特征融合阶段,对特征图实施通道和空间维度的双重注意力加权,以增强模型对关键特征的捕捉能力;采用BiFPN结构替换原有的路径聚合网络(Path Aggregation Network,PANet)结构,通过自底向上和自顶向下的双向特征流动与动态加权特征聚合,强化多尺度特征的融合效率,提升了模型对小目标和遮挡目标的检测能力;此外,构建了时空自适应特征增强机制,设计了门控式跨层注意力传递模块,引入可学习的双域门控系数实现注意力权重与特征金字塔的动态耦合,并基于能量熵的特征融合评估准则抑制低信息量特征干扰;同时,设计了轻量化推理框架,将骨干网络部署于边缘计算单元,颈部与头部网络部署于机器人嵌入式平台,通过动态管道并行降低了推理时延。使用源自晋陕蒙主要产煤区典型矿井工作面的煤矿井下多模态数据集进行试验,数据集覆盖井下巷道掘进、综采设备巡检、人员安全监控等核心场景,且遵循相关环境参数标准,能够表征实际检测难点。结果表明:改进后的YOLOv8模型在检测精度与效率上均表现优异,其准确率、召回率、F1值分别达到了96.1%、93.4%、95.8%,显著高于YOLOv8、YOLOv5等对比模型;平均检测时间为10.55 ms,优于其他对比算法;在不同环境条件下,无遮挡、正常亮度时单个目标检测时间低至2.4 ms,漏检频率仅0.5%;严重遮挡、低亮度时检测时间最长为18.7 ms,漏检频率最高为4.0%;多目标检测中,目标数为5时的检测时间为3.4 ms,目标数为25时的检测时间为24.1 ms;定位精度方面,实际坐标与定位坐标平均误差在2 mm以内,满足煤矿井下巡检需求。

       

      Abstract: To solve the problems such as low inspection efficiency and low recognition accuracy existing in traditional underground coal mine inspection robots, improve the accuracy and efficiency of underground coal mine target detection, and ensure mine safety, a target detection method for underground coal mine inspection robots based on the improved YOLOv8 is studied and proposed. Based on YOLOv8, aiming at the complex environmental characteristics such as uneven illumination, dust and fog interference, multi-scale targets and occlusion in underground coal mines, the convolutional block attention module (CBAM) is introduced. The model is optimized by CBAM and weighted bidirectional feature pyramid network (BiFPN). The CBAM attention mechanism is embedded in the neck network of YOLOv8. At the output end of the upper sampling layer, the end of the feature stitching operation, and the multi-scale feature fusion stage, dual attention weighting in the channel and spatial dimensions is implemented on the feature maps to enhance the ability of model to capture key features. The original path aggregation network (PANet) structure is replaced by the BiFPN structure. Through the bottom-up and top-down bidirectional feature flow and dynamic weighted feature aggregation, the fusion efficiency of multi-scale features is enhanced, and the detection ability of the model for small targets and occlusions is improved. Furthermore, a spatio-temporal adaptive feature enhancement mechanism is constructed, a gated cross-layer attention transfer module is designed, a learnable dual-domain gating coefficient is introduced to achieve the dynamic coupling of attention weights and feature pyramids, and the feature fusion evaluation criterion based on energy entropy is used to suppress the interference of low-information content features. Meanwhile, a lightweight reasoning framework is designed. The backbone network is deployed in the edge computing unit, and the neck and head networks are deployed in the robot embedded platform. The reasoning delay is reduced through dynamic pipeline parallelism. The research uses the multimodal data set from the working faces of typical coal mines in the main coal-producing areas of Shanxi, Shaanxi and Inner Mongolia for experiments. This data set covers core scenarios such as underground roadway excavation, comprehensive mining equipment inspection, and personnel safety monitoring, and follows relevant environmental parameter standards, which can represent the actual detection difficulties. The results show that the improved YOLOv8 model performs excellently in both detection accuracy and efficiency. Its accuracy rate, recall rate and F1 value reached 96.1%, 93.4% and 95.8% respectively, which were significantly higher than those of comparison models such as YOLOv8 and YOLOv5. The average detection time is 10.55 ms, which is superior to other comparison algorithms. Under different environmental conditions, when there is no occlusion and normal brightness, the detection time of a single target is as low as 2.4 ms, and the missed detection frequency is only 0.5%. The longest detection time is 18.7 ms when there is severe occlusion and low brightness, and the highest missed detection frequency is 4.0%. In multi-object detection, the detection time is 3.4 ms when the number of objects is 5, and 24.1 ms when the number of objects reaches 25. In terms of positioning accuracy, the average error between the actual coordinates and the positioning coordinates is within 2 mm, meeting the inspection requirements in underground coal mines.

       

    /

    返回文章
    返回