Authors :
Muhammad Mustapha Miko; Li Dequan
Volume/Issue :
Volume 11 - 2026, Issue 9 - September
Google Scholar :
https://tinyurl.com/2er2dhbn
DOI :
https://doi.org/10.38124/ijisrt/26sep483
Note : A published paper may take 4-5
working days from the publication date to appear in PlumX Metrics, Semantic Scholar, and
ResearchGate.
Abstract :
Diffusion-based detectors begin inference from noisy boxes, whereas tracking-by-detection pipelines usually
localize each video frame independently. This study examines whether propagated tracker boxes can guide a frozen
DiffusionDet detector without retraining. Confirmed tracks are propagated using their latest observation or a Kalman
prediction, mapped to the detector’s latent box space, corrupted at a selected diffusion timestep, and mixed with random
proposals for new-object discovery. On a KITTI development split, four low-noise proposals around each Kalman prediction
improve mean car/pedestrian multiple object tracking accuracy (MOTA) by 1.29 points over a matched short-timestep
random control and 0.66 points over full random initialization. Higher order tracking accuracy (HOTA) is effectively tied,
while the identity F1 score (IDF1) is lower than with full random initialization. Transferred without retuning to MOT17
half-validation, the configuration averages 53.68 HOTA, 55.88 MOTA, and 63.91 IDF1 across three inference seeds. Mean
HOTA improves by 0.99 points over full random initialization, primarily through higher recall at the cost of lower precision.
The findings support a conditional detection-recall benefit rather than a consistent improvement in identity association.
Keywords :
Multi-Object Tracking; DiffusionDet; Temporal Proposals; Kalman Prediction; ByteTrack; Tracking-by-Detection.
References :
- Chen S, Sun P, Song Y, Luo P. DiffusionDet: diffusion model for objectdetection. In: CVPR; 2023.
- Luo R, Song Z, Ma L, Wei J, Yang W, Yang M. DiffusionTrack: diffusion model for multi-object tracking. arXiv:2308.09905. 2024.
- Hashmi KA, Stricker D, Afzal MZ. Spatio-temporal learnable proposalsfor end-to-end video object detection. In: BMVC; 2022.
- Bewley A, Ge Z, Ott L, Ramos F, Upcroft B. Simple online and real-timetracking. In: ICIP; 2016.
- Zhang Y, Sun P, Jiang Y, et al. ByteTrack: multi-object tracking byassociating every detection box. In: ECCV; 2022.
- Cao J, Pang J, Weng X, Khirodkar R, Kitani K. Observation-centricSORT: rethinking SORT for robust multi-object tracking. In: CVPR; 2023.
- Zhang Y, Wang C, Wang X, Zeng W, Liu W. FairMOT: on the fairnessof detection and re-identification in multiple object tracking. International Journal of Computer Vision. 2021.
- Cai J, Xu M, Li W, Xiong Y, Xia W, Tu Z, et al. MeMOT: multi-objecttracking with memory. In: CVPR; 2022.
- Luiten J, Osep A, Dendorfer P, Torr P, Geiger A, Leal-Taixe L, et al.HOTA: a higher order metric for evaluating multi-object tracking. International Journal of Computer Vision. 2021.
- Geiger A, Lenz P, Urtasun R. Are we ready for autonomous driving?The KITTI vision benchmark suite. In: CVPR; 2012.
- Dendorfer P, Osep A, Milan A, et al. MOTChallenge: a benchmark forsingle-camera multiple object tracking. International Journal of Computer Vision. 2021.
Diffusion-based detectors begin inference from noisy boxes, whereas tracking-by-detection pipelines usually
localize each video frame independently. This study examines whether propagated tracker boxes can guide a frozen
DiffusionDet detector without retraining. Confirmed tracks are propagated using their latest observation or a Kalman
prediction, mapped to the detector’s latent box space, corrupted at a selected diffusion timestep, and mixed with random
proposals for new-object discovery. On a KITTI development split, four low-noise proposals around each Kalman prediction
improve mean car/pedestrian multiple object tracking accuracy (MOTA) by 1.29 points over a matched short-timestep
random control and 0.66 points over full random initialization. Higher order tracking accuracy (HOTA) is effectively tied,
while the identity F1 score (IDF1) is lower than with full random initialization. Transferred without retuning to MOT17
half-validation, the configuration averages 53.68 HOTA, 55.88 MOTA, and 63.91 IDF1 across three inference seeds. Mean
HOTA improves by 0.99 points over full random initialization, primarily through higher recall at the cost of lower precision.
The findings support a conditional detection-recall benefit rather than a consistent improvement in identity association.
Keywords :
Multi-Object Tracking; DiffusionDet; Temporal Proposals; Kalman Prediction; ByteTrack; Tracking-by-Detection.