Objective Floating-object detectors deployed in complex water environments must recognize small, weak-texture, and dynamically deformed targets under changes in camera location, weather, surface ripples, illumination, and occlusion. Existing deep detectors often depend on large fully annotated datasets, incur excessive computational and storage costs, and lose accuracy in unseen water scenes. Conventional domain generalization also requires multiple completely labelled source domains, whereas single-labelled domain generalization reduces annotation work but is vulnerable to source-specific bias. A double-labelled domain generalization (DLDG) method was therefore developed to balance detection accuracy, real-time efficiency, annotation cost, and generalization. The method used only two labelled source domains and several unlabelled source domains while explicitly correcting feature-extraction, classification, and localization biases.Methods A lightweight multiscale detector was constructed by coupling MobileNetV3, a Dynamic Feature Pyramid Network (DyFPN), and a Single Shot MultiBox Detector (SSD). The VGG16 backbone of SSD was replaced by MobileNetV3, and feature maps from different stages were delivered to DyFPN. Dynamic modules containing a gate, an Inception unit, and a skip connection were inserted into the lateral connections. A gating signal and Gumbel-Softmax one-hot decision determined whether each lateral convolution was executed. Useful cross-scale information was therefore selected according to the input, while redundant reflections, ripples, shoreline textures, and computation were reduced. Depthwise separable convolutions were also applied to the classification and box-regression layers. DLDG training was organized as a progressive "feature generation-category discrimination-spatial localization" process. Two labelled domains were first used to initialize the feature generator, classifier, and locator through cross-entropy and Smooth L1 losses. The remaining unlabelled domains were then used to filter three types of bias. For feature-extraction bias, cluster centers were calculated from SoftMax outputs and bounding-box regression results. Pseudo-labels were assigned according to cosine distance, and information maximization improved cluster separability and prevented concentration in a few classes. Both classification and localization information constrained pseudo-label generation, after which the feature extractor was updated using classification and localization losses from unlabelled domains. For classification bias, a conditional feature-projection network mapped unlabelled features into the discriminative space jointly defined by the two labelled domains. Consistency between true labels and pseudo-labels constrained projection, and domain-similarity attention reweighted common features across domains. For localization bias, foreground-anchor regression was trained with pseudo-labels. An adversarial regressor maximized its prediction discrepancy from the main regressor on source samples, and generalized intersection over union measured box differences. The three stages shared features and pseudo-labels and were optimized sequentially under a unified objective. Experiments were conducted using fixed-camera images from the modern water-conservancy demonstration area in Deqing County, Zhejiang Province, China. Sixty representative video sequences were sampled. After redundant frames were removed, 526 water-hyacinth images, 340 floating-weed images, and 300 plastic-bottle images were retained; training-set augmentation produced 11,623 samples. The 1,920 × 1,080-pixel images were divided into five scene domains: normal water surfaces, weather variation, wave interference, illumination variation, and target-level occlusion. The dataset was split 9:1 for training and validation, and 30 labelled-source/target combinations were generated. Experiments used PyTorch 1.10, Ubuntu 18.04, an Intel i7 processor, and an NVIDIA RTX 3080 graphics card. Stochastic gradient descent was configured with a momentum of 0.9, an initial learning rate of 0.01, a batch size of 16, and weight decay of 0.001. Performance was evaluated using mean average precision (mAP), F1 score, frames per second (FPS), floating-point operations, parameter count, and model size.Results and Discussions For double-labelled domain generalization, DLDG achieved 70.33% mAP, 71.02% F1, and 22.38 FPS on the GPU. It required 3.01 billion floating-point operations, 4.98 million parameters, and 21.24 MB of storage; CPU inference reached 7.98 FPS. The strongest compared method, domain-invariant feature enhancement domain adaptation, obtained 66.48% mAP, 66.85% F1, and 20.10 FPS. DLDG therefore improved mAP by 3.85 percentage points, F1 by 4.17 percentage points, and speed by 2.28 FPS, while using fewer parameters and less storage. Under conventional domain generalization with fully labelled source data, the method achieved 85.29% mAP, 86.28% F1, and 17.81 FPS, indicating that the progressive filtering mechanism was applicable under both limited-label and fully labelled settings. The network ablation showed that MobileNetV3-DyFPN reached 86.28% mAP, 87.33% F1, and 19.15 FPS, compared with 70.02% mAP, 71.28% F1, and 12.99 FPS for VGG16 with DyFPN. Thus, mAP increased by 16.26 percentage points and speed by 6.16 FPS. Bias-filtering ablations showed that the unfiltered model produced 62.29% mAP and 63.02% F1 at 24.37 FPS. Introducing all three filters increased mAP by 8.04 percentage points and F1 by 8.00 percentage points, while speed decreased by 1.99 FPS. Feature-extraction filtering produced the largest single-module improvement, increasing mAP by 2.43 percentage points. Localization filtering contributed more than classification filtering because adversarial regression and generalized intersection over union directly corrected box displacement. In wave-interference and low-illumination scenes, DLDG produced boxes more consistent with target boundaries, whereas FixMatch and open compound domain adaptation showed missed detections or localization shifts. Stable detections were also obtained for small plastic bottles and floating weeds whose edges were mixed with reflections and ripples.Conclusions Combining lightweight dynamic multiscale feature extraction with progressive filtering of feature, classification, and localization biases reduced dependence on extensive bounding-box annotation while maintaining real-time cross-domain detection. Two labelled domains provided model initialization, and additional unlabelled domains were exploited through classification-and-localization-constrained pseudo-labels, conditional feature projection, domain-similarity reweighting, and adversarial box regression. The method can support fixed-camera identification of floating-object accumulation, cleaning prioritization, and continuous inspection in heterogeneous water environments. The current domain definition was based mainly on visible water-surface states and object distributions because synchronized flow velocity, water level, wind speed, and rainfall were unavailable. Future integration of hydrodynamic and meteorological measurements, unmanned surface vehicles, and hydrological monitoring data could clarify the relationships among water conditions, floating-object transport, and detection performance.
更多