Unleash and Integrate the Power of Pre-Trained ViTs Via Feature Fusion for Open-Vocabulary Object Detection | AMiner