Agent-driven Cognitive Object/Motion/Environment Hybrid Vision-Language Model on CCTV Surveillance-Video Bigdata | AMiner