InternVL: Scaling Up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks | AMiner