CLIP (Contrastive Language-Image Pretraining), Predict the most relevant text snippet given an image
CLIP(对比语言-图像预训练),根据给定的图像预测与之最相关的文本片段。【此简介由AI生成】