Explore indexBack to Terms
CLIP
A model architecture that aligns cross-modal semantics through contrastive learning on large-scale image-text pairs
No public content is connected to this entity yet.
A model architecture that aligns cross-modal semantics through contrastive learning on large-scale image-text pairs
No public content is connected to this entity yet.