Posted 2025-01-06Updated 2026-03-30Notea minute read (About 197 words) visits

CLIP

https://blog.csdn.net/h661975/article/details/135116957

loss: ITC (Image Text Contrastive)

# image_encoder - ResNet or Vision Transformer 
# text_encoder - CBOW or Text Transformer 
# I[n, h, w, c] - minibatch of aligned images 
# T[n, l] - minibatch of aligned texts 
# W_i[d_i, d_e] - learned proj of image to embed 
# W_t[d_t, d_e] - learned proj of text to embed 
# t - learned temperature parameter  

# extract feature representations of each modality 
I_f = image_encoder(I) #[n, d_i] 
T_f = text_encoder(T) #[n, d_t]  

# joint multimodal embedding [n, d_e] 
I_e = l2_normalize(np.dot(I_f, W_i), axis=1) T
_e = l2_normalize(np.dot(T_f, W_t), axis=1)  

# scaled pairwise cosine similarities [n, n] 
logits = np.dot(I_e, T_e.T) * np.exp(t)  

# symmetric loss function 
labels = np.arange(n) 
loss_i = cross_entropy_loss(logits, labels, axis=0) 
loss_t = cross_entropy_loss(logits, labels, axis=1) 
loss = (loss_i + loss_t)/2

Cross_entropy_loss:

CLIP 本质上是全局图像嵌入，不利于像素对齐特征提取。

CLIP

http://chen-yulin.github.io/2025/01/06/[OBS]Reconstruct Anything-Semantic-CLIP多模态预训练模型/

Author

Chen Yulin

Posted on

2025-01-06

Updated on

2026-03-30

Licensed under

#Research-paper Image2Text CV MultiModal CLIP Contrastive-Learning VLP Image-Text

CLIP

Author

Posted on

Updated on

Licensed under

Comments

Archives

Recents

Tags