SCPT optimization boosts vision-language model accuracy for fine-grained video recognition | StreamingMeme