Y belongs to G, give it a try for yourself if you don't have spare 5 bucks and you will find out. Language on the video should not have any impact, I'd rather forcus on decription, on the other way G has done some work already on OCR with images, so I wouldn't be shocked if content of video would matter in terms of "contextual" linking.