amoon
Elite Member
- May 16, 2015
- 2,358
- 1,920
-Visual IQ test:
yes, the ones that humans take!
-OCR-free reading comprehension:
input a screenshot, scanned document, street sign, or any pixels that contain text. Reason about the contents directly without explicit OCR. This is extremely useful to unlock AI-powered apps on multimedia web pages, or “text in the wild” from real world cams.
-Multimodal chat:
have a conversation about a picture. You can even provide “follow-up” images in the middle.
-Broad visual understanding abilities,
like captioning, visual question answering, object detection, scene layout, common sense reasoning, etc.
-Audio & speech recognition:
wasn’t mentioned in Kosmos-1 paper, but Whisper is already an OpenAI API and should be fairly easy to integrate.
.
yes, the ones that humans take!
-OCR-free reading comprehension:
input a screenshot, scanned document, street sign, or any pixels that contain text. Reason about the contents directly without explicit OCR. This is extremely useful to unlock AI-powered apps on multimedia web pages, or “text in the wild” from real world cams.
-Multimodal chat:
have a conversation about a picture. You can even provide “follow-up” images in the middle.
-Broad visual understanding abilities,
like captioning, visual question answering, object detection, scene layout, common sense reasoning, etc.
-Audio & speech recognition:
wasn’t mentioned in Kosmos-1 paper, but Whisper is already an OpenAI API and should be fairly easy to integrate.
.