Microsoft unveils AI model that understands image content and solves visual puzzles

Share via:

Microsoft unveiled Kosmos-1, a multimodal model capable of analysing images for content, solving visual puzzles, performing visual text recognition, passing visual IQ tests, and understanding natural language instructions.

The researchers believe that multimodal AI—which integrates different modes of input such as text, audio, images, and video—is a critical step towards developing AGI that can perform general tasks at the level of a human. “Language Is Not All You Need: Aligning Perception with Language Models,” the researchers write in their academic paper, “is a necessity to achieve artificial general intelligence, in terms of knowledge acquisition and grounding to the real world.”

Disclaimer

We strive to uphold the highest ethical standards in all of our reporting and coverage. We StartupNews.fyi want to be transparent with our readers about any potential conflicts of interest that may arise in our work. It’s possible that some of the investors we feature may have connections to other businesses, including competitors or companies we write about. However, we want to assure our readers that this will not have any impact on the integrity or impartiality of our reporting. We are committed to delivering accurate, unbiased news and information to our audience, and we will continue to uphold our ethics and principles in all of our work. Thank you for your trust and support.

Popular

More Like this

Microsoft unveils AI model that understands image content and solves visual puzzles

Microsoft unveiled Kosmos-1, a multimodal model capable of analysing images for content, solving visual puzzles, performing visual text recognition, passing visual IQ tests, and understanding natural language instructions.

The researchers believe that multimodal AI—which integrates different modes of input such as text, audio, images, and video—is a critical step towards developing AGI that can perform general tasks at the level of a human. “Language Is Not All You Need: Aligning Perception with Language Models,” the researchers write in their academic paper, “is a necessity to achieve artificial general intelligence, in terms of knowledge acquisition and grounding to the real world.”

Disclaimer

We strive to uphold the highest ethical standards in all of our reporting and coverage. We StartupNews.fyi want to be transparent with our readers about any potential conflicts of interest that may arise in our work. It’s possible that some of the investors we feature may have connections to other businesses, including competitors or companies we write about. However, we want to assure our readers that this will not have any impact on the integrity or impartiality of our reporting. We are committed to delivering accurate, unbiased news and information to our audience, and we will continue to uphold our ethics and principles in all of our work. Thank you for your trust and support.

Website Upgradation is going on for any glitch kindly connect at office@startupnews.fyi

More like this

JioMart unlikely to enter top tier of quick commerce:...

JioMart relies on over 2,000 existing stores to...

Crypto prediction platform Polymarket to raise at $1b valuation

Polymarket allows users to wager on topics like...

Putin authorises creation of state messaging app to combat...

Russian President Vladimir Putin on Tuesday signed a...

Popular

Upcoming Events

Norwegian Deep Sea Mining Firm Plans $1.2B Bitcoin Treasury

Norwegian deep-sea mining firm Green Minerals AS says...

Exito Media Concepts Presents: The 29th Edition of BFSI...

Physical Conference on 7th of August, Mumbai  From paper trails...

JioMart unlikely to break into quick commerce top tier:...

Reliance Retail’s quick commerce service JioMart is unlikely...
ZXCVa ZXCVa ZXCVa ZXCVa ZXCVa ZXCVa ZXCVa ZXCVa ZXCVa ZXCVa ZXCVa ZXCVa ZXCVa ZXCVa ZXCVa ZXCVa ZXCVa ZXCVa ZXCVa ZXCVa ZXCVa ZXCVa ZXCVa ZXCVa ZXCVa ZXCVa ZXCVa ZXCVa ZXCVa ZXCVa ZXCVa ZXCVa ZXCVa ZXCVa