AI news story
H Company Releases NeoMME: A Family of 260M and 800M Single-Tower Multimodal Encoders That Drop the Vision Tower and Causal Decoder
We look at NeoMME, a family of 260M and 800M bidirectional encoders from H Company. Unlike ColPali-style retrievers, it processes multilingual text tokens and raw 32×32 image patches in a single Transformer, with no pretrained vision tower and no cau
Editor's take
H Company has introduced NeoMME, a new family of single-tower multimodal encoders that integrate text and image processing without a separate vision encoder or causal decoder. This approach allows for direct processing of raw image patches alongside text tokens within a unified Transformer architecture.
The significance lies in its potential to streamline multimodal model development by eliminating the need for separate, pretrained vision components like those found in some earlier models. This could lead to more efficient training and inference for tasks requiring joint understanding of visual and textual information, impacting applications in areas such as image captioning and visual question answering.
Future developments will focus on NeoMME's performance compared to established multimodal models such as CLIP or BLIP, particularly its effectiveness on downstream tasks and its scalability. The ability of these 260M and 800M parameter models to generalize across diverse visual-linguistic benchmarks will be a key indicator of their practical utility.
Signal score: 3
This event was corroborated by 55 independent sources. The signal score weighs cross-source corroboration, recency, source weight and topic salience. How we rank stories.
Original reporting
This story summarises reporting published by MarkTechPost. Read the original article at MarkTechPost.