AI news story

H Company Releases NeoMME: A Family of 260M and 800M Single-Tower Multimodal Encoders That Drop the Vision Tower and Causal Decoder

We look at NeoMME, a family of 260M and 800M bidirectional encoders from H Company. Unlike ColPali-style retrievers, it processes multilingual text tokens and raw 32×32 image patches in a single Transformer, with no pretrained vision tower and no cau

  • Generative
  • Source: MarkTechPost
  • Published: 2026-09-06
  • Signal score: 3
  • 55 sources

Editor's take

H Company has introduced NeoMME, a new family of single-tower multimodal encoders that integrate text and image processing without a separate vision encoder or causal decoder. This approach allows for direct processing of raw image patches alongside text tokens within a unified Transformer architecture.

The significance lies in its potential to streamline multimodal model development by eliminating the need for separate, pretrained vision components like those found in some earlier models. This could lead to more efficient training and inference for tasks requiring joint understanding of visual and textual information, impacting applications in areas such as image captioning and visual question answering.

Future developments will focus on NeoMME's performance compared to established multimodal models such as CLIP or BLIP, particularly its effectiveness on downstream tasks and its scalability. The ability of these 260M and 800M parameter models to generalize across diverse visual-linguistic benchmarks will be a key indicator of their practical utility.

Signal score: 3

This event was corroborated by 55 independent sources. The signal score weighs cross-source corroboration, recency, source weight and topic salience. How we rank stories.

More Generative stories

  1. Designers should not fear being replaced by AI, industry leaders say

    The Guardian AI · 2026-09-07

    Firms are more likely to use technology as ‘the intern in the office’ than as a replacement for skilled staff Professional designers should not feel “threatened” by the rapid

  2. Roland is getting into generative AI music with Melody Flip

    The Verge · 2026-09-04

    It's not quite the "push button; get song" of Suno, but Roland's new Melody Flip tool marks the company's foray into generative AI music.

  3. Mini book: Next-Gen Architecture Playbook: Insights and Patterns for the AI Era

    InfoQ · 2026-09-04

    This eMag examines how architects can lead with clarity in a rapidly evolving engineering worl

  4. Roubini Says He's Not Concerned About an AI Bubble (Video)

    Bloomberg · 2026-09-04

    Nouriel Roubini, a prominent economist known for predicting the 2008 financial crisis, has stated he is not concerned about a speculative bubble forming around artificial

  5. The sameness problem behind those unappetizing AI-generated menus

    TechCrunch · 2026-09-04

    While restaurant owners might look to generative AI as a shortcut to sprucing up their menu, customers can viscerally sense that something is wrong with the food.

  6. Comedian’s offensive mock Welcome to Country at far-right rally gets rerun in the Daily Mail | Weekly Beast

    The Guardian AI · 2026-09-04

    Daily Mail’s report on March for Australia rally quotes – in full – and embeds video of Lisa Jane Spencer’s disrespectful skit.