Jina Releases Multimodal jina-embeddings-v5-omni-small
Jina has launched jina-embeddings-v5-omni-small, a multimodal model that searches video and audio across existing text indices without requiring developers to re-embed their text databases.

AI search company Jina has released jina-embeddings-v5-omni-small, a 1.74-billion-parameter multimodal embedding model that unifies text, images, video, and audio into a shared vector space. By projecting diverse media formats into the same mathematical representation, the model enables seamless cross-modal nearest-neighbor searches. A primary advantage of this architecture is that its text embeddings remain bit-for-bit identical to those generated by the corresponding v5-text-small model under identical task settings. As a result, engineering teams can query existing text indices using audio clips or visual files without needing to re-embed their stored text collections.
Built using a GELATO frozen-tower design, the development team trained a mere 0.35 percent of the model's overall parameters across four Nvidia H100 GPUs. The model boasts a generous context length of 32k and generates 1024-dimensional output vectors by default. Through Matryoshka representation learning, developers can dynamically truncate vector sizes down to a range between 32 and 768 dimensions. This flexibility allows engineers to trade off marginal retrieval accuracy for significantly reduced storage costs and lower memory footprints. Supported file formats span .mp4, .wav, .mp3, .pdf, .jpg, and .png.
To streamline integration across common machine learning applications, jina-embeddings-v5-omni-small ships with four specialized task adapters optimized for retrieval, classification, clustering, and text-matching tasks. The retrieval adapter specifically targets asymmetric query-to-document search for retrieval-augmented generation pipelines. The model also includes native support for vLLM to simplify high-throughput production serving. Maintainers must ensure matching task configurations, query roles, and vector dimensions when combining text and multimodal pipelines. Jina offers the model under a CC BY-NC 4.0 license, requiring organizations to secure commercial authorization for business usage.
This is our own summary of reporting by AlphaSignal



