14 times faster: how we accelerated embedding generation in Manticore through ONNX

New Products

Summary

The article explains how Manticore Search sped up embedding generation by moving from SentenceTransformers/Candle to ONNX Runtime. It details benchmark gains on CPU, especially for insert workloads and batch inference, and shows how configuration changes reduced spinning and improved throughput. The piece also covers implementation choices in Rust, including session sharing, threading behavior, and padding strategy. It concludes with guidance for migrating existing models and using ONNX-based embeddings inside Manticore.

Classifications

industries
Retail
applications
Accounting and Taxes

AskAI Classifications

Labels
AI/ML Platform Developer Tools MLOps

Linked Companies

Hugging Face, Inc.
$10M to $25M
Manticore Search
$1M to $5M