14 times faster: how we accelerated embedding generation in Manticore through ONNX
Summary
The article explains how Manticore Search sped up embedding generation by moving from SentenceTransformers/Candle to ONNX Runtime. It details benchmark gains on CPU, especially for insert workloads and batch inference, and shows how configuration changes reduced spinning and improved throughput. The piece also covers implementation choices in Rust, including session sharing, threading behavior, and padding strategy. It concludes with guidance for migrating existing models and using ONNX-based embeddings inside Manticore.
Classifications
industries
Retail
applications
Accounting and Taxes
AI Classifications
Labels
AI/ML Platform
Developer Tools
MLOps
Linked Companies
Hugging Face, Inc.
$10M to $25M
Manticore Search
$1M to $5M