Fireworks AI Raises $52M Series B to Lead Industry Shift to Compound AI Systems
Summary
These features make it easier and faster for businesses to customize models and build AI applications without needing a large team of ML engineers or data scientists. Today, we serve over 100 state-of-the-art models in text, image, audio, embedding, and multimodal formats, optimized for latency, throughput, and cost per token. Cursor, for example, has used Fireworks AIs custom Llama 3-70b model to achieve 1000 tokens/sec for code generation use cases such as instant apply, smart rewrites, and cursor prediction, which boost developer productivity.We continue to enhance our platform through deep collaboration with top providers across the AI stack, including partnerships with: Meta, Mistral, and Stability AI to deliver the lowest latency on SOTA models: 0.27s for Llama 3 70b, 0.25s for Mixtral 8x22b and 7b, and 1.2s for 1024x1024 image on Stable Diffusion XL respectively.In the past three months, weve launched new features that drastically boost performance and cut costs, bridging the gap between prototyping and production. On-demand GPU deployment option, in addition to serverless and reserved cloud, for scaling companies that need reliability and speed without long-term commitments. Compound AI systems tackle tasks using various interacting parts, such as multiple models, modalities, retrievers, external tools, data, and knowledge.