The Read

The Read

Inside Fireworks: The Factory for Specialized Intelligence

Fireworks is building the platform companies use to train specialized open models on proprietary data, then serve and improve them at scale.

By Will Daney

A golfer watches a prominent fireworks display outside a modern AI data center at dusk.

Most companies rent AI through a general-purpose model API. Fireworks is built around a different idea: companies will increasingly train intelligence around their own data, then operate that intelligence at scale.

Fireworks helps companies train customized open models on proprietary data and serve them in production. Put simply, it is selling specialized intelligence as a service.

This is already the core business. Fireworks says it serves more than 40 trillion tokens per day, and 95% come from models specialized on customer data and optimized for a specific job. The company recently surpassed $1 billion in annualized revenue.

The 95% figure is the important one. Fireworks is not primarily selling access to the same models available elsewhere. Customers are using the platform to create intelligence specific to their products and businesses.

Training these models is only half the problem. Companies must choose the right model, evaluate it against their own work, improve it, and then serve it reliably at scale. The same model can perform very differently depending on how well the software, memory, and hardware are optimized.

Fireworks brings training and serving together. A customer can post-train a model and deploy it on the same platform. Fireworks then helps the customer improve quality, speed, and cost as the workload grows.

That creates a technical flywheel. Each new workload strengthens Fireworks' expertise in model selection, post-training, evaluation, and inference. Its advantage should grow as it sees more models and production workloads.

The team matters here. CEO Lin Qiao previously led PyTorch at Meta, while other founders worked on PyTorch, Meta's machine learning infrastructure, and Google Vertex AI. The ability to optimize both model training and high-performance inference is scarce. Fireworks was built by a team that has already done this work at scale.

Proprietary data is what makes the opportunity interesting. Enterprises have years of customer interactions, internal decisions, and domain expertise. Fireworks gives them a way to turn that data into intelligence they control.

Harvey offers a good example. The legal AI company post-trained Nemotron 3 Ultra against its legal benchmark and matched leading closed models on complex legal work at least 10 times lower cost per run. A customized open model does not need to be smarter at everything. It needs to be excellent at the work the customer actually performs.

Most enterprises are early in this shift. Companies are full of repetitive, high-volume workflows that do not require the world's most powerful model. As specialized models become easier to train and operate, more of this work will move away from generic APIs.

The resulting switching costs should be meaningful. Moving an optimized production model is harder than changing an API endpoint. Customers become reliant on Fireworks to retrain the model, add workloads, and keep it running reliably.

The lasting advantage will not come from serving the most tokens. It will come from helping customers create intelligence that improves with use. Fireworks is building the factory that trains, serves, and continually improves that intelligence.


Sources

  1. Fireworks Series D announcement and operating metrics
  2. Fireworks founding team
  3. Fireworks training and inference platform
  4. NVIDIA Nemotron Labs: Harvey customization results