LongCut logo

硅谷坐标 x Fireworks 联创Benny Chen:开源模型、token增速、推理优化和模型定制

By Silicon Valley Vector 硅谷坐标

Summary

Topics Covered

  • Buyer mindset, not technology, powers closed-source
  • Vertical agents are edged intelligence
  • Distillation is not a durable closed-source moat
  • Token volume is a deceptive metric
  • Open-source progress strengthens infrastructure

Full Transcript

Building a vertical-specific agent and solving the Riemann Hypothesis are likely two completely different tasks. But perhaps Frontier Lab is currently more concerned with whether they can solve the Riemann Hypothesis. We see more and more Vertical Agents, their work becoming increasingly granular, and their data cleaning processes becoming more refined. I do believe this type of work is indeed very meaningful for the customization of open-source models. In the long run, distillation is not a particularly reliable path; it might help everyone reach a decent level in the short term, though. Is there any direction that might not consume much compute right now, but where you actually see extremely fast growth in terms of trend? C-Work is definitely a major direction. I feel that everyone currently spends the vast majority of their time catering to reasoning tasks, but this is actually very counterintuitive. I actually feel that we are seeing more and more cases where the demand for world-class content is growing, such as various SaaS, web apps, and even things like video agents that have recently emerged . Many engineers are worried about losing their jobs, but I think the number of people who can actually do this work is truly minimal. If anyone is interested, you can look more into how to do it and teach your boss whether open-source can actually be adopted. You might be able to save three to five times your salary just by writing this code well and switching from closed-source to open-source. Hello Ben, welcome to Silicon Valley Coordinates. Today, we want to talk to you about your assessment of the overall situation, discuss the two businesses at Works—inference optimization and custom models—and also chat about your core competitiveness. First, let's look at the biggest core issue: open-source models versus closed-source models. We know the most significant industry variable over the past year is that open-source models are closing the gap with closed-source model capabilities at a very fast pace, at a fraction of the cost. Looking ahead, what do you think the landscape and structure between open-source and closed-source models will look like? Many of our team members come from an open-source software background; the vast majority are PyTorch contributors or users who worked at Meta before starting their own ventures. We strongly believe that in the most valuable software layers, open-source will eventually catch up. Whether it's previous verticals like operating systems, where open-source eventually caught up, or various databases where open-source also eventually caught up —we are very confident that open-source will continue to catch up. But this year, development has certainly been faster than I imagined. First off, the ARR of closed-source models is growing very, very quickly. At the same time, the capabilities of open-source models are catching up, like what we've seen recently with Qwen 2 or DeepSeek-V 2, and the latest Llama 3. Meta is also talking about open-sourcing Spark. This is quite surprising to me—the competition in open-source is so intense, and the cost-to-performance ratio of these models is so high, which really is very unexpected. Our team firmly believed that this would eventually catch up, we just didn't expect it to catch up this fast. We know some closed-source models rely on distillation to maintain their competitive edge. Do you think this approach of using distillation to stay ahead is sustainable? Do closed-source models have any way to prevent this? First of all, closed-source models will definitely patch those few APIs. They did have various issues before. Often, you’d take the output from a large model, feed it into a smaller one, and with some tweaking, you could extract its thinking tokens. That was a security gap in how closed-source models were handled previously. But I think this can be fixed very quickly; there’s no reason closed-source model providers can't handle it. So, in the long run, distillation isn't a particularly reliable path. It might help everyone reach a decent level in the short term, though. However, based on my understanding, many companies in the U.S., like Meta, might not rely heavily on distillation for their success. So, whether the success of open-source models is contingent on distillation, I’m not so sure. Actually, as long as open-source data is large-scale and clean, and data costs continue to drop, I don’t think distillation is strictly necessary for many to access high-quality models. So, I believe closed-source vendors have plenty of things they can do to maintain their competitive edge; they have many tools at their disposal. But I’m increasingly feeling that closed-source vendors don’t necessarily need to be ahead forever. I’m an Apple user myself. If you ask me about when I used Android, I found Google phones to be much more functional than iPhones, yet they still had all sorts of minor issues. Even if the price wasn't so high, I would probably still buy Apple. In the end, the bulk of the profits likely goes to the one who has always been in the lead . It could become very large, there might be no conflict at all, and the run rate will likely be largest in open source, while closed source may always claim some advantages. And for those who aren't so price-sensitive, they will claim closed source, so if you ask me where the long-term advantage lies, I think user mindset might be more important. For instance, when I was involved in procurement, we used to joke that "no one ever got fired for buying IBM." Right? Whether IBM can maintain that advantage forever, I don't know. But for a very long time, everyone was indeed inclined to be conservative and wanted to buy IBM. So, I think this mindset might not change that quickly. If people believe that continuing to purchase from Cloud or Open source providers is better, then this shift might just be a bit slower. This is likely their advantage for the next two or three years, but if you ask me about five or ten years from now, I really can't see that clearly. So do you think that open models relying on leaks to stay ahead cannot maintain a narrowing gap with closed models long-term, since this will eventually be prevented by strengthening internal security teams, but what about non-leaked categories like Meta or Mistral—do you think they can close the gap quickly? Or will this gap persist? Yes, I actually think Meta's gap is relatively smaller now. They can continue to scale up; they have enough data and enough GPUs. I see no reason for the gap to be particularly large, and you have things like reflection or thinking machines, all sorts of methods to catch up. I don't think there's any reason for the gap to persist forever; it won't be zero, but I think the gap will continue to shrink. Mm. Like I joked earlier about whether there's a gap between iPhones and Androids, I think it's just various small gaps. But if you really gave me an Android, couldn't I use it? I think I could. I don't think there's a fundamental difference, and in the end, everyone just looks at the results. For example, whether the final performance is good or not matters far more than the price, and high-quality performance is indeed where large models excel, because they are larger and were developed earlier and more meticulously , so open source does need to invest more effort to catch up. So, in the long term, how do you see the future growth of ARR for models like OpenOPIC and STA? I think ARR will definitely continue to rise. There’s one thing I’m not entirely sure about, which is their ARR accounting and how they handle downstream revenue sharing; for example , they might count 100%of revenue, but some of it actually goes to cloud providers. So, as that revenue-sharing ratio potentially grows, their ARR might keep climbing, but their GAAP revenue might not necessarily follow the same trend. I’m just not sure how everyone will analyze that narrative once they go public. So you think that in the future, the revenue-sharing ratio for STA models will be smaller, and their bargaining power will be weaker? I think it will be a bit weaker. We often look at ARR without really seeing the specifics of their revenue shares, but once they go public, there might be better visibility into those things. Also, I wanted to ask—since Harvey is one of your key benchmark customers— do you think vertical-specific application models like Harvey will come out on top, or will general-purpose models like the S-model be better for legal tasks? Which side are you on? We are definitely on the side of vertical-specific models; we truly believe that doing a vertical well requires a massive amount of effort. We are infrastructure providers ourselves, and we often joke, "Next year we're all doomed, AGI will just code us out of existence." But that’s definitely just a joke. To do a vertical well, you need extensive testing to understand user needs, and figuring out how to translate those needs into our products is a huge part of the work. Many of our customers are vertical-specific agents; for example, Harvey is a particularly famous vertical-specific agent, right? They have all kinds of verticals, and each one needs to be evaluated to determine how well the work is being done. This actually represents a very, very heavy workload. For instance, we recently have a few others, like Doximity; they built a medical agent on our platform, which is essentially a competitor to ChatGPT. Doximity is the largest GP app for doctors in North America, and it focuses more on deep research to ensure doctors see correct references, so that they don't encounter issues when making patient diagnoses. For example, outside of North America, our largest partner is a company called HiD. HiD is the largest medical scribe and medical search company outside of North America, and they’ve done a lot of fine-tuning on our platform to achieve better results. Every vertical indeed has different evaluation criteria, and for us, we certainly hope that "a thousand flowers bloom," right? We hope that vertical-specific agents across all kinds of industries can succeed because we believe that every vertical has many very specific, nuanced tasks. It is quite difficult to use a single, massive model to handle all of these numerous, detailed tasks effectively. Take the various coding, legal, and medical fields we work with; each vertical has its own unique, granular requirements, which are the "alpha" of these software companies. We believe that if one lab tried to do all of this , well, how should I put it, the costs would be extremely high, and the quality might be quite low. With so many excellent, high-quality open-source models released this year, what does the growth in token volume look like from your perspective? Yes, our volume growth has actually been quite good. We recently made some public announcements because of our funding round; I recall we disclosed about 450 million tokens per day, and this figure is likely larger than what Gemini and OpenAI have publicly disclosed. So, you could say that the volume of open-source models running on our platform has already exceeded the B2B API volume of closed-source models, which I think is a fantastic development for open-source models. What are the main applications that account for the majority of token consumption on your platform? It is definitely coding, and as I mentioned before, perhaps medical or office work; from our observation, many tasks have been reframed as coding problems. If you can identify a specific industry’s vertical specifications—where tools exist to solve specific vertical problems— those tools can be converted into API calls, and then they train these models to become proficient in using those tools to solve that vertical's problems . We initially didn't realize that PowerPoint and Excel are actually coding problems too, because once you translate how to use PowerPoint and Excel into a coding problem, all existing tools for solving coding problems can be immediately applied, creating a better vertical-specific application. So, we really see that the vast majority are being redefined as coding; if they aren't coding, they are likely search. Research and coding are likely the two biggest categories, and various smaller categories are trying to figure out if they fall under research, coding, or both, then trying to slot their vertical-specific agents and tuned models into these classifications. Is there a direction where the current consumption isn't that high, but you're actually seeing very fast token growth? Co-work is definitely a major direction . I feel like everyone is spending most of their time catering to the hard sciences, but this is actually very counterintuitive. I actually feel that now we see many examples where the demand for catering to the humanities is growing, such as various SaaS, writing tools, and even things like video generation that have emerged recently, right? Most work is ultimately geared toward serving people , while a lot of this coding stuff or what OpenAI currently loves to do— boosting math benchmarks—I'm not sure what the significance of that is, but we might end up doing a lot to serve humanities co-workers later. We've done it ourselves, like the FireworkNexus I mentioned earlier, hoping to serve co-workers well. And I think, maybe a year and a half ago, when Computer Use Agents first came out, I was truly amazed. I thought, "Wow," everyone used that term "unhobbling"—that this model wouldn't need to mess around with all kinds of messy API calls anymore; it could just use a computer directly and do any work inside it. Right? But now I think it’s reversed; a year and a half later, it’s the other way around. Computer Use Agents are relatively a smaller vertical, and the vast majority of work flows through text, so you can actually see that most of it—Successful Chinese Open Source Models, let's put it that way. Llama is indeed a multimodal model, but those Chinese Open Source Models can achieve high performance, yet they are actually pure-text models. It's true that most work has returned to pure text because text models are cheap, and to cut costs , people shift their workflows toward text. But as VM prices get cheaper and people make VMs better and better, I believe Computer Use Agents still have a very promising future. It’s just about how to optimize this path; that is likely still being explored, yes. Could you help us estimate the total spending of North American customers on open models? Wow, That's a bit difficult. I think you'd be better off asking an analyst than me . However, when it comes to token ratios, I think looking at current public leaderboards might actually skew one's judgment. That’s because the vast majority of those public leaderboards are heavily influenced by free traffic. For instance, some companies offer incentives to get people to use their models for free on public platforms, which distorts the usage ratios you see. I often talk to my team about how we’ve been doing quite well, growing from 100 million to a billion last year, but in reality, Anthropic is growing much faster than we are. Right? They might already be at 70 billion. So, looking at token consumption isn't very honest because a token from Haiku consumes the same as a token from GPT-4 Flash if you only count the quantity. But I believe looking at revenue is the most authentic metric, as people’s willingness to pay will always gravitate toward the best. If people aren’t that price-sensitive, they will always pay for the best model available. So, when we collaborate with our customers to build specialized intelligence, most of the time we aren't looking at cost-effectiveness, but simply at the absolute best performance. Whether we can outperform Opus or Haiku on a specific task is our most important work. For example, when we work with DeepSeek, we also look at whether the final results can surpass Opus. Using frontier models as a benchmark and using their revenue as a benchmark is, I think, more honest. Looking only at consumption volume feels a bit deceptive. What do you think the trend for this number— meaning the total spending or token consumption of North American customers on open-source models—will look like over the next six months? We firmly believe that open-source consumption will continue to rise. I also have great confidence in our own progress; I believe we can definitely get this right, and while this happens, I don't think people's willingness to pay for the best models will change much. More and more people will certainly want the best models, so we will work harder to help more people make customized open-source models perform extremely well, which will drive more revenue; the specific ratio depends on how many people have the conviction to customize open-source models. Because many are still afraid that customizing a model might be too costly or have a low success rate— this is something many people worry about—but I believe if we continue to compete with closed models on quality, then the traffic for open-source models will keep growing; otherwise, the traffic for closed-source models will likely continue to increase. For the North American clients you see, when they compare different open models , which ones do they tend to prefer? The four domestic ones are all about the same to me: Kimi, GLM, MiniMax, and DeepSeek—I think all four of these have a lot of usage, and they sort of take turns. It really just depends on who launches something, and then their traffic might be a bit higher. Actually , I forgot to mention Qwen earlier; yes , Qwen is also very significant, so there are probably five major domestic ones: Kimi, GLM, Qwen, MiniMax, and DeepSeek. Over here, it might be Llama or Gemma, and while it used to be Llama , nowadays Mistral might be more popular. Is there a specific ranking or order to them? I think it's hard to define a ranking because, frankly, it’s just about whoever has the newest release— whoever just launched, their traffic goes up, and once someone else launches , the older model's traffic tends to drop. It’s hard to categorize, but as of today, for example in August 2024, Kimi definitely has the highest traffic . From your perspective, how is the current penetration rate of AI coding tools? Especially for open-source model penetration, I think it’s still on the low side. You see a lot of jokes on Twitter about how people say "open source is forever," or how they’ll never run out of usage for GPT-4, and so on. Right? But based on our observations, Cloud-based coding assistants are still far ahead of open-source models, which is why our company recently launched FWORK Nexus. We hope to bring open-source models into more companies; in reality, this work is often fundamentally a Go-to-Market challenge rather than a technical one. How do you view the total addressable market (TAM) for AI coding? How big is it, and what is its ceiling? Do you think we are still in a relatively early stage? I think the TAM question might need to be split between revenue and token consumption. Token volume will definitely keep rising—I am very certain of that—but when people talk about TAM, they usually mean revenue, right? As for revenue, if the trend we observe is that models like Llama are lowering prices, or Sonnet is lowering prices—the models everyone uses daily —if they are all cutting prices, it will definitely drive more consumption, but as to which direction the revenue will go, I'm not entirely sure right now. For instance, back in the day, the most valuable company might have been General Electric, right? But even if everyone uses electricity, the final most valuable company isn't necessarily General Electric. This is something that I think is quite difficult to judge. After a more powerful foundation model is launched, have you noticed any customers on the platform whose self-built fine-tuned models have been washed out? Some may have stayed, while others were washed out. What do you think is the biggest difference between these two types of customers? Well, I think it's like this: last year , quite a few were indeed washed out, but this year I feel very few have been . I think this is also a very positive signal. To be honest, if these models kept getting washed out, we wouldn't be growing this fast. We are now discovering that the direction many verticals are moving in may indeed be a bit different from the direction Frontier Labs is taking. Building a vertical-specific agent well and solving the Riemann Hypothesis are likely two completely different tasks. But Frontier Labs is probably more concerned now with whether they can solve the Riemann Hypothesis or create new drugs, right? This is what they need to prove, and we see more and more vertical-specific agents becoming increasingly specialized. Their data cleaning processes and similar tasks are indeed, I believe, very meaningful for the customization of open-source models. So, for companies like Harvey or Even, if today's top model providers decided to pour all their resources into the legal application vertical, would it be possible for them to defeat the current Harvey and Even? Certainly, that is possible. But that wouldn't support their overall ambition , and for them, that certainly wouldn't be a smart move, would it? I firmly believe that these vertical-specific agents conduct very detailed evaluations and clean their data well, ensuring they move forward on the path they deem correct. We often joke and call this "edged intelligence," meaning they perform very well in one specific direction, but for example, could their fine-tuned model solve the Riemann Hypothesis? I think that is absolutely impossible. However, for an AGI company , if the model they build is "edged intelligence," they might only be able to serve that one customer and wouldn't be able to make a model that serves many, many people. They need to have a reasonable distribution; I can't just obsess over whether I'm competing with EvenUp or Harvey. I need to ensure I take care of every customer on my plate . It’s like, one can’t just focus on making Szechuan or Cantonese food perfectly, even though many chefs only master one dish, right? And they only dare to charge a lot of money because they can master that one dish, right? But what a company needs to do is be the McDonald's; I have to satisfy everyone's palate. So, at most, for example, when I enter the Chinese market, I adapt to the tastes of Chinese people, and when I'm in the U.S., I take care of the tastes of Americans. But it's impossible for me to change the tastes for both California and Kentucky; if I changed them all, I don't think it would be fundamentally different from In-N-Out. Sorry, this example might be a bit of a stretch, but like In-N-Out, they compete with McDonald's every day without issue, but by just catering to California and a few other states, they manage to satisfy everyone's tastes by just doing their burgers right. Maybe this isn't that comparable to LLMs, but it really is the case that trying to satisfy everyone's tastes versus catering to one vertical are two different jobs. If there really were an AI company that could cater to everyone's tastes perfectly, then I believe they would definitely make a lot of money, it's just that we haven't seen that happen yet. When enterprises consider closed-source versus open-source models, what factors do they weigh? Actually, a huge point is trust; for us , since we do Fireworks Nexus—which leans towards coding-focused, co-work type tasks—and also customized models , a big part of the trust is trusting that this ecosystem can continue to support what they need to do for the next year or two. Because when we sign contracts, we sign them for one or two years, so when they think about it, they are signing with Fireworks; they don't really care about who we are, but rather whether the models served on our platform can fulfill their requirements for the next year or two. Building this trust is, I think, a relatively long-term process; we have to communicate our success cases with more customers, communicate with them about what we do well and what we don't, and what the pros and cons of open source are—we talk about this very transparently. And I think this trust is built better by closed-source models because they can post on Twitter every day about how they've hacked some company's system, or how they've solved some Riemann hypothesis, right? That’s something they do better, so I think the go-to-market capability is actually the biggest difference. That is, the go-to-market capability of open models and open-model companies is far weaker than that of closed-model ones. The vast majority of people's mindshare is still on closed-source models. Do you think these enterprises will consider whether to deploy locally or in the cloud? We've seen many on-premise deployment cases, and now we're seeing more cloud-based ones. I think this is fundamentally a trust issue—can you ensure security? Can you manage that aspect across various virtual environments? It’s a trust issue at its core, but the best thing about cloud deployment is multi-tenancy. Many enterprises aren't necessarily able to saturate the amount of GPUs they want to rent on their own. They might not be able to saturate them, but if we can aggregate enough demand, then saturating those workloads becomes a very easy thing to do. So, if the cost of this starts taking up a larger and larger share of enterprise operations, then multi-tenancy becomes a much better approach. People will want to opt for a multi-tenant setup where they can share the costs with others, which is better for both sides. Regarding the future of enterprise AI, from POC to actual production, what will the pace of adoption look like, and what are the key variables? I think you should look more at the ARR and adoption rates of these data companies. I believe more and more companies will purchase open-source models, but they will need evaluations; they will need to convince themselves that the purchase is justified. Many lack the capability for in-house evaluation, so they need third-party data companies to come in and handle these evaluations, and they can also use them for fine-tuning. So, I think a data company's revenue diversification and run rate are very interesting indicators. You'll find that if they continue to rely on big clients—those few major labs—and they account for the bulk of their revenue, then conversely, the choices available to customers might be fewer. For example, what are the main reasons you see for failure when enterprises move from POC to production? Right, it's that the evaluation is unclear, which makes it very easy for the POC to be done poorly. We've discussed this with several people. We usually ask them what kind of KPIs they have internally, and even though it's already 2026, many are still just going by "vibes"—checking if it feels like it works well, and that's it. Even though this might already account for a significant portion of their company's cost center, they still might not be particularly rigorous. So, I personally think that while many engineers currently worry about losing their jobs , the reality is that the number of people who can actually do this work is truly minimal. If you are interested, you can look more into how to do this and teach your boss whether you can procure open-source solutions. By properly setting this up and switching from closed to open source, you could save three to five times your salary. I think this is a muscle many companies lack, and it’s arguably one of the biggest points affecting a POC. So you think there are still many jobs in this evaluation space that haven't been fulfilled yet. I think the difference between new AI companies and old software companies isn't even that significant. This might be a controversial statement, but back then, software companies had internal test suites—for example, what tests need to be met to run a SQL engine. Now, as a vertical AI company, I have a bunch of evaluations, and before delivering an agent, it has to meet certain evaluation criteria. These things all require accumulation and refinement, and a significant amount of time and money to get the distribution right, so I really don't think the difference is that big. I think it’s just that people’s work is slowly shifting from writing tests to writing evaluations, and it’s just a matter of how many people can adapt to that transition and how quickly. Fireworks has two main businesses: inference optimization and customized models. Let's look at the first business, inference optimization. How much room for growth do you think there is for your inference optimization? I think inference optimization still has a long way to go. Inference optimization involves many delicate details. For instance, there is a huge difference between how a new model performs when it first launches and its ultimate performance after a year of optimization. For us, shortening that time and ensuring it performs at its best right from the launch is what we have spent a lot of time and effort on. Also, a very important part of inference optimization is how to handle caching effectively. The essence of good caching isn't necessarily in the serving runtime itself, but whether your supporting infrastructure can handle it well. We will continue to invest more energy into this caching aspect. I think inference optimization is essentially a very, how should I say , labor-intensive job; there are just so many things that can be done. Compared to the APIs provided by model creators, what advantages does your optimization offer? Often, because we optimize so many models, the work we do tends to be a bit more detailed. For example, many 1P APIs can serve their own models well, but many API patterns are things we've observed from other models. When we talk to many open-source model providers, we point out things they might have missed or workflow patterns they haven't seen much of, which we see more often, so the distributions are different. They might want to build a larger, more comprehensive API, but for us, since we often serve customized models, the traffic patterns for those workflows are likely different. In that case, the optimization work required will also be quite different. Regarding your model routing, such as combining proprietary SOTA models with open-source models, you mentioned that this is not just about cost, but also about improving performance. Especially in your joint research with Harvey, which I saw on your official blog, it mentioned that the proprietary model is only called 0.83 times per task on average, yet the quality surpasses using the strongest model alone. Why does a division of labor between multiple models lead to a better task completion rate? How is this achieved? Because the work scenario influences the harness, and the harness determines when each model should be invoked. In the specific Harvey example, I recall Opus served as the executor and GPT-4 as the advisor, which allowed breaking down tasks with very long contexts into tasks with slightly shorter contexts. This is fundamentally because LLM performance tends to degrade as the context grows; knowing this characteristic, and knowing that the task involves long text, we can improve the harness to optimize the specific model scheduling. So, I think every task uses models differently with various nuances, and these differences dictate how we approach our optimizations. This work is sometimes a bit like the performance optimization people used to do on CPUs. Even though code might look similar before optimization, to understand the CPU's characteristics, you often need to perform various decomposition tasks, breaking it down into work suited for that CPU. This is where we need to explore with our customers—treating the underlying model as a text processor, identifying its processing characteristics, building a harness based on those traits, and then placing it in the most suitable model and location based on the harness's results . From the perspective of optimizing the cost per token, what hardware configurations do you think still have room for optimization? For example, the capacity ratio between GPU and HBM, the ratio of CPU to DRAM, the ratio of DRAM to NAND, as well as HBM bandwidth and SALP bandwidth. This question is quite specific; let me think about how to answer it. I think those interested in this topic should watch a chalkboard talk video by Dwarkesh Patel from about two or three months ago. In that video, he invited someone to explain GPU architecture and such, but there was actually a very important number mentioned, a number that has no dimension. Essentially, it's the ratio of compute to memory. The ratio itself is dimensionless, but every piece of hardware has a different ratio, which influences how each model is designed. Before training a model, one would think, "What hardware will this model be served on?" Once I have determined the likely hardware for serving, I will modify my model architecture accordingly. So, most of the time when we ask how hardware can be improved, we overlook the fact that the vast majority of models are trained specifically to be served on existing Nvidia or AMD hardware. Since the ratios between AMD and Nvidia are quite similar, you can often quickly move an Nvidia-served model to AMD, but moving it to other hardware might require significant effort because the ratios are different, and the hardware characteristics are simply not the same . Conversely, if someone could find extremely cheap hardware that provides a similar ratio—or even a different one, that’s okay—and then incentivize model developers to overfit to their specific cards, then the nature of this problem would be completely flipped. The question is how I can provide enough incentive for everyone to overfit to my hardware; first, the card's volume must be large, and second, the card's ecosystem must be strong. So, I may not have answered your question directly regarding how this hardware should be modified or where there is room for optimization. I think a point Jensen Huang made earlier was excellent; he joked that if you have electricity now and don't buy my cards, you'll actually lose money. The point is that buying this card is essentially cost-neutral because it can earn the money back very quickly. Other cards might lack this characteristic, which means if I overfit to those other cards, I might not be able to earn that money back. In that case, the question might not be what my scale-up ratio should be, or whether I should use HBM or not—NVIDIA might not even use HBM itself, right? The question to ask might be whether I can find a card that integrates most of the reasonably priced upstream components, and through various incentive programs and a well-built software ecosystem, encourages people to overfit to my card , especially during serving. Whoever can provide this incentive, I believe they will succeed in the end, no matter how strange their final ratio might be. But for someone who cannot provide this incentive, even if their ratio looks exactly like NVIDIA's, they will still fail. Next, we want to talk about customized models: which companies need them, which ones don't, and how is this dividing line determined? I think vertical SaaS companies might be better off customizing, while we don't recommend specific consumer businesses to do so. For example, as we mentioned earlier with Doximity or Hidy , they engage with hospitals and know the needs of, say, 100,000 hospitals, so they can do this well and ensure they cover various cases; they have the ability to build a foundation model, or a model that supports various law firms , covering all sorts of cases, making it quite comprehensive. But when it comes down to each individual hospital or each legal practice itself, I think building customized models is still extremely difficult. They might be better suited to solidifying their workflows into a skill, writing internal tools, or using agents; that might be more appropriate. In these types of enterprise projects for customized LLMs, which part is the most time-consuming and expensive? GPU consumption is certain, and another part is data. During the process, we need to work with our customers to check if the data has reward-hacking behaviors, evaluate the environment stability, and look at how to scale up the environment. We review all these aspects with our customers. So, for the whole customization process, what kind of data volume is required? Perhaps a few thousand environments would be sufficient. So only a few thousand are needed? Why is that? Why isn't a larger volume of data necessary? Because during the entire RL process, the model constantly explores, so the amount it generates is far greater than those few thousand entries. For example, if every step has a rollout, say 128, then 1,000 entries could generate 128K of data, right? This volume actually grows with the environments; it may not need that many environments, but it still has a huge amount of data to teach the model what is correct and what is wrong. For instance, after a customer's model goes live, does it need to be retrained in the future? Retraining is necessary; for example, once a new model is launched, we help them rebase it. And this rebase process might not take as long; as long as the evaluation setup is fixed, it won't be that long. But customers always want to add new things, or they have new workloads coming in, and they want to push for harder tasks, right? So in that case, some customization work might still be needed, but once the evaluation is solidified, rebasing is a very fast process. What is the general rhythm for this retraining? It depends on the open-source model release cycle, yes. Next, I would like to discuss the core competitiveness of Fireworks. Helping enterprises deploy open models requires three types of capabilities: hardware compute, like companies such as DBRX or Groq; software optimization companies like yours; and companies with client relationships like Palantir. We are now seeing some hardware compute companies acquiring software capabilities, for example, companies like OSS Service that specialize in inference optimization. How do you maintain your moat? Will you become an asset-heavy company and buy your own GPUs? I think there are a few points; first, regarding whether Fireworks will make its own hardware, we definitely don't have that intention right now. Our setup is more like a secondary landlord, so we still want to build out the software layer. We truly believe there's value in optimizing the compute layer, and the difference between us and other companies is that we genuinely believe in specialized intelligence, and we are moving toward customized models. The vast majority of our traffic isn't vanilla-based models; it's all customized models. This might differ from other companies—for example, Nebus or other cloud providers might see most of their traffic in base models, and I find the competition for base model business to be quite intense . However, we firmly believe that the bulk of the profit lies in high-performing models. With high-performing models, our customers are able to charge more because they can deliver these results, and in turn, we can get a larger share of that revenue. That’s why we firmly believe in customized intelligence. We firmly believe that by doing training well, our clients can make money from their inference. What makes us different from other AI service companies is that many of them only do training or only do consulting. Their incentives aren't as aligned with the customer's success as ours; for training, they might just want the model to perform well without needing to earn much more, whereas we want their inference traffic to consistently grow. So, our incentive aligns with theirs in that they make money through inference, and we get a share of the money they earn. If it's a consulting firm or a pure training infrastructure company, I think their reach is limited because their incentives aren't as aligned with customer success—we align toward their inference traffic, while they align toward their training traffic. So , I think the real competitors aren't NE Cloud or the current providers, but rather the CSPs. CSPs have the capability, as I mentioned before, to build the entire chain; they can provide AI services and manage inference well. Therefore, we are actually looking more at the capabilities of CSPs and how to partner with them. I think they are, how should I put it, the final bosses to beat. It’s not necessarily that we are against NABS or New Cloud specifically; it's just that given their limited manpower, I think it would be very difficult for them to build out the entire chain themselves. Your relationship with CSPs is one of both competition and cooperation—a " coopetition" of sorts—so how do you maintain your niche? That question might be a bit hard to answer, but it's true that we work most closely with [Microsoft]. We are a 1P on Azure, meaning you can buy credits —Microsoft credits—and get our FWK services directly. So we do work very closely with them, and we also have partnerships with GCP as well. So, you’re saying your inference optimization is actually better than what the CSPs 'own teams can do. At least for now, yes. But honestly, I’m not that worried about whether the specific inference optimization is better or worse. If you ask me, CSPs have very strong capabilities across their entire chain because they’ve done this work before. What we care more about is completing the entire ML pipeline cycle; we want a better developer lifecycle experience. I think that’s something CSPs are better at, whereas current new clouds might not be as comprehensive. So what is your core competitiveness? Right, perhaps where we do better than CSPs is in providing a superior developer lifecycle, and compared to new clouds, our chain is more complete, while they might focus more on base models. For example, regarding the recursive superintelligence we’re discussing—if optimization, inference , training, and post-training could all be fully automated, where would your core competitiveness lie? AGI is extremely hard to predict, but if there really were an AGI that could do all those jobs well, I don't think there would be much core competitiveness left. I think the difficulty lies in the fact that even with AGI, you still need to talk to customers about their needs. It’s very difficult for it to figure out all customer needs from within a lab. If those needs can't be obtained through RSI, I think the growth of RSI might be limited. It might indeed be able to solve Riemann's hypothesis, but I’m not sure it can solve customer requirements. For Fireworks, looking from now to the next year or two, what do you think is your biggest challenge? Over the next year or two, it will still be compute. I think the challenge is whether we can get enough compute at a reasonable price in the market, because compute prices are currently very, very high, and since our revenue is ultimately compute-bound, acquiring enough compute will definitely be difficult. Since this summer, a major debate in the capital markets is that the quality of open-source models is catching up, which puts pressure on the revenue side of proprietary models. The previous trend might not continue, which in turn could pressure their future AI investments, hence why we see a significant pullback in hardware stocks and the semiconductor industry as a whole. With open models continuously catching up, how do you view this narrative in the capital markets? I think the point isn't whether the ROI is high or not; I think if you look at it the other way, suppose these big players miss this wave, even if it's highly profitable, they’ve missed out . Then they might never be able to get back to the table again. So their perspective is that even if I overspend , I won't face much of a downside. Right? At worst, my books look a bit ugly for a couple of years, and I might need five to ten years of depreciation to digest this. That's not an existential risk, is it? They have 20 or 30 years ahead to digest this; it's not a problem. But conversely, if they miss this wave, they can't fix it. So when Google or Meta run these numbers, they aren't just looking at the ROI; they're thinking, what if this thing is huge and I missed it because I didn't make the investment? Then I might be finished. For the CEO, I think they need to be more honest about just how high the investment costs actually are. However, based on our observations, in the vast majority of scenarios, it is still compute-constrained. Inference is definitely compute-constrained, and we certainly hope to find more compute, so I see a compute-constrained setup for at least the next year. Predicting beyond a year is truly very, very difficult, but they do have to realize that once the capital is deployed, it might take three years to break even, right? So this question is indeed open-ended, and I don't think it can be answered. But I think one thing we can look at is this: We can look at the debt of each company. Their bond ratings are very transparent, right? And if some companies have project delays or similar issues, their bond ratings will drop. So, I think the point here to test the market is: Suppose some player runs into bond issues over the next three to six months and can't borrow more money, requiring some restructuring; how many people in this market actually have the cash to bail them out? I think that’s a "watershed moment." Before this watershed moment, I think it’s hard to test this because we aren't sure how much cash everyone still has in their pockets to bail someone out. As long as they can be bailed out, I don't think it's a big problem, and everyone can keep building; but if they can't be bailed out, then I think it gets really difficult. Regarding open-source models catching up leading to a drop in hardware prices, I really can't figure out that correlation. But if you ask me , I think open-source models catching up is definitely a favor to infrastructure providers. Because it weakens the bargaining power of the model layer, right? So conversely, if the money remains the same, it must favor those who provide the infrastructure, right? Therefore, the share prices of those closer to the infrastructure should be rising, not falling. For example, when DeepSeek first came out last year, their stock price dropped too. I really don't understand what those people were thinking. But looking at it from another angle, open source definitely favors these hardware providers. To be honest, I'm not cut out for stock trading because I can't sort out the logic behind it. If you ask me, I think it definitely favors the infrastructure providers. Thank you so much, BNY, for sharing with us today; it was truly very inspiring. Thank you, thanks.

Loading...

Loading video analysis...