OTHER

Has the Search for AI Compute Uncovered the Next Cerebras?

The escalating demand for computers capable of executing AI models has grown more urgent, yet two major hurdles must be navigated by industry players: procuring the right chips and ensuring they can be effectively utilized in data centers to start bringing in revenue.

General Compute, an emerging neocloud focused on renting AI processing power during the operational phase—where models actively engage with users instead of being trained—provides valuable insights into the evolving AI landscape. These insights contributed to a $15 million seed round at a $60 million post-money valuation, spearheaded by FUSE VC, with additional support from Carya Venture Partners and Village Global Ventures.

So, what defines the right chip? The demand for GPUs has soared, yet it’s increasingly recognized that they may not be the optimal choice for running AI models post-training. The requirements for a model generating responses significantly differ from those during training, leading to the creation of a new class of chips designed specifically for this task. Nvidia’s $20 billion acquisition of Groq last December and Cerebras’ recent IPO valued at $57 billion exemplify this trend.

Due to capacity constraints at these companies, General Compute’s co-founders, CEO Finn Puklowski and CTO Jason Goodison, sought alternatives by turning to specialized chips from SambaNova, an Intel-supported chipmaker that focuses on inference and has somewhat stepped back from the Silicon Valley limelight.

This narrative might change when SambaNova unveils its new chips this year. The architecture is more flexible, employing additional memory to store context during inference computations. SambaNova asserts that its chips outpace not just GPUs but also specialized chips from firms like Groq and Cerebras. Puklowski notes that the new chips will achieve 600 to 700 tokens per second, compared to around 250 tokens per second for GPUs.

General Compute has placed an order for $300 million worth of the SN50 chips and claims it will be the first neocloud to implement them.

These chips also address the second major concern—where to install them—for General Compute: they are air-cooled rather than water-cooled and consume less power, enabling installation in existing data centers without substantial infrastructure upgrades.

Puklowski is exploring colocation arrangements—where General Compute installs its hardware in third-party facilities—not only with data center providers but also with cryptocurrency miners looking to repurpose their facilities, as bitcoin production costs have often exceeded market prices.

General Compute launched its cloud service last week, claiming to be the fastest at running MiniMax 2.7, a robust open-source large language model (LLM).

Venture investor Joe Hasselmann, who got involved early in the inference trend by investing in Groq in 2021, has established a new fund this year, Evercrest Capital Partners, focusing on the AI sector, with General Compute as his first investment. Hasselmann observes parallels between SambaNova’s partnership with General Compute and Coreweave’s collaboration with Nvidia, as well as Groq’s chip manufacturing alongside its previous cloud offerings.

“They require a diverse mix of customers willing to utilize their chips in high-growth environments,” Hasselmann stated. “Just as General Compute is betting on SambaNova, SambaNova is equally investing in General Compute.”

The critical question persists: which computer architecture will yield the highest value in the future of AI? Inference clouds represent implicit bets on a diverse landscape of multiple models and agents, where no single provider dominates, making the speed and cost of inference essential competitive elements. The recent $113 million Series B round raised by OpenRouter highlights this, showcasing the company’s ability to grant access to various models for optimal token expenditure.

Speed is paramount in this context, influencing both cost and capabilities. Puklowski intends to reduce coding tasks that currently take an hour down to just five or ten minutes, and to enhance the efficiency of audio agents in customer service, necessitating rapid inference for effective dialogue. “Using ChatGPT at 50 tokens per second is considerably faster than our reading capability,” Puklowski explained to TechCrunch. “As interactions shift to agent-to-agent communication, where agents conduct readings or database queries on our behalf, they must operate at an increased pace.”

When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.