OTHER

Kog Investigates Advanced Methods to Enhance GPU Inference Performance

The race for advanced AI inference capabilities is heating up, with Cerebras making waves following its IPO in May, credited to its specialized chips. In contrast, the French startup Kog believes there is still substantial power to be tapped from traditional GPUs.

In May, Kog attracted significant attention by appearing on the front page of Hacker News, unveiling a technology preview that showcased the feasibility of “extremely fast single-request decoding on standard datacenter GPUs that companies already possess” — namely, the AMD MI300X and Nvidia H200 GPUs highlighted in their demonstration.

Some were disappointed to find out that this capability doesn’t extend to personal laptop GPUs, yet many recognized its potential. Given the current challenges posed by inference speed and costs, Kog’s commitment to unlocking new capabilities on existing infrastructure through software optimization has garnered substantial interest. “We had 200 tangible business leads,” noted CEO Gaël Delalleau in an interview with TechCrunch.

According to initial feedback, the sole founder anticipates that software engineering will serve as the primary initial use case. Experienced users of Claude Code often face lengthy wait times for results. Anthropic understands that speed equates to value, which is why it charges a premium for Claude’s Fast Mode.

Kog targets clients who have been put off by such delays, especially those dependent on AI workflows for their professional tasks. The startup also partners with design collaborators to enable users to create games and applications through prompts, where faster outcomes from the Kog Inference Engine (KIE) could result in increased revenue, according to Delalleau.

The company recognizes that the market remains dynamic. While gauging demand, Kog found that its potential customers are typically not inclined to tune smaller models. “This is why, since our launch, we have been focused solely on accelerating the development of larger models to meet the demand we’ve identified.”

This sets a daunting challenge for Kog to deliver on its promise of “30x faster LLM inference.” Their demonstration revealed an impressive throughput of 3,000 tokens per second (TPS) — achieved with a specialized, smaller model featuring approximately 2 billion parameters, specifically the now open-sourced Laneformer 2B.

Ignoring the skeptics, Delalleau remains confident that the same principles can be successfully applied to larger LLMs, despite their size posing challenges for inference chips. “GPUs have significant potential,” he asserted, emphasizing that the belief that GPUs are ineffective for decoding is misguided; modern GPUs now provide enhanced memory bandwidth that needs to be effectively utilized.

Kog is not the only company that believes software optimization can expand the capabilities of GPUs beyond their advertised limits. The French company ZML has also launched hardware-agnostic software that bypasses Nvidia’s CUDA, facilitating swift inference across competing chips. However, Delalleau noted that Kog’s approach is more aligned with Stanford’s Hazy Research lab, focusing on advanced GPU acceleration techniques.

Delalleau himself does not come from a research background; his first startup, Stribe, which participated in TechCrunch50 in 2009, is unrelated to Kog, aside from his former co-founder Kamel Zeroual, now a VC whose firm Varsity VC co-led Kog’s seed funding round. Nevertheless, Kog’s focused strategy stems from Delalleau’s unique background.

Having studied solid-state physics at France’s École Polytechnique, he later moved into offensive cybersecurity — often referred to as white hat hacking. Delalleau mentioned that this experience has shaped the mindset he fosters within his team. Scientifically, “there’s a mindset of understanding the laws of physics and the GPU to maximize their potential.”

Regarding hacking, the four-time DEFCON CTF finalist explained that it taught him “to reverse-engineer components at a fundamental level — down to assembly language and binary code — to understand their operation and utilize them for purposes for which they weren’t initially intended.”

The downside of this approach, however, is its labor-intensive and time-consuming nature. “For each new GPU, we will dedicate several weeks or even months to thoroughly examine the specifics and conduct GPU engineering research on that hardware.” With a team of 11, this limits the number of chips that Kog can work with in the near term.

Looking ahead, Kog aims to apply its methodology to agent-based pipelines, allowing for broader support across various chips and models. As Europe seeks to enhance its capabilities in both sectors, this could strengthen Kog’s position, already supported by Scaleway and backed by France’s Bpifrance and the French Tech 2030 initiative.

At this juncture, Kog needs to validate the efficacy of its approach on LLMs. This validation is essential for securing further funding. “Once we roll out our first major model at 10x speed, which I expect will happen in September, we’ll be able to showcase customer traction and subsequently pursue our Series A funding,” Delalleau concluded.

When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.