IT-PUB NEWS

Kog bets on faster AI inference on existing GPUs

15.08.2026 11:03 • Author: IT-PUB

Kog bets on faster AI inference on existing GPUs

The French startup says software tuning can speed up datacenter GPUs for AI inference, but it still has to prove the method on larger models.

French startup Kog is trying to show that AI inference can be sped up without waiting for new hardware. The company says it can pull much more performance from standard datacenter GPUs already used by enterprises, at a time when inference speed and cost have become major bottlenecks. That pitch has drawn attention, but it comes with an obvious caveat. Kog still has to prove the approach works on large language models, not just on a smaller demo.

Kog is pitching faster AI on hardware companies already use

Kog first drew wider attention in May, when a tech preview reached the front page of Hacker News. The startup said it wanted to demonstrate that “extremely fast single-request decoding” is possible on GPUs companies already own, including AMD MI300X and Nvidia H200 chips used in its demo.

The idea is straightforward, and timely. Many businesses are already running AI workloads on existing infrastructure and want to make those systems faster without replacing hardware. Kog’s argument is that software optimization alone may unlock more speed from GPUs that are already in use.

CEO Gaël Delalleau said the preview generated more than interest. According to IT-PUB News, the company received 200 tangible business leads after the launch.

Early demand is coming from software teams

From the early feedback, Kog expects software engineering to become its first major use case. That lines up with a familiar problem for users of AI coding tools: delays long enough to interrupt work.

Delalleau pointed to Claude Code as an example of a product where speed has direct business value. Anthropic, he noted, charges a higher price for Claude’s Fast Mode, suggesting that faster responses can justify higher costs for professional users.

Kog is targeting customers who are put off by delays because they depend on AI workflows for work. The startup also has design partners building games and apps from prompts, where faster results could translate into more revenue.

At the same time, Delalleau said the market is still not fully mature. Kog found that many potential customers are not ready to fine-tune small models, which pushed the company to focus on accelerating larger models instead.

The demo showed speed, but not yet on large LLMs

Kog’s headline promise is “30x faster LLM inference,” but the company still has a significant gap to close before showing that on the kind of models most people mean when they say LLMs.

Its demo delivered 3,000 per-request tokens per second, but that result came from a purpose-built model with about 2 billion parameters. The model, Laneformer 2B, is now open sourced. That makes the demo notable, but also narrow: it does not yet show the same performance on larger models, which are harder to optimize.

Delalleau said he believes the same method can work on LLMs as well. He argued that the idea GPUs are not suited for decoding is a misconception, and said newer GPUs have growing memory bandwidth that can be unlocked through software.

The next step is clear. Kog now has to move from a strong technical demo to a result that matters commercially, and the company says it plans to show its first major model running at 10x speed in September.

Kog’s low-level GPU work also limits how fast it can scale

Kog is not alone in trying to improve inference through software. Another French startup, ZML, has released hardware-agnostic software that bypasses Nvidia’s CUDA to support fast inference across competing chips.

Delalleau said Kog is closer in spirit to Stanford University lab Hazy Research, with a deeper focus on GPU acceleration. It is a highly technical, hands-on approach, and his background fits that work.

Delalleau studied solid-state physics at France’s École Polytechnique and later worked in offensive cybersecurity, also known as white hat hacking. He said both experiences shaped how he thinks about the problem: understanding the “laws of physics” and the “laws of the GPU,” then pushing the hardware as far as possible.

He also said his hacking background taught him to reverse-engineer systems at a very low level, down to assembly language and binary code, to use them for goals they were not originally designed for.

That kind of work is slow by nature. Delalleau said that for every new GPU, Kog may spend several weeks or even months digging into the details and doing engineering research on that hardware. With only 11 people on the team, that limits how many chips the company can support for now.

European backing may help, but proof still comes first

Kog’s long-term plan is to turn its methodology into agent-based pipelines that would let it support more chips and more models. That could matter beyond the startup itself, as Europe tries to build more of its own capability in those areas.

The company already has support from Scaleway and backing from France’s Bpifrance and the French Tech 2030 program. Delalleau suggested that this broader push could create favorable conditions for Kog as it tries to scale.

For now, though, the company still has to show that its method works on large language models. That step matters not just for technical credibility, but also for fundraising. Delalleau said that once Kog has implemented its first major model at 10x speed, it expects to show customer traction and then raise a Series A.

Until that happens, Kog remains a bet on a simple but ambitious idea: that the GPUs companies already own may still have more performance left than the market assumes.


Improve SEO for a small/medium business website for $50