WORK / ArchiveKelly Personal Marketing Intelligence OS
阅读READ
每日简报Daily Brief市场情报Market Intelligence品牌案例库Brand Casebook公司研究Company Dossier
收听与学习LISTEN & LEARN
播客Podcasts商务英语Business English
创作CREATE
创意工作室Creative Studio视觉素材库Visual Library作品集Portfolio
职业CAREER
面试题库Interview Bank营销工具箱Marketing Toolkit
资料库LIBRARY
收藏集Collections观察名单Watchlists来源体系Sources
我的Profile设置Settings
⌘K
更新于 —KKelly
今日情报播客来源我的
WORK / ArchiveKelly Personal Marketing Intelligence OS
阅读READ
每日简报Daily Brief市场情报Market Intelligence品牌案例库Brand Casebook公司研究Company Dossier
收听与学习LISTEN & LEARN
播客Podcasts商务英语Business English
创作CREATE
创意工作室Creative Studio视觉素材库Visual Library作品集Portfolio
职业CAREER
面试题库Interview Bank营销工具箱Marketing Toolkit
资料库LIBRARY
收藏集Collections观察名单Watchlists来源体系Sources
我的Profile设置Settings
⌘K
更新于 —KKelly
WORK / ArchiveKelly Personal Marketing Intelligence OS
阅读READ
每日简报Daily Brief市场情报Market Intelligence品牌案例库Brand Casebook公司研究Company Dossier
收听与学习LISTEN & LEARN
播客Podcasts商务英语Business English
创作CREATE
创意工作室Creative Studio视觉素材库Visual Library作品集Portfolio
职业CAREER
面试题库Interview Bank营销工具箱Marketing Toolkit
资料库LIBRARY
收藏集Collections观察名单Watchlists来源体系Sources
我的Profile设置Settings
⌘K
更新于 —KKelly
Market Intelligence/TechCrunch

Kog is going deeper to squeeze more inference out of GPUs

The idea that GPUs are poorly suited for agentic workflows may be a misconception, according to French startup Kog.

Anna Heim·2026.08.14EN
档案整理中本篇暂以摘要模式呈现,完整解析待补充。可点击右侧「阅读原文」查看来源。
事件背景基于真实抓取数据整理

本条来自 TechCrunch(AI / 创投),聚焦 consumer。 Kog is going deeper to squeeze more inference out of GPUs

Original Intelligence基于真实抓取数据整理

The idea that GPUs are poorly suited for agentic workflows may be a misconception, according to French startup Kog

  • Media & Entertainment
  • TechCrunch Brand Studio
  • Anna Heim
  • 7:50 AM PDT · August 14, 2026
  • Kog is going deeper to squeeze more inference out of GPUs

The idea that GPUs are poorly suited for agentic workflows may be a misconception, according to French startup Kog

Media & Entertainment

TechCrunch Brand Studio

Kog is going deeper to squeeze more inference out of GPUs

Anna Heim

7:50 AM PDT · August 14, 2026

The race for faster AI inference is on, and markets gave Cerebras and its purpose-built chips a warm welcome in its IPO debut in May. But French startup Kog is betting that there’s a lot more power to be squeezed out of conventional GPUs.

The startup hit the front page of Hacker News in May with a tech preview aimed at proving that “extremely fast single-request decoding is possible on the standard datacenter GPUs enterprises already own” — such as the AMD MI300X and Nvidia H200 GPUs it used for its demo.

Some were disappointed to hear this didn’t extend to GPUs in our laptops, but others saw the potential. With inference speed and costs now being a critical bottleneck, Kog’s promise to unlock new capabilities on existing hardware with software optimization attracted more than onlookers . “We had 200 tangible business leads,” CEO Gaël Delalleau told TechCrunch.

Based on early feedback, the solo founder expects software engineering to be the first use case. Veteran Claude Code users are well aware that they sometimes have to wait hours to get results. Anthropic itself understands that speed is worth money, and charges a price multiple for Claude’s Fast Mode.

Kog is hoping to target customers put off by those delays, usually because they rely on AI workflows for professional tasks. But the startup also has design partners that let users generate games and apps with a prompt, and for whom a faster outcome thanks to the Kog Inference Engine (KIE) would mean more revenue, Delalleau said.

The company realizes this market is not quite mature yet. While observing demand, Kog learned that its prospective customers aren’t prepared to fine-tune small models. “And that’s why since the launch, we’ve been fully focused on accelerating the development of larger models to meet the demand we’ve seen.”

This leaves Kog with a huge leap to make to deliver on its promise of “30x faster LLM inference.” Its demo showed an impressive 3,000 per-request tokens per second (TPS) — but with a purpose-built small model with only some 2 billion parameters, the now open sourced Laneformer 2B .

Contradicting skeptics, Delalleau is confident the same approach can work just as well with LLMs, whose size can be a challenge for inference chips. “GPUs have a bright future,” he said. For Kog’s CEO, the idea that they aren’t well suited for decoding has become a misconception; newer GPUs have more and more memory bandwidth that only begs to be unlocked.

Kog isn’t alone in thinking that software optimization can help GPUs do more than it says on the box. ZML , also from France, released hardware-agnostic software that bypasses Nvidia’s CUDA to support fast inference across competing chips. But Delalleau said Kog is more akin to Stanford University lab Hazy Research , with an even deeper-level focus on GPU acceleration.

Delalleau himself is not a researcher, and his first startup, TechCrunch50 2009 alum Stribe , has nothing to do with his new one — other than his former co-founder turned VC Kamel Zeroual, whose firm Varsity VC co-led Kog’s seed round. But the startup’s deep-level focus stems from his unique background.

Having studied solid-state physics at France’s École Polytechnique, he went on to work in offensive cybersecurity — also known as white hat hacking. According to Delalleau, this shaped the mindset he is now encouraging his team to adopt. On the science side, “there’s this mindset of understanding the laws of physics, and the laws of the GPU in order to make the most of them.”

As for hacking, the four-time finalist at DEFCON’s CTF tournament said it taught him “to reverse-engineer things at a very low level — down to assembly language and binary code — to understand how it works, and to try to use it to achieve a goal for which it wasn’t necessarily designed.”

The downside of this approach is that it is very hands-on and time-consuming. “For every new GPU, we’ll dedicate several weeks or even months, to really dig into the details and conduct GPU engineering research on that hardware.” With a team of 11 people, this puts a limit to the number of chips that Kog can work with, at least for the foreseeable future.

In the longer run, Kog hopes to feed its methodology into agent-based pipelines that will let it support more chips and models. As Europe seeks to build its own capability on those two fronts, this could add sovereignty tailwinds for the startup, which is already supported by Scaleway and backed by France’s Bpifrance and French Tech 2030 ’s program.

For now, though, Kog needs to prove to the world that its approach works on LLMs. This will also be key to securing more funding. “Once we’ve implemented our first major model at 10x speed, which I think will be in September, we’ll be able to start demonstrating customer traction and from there, raise our Series A,” Delalleau said.

When you purchase through links in our articles, we may earn a small commission . This doesn’t affect our editorial independence.

Anna Heim

Freelance Reporter

Anna Heim is a writer and editorial consultant.

You can contact or verify outreach from Anna by emailing annatechcrunch [at] gmail.com.

As a freelance reporter at TechCrunch since 2021, she has covered a large range of startup-related topics including AI, fintech & insurtech, SaaS & pricing, and global venture capital trends.

As of May 2025, her reporting for TechCrunch focuses on Europe’s most interesting startup stories.

Anna has moderated panels and conducted onstage interviews at industry events of all sizes, including major tech conferences such as TechCrunch Disrupt, 4YFN, South Summit, TNW Conference, VivaTech, and many more.

A former LATAM & Media Editor at The Next Web, startup founder and Sciences Po Paris alum, she’s fluent in multiple languages, including French, English, Spanish and Brazilian Portuguese.

October 13 – 15

San Francisco

Scale faster. Grow your portfolio. Gain practical expertise. No matter your goal, Disrupt can empower you.

Save up to $300 toda y!

Most Popular

Some Claude users are mad that Anthropic’s new watermarks will catch them using it at their jobs, classes

Lucas Ropek

Phoebe Gates and Sophia Kianni reportedly knew Phia was ‘cookie stuffing’ for months

Dominic-Madori Davis

Delta investigating after someone set up fake Wi-Fi network mid-flight

Lorenzo Franceschi-Bicchierai

Anthropic says it will watermark text generated by its AI models

Ivan Mehta

Mark Zuckerberg’s AI manifesto is exactly why people don’t like AI

Russell Brandom

YouTube now requires creators to have twice as many watch hours to start earning money

Aisha Malik

This ‘adversarial’ pattern can prevent surveillance cameras from detecting you

Zack Whittaker

X

LinkedIn

Facebook

Instagram

youTube

Mastodon

Threads

Bluesky

TechCrunch Staff Contact Us Advertise Site Map

Terms of Service Privacy Policy RSS Terms of Use Code of Conduct

Made by Google Phia Blacksmith Pixel Tag Disrupt 2026 Tech Layoffs ChatGPT

© 2026 TechCrunch Media LLC.

❧
Industry Analysis规则派生 · 可核对

本条目归入「Consumer Trends」垂直,涉及真实话题:consumer。

· 市场:关注 consumer 对相关品类与竞争格局的潜在影响。

· 消费者:受众行为与偏好变化值得追踪。

· 品牌:本动向对品牌资产建设的启示。

· 渠道:内容分发与触点组合(社媒 / 电商 / 线下)的协同值得复盘。

Marketing Insight规则派生 · 可核对

· 核心话题:consumer。

· 可思考:如何把「consumer」的洞察,转化为可衡量的内容与增长动作?

Career Usage规则派生 · 可核对

面试中可引用「Kog is going deeper to squeeze more inference out of GPUs」:围绕 consumer,说明你对行业动向的判断与可落地动作。

本条目相关英文术语可在「商务英语」模块按话题检索,用于外企面试表达训练。

关联播客真实 RSS 单集
Caracas under pressure: democracy in Venezuela
The Intelligence · 2026.08.13
Crude retreats but fuel prices stay high
FT News Briefing · 2026.08.07
The new rules of brand building, with Wieden+Kennedy CEO
Masters of Scale · 2026.08.04
延伸信源A / B 级权威来源 · 供深挖
Marketing BrewACampaignAThe DrumAWARCAAdweekADigidayA
Business English提取正文真实商业词汇
revenueaillmsaasipocreator
revenue

Kog is hoping to target customers put off by those delays, usually because they rely on AI workflows for professional tasks. But the startup also has design partners that let users…

ai

The race for faster AI inference is on, and markets gave Cerebras and its purpose-built chips a warm welcome in its IPO debut in May. But French startup Kog is betting that there’s…

llm

This leaves Kog with a huge leap to make to deliver on its promise of “30x faster LLM inference.” Its demo showed an impressive 3,000 per-request tokens per second (TPS) — but with…

系统商务英语 →
关联 English Brief
EXCLUSIVE: Reese Cooper Is Putting Down Roots in L.A.
WWD
造车狂奔十年,烂摊子全留给了车主
虎嗅
来源
阅读原文 · TechCrunch ↗
发布:2026.08.14
类型:AI / 创投
话题:consumer
相关阅读
TechCrunch/2026.08.15
Talks to sell PayPal to Stripe and Advent are heating up
TechCrunch/2026.08.15
Unforgetful is a new reminders app for people who can’t stop hitting snooze
TechCrunch/2026.08.14
Apple proposes to take a 15% cut of purchases made outside the App Store
TechCrunch/2026.08.14
Hyperscalers might regret embracing natural gas if new forecast proves correct
TechCrunch/2026.08.14
Anthropic set AI agents loose on the same task. They started a turf war.
个人笔记
自动同步到云端