Uploaded July 2026 | Updated September 2026, 3 weeks ago
The constraint on edge AI is not compute, it is RAM, and it is getting worse: phone makers are shipping less of it this year, and a 6GB Raspberry Pi costs 2.5 times what it did at launch. So Cormac Brick's team at Google AI Edge spends its effort making models small enough to fit. A 2 billion parameter Gemma, quantized to 2.9 bits per weight, runs on a Raspberry Pi at about 8 tokens per second and on a Qualcomm NPU fast enough for a few frames of vision a second.
Below that sit tiny models, from 500 million parameters down to 50, that reach the older laptops and cheap devices where even a small model will not fit. They usually need fine tuning rather than prompting, but the payoff is real: a fine tuned Gemma turns free text into the right function call across ten actions at over 86% reliability, and putting a speech model in front gives you voice to function calling. One shipped example is an offline voice dictation app with no subscription, built on two sub billion Gemma models that also strip your ums and ahs.
Speaker info:
- https://x.com/cormacb
- linkedin.com/in/cbrick
- github.com/google-ai-edge/gallery
Timestamps:
0:00 - Why intelligence at scale needs tiny models
1:17 - The Google AI Edge team and its open source stack
2:35 - Why run on the edge at all
3:25 - The real constraint: DRAM cost
4:40 - Small models: 1 to 4 billion parameters
6:08 - Shrinking Gemma to 2.9 bits per weight
7:36 - Decode speeds across Raspberry Pi, Jetson, and NPUs
9:30 - Try it yourself: AI Edge Gallery and a hobby robot
12:07 - When small is still too big: tiny models
13:24 - Off the shelf tiny models: ASR, vision, embeddings
14:28 - Fine tuning for voice to function calling
17:50 - In production: offline voice dictation
19:30 - Takeaways and Q&A
The constraint on edge AI is not compute, it is RAM, and it is getting worse: phone makers are shipping less of it this year, and a 6GB Raspberry Pi costs 2.5 times what it did at launch. So Cormac Brick's team at Google AI Edge spends its effort making models small enough to fit. A 2 billion parameter Gemma, quantized to 2.9 bits per weight, runs on a Raspberry Pi at about 8 tokens per second and on a Qualcomm NPU fast enough for a few frames of vision a second.
Below that sit tiny models, from 500 million parameters down to 50, that reach the older laptops and cheap devices where even a small model will not fit. They usually need fine tuning rather than prompting, but the payoff is real: a fine tuned Gemma turns free text into the right function call across ten actions at over 86% reliability, and putting a speech model in front gives you voice to function calling. One shipped example is an offline voice dictation app with no subscription, built on two sub billion Gemma models that also strip your ums and ahs.
Speaker info:
- https://x.com/cormacb
- linkedin.com/in/cbrick
- github.com/google-ai-edge/gallery
Timestamps:
0:00 - Why intelligence at scale needs tiny models
1:17 - The Google AI Edge team and its open source stack
2:35 - Why run on the edge at all
3:25 - The real constraint: DRAM cost
4:40 - Small models: 1 to 4 billion parameters
6:08 - Shrinking Gemma to 2.9 bits per weight
7:36 - Decode speeds across Raspberry Pi, Jetson, and NPUs
9:30 - Try it yourself: AI Edge Gallery and a hobby robot
12:07 - When small is still too big: tiny models
13:24 - Off the shelf tiny models: ASR, vision, embeddings
14:28 - Fine tuning for voice to function calling
17:50 - In production: offline voice dictation
19:30 - Takeaways and Q&A







![Building Turbopuffer: Gergely Orosz (@pragmaticengineer ) × Simon Eskildsen (CEO)
This fireside chat between Gergely Orosz and Simon Eskildsen explores the technical journey and engineering philosophy behind the database company Turbopuffer.
Video Timestamps
0:00 Introduction and Simon’s early history with computers
3:02 The International Olympiad in Informatics and early competitive programming
4:13 How Simon was recruited by Shopify while still in high school
8:46 Engineering challenges and scaling infrastructure at Shopify
14:56 Decision to leave Shopify and the creation of the napkin math project
20:40 The origin and technical motivations behind Turbopuffer
24:46 Design challenges of building a database on top of S3
28:41 Cursor becoming the first major customer
35:36 The meeting with Jensen Huang and Nvidia’s push for GPUs
39:01 The competitive reality of cloud infrastructure and CPU scarcity
43:06 Philosophical perspective on venture capital and funding
51:45 Building a remote-first culture with the campfire concept
Quotes
(19:49) Because you batch. So an f-sync happens on usually a 4K... its not intuitive. Its actually—I got caught—I just got obsessed with this question.
(30:33) Yeah, you could do a million vectors for a dollar. And before that, I think the cheapest was maybe $100 per million for something that actually worked.
(36:56) [Jensen Huang] said, Judging by your slide, maybe you should [pivot into vapes].
(43:53) I promised Cursor that Justine and I could get their bill to 4K a month... thats the pricing we ship with.
(49:54) The third reason to raise capital is for the founders ego... I wish that it was more talked about because youre diluting all of your employees when you do it.
## Speakers
### Gergely Orosz
Author / Founder, The Pragmatic Engineer · The Pragmatic Engineer
[X/Twitter](https://twitter.com/gergelyorosz) · [LinkedIn](https://www.linkedin.com/in/gergelyorosz/) · [Website](https://pragmaticengineer.com)
Software engineer, engineering leader, and author of The Software Engineers Guidebook; best known for The Pragmatic Engineer newsletter and blog covering software engineering practices, engineering leadership, and the tech industry. Previously held engineering leadership roles at Uber and worked at companies including Skype and Skyscanner.
### Simon Eskildsen
CEO and co-founder · turbopuffer
[X/Twitter](https://x.com/Sirupsen) · [LinkedIn](https://www.linkedin.com/in/sirupsen/) · [Website](https://sirupsen.com) · [Blog](https://sirupsen.com/napkin)
Co-founder and CEO at turbopuffer. Formerly Principal Engineer at Shopify, where he helped scale infra from 1K → 1M RPS.
— [View on the schedule](https://www.ai.engineer/worldsfair/schedule?session=asn_slot_2026_06_30_main_stage_1230_2026_06_25t07_57_06_000z) Building Turbopuffer: Gergely Orosz (@pragmaticengineer ) × Simon Eskildsen (CEO)](https://i.ytimg.com/vi/jQDXzEVHMSE/mqdefault.jpg)


