Gerhard Lazu
360˚ with Docker @ KubeCon EU 2024, Paris - Day 2
updated
Nabeel Sulieman, Staff Software Engineer at GitHub, joins us for a makeitwork.club chat to talk about his 3‑month passion project that involves turning a very normal, very analog garage door into something that’s fully automated, observable, and Home Assistant–native.
In this video we cover:
* The initial plan and how the project evolved over 3 months
* The background context: why even bother upgrading a working garage door?
* Early ideas (including a few that didn’t make the cut) and what Nabeel learned from each
* Choosing “the brains” of the setup
* Microcontrollers, sensors, relays, and how everything talks to Home Assistant
* Figuring out the physics
* How do you safely and reliably sense door position, movement, and state?
* Dealing with power, wiring runs, and a not‑so‑forgiving physical environment
* What an optocoupler is, and why it matters for this build
* Installing ESPHome on a chip and getting it onto your Wi‑Fi + into Home Assistant
* Putting it all together into a reliable, testable system
* How to control two doors with a single microcontroller
* Mounting everything permanently so it’s safe, tidy, and spouse‑approved
* ESPHome tips, YAML gotchas, and debugging strategies
* The big question: was this actually worth it?
* Reflections on “shipping the scrappy v1” vs. endlessly polishing
* What surprised Nabeel the most during the build
* Where he wants to take the project next (including a GitHub Action that ties into the setup)
If you’re into Home Assistant, ESPHome, microcontrollers, or just love seeing a software engineer tackle a real‑world hardware problem, you’ll get a ton of practical ideas from this conversation—plus a realistic look at the tradeoffs, dead ends, and aha moments along the way.
If you enjoyed this and want to explore more:
- Small group discussions like this one in 🪩 https://makeitwork.club (the next one is on Nov. 21, 2025 where I cover my experience of KubeCon NA 2025)
- Become a member at 📺 makeitwork.tv for the full length content (includes Jellyfin media server access)
- Tune into 🎧 https://makeitwork.fm for the podcast
00:00 - The plan for today
00:51 - Intro
02:09 - The Background
04:19 - Idea 1: Replace the whole system
05:12 - Idea 2: Use a pre-made kit
05:47 - Idea 3: Build it yourself!
07:30 - The brains of the setup
09:17 - Figuring out the physics
13:18 - What is an optocoupler?
15:35 - Installing ESPHome on a chip
18:28 - Putting it all together
20:43 - How can you connect both doors with one microcontroller?
22:05 - How to mount everything permanently?
27:15 - ESPHome
29:25 - Was this worth it?
32:01 - I am wondering...
33:05 - Ship the scrappy v1
33:37 - I forgot to mention...
36:02 - What else surprised you during this build?
38:56 - How are you thinking of continuing this?
40:57 - A GitHub Action that...
What did you think of Nabeel's passion project? Let us know in the comments below!
After 3 years of tinkering, Matt Johnson, Senior Site Reliability Engineer, shares with us:
- his motivations behind building a local, private environment for experimentation
- the values that guided his homelab design
- why he started with Incus - linuxcontainers.org/incus - and deferred Kubernetes for 3 years
- how OpenTofu & Ansible are used to bootstrap hosts
- the case for PowerDNS & WireGuard
In the second half, Matt shows us how he spins up a new Incus host, and what are some of the considerations for the entire homelab as this new host takes shape.
If you enjoyed this, you are likely to also enjoy joining https://makeitwork.club where we discuss about topics like this one every 2 weeks.
Homelab stats:
- 64 vCPUs
- 240GB RAM
- 10.75TB storage
- $5,000 over a 3 year period
- $25/month running costs
00:00 Intro
00:44 Why I started a homelab
02:47 What do I value in a homelab?
05:26 What am I running in my homelab?
09:09 How I settled on Incus?
11:54 Could the kids help with the homelab?
14:31 When do you have time to work on this?
15:45 Does Incus provide scheduling & orchestration?
17:27 What services do others depend on in your homelab?
19:59 PowerDNS & WireGuard
21:48 What do you use for observability?
23:47 What am I working on now?
27:26 A Christmas present for Gerhard
28:47 Key learnings
31:12 How much would 64 vCPUs cost in AWS?
32:26 How do you install the host OS?
34:41 What is the easiest way to consume these learnings?
39:06 Let's spin up a new Incus host
46:37 What would help understand this setup?
48:44 Infrastructure diagram example
51:41 Software for infrastructure diagrams
55:20 What did you learn from your homelab?
58:56 Wrap-up
Have you wondered how fast is Varnish HTTP Cache (soon to be renamed to Vynil) when CPU & network are not a bottleneck? FWIW, this HTTP benchmark is for the open-source CDN that powers changelog.com: github.com/thechangelog/pipely
My 10Gbit connection was the bottleneck last time, so I went to work:
- Ryzen 7 5800X for the client, github.com/hatoo/oha
- Threadripper 9970X for the server, Varnish 7.7
- 2 x 100gbit NVIDIA Mellanox ConnectX-5 connecting both hosts via QSFP28 DAC
- Adam Stacoviak & Jerod Santo as my "partners in crime"
If you enjoyed this, there's more to it:
- Full conversation: youtube.com/watch?v=hYahjUGoy5Y
- Join the most loyal members for small group discussions every 2 weeks 🪩 https://makeitwork.club
- Subscribe at 📺 makeitwork.tv for the full length content
- Tune into 🎧 https://makeitwork.fm for the podcast
00:00 - Intro
00:33 - 64vCPU @ 5.4GHz max boost
01:36 - 16vCPU @ 4.7GHz max boost
02:26 - MCX556A-ECAT
04:05 - 20Gbps HTTP benchmark
06:16 - 40Gbps HTTP benchmark
08:24 - 80Gbps HTTP benchmark
10:01 - changelog.com on Vynil on Fly.io
11:11 - What's next?
14:45 - One more thing...
For my first pair of 100gbit network cards, I didn't know what to expect. I installed drivers the WRONG way and hit an issue on the first kernel upgrade. That bend was also driving me crazy! What I didn't think of is to check just how hot this network card actually runs.
Until the next https://makeitwork.club when I will dig into all the improvements - SPOILER(3D printed tall bracket, LACP bond, 40mm Noctua fans, CDN benchmarks) - this is a quick peek into what is coming.
Thank you for all the great comments, and especially those that addressed the lack of active cooling on the Mellanox ConnectX-5 card, and suggested a temporary fix until I get a 3D printed tall bracket.
00:00:00 Intro
00:00:29 I was doing MLNX_EN wrong
00:01:34 _jodi33 has a solution for the bend
00:02:41 WillFul noticed that there is no active cooling
00:04:37 99Gbits/s for Marcos
00:07:31 Does the storage support 100Gbit/s?
00:09:46 Hopefully NO MORE Kubernetes
Rio Kierkels, Head of Platform Engineering for Tribe.ai, shares his expert setup, from a lightning-fast shell prompt (you know it as starship.rs) to reproducible dev containers and seamless dotfile management.
🛠️ TOOLS MENTIONED IN THIS VIDEO
1. starship.rs - the minimal, blazing-fast, and infinitely customizable prompt for any shell!
Learn more about Starship
2. Bluefin OS - a custom image of Fedora Silverblue, designed for developers who need a reliable and powerful workstation.
3. Distrobox - use any Linux distribution inside your terminal. Enable both backward and forward compatibility with software and freedom to use whatever you want.
4. Devpod - create reproducible developer environments with a single command.
5. chezmoi - a powerful dotfile manager that helps you manage your personal configuration files across multiple machines.
6. mise - a next-generation version manager for your development tools, replacing asdf, nvm, and more.
7. Bat - a cat clone with syntax highlighting and Git integration.
If you enjoyed this and want to explore more:
- Become a member at 📺 makeitwork.tv for the full length content
- Tune into 🎧 https://makeitwork.fm for the podcast
- Join the most loyal members for small group discussions every 2 weeks 🪩 https://makeitwork.club
00:00:00 - Intro
00:00:37 - starship.rs
00:02:10 - Bluefin OS
00:03:55 - Distrobox
00:08:41 - xeyes & cmatrix
00:10:38 - devpod.sh
00:12:22 - chezmoi
00:15:42 - mise
00:18:27 - setup.sh
00:23:28 - How do you run databases in dev?
00:24:48 - What do you use for live-reloading K8s in dev?
00:26:28 - Let's build a make-it-work devpod
00:27:25 - What do teams often miss?
00:28:50 - What is good team communication?
00:32:10 - Thank you Rio for sharing all these tips with us
#cli #unix #linux #terminal #dev #devlife #devops
Have you ever looked at your shiny new, maxed-out MacBook Pro and thought: "This is great, but where's my 100Gbit dongle?" Yeah, me neither.
My first goal at the new startup was to build a 100Gbit homelab. I needed to benchmark high-performance software for live, cross-host process migrations (memory & all), remote GPUs and a bunch of other things which are as ridiculous and awesome as they sound.
In this https://makeitwork.club conversation, I take you on my journey to build a home network that's 2x as fast as the fastest MacBook Pro NVMe available today. We'll cover:
- 100Gbit dongles
- Why most consumer grade motherboards are not enough
- The great Linux distro debate (the answer might surprise you!)
- Jumbo frames (MTU 9000) and a few more tuning tips
- A live unpeeling that was definitely not what it looked like
If you are a networking nerd, a homelab enthusiast, or just someone who enjoys watching nerds getting way too excited about PCIe lanes, then this video is for you.
This is about half of the full conversation that we had in https://makeitwork.club, a place where we started meeting every 2 weeks. Everyone is welcome to join!
There is more to this:
- Become a member at 📺 makeitwork.tv for the full length content
- Tune into 🎧 https://makeitwork.fm for the podcast
00:00:00 Intro
00:00:45 New company = new Apple hardware
00:02:20 Where is my 100gbit dongle?
00:02:59 Would this hardware work for 100gbit?
00:05:38 The second half of the homelab
00:07:44 Which Linux distro did I pick?
00:09:47 Check network card PCIe speed
00:12:05 netplan
00:13:10 jumbo frames
00:15:13 Why only 95Gbps?
00:17:10 mlnx_tune
00:18:00 How much is 98Gbps?
00:18:52 What will this new host be used for?
00:20:57 100Gbit networking
00:24:58 Thunderbolt 5
00:27:15 OWC 10Gbps dock
00:28:09 One more thing...
00:29:47 Intro to next meeting's topic
What started as a fascinating "Exit Interview" blog post - fly.io/blog/the-exit-interview-jp - uncovering the human element behind a massive piece of technology, turned into this in-depth conversation about building and scaling a global, multi-tenant system.
In this conversation, JP shares his reflections on the evolution of Fly.io Machines, the critical design decisions that shaped the platform, and the trade-offs made along the way. We explore the journey from a Nomad-based architecture to the custom-built Fly.io Machines, the reasoning behind choosing BoltDB over SQLite, and the crucial role of fly-proxy in handling traffic at a massive scale.
We also take a look at some of the open-source components that power Fly.io, including a code walkthrough of the Finite State Machine (FSM) that manages the lifecycle of machines and volumes. JP explains how this simple yet powerful component ensures reliability and simplifies operations.
Whether you're a software engineer, a DevOps professional, or just curious about how large-scale distributed systems are built, this conversation is packed with valuable insights and hard-earned lessons.
If you enjoy this, there is more:
- Join 🪩 https://makeitwork.club for regular members-only discussions
- Tune into 🎧 https://makeitwork.fm for the podcast
- Sign-up at 📺 makeitwork.tv for more content
00:00:00 — It all started with The Exit Interview
00:03:24 — Have any of your thoughts refined since the blog post?
00:09:12 — How did Fly.io Machines start?
00:17:20 — Initial design considerations
00:28:21 — SQLite vs BoltDB
00:36:07 — Why no database for flaps?
00:38:50 — fly-proxy
00:46:21 — Open-source components
00:54:09 — Let’s look at the FSM code
01:06:57 — Something simple that works as designed
01:11:07 — Our take aways
Whether you're new to Talos or looking for a streamlined way to run K8s on OCI, this video is for you!
This is about half of the full conversation that we had in https://makeitwork.club, a place where we started meeting every 2 weeks. Everyone is welcome to join!
There is more to this:
- Become a member at 📺 makeitwork.tv for the full length content
- Tune into 🎧 https://makeitwork.fm for the podcast
00:00 - Introduction
00:30 - Context for Talos Linux on Oracle Cloud
02:25 - Why Kubernetes?
04:42 - Why Talos?
05:34 - What's the first thing that you like about Talos?
07:05 - A look at the Talos Linux docs
08:11 - Improving the single-node Talos install
09:45 - Creating an Oracle Cloud Image with factory.talos.dev
12:48 - The Oracle Cloud Image (OCI)
13:49 - Let's create an instance
15:15 - Configuring the new Talos instance
18:11 - Using talosctl gen config
18:44 - Setting up a single-node Talos cluster
20:48 - Exploring the talosctl dashboard
21:12 - Where do you keep configs with credentials?
22:54 - Generating the talosctl kubeconfig
23:11 - Did it work?
25:15 - Let's start deploying stuff!
27:05 - Who would like to see part 2?
29:00 - Troubleshooting connection issues
31:09 - Key takeaways
32:51 - What's next?
Marcos Nils shows us how to use it to debug all TLS requests from Docker to any container image repository.
eCapture also supports:
- openssi, libressl, boringssl, gnutls, gotls & nspr(nss)
- capturing bash & zsh commands for Host Security Audit.
- sql queries from mysald 5.6 & postgres 10
- pcap mode
There is more to this:
- Sign-up at 📺 makeitwork.tv for the full length content
- Tune into 🎧 https://makeitwork.fm for the podcast
- Join 🪩 https://makeitwork.club for regular members-only discussions
If you found this video helpful, please give it a thumbs up, and don't forget to subscribe for more content on cloud infrastructure, DevOps, and all things tech! Let us know in the comments if you have any questions or what you'd like to see next.
There is more to this:
- Sign-up at 📺 makeitwork.tv for the full length content
- Tune into 🎧 https://makeitwork.fm for the podcast
- Join 🪩 https://makeitwork.club for regular members-only discussions
00:00:00 Intro
00:00:33 Serverless talk while we deploy a new AWS CloudFront distribution
00:02:48 Let's test this AWS CloudFront
00:05:06 AWS Route53 Geolocation records
00:06:05 Let's test everything by selectively resolving registry.dagger.io to the new setup
00:08:06 So what did we do today?
00:09:10 How do you want to continue with this?
00:12:03 This is a huge milestone
While ghcr.io provides a generous free tier that served us well for years, the existing analytics on ghcr.io offer limited insights, latency is sometimes poor, and unavailability hits hard. A custom domain allows us to gather more granular data, makes requests consistently faster & more reliable anywhere in the world, and we have the option of switching container registries without disrupting users.
A significant challenge in this project was understanding the HTTPS traffic between the user's machine and the CloudFront endpoint. To do this without the complexity of a man-in-the-middle proxy, we used eCapture, a powerful tool that leverages eBPF (Extended Berkeley Packet Filter) to capture SSL/TLS traffic in plaintext. This allowed us to inspect the requests and responses, understand the authentication flow, and identify which requests to cache.
By using eCapture, we demonstrate that the new setup works as expected. We show how the Docker CLI interacts with the CloudFront endpoint, how the token is fetched, and how the manifest and blobs are pulled. We also demonstrate that the caching is working correctly, with subsequent requests for the same image being served from the CloudFront cache, resulting in a significant performance improvement - it's 100% faster.
We plan to continue optimizing this setup including:
- Caching the token: To limit the impact of ghcr.io service disruption
- Fine-tuning the caching policies: To further improve performance and reduce costs
- Origin shield: To reduce the load on ghcr.io
There is more to this:
- Sign-up at 📺 makeitwork.tv for the full length content
- Tune into 🎧 https://makeitwork.fm for the podcast
- Join 🪩 https://makeitwork.club for regular members-only discussions
00:00 Intro
00:36 What are we trying to achieve?
01:44 Why custom domain for ghcr.io?
03:34 registry-redirect
05:48 vector.dev
07:45 ecapture
14:44 ecapture new image from registry.dagger.io
21:56 Does it work?
29:47 How can we improve on this?
31:56 Is it faster than ghcr.io?
36:59 How to set this up
42:58 Follow-up optimisations
In this demo, Shivansh Vij, CEO of Loophole Labs, introduces us to architect.run, a groundbreaking live migration technology in action.
We start with a standard, off-the-shelf PostgreSQL OCl image inside a virtual machine. We then live-migrate it between two different AWS instances - back and forth, every three seconds.
Watch our real-time dashboards as we track:
- The VM's location bouncing between hosts.
- Continuous, successful read/write operations from a single psql client.
- Stable, persistent network connections that survive every migration.
- The most incredible part? The client application is completely unaware of the chaos. It connects to the exact same IP address the entire time, thanks to our intelligent network plane that handles all the magic behind the scenes. No custom database builds, no application code changes required.
Features worth emphasizing:
1. Zero-Downtime Migration: We move a stateful PostgreSQL database between hosts in seconds, demonstrating true, seamless live migration.
2. Complete State Preservation: In-memory data, disk state, and active network connections are all perfectly preserved and transferred during every migration.
3. Cross-Cloud Ready: While this demo uses two AWS instances, our custom kernel is built to normalize CPU features, enabling migrations between different cloud providers (like AWS and GCP) and even different CPU architectures-a problem that makes traditional snapshots non-portable.
4. Instant Fault Tolerance: We discuss how this technology provides incredible high availability. If a host machine were to disappear without warning, the workload could be restarted from its last snapshot (in this case, from just 3 seconds prior) on another machine.
00:00 - Introduction: The GCP Kernel Update Story
00:58 - Explaining the Live Demo Dashboard Setup
02:06 - The "Wow" Moment: Migrating a Live Database Every 3 Seconds
02:21 - How Network Connections Survive the Migration
03:04 - Proof: Real-Time Read/Write Operations to PostgreSQL
04:36 - The Magic: How the Client Connects to a Single, Unchanging IP
05:14 - No Custom Builds: Using an Official, Off-the-Shelf PostgreSQL Image
06:17 - The Secret to Multi-Cloud: How We Handle Different CPU Architectures
07:20 - What Happens if a Host Fails? (High Availability & Fault Tolerance)
07:48 - How Snapshots Enable Near-Instant Recovery
10:08 - Visualizing the Migration Time
#LiveMigration #CloudComputing #HighAvailability #PostgreSQL #AWS #DevOps #ZeroDowntime #CloudMigration #Virtualization
18 months ago with a simple question: "Can we do this?" Watch as we culminate months of work by migrating our entire production traffic from Fastly to our new, open-source CDN, "Pipely," live on stage.
We started with a cache hit ratio of just 17.93%, meaning over 80% of our requests were slow and prone to failure. Follow along as we debug a last-minute memory issue, push the fix to production in under three minutes, and perform a live DNS switchover. The result? A staggering 99.1% cache hit rate!
This is a real-world demonstration of continuous improvement, embracing public learning, and the power of building with friends.
00:00 - The story of Pipely: How it all began 18 months ago.
01:12 - Welcome to Kaizen 20: The first in-person meetup after a decade of remote collaboration.
02:20 - All our work is open source! Find the code and discussions on GitHub.
04:10 - The Problem: Our abysmal 17.93% cache hit ratio.
06:21 - Let's fix it live! Debugging an out-of-memory crash that happened the night before.
07:04 - Pushing the fix to production with a simple Git tag.
08:41 - Watching the live, zero-downtime, blue-green deployment across our global infrastructure on Fly.io.
11:47 - The Moment of Truth: Switching DNS from Fastly to our new CDN, Pipely.
14:20 - The Result: From 17% to a 99.1% cache hit rate!
15:05 - Verifying the DNS change has propagated worldwide.
Learn more about the project: https://pipely.tech/
Dive into the code and technical details (Discussion #546): github.com/changelog/changelog.com/discussions/546
A huge thank you to our friends and contributors who made this possible, and to Fly.io and DNSimple for their amazing platforms.
#Kaizen #LiveCoding #DevOps #CDN #Varnish #CloudNative #OpenSource #WebPerf
Follow the journey as we:
- Start with a completely fresh Windows virtual machine to hunt down the root cause.
- Discover how containerizing our entire local environment with Dagger was a game-changer.
- Uncover the subtle but critical differences between how PowerShell and WSL handle home directories, leading to our "aha!" moment.
- Make the tough call to simplify, rather than over-engineer, by standardizing our workflow on WSL for everyone.
This isn't just about fixing a bug; it's a story about simplifying complexity, the power of good documentation, and the satisfaction of deleting code that's no longer needed. We improved the developer experience, cleaned up our repository, and made Pipely better ever.
What is your experience running just.systems in Windows? Let us know in the comments below!
makeitwork.tv 👈 full version shipping any week now
0:00 Intro
0:43 Windows dev env
1:35 Update docs for local dev tests
2:27 Just home_dir() does not work correctly in Windows
5:22 Read your own PRs
8:12 tmux is not the easiest
9:38 Just in Powershell vs WSL
11:15 Another issue...
11:59 Working from home FTW
See the context: youtube.com/watch?v=rFcC8cAbW-M
Watch the full pairing session: makeitwork.tv
From outdated docs, to outdated forks, to just & dagger. It's a small excerpt from the final version that will ship any week now on makeitwork.tv
WDYT, are the dev docs better now? github.com/thechangelog/pipely/blob/v1.0.0-rc.5/docs/local_dev.md
We implement the following one-liner bash script the Elixir Way™:
dig cdn-2025-02-25.internal AAAA +short | while read ipv6; do curl -so /dev/null -w "%{http_code} %{method} %{url_effective}\n" -X PURGE "http://[$ipv6]:9000/"; done
Find the code and the conversation in this pull request 🐙 github.com/thechangelog/changelog.com/pull/549
🤔 THE PROBLEM: CDN is serving stale content, with some entries expired for over 4 minutes. We need a way to explicitly purge cache across all CDN instances rather than waiting for user requests to trigger refreshes.
🙇♂️ THE SOLUTION: DNS lookups to discover all instances and coordinate purges across them.
🧑🔧 THE DNS THING (02:26): Using Fly.io's internal DNS system, we implement IPv6 lookups to discover all CDN instances. Claude Code helps. It works the first time 🎉
🐞 BUGS & PREDATORS (3:00): Undefined function errors, mock testing issues, and the classic double-slash URL bug. No Claude Code, just experience.
🧐 DID IT WORK? (6:56): Let's force push straight into production and see what happens 😱
This session isn't just a tutorial - it's an authentic look at collaborative coding, where DNS magic meets Elixir elegance to conquer distributed caching.
If you're into DevOps, CDNs, or just love seeing bugs squashed in real time, this gives you the tools to understand (and maybe implement) your own purge system.
00:00 The Problem
00:46 The Solution
02:26 The DNS thing
03:00 Bugs & Predators
04:43 Let's fire this up
06:56 Let's force push it
07:25 Is it working?
08:53 No no no
Open-source Varnish HTTP Cache does not support TLS backends. What does the simplest solution look like? Why are we even doing this? And "Why no SSL?"
We go all the way back to 2011 to understand the reasons behind this, learn about PHK, The Bikeshed and conclude what's next for Pipely, "the 20-line Varnish config that we deploy around the world and: Look at our CDN!"
This is a follow-up to youtube.com/watch?v=8bDgWvyglno
00:00 What is Pipely?
00:34 Three years ago
00:57 Resolving the Varnish TLS issue
03:41 Hitch
05:44 Why no SSL?
07:17 Who is PHK?
09:10 The Bikeshed
But here's the thing: technical expertise alone won't make you remarkable. It's your dedication to perfecting your craft that creates true mastery. The greatest experts, like Jiro who's been making sushi daily for 80 years straight, understand something most technologists never grasp.
This is an excerpt from the 🍱 DevOps Sushi - Part 1 🎬 movie that is now available on https://makeitwork.tv. Join as a member & put this conversation into context by watching the full-length movie.
☸️ 4-node Kubernetes cluster built with affordable refurbished hardware
💻 Control plane running on a ThinkPad T430
🗃️ Three worker nodes using HP EliteDesks (~€170 each)
🐧 Minimalist Arch Linux setup with Wayland and Hyperland
🏡 Homepage as his central dashboard for all services
🦘 Wallabag for saving and storing articles
👔 Linkding for browser-agnostic bookmark management
🐘 PostgreSQL databases for data persistence
This is an excerpt from the 🍱 DevOps Sushi - Part 1 🎬 movie that is now available on https://makeitwork.tv. Join as a member & put this conversation into context by watching the full-length movie.
When Mischa van den Burg created his first Kubernetes deployment with three (3) replicas, killed a container, and watched it automatically resurrect itself, he knew that this was the setup he really wanted. That feeling of seeing a rolling update happen flawlessly is something that you can't understand until you experience it first-hand. The automatic rollback when something goes wrong is the true epiphany.
If you want a homelab that runs as reliably as professional cloud infrastructure, you need to consider what begins as overkill, but starts making more sense after some hands-on experience.
This is an excerpt from the 🍱 DevOps Sushi - Part 1 🎬 movie that is now available on https://makeitwork.tv. Join as a member & put this conversation into context by watching the full-length movie.
Hugo Santos shows us Foundation, an open-source Kubernetes application platform that powers Namespace's developer-optimized compute platform. We see a testing framework that has a:
- Service-Oriented Architecture: The system manages how services are written, built, tested, deployed, and monitored in production
- Container-Centric Approach: Everything runs in containers within Kubernetes clusters
- Fast Test Execution: Tests that would normally be slow are dramatically accelerated through parallelization
- Complete Isolation: Each test runs in its own dedicated Kubernetes cluster for complete isolation
TECHNICAL INFRASTRUCTURE:
1. Kubernetes Distribution: Uses customized k3s for speed
2. Container Runtime: Custom-managed ContainerD (version 1.7)
3. Operating System: Wolfi-based with custom kernel (6.7)
4. Deployment Speed: New Kubernetes clusters become ready in under 2 seconds
5. Scale: Can easily manage dozens of parallel test environments
If you want to watch the full conversation, as a 4k movie, go to 📺 makeitwork.tv/fast-infrastructure and join as a member.
Alternatively, you can listen to the conversation at 🎧 https://makeitwork.fm/12
There are 2 other videos that go well with this one:
youtube.com/watch?v=zDJ-EE1f82w
youtube.com/watch?v=BSD-E8I7JNE
- First, he runs a local build using yarn (takes ~10 seconds)
- Then packages the build into a container using his local Docker setup
- He then compares this with a remote Docker build on Namespace's infrastructure.
Gerhard then showcases Namespace's preview environment capabilities:
- Using nsc run creates temporary deployment environments in seconds (~7 seconds)
- These environments get unique URLs for testing
- By default, previews require authentication, but can be made public with --ingress *=noauth
- Preview instances automatically get cleaned up after a default duration (4 hours)
Some highlights discussed:
- Extremely fast container builds in the cloud
- Quick preview environment creation
- Namespace's infrastructure optimizes for throughput
- Upcoming features like deferred starting (scale-to-zero) that only spins up when someone accesses the URL
The conversation includes interesting implementation details:
- Namespace uses advanced caching for fast instance creation (~50-100ms for cloning cache volumes)
- Deployments use global capacity planning rather than region selection
- Preview environments support various protocols including TCP, TLS passthrough, and TLS termination
If you want to watch the full conversation, as a 4k movie, go to 📺 makeitwork.tv/fast-infrastructure and join as a member.
Alternatively, you can listen to the conversation at 🎧 https://makeitwork.fm/12
There are 2 other videos that go well with this one:
youtube.com/watch?v=Um77B1RiLaI
youtube.com/watch?v=BSD-E8I7JNE
If you want to watch the full conversation, as a 4k movie, go to 📺 makeitwork.tv/fast-infrastructure and join as a member.
Alternatively, you can listen to the conversation at 🎧 https://makeitwork.fm/12
There are 2 other videos that go well with this one:
youtube.com/watch?v=Um77B1RiLaI
youtube.com/watch?v=zDJ-EE1f82w
1. Set up Keep with a MySQL database using Docker
2. Create a workflow that:
- Receives alerts
- Queries a MySQL database
- Sends webhooks based on query results
Tal Borenstein sprinkles a high level perspective, and helps fix a misconfiguration.
Key features highlighted:
- Workflows can be defined in YAML (similar to GitHub Actions syntax)
- Supports multiple trigger types: alerts, incidents, manual triggers, and intervals
- Includes deduplication and correlation of alerts
- Integrates with external systems via "providers" (MySQL, HTTP webhooks, etc.)
Orbstack & webhook.site are worth exploring on your own.
Listen to the conversation before & after this demo 🎧 https://makeitwork.fm/11
If you enjoyed this content and want to watch the full movie in 4k, go to 📺 https://makeitwork.tv. Offline download is available.
00:16 Setup Keep locally
01:21 Orbstack
02:20 Our first alert
03:07 What makes Keep different
04:30 Let me try to understand this
06:21 Our first workflow
13:29 GitHub Actions for monitoring tools
19:50 I can't believe it worked
22:43 Something bad happened
28:01 Reporting capabilities
Most of this is captured in 🐙 github.com/thechangelog/changelog.com/pull/518
We made IAD the primary region (it's the closest one to the origin), discussed VARNISH_SIZE and exposed environment variables as custom HTTP headers. We then struggle with IPv6 DNS resolution. At the end, we are surprised when this starts working all of a sudden.
Listen to the conversation before the screen share 🎧 https://makeitwork.fm/10
If you enjoyed this content and want to watch the full movie in 4k, go to 📺 https://makeitwork.tv. Offline download is available.
00:00 Intro
00:57 Who is James?
02:12 Who is Matt?
03:23 What are we trying to achieve today?
04:26 GitHub Pull Request 518
05:53 Make IAD the primary region
06:31 How do we think about VARNISH_SIZE?
09:43 Expose env var as custom HTTP header
12:15 503 backend fetch failed
22:59 We could create a GitHub issue...
24:27 Why does this work?
We look at Go code, discuss procedural (imperative) vs. declarative, spend some time on state management & introduce the concept of Ninjas in the context of infrastructure: move fast & break nothing.
In the second half, Matias uses diagrams to talk through different ideas of rolling this out into production. Which of the two approaches would you choose?
Testing Pulumi programs: pulumi.com/docs/iac/concepts/testing
Listen to the audio version 🎧 https://makeitwork.fm/9
The full content experience 📺 makeitwork.tv
00:00 Why Pulumi instead of Terraform?
02:40 Procedural or declarative?
07:19 What is Ninjastructure?
08:47 First thing that gets provisioned in an AWS account
11:18 How does the network module work?
14:29 Biggest advantage to using Pulumi over Terraform
17:02 Stacks = different environments
18:20 Where is state stored?
20:18 Where did you choose to store the state?
21:46 How to use this in production?
24:32 The GitOps approach
29:17 Outro
Listen to 3 other conversations from TalosCon 2024 🎙 https://makeitwork.fm/8
00:00 Intro
01:06 How this started
01:55 What are the two things wrong?
02:48 LattePanga Sigma
03:29 Portable homelab
04:15 What is Omni?
04:42 How do I bootstrap my homelab?
06:20 What is running on this homelab?
08:15 Daggerverse on the Tailnet
09:36 How many requests per second?
13:40 Can we squeeze more requests per second?
14:45 300B requests per month on a K8s phone
15:09 Why Single Node?
16:33 Why Kubernetes?
17:29 Let's scale this homelab into production
19:02 What is Dagger?
20:55 The digitalocean function
24:22 dagger call digitalocean
28:38 We have a problem
29:49 How to bootstrap db from backup
30:59 Production network benchmark
31:34 Takeaways
34:22 Outro
35:53 Questions from the audience
40:39 One more thing


