Uploaded July 2026 | Updated September 2026, 2 hours ago
AI is pushing more code into delivery workflows, but review capacity is not moving at the same speed. The question is what teams need between a plausible diff and a merge decision: proof from the running app.
Watch the recording of a webinar, in which Callstack’s Michał Pierzchała, Principal Engineer, and Piotr Miłkowski, Senior AI System Engineer, break down why code review still misses runtime behavior, why agentic code review remains too close to code, and why end-to-end suites often do not answer the exact question attached to a pull request.
They also show what behavior review can look like in practice: a bounded QA agent that checks the flow touched by a PR, inspects the app at runtime, and returns evidence a reviewer can use. The session stays grounded in real delivery work, with Agent Device presented as the infrastructure layer that gives the agent access to the running app.
Once you watch it, you leave with a clearer view of where this fits in an engineering workflow, what makes it trustworthy, and where to start if your team wants to test it without redesigning the whole QA stack.
What you'll learn from the webinar:
🔹 why QA becomes the bottleneck when AI increases code output
🔹 why code review and E2E tests still leave blind spots at runtime
🔹 what a bounded QA agent should verify on each PR
🔹 what makes behavior review reliable enough for production use
🔹 how to build your first own agent for app behavior review on every PR
Explore AI-Driven QA & Testing: clstk.com/3RaIexP
Check out Agent Device: github.com/callstackincubator/agent-device
Chapters
00:00 Meet the speakers and the webinar focus
04:20 When a green PR breaks on a real device
06:35 Why code review cannot keep up alone
09:09 Why agentic review still misses the product
11:04 Behavior review: a new PR review surface
12:26 Building a bounded QA agent
25:19 Live pull-request behavior-review demo
33:39 Five rules for trustworthy behavior review
38:44 Start with one workflow and one product area
40:10 Q&A: adoption, models, cost, local setup, and real devices
AI is pushing more code into delivery workflows, but review capacity is not moving at the same speed. The question is what teams need between a plausible diff and a merge decision: proof from the running app.
Watch the recording of a webinar, in which Callstack’s Michał Pierzchała, Principal Engineer, and Piotr Miłkowski, Senior AI System Engineer, break down why code review still misses runtime behavior, why agentic code review remains too close to code, and why end-to-end suites often do not answer the exact question attached to a pull request.
They also show what behavior review can look like in practice: a bounded QA agent that checks the flow touched by a PR, inspects the app at runtime, and returns evidence a reviewer can use. The session stays grounded in real delivery work, with Agent Device presented as the infrastructure layer that gives the agent access to the running app.
Once you watch it, you leave with a clearer view of where this fits in an engineering workflow, what makes it trustworthy, and where to start if your team wants to test it without redesigning the whole QA stack.
What you'll learn from the webinar:
🔹 why QA becomes the bottleneck when AI increases code output
🔹 why code review and E2E tests still leave blind spots at runtime
🔹 what a bounded QA agent should verify on each PR
🔹 what makes behavior review reliable enough for production use
🔹 how to build your first own agent for app behavior review on every PR
Explore AI-Driven QA & Testing: clstk.com/3RaIexP
Check out Agent Device: github.com/callstackincubator/agent-device
Chapters
00:00 Meet the speakers and the webinar focus
04:20 When a green PR breaks on a real device
06:35 Why code review cannot keep up alone
09:09 Why agentic review still misses the product
11:04 Behavior review: a new PR review surface
12:26 Building a bounded QA agent
25:19 Live pull-request behavior-review demo
33:39 Five rules for trustworthy behavior review
38:44 Start with one workflow and one product area
40:10 Q&A: adoption, models, cost, local setup, and real devices





